Autonomous search method combined with semantics in large-scale unknown environment and related equipment
By constructing a semantic octree and a Gaussian mixture model combined with a large language model to optimize the search path, the problem of inefficient target search in large-scale unknown environments is solved, adaptive and efficient target positioning is achieved, and the security and continuity of the search process are ensured.
Patent Information
- Application Number
- CN202510684881.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies are inefficient in target search in large-scale unknown environments and have difficulty coping with diverse tasks and dynamic environmental changes. Especially in outdoor scenarios, the correlation between targets and the environment is ambiguous, and existing methods lack security and continuity.
A semantic octree is constructed by combining the Bayesian probability method, and a Gaussian mixture model is used to analyze the panoramic image to determine the target location probability. The search path is optimized through a large language model and task probability graph, and the robot search trajectory is generated. The semantic octree and multi-layer task probability graph are combined to optimize the search path trajectory.
It achieves adaptive and efficient target search in large-scale unknown environments, improves the real-time and efficiency of robot search, and ensures the safety and continuity of the search process.
Smart Images

Figure CN120635858A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot navigation technology, in particular to an autonomous search method and related equipment combined with semantics in a large-scale unknown environment. Background Art
[0002] While existing technologies have made some progress in object search tasks, they still have significant shortcomings. Traditional methods rely on pre-designed search strategies, making them difficult to adapt to diverse tasks and dynamic environments, resulting in poor performance in complex scenarios. While neural network and reinforcement learning-based methods show promise, their generalization is hampered by limited training data, hindering their efficient application in large-scale environments. Large language models (LLMs) excel in reasoning and decision-making, but they still have significant limitations in task-specific training, leveraging historical experience, and adapting to dynamic environments. For example, task-specific training requires extensive human input, which is time-consuming and labor-intensive. LLMs also have limited ability to store and access historical experience, significantly reducing their efficiency when performing repeated similar search tasks. As the scale of environments increases, especially in outdoor scenarios, the correlation between the object and its surroundings or topological semantics becomes blurred, reducing the efficiency of LLM-based search methods. Furthermore, in dynamic environments, robots must navigate obstacles in real time, a challenge for existing methods, making it difficult to ensure a safe and continuous search process. These shortcomings limit the practical application of existing technologies in large-scale and unknown environments. Summary of the Invention
[0003] In view of this, embodiments of the present application provide an autonomous search method and related equipment combined with semantics in a large-scale unknown environment to achieve adaptive and efficient target search.
[0004] One aspect of an embodiment of the present application provides a semantically-integrated autonomous search method in a large-scale unknown environment, the method comprising the following steps:
[0005] Constructing a semantic octree based on a Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and the semantic labels of the objects in the panoramic image;
[0006] Analyzing the panoramic image using a Gaussian mixture model to determine the probability of a target location, and then constructing a multi-layer task probability map based on the probability of the target location;
[0007] Obtaining a task description and determining whether the search task corresponding to the task description is a repeated task;
[0008] If not, a large language model is used to determine a search direction based on the task description and the panoramic image, and then determine a local target point;
[0009] If so, determining the search direction according to the semantic octree and the multi-layer task probability graph, and then determining the local target point;
[0010] generating a search path for the local target point, and optimizing a trajectory of the search path according to the semantic octree and the multi-layer task probability graph;
[0011] The robot is controlled to move and search according to the trajectory.
[0012] In some embodiments, constructing a semantic octree based on a Bayesian probability method comprises the following steps:
[0013] Dynamically acquiring environmental information; wherein the environmental information includes the panoramic image and the point cloud data;
[0014] Segmenting the panoramic image and generating semantic masks, and assigning semantic labels to various objects in the panoramic image;
[0015] After the panorama segmentation is completed, the point cloud data is converted and projected onto an equidistant cylindrical projection aligned with the panorama coordinate system to align the point cloud data with the semantic labels, thereby generating a semantic point cloud;
[0016] The semantic point cloud is gradually mapped to the semantic octree.
[0017] In some embodiments, analyzing the panoramic image using a Gaussian mixture model to determine the probability of the target location includes the following steps:
[0018] Determining a Gaussian distribution of each Gaussian component; wherein each Gaussian component represents a probability of finding a target object at a target position in the panoramic image;
[0019] The probability density function of each Gaussian component is defined as:
[0020] ;
[0021] in, represents the spatial coordinates, It is The mean of the Gaussian components represents the observed object position, is the covariance matrix, quantifying the uncertainty of the target object’s position;
[0022] Constructing the Gaussian mixture model according to the weighted sum of the probabilities provided by the respective Gaussian components;
[0023] The weighted sum is:
[0024] ;
[0025] in, Indicates the The weights of the Gaussian components.
[0026] In some embodiments, the method further comprises the following steps:
[0027] When the search task is successfully completed and the location When the target object is found, the Gaussian mixture model is updated by the following steps:
[0028] Add a new Gaussian component to the Gaussian mixture model , the parameters are defined as:
[0029] ;
[0030] in, is the number of Gaussian components present in the Gaussian mixture model, Proportional to the size of the target object, indicating the uncertainty of the target object's position;
[0031] If the new Gaussian component With existing weight Euclidean distance between If it is less than the sum of the corresponding standard deviations, then merge and ; The combined Gaussian components The parameters are:
[0032] ;
[0033] After merging, the new Gaussian components discarded;
[0034] After adding a new Gaussian component or updating the Gaussian mixture model, the weights of all the Gaussian components are normalized as follows:
[0035] .
[0036] In some embodiments, the step of using a large language model to determine a search direction based on the task description and the panoramic image includes the following steps:
[0037] Dividing the panoramic image into multiple sub-images, each sub-image corresponding to a candidate search direction;
[0038] Outputting a list of candidate propositions based on the task description and each subgraph using the large language model; wherein the list of candidate propositions includes corresponding decisions;
[0039] Calculating the cost function value of each of the decisions;
[0040] The decision with the smallest cost function value is selected as the target search direction.
[0041] In some embodiments, calculating the cost function value of each candidate search direction includes the following steps:
[0042] Calculate the safety cost, invalid exploration cost and direction change cost of each decision in turn;
[0043] The security cost is:
[0044] ;
[0045] in, is the security cost, is a predefined safety distance, is the shortest distance to the obstacle;
[0046] The invalid exploration cost is:
[0047] ;
[0048] in, is the ineffective exploration cost, and Represent the area of the effective exploration coverage area and the total area covered by the sensor respectively;
[0049] The direction change cost is:
[0050] ;
[0051] in, is the direction change cost, It is a decision The corresponding direction, is the robot’s current moving direction;
[0052] Determining the cost function value according to the safety cost, the invalid exploration cost, and the direction change cost;
[0053] The expression of the cost function value is:
[0054] ;
[0055] in, is the cost function value, is the weight corresponding to each cost; Representing a decision The candidate proposition list The original index in .
[0056] In some embodiments, determining the search direction according to the semantic octree and the multi-layer task probability map, and then determining the local target point, includes the following steps:
[0057] The problem of determining the local target point in the multi-layer task probability graph is expressed as a first traveling salesman problem, and the local target point in the multi-layer task probability graph is obtained by solving the first traveling salesman problem;
[0058] The first traveling salesman problem is:
[0059] ;
[0060] Among them, multiple potential local target points are represented as , Indicates the current position of the robot , ( ) indicates the The spatial coordinates of the Gaussian components, Represents the Euclidean distance between two local target points, evaluated using the A* algorithm in the semantic octogram; It is The weights of the Gaussian components, is the scaling factor;
[0061] Expressing the problem of determining the local target point in the semantic octree as a second traveling salesman problem, and solving the second traveling salesman problem to obtain the local target point in the semantic octree;
[0062] The second traveling salesman problem is:
[0063] .
[0064] Another aspect of the present application further provides an autonomous search device incorporating semantics in a large-scale unknown environment, the device comprising:
[0065] A semantic octree construction unit, configured to construct a semantic octree based on a Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and semantic labels of objects in the panoramic image;
[0066] a multi-layer task probability map construction unit, configured to analyze the panoramic image using a Gaussian mixture model to determine the probability of a target location, and then construct a multi-layer task probability map based on the probability of the target location;
[0067] A repeated task determination unit, configured to obtain a task description and determine whether the search task corresponding to the task description is a repeated task;
[0068] a first search unit, configured to, if not, determine a search direction using a large language model according to the task description and the panoramic image, and further determine a local target point;
[0069] a repeated search unit, configured to, if yes, determine the search direction according to the semantic octree and the multi-layer task probability map, and further determine the local target point;
[0070] a trajectory generation and optimization unit, configured to generate a search path for the local target point and optimize the trajectory of the search path according to the semantic octree and the multi-layer task probability graph;
[0071] The search control unit is used to control the robot to move and search according to the trajectory.
[0072] Another aspect of the embodiments of the present application further provides an electronic device, including a processor and a memory;
[0073] The memory is used to store programs;
[0074] The processor executes the program to implement any of the above methods.
[0075] Another aspect of the embodiments of the present application further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement any of the above methods.
[0076] This application has at least the following beneficial effects:
[0077] The present application can construct a semantic octree based on the Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and the semantic labels of the objects in the panoramic image; the panoramic image is analyzed using a Gaussian mixture model to determine the probability of the target position, and then a multi-layer task probability map is constructed based on the probability of the target position; the task description is obtained and it is determined whether the search task corresponding to the task description is a repeated task; if not, the search direction is determined based on the task description and the panoramic image using a large language model, and then the local target point is determined; if so, the search direction is determined based on the semantic octree and the multi-layer task probability map, and then the local target point is determined; a search path for the local target point is generated, and the trajectory of the search path is optimized based on the semantic octree and the multi-layer task probability map; the robot is controlled to move and search based on the trajectory. The present application, combined with the semantic octree and the multi-layer task probability map, can quickly and accurately determine the robot's search trajectory, which can improve real-time performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0079] Figure 1 A flowchart of a method for autonomous search in a large-scale unknown environment combined with semantics provided by an embodiment of the present application;
[0080] Figure 2 An example flow chart of an autonomous search method combining semantics in a large-scale unknown environment provided by an embodiment of the present application;
[0081] Figure 3 This is an example diagram of the reasoning analysis of a panoramic image captured by a robot during an exploration process and the corresponding LLM provided in an embodiment of the present application;
[0082] Figure 4 This is an example diagram of the reasoning analysis of the panoramic image captured by the robot and the LLM corresponding to another exploration process provided by an embodiment of the present application;
[0083] Figure 5 This is a structural block diagram of an autonomous search device combining semantics in a large-scale unknown environment provided by an embodiment of the present application. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0085] Before describing the embodiments of the present application in detail, some of the related technologies involved in the embodiments of the present application are first described as follows:
[0086] Large Language Model (LLM).
[0087] Gaussian Mixture Models (GMM).
[0088] Traveling Salesman Problem (TSP).
[0089] Object search in an unknown environment refers to the process of a robot using autonomous navigation and perception technologies to find and locate a specific target, without prior knowledge or complete understanding of the environment. Such tasks often involve complex environmental dynamics, uncertainty, and ambiguity in the target's location. Examples include searching for survivors in disaster relief, finding specific items in a warehouse, or locating a specific location in an outdoor environment. The core challenge of object search lies in how to complete the task efficiently and accurately within limited information and resources.
[0090] In target search tasks, traditional methods mainly rely on pre-designed search strategies, neural network-based methods, and reinforcement learning algorithms. Pre-designed search strategies are usually based on rules and heuristic algorithms, such as breadth-first search (BFS) or depth-first search (DFS), which are suitable for structured environments but have limited performance in complex and dynamic scenarios. Neural network-based methods train models with large amounts of data and can adapt to a variety of tasks to a certain extent, but their generalization ability is limited by the quality and scale of the training data. Reinforcement learning algorithms continuously optimize strategies through interaction with the environment and perform well in specific tasks, but face challenges in computational complexity and convergence speed in large-scale and unknown environments. Overall, traditional methods often exhibit problems such as insufficient adaptability, low efficiency, and limited decision-making accuracy when dealing with complex and dynamic environments.
[0091] The emergence of large language models (LLMs) provides a new solution for target search tasks. LLMs can combine multimodal inputs (such as images, speech, and sensor data) to achieve a deep understanding of complex environments. For example, by analyzing environmental images and contextual information, LLMs can infer the location of potential targets and dynamically adjust search strategies. In addition, LLMs can dynamically optimize decisions based on real-time environmental information to improve search efficiency. In dynamic environments, LLMs can plan paths in real time, avoid obstacles, and ensure the safety and continuity of the search process. At the same time, by storing and accessing historical experience, LLMs can quickly learn and optimize strategies when repeatedly performing similar tasks, significantly improving the efficiency and adaptability of robots in target search tasks. Overall, LLMs provide more powerful technical support for target search tasks, especially showing significant advantages in complex and unknown environments.
[0092] Reference Figure 1 The embodiment of the present application provides an autonomous search method combining semantics in a large-scale unknown environment, specifically comprising the following steps S100 to S160:
[0093] S100: Constructing a semantic octree based on a Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and semantic labels of objects in the panoramic image;
[0094] S110: Analyzing the panoramic image using a Gaussian mixture model to determine the probability of a target location, and then constructing a multi-layer task probability map based on the probability of the target location;
[0095] S120: Obtain a task description and determine whether the search task corresponding to the task description is a repeated task;
[0096] S130: If not, using a large language model to determine a search direction according to the task description and the panoramic image, and then determine a local target point;
[0097] S140: If yes, determining the search direction according to the semantic octree and the multi-layer task probability map, and then determining the local target point;
[0098] S150: Generate a search path for the local target point, and optimize the trajectory of the search path according to the semantic octree and the multi-layer task probability graph;
[0099] S160: Control the robot to move and search according to the trajectory.
[0100] Optionally, constructing a semantic octree based on a Bayesian probability method comprises the following steps:
[0101] Dynamically acquiring environmental information; wherein the environmental information includes the panoramic image and the point cloud data;
[0102] Segmenting the panoramic image and generating semantic masks, and assigning semantic labels to various objects in the panoramic image;
[0103] After the panorama segmentation is completed, the point cloud data is converted and projected onto an equidistant cylindrical projection aligned with the panorama coordinate system to align the point cloud data with the semantic labels, thereby generating a semantic point cloud;
[0104] The semantic point cloud is gradually mapped to the semantic octree.
[0105] Optionally, analyzing the panoramic image using a Gaussian mixture model to determine the probability of the target location includes the following steps:
[0106] Determining a Gaussian distribution of each Gaussian component; wherein each Gaussian component represents a probability of finding a target object at a target position in the panoramic image;
[0107] The probability density function of each Gaussian component is defined as:
[0108] ;
[0109] in, represents the spatial coordinates, It is The mean of the Gaussian components represents the observed object position, is the covariance matrix, quantifying the uncertainty of the target object’s position;
[0110] Constructing the Gaussian mixture model according to the weighted sum of the probabilities provided by the respective Gaussian components;
[0111] The weighted sum is:
[0112] ;
[0113] in, Indicates the The weights of the Gaussian components.
[0114] Optionally, the method further comprises the following steps:
[0115] When the search task is successfully completed and the location When the target object is found, the Gaussian mixture model is updated by the following steps:
[0116] Add a new Gaussian component to the Gaussian mixture model , the parameters are defined as:
[0117] ;
[0118] in, is the number of Gaussian components present in the Gaussian mixture model, Proportional to the size of the target object, indicating the uncertainty of the target object's position;
[0119] If the new Gaussian component With existing weight Euclidean distance between If it is less than the sum of the corresponding standard deviations, then merge and ; The combined Gaussian components The parameters are:
[0120] ;
[0121] After merging, the new Gaussian components discarded;
[0122] After adding a new Gaussian component or updating the Gaussian mixture model, the weights of all the Gaussian components are normalized as follows:
[0123] .
[0124] Optionally, the using a large language model to determine a search direction according to the task description and the panoramic image includes the following steps:
[0125] Dividing the panoramic image into multiple sub-images, each sub-image corresponding to a candidate search direction;
[0126] Outputting a list of candidate propositions based on the task description and each subgraph using the large language model; wherein the list of candidate propositions includes corresponding decisions;
[0127] Calculating the cost function value of each of the decisions;
[0128] The decision with the smallest cost function value is selected as the target search direction.
[0129] Optionally, calculating the cost function value of each candidate search direction includes the following steps:
[0130] Calculate the safety cost, invalid exploration cost and direction change cost of each decision in turn;
[0131] The security cost is:
[0132] ;
[0133] in, is the security cost, is a predefined safety distance, is the shortest distance to the obstacle;
[0134] The invalid exploration cost is:
[0135] ;
[0136] in, is the ineffective exploration cost, and Represent the area of the effective exploration coverage area and the total area covered by the sensor respectively;
[0137] The direction change cost is:
[0138] ;
[0139] in, is the direction change cost, It is a decision The corresponding direction, is the current moving direction of the robot;
[0140] Determining the cost function value according to the safety cost, the invalid exploration cost, and the direction change cost;
[0141] The expression of the cost function value is:
[0142] ;
[0143] in, is the cost function value, is the weight corresponding to each cost; Representing a decision The candidate proposition list The original index in .
[0144] Optionally, determining the search direction according to the semantic octree and the multi-layer task probability map, and then determining the local target point, includes the following steps:
[0145] The problem of determining the local target point in the multi-layer task probability graph is expressed as a first traveling salesman problem, and the local target point in the multi-layer task probability graph is obtained by solving the first traveling salesman problem;
[0146] The first traveling salesman problem is:
[0147] ;
[0148] Among them, multiple potential local target points are represented as , Indicates the current position of the robot , ( ) indicates the The spatial coordinates of the Gaussian components, Represents the Euclidean distance between two local target points, evaluated using the A* algorithm in the semantic octogram; It is The weights of the Gaussian components, is the scaling factor;
[0149] Expressing the problem of determining the local target point in the semantic octree as a second traveling salesman problem, and solving the second traveling salesman problem to obtain the local target point in the semantic octree;
[0150] The second traveling salesman problem is:
[0151] .
[0152] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.
[0153] The target search framework GET proposed in this embodiment is mainly composed of four modules: perception, decision-making, trajectory refinement and memory storage. Each module works together to achieve adaptive and efficient target search. The solution structure diagram is shown in the figure below. Figure 2 shown.
[0154] Figure 2 Schematic diagram of the GET solution, which integrates perception, memory, decision-making, and trajectory optimization modules. The perception module processes panoramic images and point clouds into semantic point clouds, providing a structured understanding of the environment. The memory model consists of a semantic octree and a multi-level task probability graph that are updated in real time, explicitly recording environmental data and historical task experience. The decision module operates in two modes: reasoning-based search and experience-based search. In reasoning-based search, the DoUT module serves as the core, inferring potential target locations from task descriptions and environmental panoramas, while providing feedback to the LLM through autonomous learning to improve reasoning accuracy. In experience-based search, the robot locates the target using historical data and the semantic octree, and plans the search route by constructing and solving the TSP problem. Finally, the trajectory optimization module optimizes the robot's path to achieve smooth and efficient navigation, seamlessly integrating real-time perception, reasoning, and motion planning.
[0155] The perception module uses panoramic cameras and lidar to acquire real-time environmental information, including panoramic images and point cloud data. After semantic segmentation, the panoramic image is semantically mapped to the point cloud to generate a semantic point cloud. This is efficiently stored in an octree map, supporting incremental updates and fast queries. The panoramic image is also used to determine whether the target has been found. Once a target is found, its location information is updated in the multi-layer task probability map, accelerating the execution of repetitive tasks.
[0156] The decision module is the core of GET. It is categorized into two types: first-time search and repeated search, depending on the nature of the search task. For first-time search tasks, a large language model comprehensively considers the input task description and the real-time panoramic image, inferring the potential location of the target object and generating a series of candidate proposals. A scoring mechanism is then designed to score and rank these proposals, generating ranked proposals that provide the robot with reference options for dynamic decision-making. For repeated search tasks, historical experience is retrieved from the memory storage module to construct a Transitional Plan (TSP), guiding the robot through possible locations of the target, further accelerating the search process.
[0157] Based on the output of the decision module and the distribution of the robot's surrounding environment, the trajectory refinement module searches for and optimizes feasible paths, generating safe, smooth, and consistent search trajectories that conform to the robot's dynamic model, ensuring the continuity and stability of the movement process. This module also responds to dynamic obstacles (such as vehicles and pedestrians) in real time, dynamically replanning the path to ensure the robot's safety.
[0158] The memory storage module consists of a real-time updated semantic octree map and a multi-layered task probability map. The semantic octree map records environmental data, providing critical support for the evaluation process; the task probability map, on the other hand, stores successful search experiences, with each layer dedicated to a specific target object. By combining the semantic octree map and the task probability map, the robot can fully leverage current environmental data and historical experience to more quickly locate targets in repetitive tasks, and continuously improve search efficiency as task experience accumulates.
[0159] More specifically, the scheme of this embodiment is as follows:
[0160] 1. Semantic octree.
[0161] To construct an accurate semantic octree, the environment panorama and the lidar point cloud are first synchronized in real time. Subsequently, Grounded-SAM combined with RAM is used to semantically segment the panorama, automatically generating semantic masks and assigning labels to various objects in the environment. After segmentation, the point cloud is transformed and projected onto an equidistant cylindrical projection aligned with the panorama coordinate system, ensuring precise alignment of the point cloud and semantic labels, thereby generating a semantic point cloud. Finally, the semantic point cloud is gradually mapped into the semantic octree.
[0162] Taking into account sensor noise and potential errors in the automatic annotation process, the semantic octree is constructed using a Bayesian probability-based approach to update the occupancy status and semantic labels of the octree cells. This approach supports incremental construction and updating of the map, ensuring that the environmental representation becomes more accurate over time.
[0163] 2. Task probability map.
[0164] Because objects are often located at specific locations, a probabilistic approach is advantageous for representing the likelihood of an object's location based on historical search tasks. In this embodiment, a Gaussian mixture model (GMM) is used to analyze past experience and capture the probability of a target's location. Each Gaussian component represents the probability of finding the object at a specific location. This model is favored because it supports incremental updates, ensuring real-time data without requiring extensive reprocessing, and its compact structure significantly reduces storage requirements.
[0165] 2.1. Gaussian mixture model.
[0166] The Gaussian mixture model is a probability model that represents a mixture of multiple Gaussian distributions. Each component corresponds to a Gaussian distribution and represents the probability of encountering a target at a specific location. The probability density function of each component is defined as:
[0167] (1.1)
[0168] in represents the spatial coordinates, It is The mean of the Gaussian components represents the observed position of the object, is the covariance matrix, quantifying the uncertainty of the object's position.
[0169] For the inclusion Gaussian mixture model of components, at a specific location The probability of finding an object is the weighted sum of the probabilities provided by each Gaussian component:
[0170] (1.2)
[0171] in Indicates the The weight of the components.
[0172] 2.2. Create and update Gaussian components.
[0173] When the search task is successfully completed and the location When the target object is found, the Gaussian mixture model is updated in the following two steps:
[0174] 1) Add a new Gaussian component:
[0175] Add a new Gaussian component to the model , whose parameters are defined as:
[0176] (1.3)
[0177] in is the number of existing components in the Gaussian mixture model, Proportional to the size of the object, it represents the uncertainty of its position.
[0178] 2.3. Merge Gaussian components.
[0179] If the new With existing weight Euclidean distance between If the sum of their standard deviations is less than the sum of their standard deviations, they are merged. The merged Gaussian components The parameters are:
[0180] (1.4)
[0181] After merging, the newly created component This merging process combines historical and recent data, providing a more accurate reflection of the probability distribution of objects in the search space.
[0182] 2.4. Gaussian component weight adjustment.
[0183] The weights of the Gaussian components reflect their importance in the overall Gaussian mixture model. Components with higher weights indicate a higher probability of encountering the target at those locations. After adding a new component or updating the Gaussian mixture model, the weights of all components are normalized as follows:
[0184] (1.5)
[0185] This normalization ensures that the total weight of all components remains equal to 1, maintaining the consistency of the model representation. The incremental update process enables the Gaussian mixture model to dynamically adapt to changes in the environment. With each new recorded experience, the Gaussian mixture model reflects the latest information about the location of the target object while gradually weakening the influence of outdated data. This adaptive method achieves an efficient and compact representation of the search experience using parameters , thereby reducing the maintenance cost of large-scale environment maps.
[0186] 3. Goal-oriented search strategy.
[0187] 3.1. Reasoning Search
[0188] At each time step , robot position The panoramic camera captures a panoramic view of the environment Panorama is divided into multiple segments, denoted as ,like Figure 2 Each segment represents a potential direction of motion, which facilitates LLM reasoning and improves the interpretability of its decisions. To avoid oscillatory behavior in the decision-making process, it is recommended ,in This configuration ensures that the robot's current orientation is in the middle segment. On the contrary, when When it is an even number, the robot's direction falls between the two middle segments and The junction of the two segments can cause oscillations between them and disrupt smooth motion. Furthermore, because the environment panorama is captured in the robot's local coordinate system, the LLM can reason directly based on the current state without complex coordinate transformations. This intuitive observation format improves the LLM's reasoning accuracy and simplifies the decision-making process.
[0189] For example, Figure 3 and Figure 4 The following are the panoramas captured by the robot during the two exploration processes and the corresponding reasoning analysis of LLM. Among them, the candidate proposal list of LLM reasoning is LLM preferred and , because there are vehicles in these areas, it is inferred that they may be parking areas; the candidate proposal list of LLM reasoning is LLM recommends exploring intersections first , because there are no areas within the current boundary and field of view that may contain vehicles.
[0190] The LLM then uses its reasoning ability to infer the direction that is most likely to contain the search object and outputs a list of candidate propositions. , the basis of LLM reasoning is as follows Figure 2 However, due to the possible errors in LLM reasoning and the distortion of the environment panorama under equirectangular projection, it is difficult for LLM to intuitively interpret the distance to the surrounding obstacles. The candidate proposition list output by LLM is While feasible, this approach still may contain incorrect sorting. If the robot were to execute the task directly according to this proposal, a safety incident could occur. Therefore, this embodiment designs an evaluation mechanism that interfaces with the LLM. Combined with the constructed environment octree semantic map, this mechanism performs a secondary evaluation of the candidate propositions output by the LLM and selects the optimal decision. The secondary evaluation process considers the following costs:
[0191] 1) Security cost:
[0192] This standard is based on the robot along the proposition When moving, the robot evaluates the distance from the obstacle. The closer the robot is to the obstacle, the higher the cost. To accurately evaluate the cost, the point cloud obtained by the robot from the lidar sensor is first converted to a spherical coordinate system. , and then the azimuth According to the same discretization into corresponding regions, each of which corresponds to a region after the panorama is segmented. The minimum distance of the point cloud in each region is then calculated. The distance from the area to the obstacle On this basis, the decision The safety cost Defined as:
[0193] (1.6)
[0194] in is a predefined safety distance. If , no penalty is applied. Otherwise, the penalty grows quadratically as the distance to the nearest obstacle decreases.
[0195] 2) Ineffective exploration cost:
[0196] When the robot searches for a target, if a new unknown area is covered by the sensor's sensing range, it is called effective exploration. On the contrary, if the sensor covers an area that has already been explored, it is called ineffective exploration. In order for the robot to search for the target as quickly as possible, it is necessary to allow the robot to conduct effective exploration as much as possible. Therefore, this item is used to penalize the robot's ineffective exploration. The cost is defined as:
[0197] (1.7)
[0198] in and They represent the area of the effective exploration coverage area and the total area covered by the sensors, respectively.
[0199] 3) Direction change cost:
[0200] During the robot's search for a target, frequent changes in direction will affect the robot's search efficiency and increase energy consumption. Therefore, this item is used to select decisions with smaller direction changes. This item is defined as:
[0201] (1.8)
[0202] in It is a decision The corresponding direction, is the current moving direction of the robot.
[0203] In summary, each candidate proposition The cost function of the allocation is defined as:
[0204] (1.9)
[0205] in is the weight that balances the impact of each criterion. express List of candidate propositions obtained in LLM reasoning The original index in encourages priority based on its initial order. Finally, calculate all The corresponding cost is re-arranged from small to large to obtain the optimal decision sequence , the robot will choose the strategy with the lowest cost as the next decision. If the decision cannot be made due to dynamic obstacles, the robot will choose the strategy with the second lowest cost as the decision, and so on.
[0206] Once the local goal is determined, the robot evaluates the feasibility of the direct path. If the direct path is blocked, the robot dynamically searches for an alternative path in the semantic octogram using the A* algorithm.
[0207] 3.2. Experience search.
[0208] When performing repeated search tasks in the same environment, semantic octograms and multi-layered task probability graphs are used to improve search efficiency. During the search process, the GET framework queries these maps to identify potential target locations. The problem is formulated as a traveling salesman problem (TSP), where the goal is to plan an efficient search path through multiple potential target areas.
[0209] 1) Target location in the multi-layer task probability map.
[0210] If the multi-layer task probability map contains the target object, the components of the corresponding GMM are used to generate a set of potential target locations, denoted as ,in Indicates the current position of the robot , ( ) indicates the The spatial coordinates of the Gaussian components. Then, TSP is expressed as:
[0211] (1.10)
[0212] in Represents the Euclidean distance between two target locations, evaluated using A* in the semantic octogram. It is the first The weights of the Gaussian components, is the scaling factor.
[0213] 2) Target location in the semantic octogram.
[0214] If the multi-layer task probability map does not contain the target object, then define the potential target location ,in Indicates the current position of the robot , ( ) is identified directly from the semantic octogram by querying regions labeled with the target category. In this case, TSP is formulated as:
[0215] (1.11)
[0216] Once the TSP is solved, the optimal search path is obtained Then, the robot can search for a path that passes through the semantic octogram. The robot then searches for the target location in the image, guiding it through the potential target area for an efficient search. If the target is not found after completing the path, the robot transitions to the inference search phase and repeats the process until it finds the target.
[0217] In summary, this embodiment has the following technical effects:
[0218] 1. Perception and semantic modeling: Obtain environmental information through panoramic cameras and lidar, combine semantic segmentation and semantic mapping to generate semantic point clouds, and use semantic octree maps for efficient storage and incremental updates.
[0219] 2. Dynamic decision-making mechanism: Exploration tasks are divided into two types: initial exploration and repeated exploration. In the initial search task, a large language model (LLM) is used to generate candidate proposals, and dynamic decisions are made through a scoring mechanism. In the repeated search task, historical experience and a multi-layer task probability graph are used to construct the traveling salesman problem (TSP) to plan the optimal search path.
[0220] 3. Utilize historical experience, explicitly record the probability distribution of target locations through multi-layer task probability graphs and Gaussian mixture models (GMM), support incremental updates and dynamic adjustments, and improve the efficiency of repeated searches.
[0221] In summary, this embodiment includes the following technical solutions:
[0222] 1. Multi-layer task probability graph and Gaussian mixture model, protecting the method of using Gaussian mixture model (GMM) to represent the probability distribution of target positions, as well as the creation and update mechanism of multi-layer task probability graph.
[0223] 2. Dynamic decision-making and scoring mechanism, protecting the LLM method of generating candidate proposals in the first search task, and the scoring mechanism that combines safety cost, invalid exploration cost and direction change cost.
[0224] 3. Utilization of historical experience and TSP planning, protecting the method of using historical experience and semantic octree maps to construct the traveling salesman problem (TSP), as well as the optimal search path planning technology.
[0225] The beneficial effects of this embodiment include:
[0226] 1. Through the large language model (LLM), comprehensive reasoning is performed on the task description and the environment panorama to generate candidate proposals, and a scoring mechanism is combined to make dynamic decisions, significantly improving the efficiency of the first search.
[0227] 2. Leveraging historical experience and multi-layer task probability graphs, we construct a traveling salesman problem (TSP) to plan the optimal search path, enabling the robot to quickly locate the target in repetitive tasks and reduce search time.
[0228] 3. Through real-time updates of the semantic octree map and multi-layer task probability map, the robot can dynamically adapt to environmental changes and ensure the continuity and stability of the search process.
[0229] 4. A scoring mechanism was designed to conduct a secondary evaluation of candidate proposals based on safety cost, invalid exploration cost, and direction change cost, avoiding LLM proposals with security risks and ensuring the safety and efficiency of decision-making.
[0230] Reference Figure 5 The embodiment of the present application provides an autonomous search device combining semantics in a large-scale unknown environment, including:
[0231] A semantic octree construction unit, configured to construct a semantic octree based on a Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and semantic labels of objects in the panoramic image;
[0232] a multi-layer task probability map construction unit, configured to analyze the panoramic image using a Gaussian mixture model to determine the probability of a target location, and then construct a multi-layer task probability map based on the probability of the target location;
[0233] A repeated task determination unit, configured to obtain a task description and determine whether the search task corresponding to the task description is a repeated task;
[0234] a first search unit, configured to, if not, determine a search direction using a large language model according to the task description and the panoramic image, and further determine a local target point;
[0235] a repeated search unit, configured to, if yes, determine the search direction according to the semantic octree and the multi-layer task probability map, and further determine the local target point;
[0236] a trajectory generation and optimization unit, configured to generate a search path for the local target point and optimize the trajectory of the search path according to the semantic octree and the multi-layer task probability graph;
[0237] The search control unit is used to control the robot to move and search according to the trajectory.
[0238] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0239] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0240] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present application as set forth in the claims using ordinary techniques without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0241] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0242] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0243] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0244] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0245] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0246] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.
[0247] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application, and these equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. An autonomous search method combining semantics in a large-scale unknown environment, characterized by: The method comprises the following steps: Constructing a semantic octree based on a Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and the semantic labels of the objects in the panoramic image; Analyzing the panoramic image using a Gaussian mixture model to determine the probability of a target location, and then constructing a multi-layer task probability map based on the probability of the target location; Obtaining a task description and determining whether the search task corresponding to the task description is a repeated task; If not, a large language model is used to determine a search direction based on the task description and the panoramic image, and then determine a local target point; If so, determining the search direction according to the semantic octree and the multi-layer task probability graph, and then determining the local target point; generating a search path for the local target point, and optimizing a trajectory of the search path according to the semantic octree and the multi-layer task probability graph; The robot is controlled to move and search according to the trajectory.
2. The autonomous search method combining semantics in a large-scale unknown environment according to claim 1 is characterized in that: The method of constructing a semantic octree based on the Bayesian probability method includes the following steps: Dynamically acquiring environmental information; wherein the environmental information includes the panoramic image and the point cloud data; Segmenting the panoramic image and generating semantic masks, and assigning semantic labels to various objects in the panoramic image; After the panorama segmentation is completed, the point cloud data is converted and projected onto an equidistant cylindrical projection aligned with the panorama coordinate system to align the point cloud data with the semantic labels, thereby generating a semantic point cloud; The semantic point cloud is gradually mapped to the semantic octree.
3. The autonomous search method combining semantics in a large-scale unknown environment according to claim 1 is characterized in that: The method of analyzing the panoramic image using a Gaussian mixture model to determine the probability of the target location includes the following steps: Determining a Gaussian distribution of each Gaussian component; wherein each Gaussian component represents a probability of finding a target object at a target position in the panoramic image; The probability density function of each Gaussian component is defined as: ; in, represents the spatial coordinates, It is The mean of the Gaussian components represents the observed position of the object, is the covariance matrix, quantifying the uncertainty of the target object’s position; Constructing the Gaussian mixture model according to the weighted sum of the probabilities provided by the respective Gaussian components; The weighted sum is: ; in, Indicates the The weights of the Gaussian components.
4. The autonomous search method combining semantics in a large-scale unknown environment according to claim 3 is characterized in that: The method further comprises the following steps: When the search task is successfully completed and the location When the target object is found, the Gaussian mixture model is updated by the following steps: Add a new Gaussian component to the Gaussian mixture model , the parameters are defined as: ; in, is the number of Gaussian components present in the Gaussian mixture model, Proportional to the size of the target object, indicating the uncertainty of the target object's position; If the new Gaussian component With existing weight Euclidean distance between If it is less than the sum of the corresponding standard deviations, then merge and ; The combined Gaussian components The parameters are: ; After merging, the new Gaussian components discarded; After adding a new Gaussian component or updating the Gaussian mixture model, the weights of all the Gaussian components are normalized as follows: 。 5. The autonomous search method combining semantics in a large-scale unknown environment according to claim 1 is characterized in that: The method of using a large language model to determine a search direction according to the task description and the panoramic view includes the following steps: Dividing the panoramic image into multiple sub-images, each sub-image corresponding to a candidate search direction; Outputting a list of candidate propositions based on the task description and each subgraph using the large language model; wherein the list of candidate propositions includes corresponding decisions; Calculating the cost function value of each of the decisions; The decision with the smallest cost function value is selected as the target search direction.
6. The autonomous search method combining semantics in a large-scale unknown environment according to claim 5 is characterized in that: Calculating the cost function value of each candidate search direction includes the following steps: Calculate the safety cost, invalid exploration cost and direction change cost of each decision in turn; The security cost is: ; in, is the security cost, is a predefined safety distance, is the shortest distance to the obstacle; The invalid exploration cost is: ; in, is the ineffective exploration cost, and Represent the area of the effective exploration coverage area and the total area covered by the sensor respectively; The direction change cost is: ; in, is the direction change cost, It is a decision The corresponding direction, is the current moving direction of the robot; Determining the cost function value according to the safety cost, the invalid exploration cost, and the direction change cost; The expression of the cost function value is: ; in, is the cost function value, is the weight corresponding to each cost; Representing a decision The candidate proposition list The original index in .
7. The autonomous search method combining semantics in a large-scale unknown environment according to any one of claims 1 to 6, characterized in that: Determining the search direction according to the semantic octree and the multi-layer task probability map, and then determining the local target point, includes the following steps: The problem of determining the local target point in the multi-layer task probability graph is expressed as a first traveling salesman problem, and the local target point in the multi-layer task probability graph is obtained by solving the first traveling salesman problem; The first traveling salesman problem is: ; Among them, multiple potential local target points are represented as , Indicates the current position of the robot , ( ) indicates the The spatial coordinates of the Gaussian components, Represents the Euclidean distance between two local target points, evaluated using the A* algorithm in the semantic octogram; It is The weights of the Gaussian components, is the scaling factor; Expressing the problem of determining the local target point in the semantic octree as a second traveling salesman problem, and solving the second traveling salesman problem to obtain the local target point in the semantic octree; The second traveling salesman problem is: 。 8. An autonomous search device combining semantics in a large-scale unknown environment, characterized by: The device comprises: A semantic octree construction unit, configured to construct a semantic octree based on a Bayesian probability method; wherein the semantic octree includes the aligned point cloud data and semantic labels of objects in the panoramic image; a multi-layer task probability map construction unit, configured to analyze the panoramic image using a Gaussian mixture model to determine the probability of a target location, and then construct a multi-layer task probability map based on the probability of the target location; A repeated task determination unit, configured to obtain a task description and determine whether the search task corresponding to the task description is a repeated task; a first search unit, configured to, if not, determine a search direction using a large language model according to the task description and the panoramic image, and further determine a local target point; a repeated search unit, configured to, if yes, determine the search direction according to the semantic octree and the multi-layer task probability map, and further determine the local target point; a trajectory generation and optimization unit, configured to generate a search path for the local target point and optimize the trajectory of the search path according to the semantic octree and the multi-layer task probability graph; The search control unit is used to control the robot to move and search according to the trajectory.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.