Self-adaptive robot sensing, guiding and controlling integrated system facing strange environment

By integrating visual language models, large language models and visual basic models in autonomous mobile robot systems, the robot's autonomous perception, exploration and control in unfamiliar environments is realized, which solves the problem of poor task completion results in traditional systems in complex environments, and improves the system's adaptability and collaboration capabilities.

CN120085659AActive Publication Date: 2025-06-03BEIJING INST OF TECH

Patent Information

Application Number
CN202510559983.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

When traditional autonomous mobile robot systems face unfamiliar, unstructured, and terrain-changing environments, their perception and control systems are limited, resulting in poor task completion results and inability to effectively adapt to diversified and abstract task needs.

Method used

An integrated robot sensing and control system adapted to strange environments is designed, combining a semantic map builder based on visual language model, a semantic planner based on large language model, and a terrain adaptation controller based on visual basic model to realize robot active perception, independent exploration and efficient control.

Benefits of technology

It enhances the robot's task planning and real-time adjustment capabilities in unfamiliar environments, improves the trajectory tracking performance under complex terrain and dynamic interference, and realizes the deep coordination of robot perception, navigation and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085659A_ABST
    Figure CN120085659A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive robot sensing, guiding and controlling integrated system for a strange environment, and belongs to the field of intelligent unmanned systems. Each robot in the system comprises a sensing device, a sensing module, a navigation module and a control module. The sensing module is a VLM-based semantic mapping device, performs ground feature segmentation and semantic recognition on the environment information, and generates a global semantic map and a ground image embedded with a ground feature image semantic descriptor; the navigation module is a semantic planner based on LLM, adopts LLM to perform task understanding on natural language task description, and reasones a task sequence of a current environment; the control module is a terrain adaptation controller based on a VFM, the VFM is adopted to recognize terrain features from ground images, the terrain features and the real-time state of the robot are mapped into a driving matrix, and robot motion control is achieved. According to the invention, deep cooperation of active perception, autonomous exploration and efficient control of the robot is realized, and task planning and real-time adjustment capabilities of the robot in an unfamiliar environment are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to advanced perception of mobile robots, autonomous task planning, navigation, and terrain adaptation control technologies, belonging to the field of intelligent unmanned systems. Specifically, it relates to a robot integrated sensing, navigation, and control system adapted to unfamiliar environments, which can be used in multiple scenarios such as indoor service, warehousing logistics, search and rescue patrol, etc. Background Art

[0002] With the development of mobile robot technology, more and more robots are deployed in application scenarios such as large conference venue guiding or outdoor search and rescue patrol. The traditional autonomous mobile robot system uses relatively separate perception, navigation, and control modules to achieve the established task requirements. In various application examples, it is found that the autonomous mobile robot system limited by the traditional framework usually needs to design customized solutions for specific tasks and cannot accept diverse and abstract task requirements. At the same time, in the face of unstructured, terrain-changing, unknown or partially known environments, the perception and control systems of the robot are greatly affected, and the task completion effect is poor.

[0003] In the traditional autonomous mobile robot system framework, the perception module provides the environment map and robot positioning for the navigation module. Although the semantic map can provide the robot with high-level information about the surrounding ground objects, the utilization of this type of information by the navigation module is usually only limited to avoiding dynamic and static obstacles and cannot have a more favorable upper-level impact on real-time path planning in the face of autonomous exploration tasks in unfamiliar environments.

[0004] With the rapid development of the field of machine learning and the continuous improvement of hardware computing power, the application of large language models (LLMs) in robot navigation systems has gradually become an emerging research direction. Natural language has become an effective means of assigning tasks, and LLM-based planners are widely used in multiple fields, including mobile manipulation, service robots, autonomous driving, navigation, and fault detection. The current mainstream method configures the LLMs through context prompts, enabling them to apply common sense to specific domains without fine-tuning the model. At the task level, most existing research focuses on explicit task descriptions, and some teams are committed to developing LLM planners that can transform task descriptions into formal languages (such as linear temporal logic LTL or planning domain definition language PDDL). Although formal languages can clearly elaborate subtasks and semantic relationships, they are very complex. Other studies have considered natural language task instructions with non-explicit task descriptions (such as "Get me something to drink"), but still assume that the working scenario has been pre-mapped or has relatively clear structured features. At the environmental level, some scholars have tried to relax the dependence on pre-built semantic maps by introducing feedback from the perception system or specifying semantics at runtime. However, such perception methods are often limited to object detection and usually only applicable to small rooms or enclosed environments that can utilize explicit hierarchical structures and natural boundaries of the environment.

[0005] Generally speaking, most existing research assumes the existence of pre-built maps or structured indoor environments, or focuses on specific tasks such as object search and goal navigation. These assumptions often cannot be effectively applied in many practical application scenarios, especially in outdoor environments or tasks where users only provide high-level requirement descriptions.

[0006] In recent years, there have been relatively successful examples of the combination of the perception module and the control module of autonomous mobile robots. In terms of perception, taking the original image as input, the Visual Foundation Model (VFM) obtained through self-supervised learning has the potential to learn general visual features when the amount of pre-training data is large enough. By adapting, fine-tuning or further model construction, it can help the robot achieve specific tasks. This method has been applied to multiple tasks in the field of robotics, including image semantic segmentation, traversability estimation, and manipulation tasks. Its key advantages lie in the robustness to light changes and occlusions, as well as the good generalization ability between different images in the same scene. In terms of control, Vision-based Reinforcement Learning (VRL) has demonstrated the ability to control agents in simulated environments and has also achieved success in robot control in real environments. However, the policies generated by VRL often lack interpretability and do not have guarantees of safety and robustness. To solve this problem, some scholars use the Inverse Reinforcement Learning (IRL) method to interpret the terrain traversability as a reward map, thereby enhancing the robot's understanding of the environment. Although certain progress has been made in the research combining vision and reinforcement learning, how to incorporate appropriate different terrain models into the control strategy while ensuring theoretical feasibility and safety remains an urgent problem to be solved. Summary of the Invention

[0007] In view of this, the present invention provides a robot perception, navigation and control integrated system adaptable to unfamiliar environments, realizing the deep coordination of active perception, autonomous exploration and efficient control of the robot, and enhancing the robot's task planning and real-time adjustment capabilities in unfamiliar environments.

[0008] To solve the above technical problems, the present invention is implemented as follows.

[0009] A robot perception, navigation and control integrated system adaptable to unfamiliar environments, each robot in the system includes a sensing device, a perception module, a navigation module and a control module; The perception module is a semantic mapper based on the Visual Language Model (VLM), which is used to perform ground object segmentation and VLM-based semantic recognition on environmental information, and generate a global semantic map embedding the semantic descriptors of ground object images; the global semantic map is provided to the navigation module, and the segmented ground image is provided to the control module; The navigation module is a semantic planner based on the large language model (LLM), which is used to understand the input natural language task description by using the LLM, and combine with the global semantic map provided by the perception module to infer the task sequence in the current environment; the task sequence is updated and optimized in a rolling manner as the robot moves and the global semantic map provided by the perception module is updated; generate an expected trajectory according to the rolling optimized task sequence and send it to the control module; The control module is a terrain adaptation controller based on the vision foundation model (VFM). The VFM is used as a general terrain understanding model to identify terrain features from the segmented ground images. The terrain features and the real-time state of the robot are mapped to a driving matrix through a robot-specific terrain model; generate the control input of the robot based on the driving matrix and the expected trajectory to achieve the motion control of the robot.

[0010] Preferably, the perception module includes: a mapping module, a data association module, and a pose graph optimization module; The mapping module obtains the image information and point cloud information of the environment, as well as the acceleration information of the robot from the sensing device, and generates a local sub-map embedding the semantic descriptor of the ground object image based on the image segmentation algorithm and VLM; local sub-maps will be continuously generated during the robot's driving process; The data association module receives the local sub-maps sent by other robots through distributed communication and merges them into a global semantic map; calculates the relative transformation matrix of the current pose of each robot according to the global semantic map and provides it to the pose graph optimization module; The pose graph optimization module determines the global pose of each robot according to the odometry information obtained by using the acceleration information and the relative transformation matrix of the poses between robots.

[0011] Preferably, the mapping module includes an image segmentation module, a vision language module, a visual odometry calculation module, a local sub-map generation module, and a communication module; The image segmentation module segments the ground objects in the surrounding environment from the image based on the image information of the environment, sends the picture of the ground object image to the vision language module, and sends the segmented ground image to the control module; The vision language module recognizes the picture of each ground object and generates the corresponding semantic descriptor, and sends it to the local sub-map generation module; The visual odometry calculation module generates odometry information based on the image information of the environment and the acceleration information of the robot, and sends it to the local sub-map generation module; The local sub-map generation module receives the point cloud information of the environment, embeds the semantic descriptor into the point cloud, and combines it with the odometry information to generate a local sub-map; The communication module sends the local sub-map of the robot where it is located to other robots through distributed communication.

[0012] Preferably, the sensing device includes a depth camera and an Inertial Measurement Unit (IMU); the depth camera acquires image information and point cloud information of the environment; the IMU acquires the acceleration information of the robot.

[0013] Preferably, the global semantic map uses a list to represent the information of three-dimensional features, including the semantic information, centroid position, shape, and size of each feature object in the explored area.

[0014] Preferably, the navigation module includes a decision-making generation module, a decision-making verification module, and a path planning module; The decision-making generation module includes a Large Language Model (LLM) and a behavior library. The LLM analyzes the natural language task description issued by the user, combines the global semantic map and the global pose of the robot, and calls the preset atomic behaviors in the behavior library to generate a task sequence composed of atomic behaviors; the task sequence is transmitted to the decision-making verification module for verification; the atomic actions include three categories: navigation, active perception, and user interaction; The decision-making verification module is used to verify the task sequence; for the task sequence that passes the verification, it determines the navigation target point and sends it to the path planning module; The path planning module is used to plan the desired trajectory from the current position to the navigation target point and send it to the control module.

[0015] Preferably, the verification of the decision-making verification module includes semantic verification and spatial verification; the semantic verification checks whether the semantics of the task conform to the predetermined rules. If there are semantic errors or unreachable targets, it feeds back to the LLM of the decision-making generation module for adjustment; the spatial verification adopts a frontier exploration method to ensure that all spatial paths of the tasks are feasible.

[0016] Preferably, the control module includes a three-layer structure, namely a general terrain understanding layer implemented by VFM, a robot-specific terrain model layer implemented by a neural network, and a real-time disturbance adaptive layer; The VFM is trained offline based on an image set to achieve the basic modeling of terrain recognition; The neural network inputs information that combines terrain features and the real-time state of the robot; this neural network adopts meta-learning technology and is trained based on the segment driving data of the robot where it is located to capture the unmodeled dynamics during the interaction between the robot and the terrain; The real-time disturbance adaptive layer online adjusts the weights of the last layer of the deep neural network to compensate for the dynamic disturbances that were not captured during the offline training process.

[0017] Preferably, the control module includes VFM, neural network, and composite adaptive controller; The ground image generated by the perception module segmentation is input into VFM, and VFM outputs terrain features; The terrain features and the current state of the robot feedback are spliced ​​and input into the neural network, and the neural network outputs the driving matrix K; The weights of the last layer of the neural network is an online adaptive time-varying linear parameter vector, which is transmitted to the composite adaptive controller together with the driving matrix K; During online operation, the composite adaptive controller updates the time-varying linear parameter vector of the neural network in real time according to the set rules, and calculates the control input of the robot at the current moment in combination with the expected trajectory and the drive matrix K provided by the navigation module.

[0018] Preferably, the neural network adopts a lightweight deep neural network, including an input layer, 2 hidden layers and an output layer.

[0019] Beneficial effects: (1) The perception module provided by the present invention adopts a semantic map builder based on a visual language model (VLM), which realizes multimodal perception and semantic map construction of complex environments. The VLM does not need to be pre-trained for the recognition of specific objects, and is less affected by changes in illumination and perspective, and has strong robustness and versatility. In a preferred embodiment, the semantic map adopts a list format, so that the robot can efficiently store and transmit local semantic maps, and the low-dimensional data represented by the list facilitates the semantic understanding of the navigation module.

[0020] (2) The navigation module provided by the present invention adopts a semantic planner based on LLM. In the prior art, the navigation module mainly performs path planning based on a given location point. However, the present invention inputs the operator's natural language task description to the navigation module. The semantic planner understands the task description with the help of LLM and has the ability to receive high-level tasks described in natural language. Combined with the real-time updated semantic map, it generates and verifies dynamic task sequences online through recursion, which significantly enhances the robot's task planning and real-time adjustment capabilities in unfamiliar environments.

[0021] (3) The control module provided by the present invention realizes real-time disturbance adaptation capability under complex terrain through a terrain adaptation controller based on VFM, combined with deep neural network and adaptive control, effectively compensates for the dynamic interaction not captured by the offline model, and improves the trajectory tracking performance of the autonomous mobile robot under different terrains and dynamic disturbances.

[0022] (5) The control module adopts a three - layer structure: the general terrain understanding layer, the robot - specific terrain model layer, and the real - time disturbance adaptation layer; the general terrain understanding layer is pre - trained on a large - scale image set and can extract features with high semantic relevance from terrain images; the robot - specific terrain model layer is trained on a small - scale data set obtained based on robot driving to capture complex dynamics not modeled during the interaction between the robot and the terrain; the real - time disturbance adaptation layer online adjusts the parameters of the last layer of the deep neural network through adaptive control to compensate for dynamic disturbances not captured during offline training.

[0023] (6) Through the information sharing and integrated design among the perception, navigation, and control modules, the present invention breaks through the limitations of traditional independent modular systems, realizes the deep coordination of active perception, autonomous exploration, and efficient control of the robot, and provides a high - value solution for complex tasks such as post - disaster rescue, on - site exploration, and unmanned inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a system framework diagram of the embodiment of the present invention.

[0025] Figure 2 It is a framework diagram of the perception module of the embodiment of the present invention.

[0026] Figure 3 It is a flow chart of the perception module of the embodiment of the present invention.

[0027] Figure 4 It is a framework diagram of the navigation module of the embodiment of the present invention.

[0028] Figure 5 It is a flow chart of the navigation module of the embodiment of the present invention.

[0029] Figure 6 It is a framework diagram of the control module of the embodiment of the present invention.

[0030] Figure 7 It is a flow chart of the control module of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.

[0032] The present invention provides a robot integrated sensing, navigation, and control system for adaptation to unfamiliar environments. As Figure 1 shown, the system includes multiple robots that can communicate wirelessly with each other. Each robot includes a sensing device, a perception module, a navigation module, and a control module.

[0033] In this embodiment, the sensing devices configured on the robot include a depth camera and an IMU. The depth camera is used to obtain image information and point cloud information, and the IMU is used to obtain acceleration and angular velocity information. In practice, a radar can also be used to obtain point cloud information, and the depth camera can be an Intel Realsense D456 depth camera.

[0034] The perception module is mainly embodied as an open-set semantic mapper based on a vision-language model (VLM). The perception module uses the depth camera to obtain environmental information, performs ground object segmentation and VLM-based semantic recognition on the environmental information, and generates a global semantic map embedding semantic descriptors of ground object images; the global semantic map is provided to the navigation module, and the segmented ground image is provided to the control module.

[0035] The navigation module is mainly embodied as a semantic planner based on a large language model (LLM). In the prior art, the navigation module mainly performs path planning based on given position points. However, in the present invention, the input to the navigation module is a natural language task description of the operator. The navigation module uses the LLM to understand the task description, and combines the perception information provided by the perception module to infer the task sequence in the current environment; a real-time rolling optimized task sequence is generated as the perception information is updated. A desired trajectory is generated according to the task sequence and sent to the control module.

[0036] The control module is mainly embodied as a terrain adaptation controller based on a vision foundation model (VFM). The control module identifies terrain features from the ground image segmented by the perception module; uses a neural network to map the terrain features and the real-time state of the robot to a drive matrix; generates a control input for the robot based on the drive matrix and the desired trajectory provided by the navigation module to achieve motion control of the robot. The neural network preferably uses a deep neural network (DNN).

[0037] It can be seen that the present invention combines semantic description, LLM, and VFM, enabling the system to accept high-level tasks described in natural language, analyze and decompose task-related elements, actively perceive and autonomously explore in an unfamiliar environment, and achieve complex functions such as post-disaster rescue, on-site exploration, and unmanned patrol.

[0038] The following will describe the specific implementations of the perception module, navigation module, and control module in detail.

[0039] (1) Perception module The perception module is used to construct a global semantic map inserting semantic descriptors based on high-level features detected by an open-set image segmentation algorithm in cooperation with a vision-language model (VLM).

[0040] In this embodiment, the input data of the perception module includes RGB-D depth camera images (including image information and point cloud information) and the acceleration and angular velocity information provided by an Inertial Measurement Unit (IMU).

[0041] See Figure 2 , the perception module includes: a mapping module, a data association module, and a pose graph optimization module.

[0042] The mapping module obtains the image information and point cloud information of the environment, as well as the acceleration information of the robot from the sensing device, and generates a local sub-map embedding the semantic descriptors of the ground objects based on the image segmentation algorithm and VLM; local sub-maps will be continuously generated during the robot's movement.

[0043] The data association module receives the local sub-maps sent by other robots through distributed communication; based on the data association algorithm of the consistency graph, combined with prior information such as gravity, it realizes the alignment of the local sub-maps, and the aligned local sub-maps are merged into a global semantic map; according to the global semantic map, it calculates the relative transformation matrix of the current pose of each robot and provides it to the pose graph optimization module.

[0044] The pose graph optimization module determines the global pose of each robot according to the odometry information obtained using the acceleration information and the relative transformation matrix of the poses between robots.

[0045] Finally, the perception module outputs the global semantic map and the pose information of each robot. The global semantic map is preferably a list, including the semantic information, centroid position, shape, size and other attributes of each ground object in the explored area. This representation enables the robot to efficiently store and transmit the local semantic map, and the low-dimensional data represented by the list is convenient for the semantic understanding of the navigation module. This perception module does not require pre-training for the recognition of specific objects and has strong robustness and versatility.

[0046] In a preferred embodiment, the mapping module includes an image segmentation module, a vision-language module, a visual odometry calculation module, a local sub-map generation module, and a communication module. Among them, the image segmentation module segments the ground objects in the surrounding environment from the image based on the image information of the environment, and sends the picture of the ground object image to the vision-language module, and sends the segmented ground image to the control module. The vision-language module recognizes the picture of each ground object, generates a corresponding semantic descriptor, and sends it to the local sub-map generation module. The visual odometry calculation module generates odometry information based on the image information of the environment and the acceleration information of the robot, and sends it to the local sub-map generation module; the local sub-map generation module receives the point cloud information of the environment, embeds the semantic descriptor into the point cloud, and combines it with the odometry information to generate a local sub-map. The communication module (not shown in the figure) sends the local sub-map of the robot to other robots through distributed communication. During the driving process of the robot, local sub-maps will be continuously generated, and through distributed communication, a single robot receives the local sub-maps sent by other robots.

[0047] Figure 3 The working process of the perception module is shown as Figure 3 shown below, and the specific steps are as follows: Step 11: Receive the RGB-D depth camera image input and the IMU sensor input, and use the YOLOv7 algorithm to filter out the dynamic obstacles in the surrounding environment. Generate odometry information using the visual odometry algorithm.

[0048] Step 12: Use the FastSAM image segmentation algorithm to segment the ground objects in the surrounding environment for the construction of the semantic map, and send the segmented ground image to the control module.

[0049] Step 13: Use the BLIP algorithm as the vision-language model to generate semantic descriptors for the ground objects. Among them, the BLIP algorithm proposes a multi-modal hybrid architecture - MED (Multimodal mixture of Encoder-Decoder), which combines the advantages of the encoder and decoder and can handle both understanding and generation tasks simultaneously.

[0050] Step 14: Project the segmented ground objects onto the point cloud, complete the merging of inter-frame data by calculating and comparing the overlap degree (Intersection over Union, IOU) of the point cloud, and construct a local semantic map, that is, a local sub-map, in combination with the odometry information provided by the visual odometry calculation method.

[0051] Step 15: Implement the transmission of the local semantic map between robots through the ROS-based remote topic manager and the AP bridge. A single robot receives the local sub-maps sent by other robots.

[0052] Step 16: Achieve sub-map alignment and closed-loop detection between robots through a data association method based on a consistency graph. At the same time, combine the global semantic map and odometer information to determine the global pose of each robot, and continuously update the global semantic map and the global pose of the robot during the movement of the robot, and send them to the navigation module.

[0053] (2) Navigation module The navigation module is a semantic planner based on the LLM. It uses the LLM combined with an action library and a verification module to implement a navigation algorithm for task planning. This algorithm, through a recursive method, receives user tasks and prior information from the semantic mapper, and generates a series of subtasks according to semantic and syntactic rules. The semantic planner can refine these subtasks online to adapt to the dynamically changing environment, and ensure the effectiveness and executability of the task sequence through the verification module. In each planning iteration, the semantic mapper provides the global semantic map and the updated robot pose. The LLM uses these updates to adjust the task sequence and calls the atomic actions in the action library to implement specific tasks. The action library contains atomic actions such as navigation, active perception, and user interaction. The semantic planner generates a task sequence according to the current task and conducts a review to ensure that the generated task semantics are correct and verify the reachability and spatial effectiveness of the task.

[0054] See Figure 4 , the navigation module includes a decision generation module, a decision verification module, and a path planning module.

[0055] The decision generation module includes the LLM and a behavior library. The LLM analyzes the high-level language instructions issued by the user, reasons about relevant elements, combines the global semantic map and the robot pose, and calls the preset atomic behaviors in the behavior library to generate a task sequence composed of atomic behaviors (including the parameters corresponding to each task); the task sequence is passed to the decision verification module for verification. For example, if the instruction provided by the user is: "I sent a robot to collect supplies from an incoming ship. I haven't received a reply yet. What's going on?" At this time, the LLM should analyze that it is looking for a robot in combination with the global semantic map given by the perception module, and use the provided context information to infer that the robot is likely to be near the dock in the scene. There is no direct path to the dock on the map, so exploration must be carried out to reach the dock and find the task robot. Many atomic behaviors are preset in the behavior library, covering three categories: navigation, active perception, and user interaction. These behaviors can be called through the API interface configured by the system, and the behavior library is used for task reasoning. In each iteration, the LLM generates a task sequence, which is optimized incrementally as the robot moves and the global semantic map and the updated robot pose provided by the perception module. The task sequence generated by the decision generation module is passed to the decision verification module for verification.

[0056] In a preferred embodiment, the LLM can also maintain all input information and task history in the form of context through an API, forming a context learning database. In this way, the LLM can perform more reasonable reasoning based on the context information during inference.

[0057] A decision verification module, used to verify the task sequence; determine the navigation target point of the task sequence after passing the verification, and send it to the path planning module.

[0058] In a preferred manner, the verification of the decision verification module includes semantic verification and spatial verification. To ensure that the generated tasks are executable, the decision verification module first checks whether the tasks meet semantic correctness, and then checks the spatial reachability of the given navigation target. If the task contains unreachable targets or actions that do not conform to the predetermined rule constraints, the decision verification module will feedback to the LLM of the decision generation module to adjust the task sequence. For example, if the LLM attempts to call the goto behavior to go to an unreachable location x, the validator will provide the following feedback: "Location x is unreachable from your current position. Please consider exploring first and update your plan accordingly." Spatial verification uses a frontier exploration method to ensure that the spatial paths of all tasks are feasible, avoiding the robot from performing dangerous or unachievable tasks. After passing the verification, the coordinates of the navigation target point are sent to the path planning module.

[0059] A path planning module, used to plan the desired trajectory from the current position to the navigation target point, and send it to the control module. When the robot performs active perception actions, it will explore unknown areas in the map. At this time, the perception module will send the real-time updated semantic map to the navigation module.

[0060] Figure 5 Shows the working process of the navigation module, as Figure 5 shown, the specific steps are as follows: Step 21: The user issues a natural language task description, and the LLM receives this task description, understands the task objectives, environmental constraints, and analyzes relevant elements.

[0061] Step 22: According to the global semantic map provided by the semantic mapper and the robot pose, the semantic planner calls the API interface such as ChatGPT 4o to infer the task composition in the current environment. During this process, the semantic mapper provides real-time updates of the global semantic map and the robot pose to ensure that the task sequence generated by the LLM conforms to the current environmental conditions.

[0062] Step 23: Gradually reason through the LLM and generate a task sequence using the available action library. All input information and action history are maintained in the form of context through the API. Atomic actions include three categories: navigation, active perception, and user interaction. The semantic planner automatically selects and combines corresponding task actions according to the current requirements.

[0063] Step 24: The generated task sequence is sent to the decision verification module for review. Among them, semantic verification checks the correctness of the tasks to ensure that the semantics of all tasks conform to the predetermined rules. If there are semantic errors or unreachable goals in the tasks, the decision verification module will feedback to the LLM for adjustment. Spatial verification adopts a frontier exploration method to ensure that all spatial paths of the tasks are feasible. During the spatial path verification of the tasks, the decision verification module checks whether the robot can safely execute the tasks in the existing environment and avoids executing dangerous or unachievable tasks.

[0064] Step 25: After verification, using the navigation target point coordinates of the optimized task sequence, plan a desired trajectory from the current position to the navigation target point and finally send it to the control module. The control module controls the robot to perform corresponding operations according to the instructions of the task sequence.

[0065] (3) Control module The control module is a terrain adaptation controller based on VFM. In this embodiment, it is an algorithm that uses a visual foundation model (VFM) combined with a deep neural network (DNN) and adaptive control to achieve vehicle motion control.

[0066] The input data of the control module includes the ground image segmented by the perception module and the robot state information. After processing the ground image through VFM, general terrain features are extracted, and combined with the robot-specific terrain model to achieve efficient adaptation to complex terrains and dynamic disturbances. The algorithm in the control module includes a three-layer structure: a general terrain understanding layer, a robot-specific terrain model layer, and a real-time disturbance adaptation layer. The general terrain understanding layer uses VFM, the robot-specific terrain model layer uses a deep neural network (DNN), and the real-time disturbance adaptation layer uses a composite adaptive controller.

[0067] First, understand the terrain features through VFM. VFM is pre-trained on an image set of approximately 1.2M images and can extract features with high semantic relevance from terrain images. These features are further combined with the robot state information through a deep neural network to generate a driving matrix K for a specific terrain. The deep neural network uses meta-learning techniques and is trained on a small-scale dataset (such as 20 minutes of driving data) to capture complex dynamics not modeled during the interaction between the robot and the terrain (such as track slippage or wheel degradation). During real-time operation, the control module online adjusts the parameters of the last layer of the deep neural network through adaptive control to compensate for dynamic disturbances not captured during offline training. The control module generates a driving matrix based on the current terrain features and robot state, thereby quickly adjusting the control strategy to ensure the trajectory tracking performance of the robot in complex terrains.

[0068] See Figure 6 , which shows an implementation manner where the control module includes VFM, DNN, and a composite adaptive controller. The ground image generated by the segmentation of the perception module is input into VFM, and VFM outputs a terrain feature vector ; the terrain feature vector and the current state (including the current linear velocity and angular velocity) fed back by the robot are concatenated and then input into DNN. This DNN has two hidden layers, and the network structure itself serves as the basis function required for adaptive control. Its output can be deformed to obtain the driving matrix K. The weights of the last layer of DNN are time-varying linear parameter vectors for online adaptation, which are transmitted to the composite adaptive controller together with the driving matrix K. During online operation, the composite adaptive controller can update the time-varying linear parameter vector in real time according to the formulated adaptive control rules, and combine the desired trajectory provided by the navigation module (i.e., the desired linear velocity and angular velocity at the current position), the driving matrix K, and the time-varying linear parameter vector to calculate the control input u (linear velocity and angular velocity) at the current moment to achieve the control of the robot motion platform. Among them, the specific update and u calculation can be implemented according to the designed adaptive control rules as needed, which will not be elaborated here.

[0069] Figure 7 shows the working process of the control module. As Figure 7 shown, the specific steps are as follows: Step 31: The control module receives the ground image input from the perception module and obtains the robot state information.

[0070] Step 32: Use the visual foundation model pre-trained on 1.2M images to extract the terrain feature vector .

[0071] Step 33, terrain feature vector The terrain feature vector and the robot state V are input into a lightweight deep neural network to capture the complex dynamics that are not modeled during the interaction between the robot and the terrain. The lightweight deep neural network outputs a driving matrix K, which is transmitted to the composite adaptive controller together with the weights of the last layer of the deep neural network.

[0072] Step 34, when the robot is running in real time, a meta-learning technique is adopted to online adjust the weights of the last layer of the deep neural network through a composite adaptive control mechanism. To compensate for the dynamic disturbances that cannot be captured during the offline training process and ensure the motion control accuracy of the robot in complex environments.

[0073] Step 35, the composite adaptive controller combines the desired trajectory provided by the navigation module, the driving matrix K, and the time-varying linear parameter vector to calculate the control input of the robot at the current moment; based on the current terrain features and the robot state, the driving matrix is updated in real time, and then the motion state of the robot is adjusted. Ensure that the robot can stably track the target trajectory in complex terrains, overcoming dynamic disturbances and environmental changes.

[0074] The dynamic system model of the control module is expressed as: (1) Where represents the system state, represents the control input, t represents time, is the nominal dynamics model, represents the unknown disturbance; The terrain feature vector generated by VFM, through the feature mapping of the deep neural network, models the unknown disturbance (2) Where is the time-varying linear parameter vector, is the representation error; The feature mapping is obtained through offline training of the deep neural network, serves as the weights of the last layer of the deep neural network and is continuously adaptively adjusted during operation.

[0075] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not restricted. Therefore, those skilled in the art of the present invention can modify or make equivalent replacements to the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the gist and technical solutions of the present invention, and shall all fall within the protection scope of the present invention.

Claims

1. A robot sensing, guidance and control integrated system that is adaptive to unfamiliar environments, wherein each robot in the system includes a sensing device, a perception module, a navigation module and a control module, and is characterized by: The perception module is a semantic map builder based on a visual language model (VLM), which is used to perform object segmentation and VLM-based semantic recognition on environmental information, and generate a global semantic map embedded with semantic descriptors of object images; The global semantic map is provided to the navigation module, and the segmented ground image is provided to the control module; The navigation module is a semantic planner based on the large language model (LLM), which is used to use the LLM to understand the input natural language task description, and to infer the task sequence in the current environment in combination with the global semantic map provided by the perception module; the task sequence is updated and optimized with the movement of the robot and the global semantic map provided by the perception module; the expected trajectory is generated according to the rolling optimized task sequence and sent to the control module; The control module is a terrain adaptation controller based on the visual fundamental model VFM. It uses VFM as a general model for terrain understanding to identify terrain features from the segmented ground image. The terrain features and the real-time state of the robot are mapped into a drive matrix through a robot-specific terrain model. The control input of the robot is generated based on the drive matrix and the expected trajectory to achieve robot motion control.

2. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 1, characterized in that: The perception module includes: a mapping module, a data association module, and a pose graph optimization module; The mapping module obtains image information and point cloud information of the environment and acceleration information of the robot from the sensor device, and generates a local sub-map embedded with semantic descriptors of the ground feature image based on the image segmentation algorithm and VLM; the local sub-map will be continuously generated during the robot's driving process; The data association module receives local sub-maps sent by other robots through distributed communication and merges them into a global semantic map; calculates the current relative transformation matrix of each robot according to the global semantic map and provides it to the posture graph optimization module; The pose graph optimization module determines the global pose of each robot based on odometer information obtained using acceleration information and a relative pose transformation matrix between robots.

3. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 2, characterized in that: The mapping module includes an image segmentation module, a visual language module, a visual mileage calculation module, a local sub-map generation module, and a communication module; The image segmentation module segments the ground objects in the surrounding environment from the image based on the image information of the environment, sends the picture of the ground object image to the visual language module, and sends the segmented ground image to the control module; The visual language module identifies the image of each ground object, generates a corresponding semantic descriptor, and sends it to the local sub-map generation module; The visual odometer calculation module generates odometer information based on the image information of the environment and the acceleration information of the robot, and sends it to the local sub-map generation module; The local sub-map generation module receives point cloud information of the environment, embeds the semantic descriptor into the point cloud, and combines it with the odometer information to generate a local sub-map; The communication module sends the local submap of the robot to other robots through distributed communication.

4. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 1, characterized in that: The sensing device includes a depth camera and an inertial measurement device IMU; the depth camera obtains image information and point cloud information of the environment; the IMU obtains acceleration information of the robot.

5. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as described in any one of claims 1 to 4, characterized in that: The global semantic map uses a list to represent the information of three-dimensional objects, including the semantic information, centroid position, shape and size of each object in the explored area.

6. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 1, characterized in that: The navigation module includes a decision generation module, a decision verification module and a path planning module; The decision generation module includes LLM and behavior library. LLM analyzes the natural language task description issued by the user, combines the global semantic map and the global posture of the robot, calls the preset atomic behaviors in the behavior library, and generates a task sequence composed of atomic behaviors; the task sequence is passed to the decision verification module for verification; the atomic actions include three categories: navigation, active perception and user interaction; The decision verification module is used to verify the task sequence; determine the navigation target point for the verified task sequence and send it to the path planning module; The path planning module is used to plan the expected trajectory from the current position to the navigation target point and send it to the control module.

7. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 6, characterized in that: The verification of the decision verification module includes semantic verification and spatial verification; the semantic verification checks whether the semantics of the task conforms to the predetermined rules. If there are semantic errors or unreachable targets, the LLM feedback is provided to the decision generation module for adjustment; the spatial verification adopts a frontier exploration method to ensure that the spatial paths of all tasks are feasible.

8. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 1, characterized in that: The control module includes a three-layer structure, namely a general terrain understanding layer implemented by VFM, a robot-specific terrain model layer implemented by a neural network, and a real-time disturbance adaptation layer; The VFM is trained offline based on the image set to achieve basic modeling for terrain recognition; The neural network input information combines terrain features and the real-time status of the robot; the neural network uses meta-learning technology and is trained based on fragmented driving data of the robot to capture the unmodeled dynamics of the robot's interaction with the terrain; The real-time perturbation adaptation layer adjusts the weights of the last layer of the deep neural network online , to compensate for the dynamic disturbances that were not captured during offline training.

9. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 1, characterized in that: The control module includes VFM, neural network, and composite adaptive controller; The ground image generated by the perception module segmentation is input into VFM, and VFM outputs terrain features; The terrain features and the current state of the robot feedback are spliced ​​and input into the neural network, and the neural network outputs the driving matrix K; The weights of the last layer of the neural network is an online adaptive time-varying linear parameter vector, which is transmitted to the composite adaptive controller together with the driving matrix K; During online operation, the composite adaptive controller updates the time-varying linear parameter vector of the neural network in real time according to the set rules, and calculates the control input of the robot at the current moment in combination with the expected trajectory and the drive matrix K provided by the navigation module.

10. The robot sensing, guidance and control integrated system for adapting to unfamiliar environments as claimed in claim 8 or 9, characterized in that: The neural network adopts a lightweight deep neural network, including an input layer, 2 hidden layers and an output layer.

Citation Information

Patent Citations

  • Target navigation method and system based on hierarchical semantic map

    CN118189961A

  • Task execution method and device, equipment, medium and program product

    CN118551057A

  • Robot task reasoning method and system fusing multi-modal information and ontology knowledge

    CN118657216A

  • Visual language navigation planning method and equipment based on topological semantic map prompt

    CN118999554A

  • Exhibition hall robot visual language navigation method based on large model

    CN119309580A

Cited By

  • Self-adaptive control method and system of autonomous mobile platform

    CN120848181A

  • Adaptive control method and system for autonomous mobile platform

    CN120848181B

  • Wheeled robot trafficability prediction method and system in complex scene

    CN120993920A

  • Humanoid robot indoor action planning and action control method and system and robot

    CN121187142A

  • Multi-mode blind-assisting indoor navigation method and system

    CN121612280A