Robot control method, device, apparatus, and storage medium

CN122546986APending Publication Date: 2026-08-11UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,视觉语言模型通常局限于二维特征空间,导致定位精度不足,难以实现物理空间中的精准导航

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122546986A_ABST
    Figure CN122546986A_ABST
Patent Text Reader

Abstract

This application provides a robot control method, apparatus, device, and storage medium. The method includes: determining the robot's pose based on images and depth information acquired by the robot; mapping the pixel features of multiple pixels in the image to multiple cells in a horizontal grid map of the robot's environment to obtain cells mapped with pixel features; determining the robot's target position based on a first correlation between the robot's control command text and the pixel features mapped to each cell, and generating an obstacle map based on a second correlation between obstacle description text in the environment and the pixel features mapped to each cell; performing path planning based on the pose, target position, and obstacle map to obtain the robot's motion path; generating multiple waypoints for the robot based on the motion path, and controlling the robot's motion based on the multiple waypoints. This application improves the accuracy of robot control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a robot control method, apparatus, device, and storage medium. Background Technology

[0002] Utilizing natural language commands to drive robot navigation is a significant area of ​​current development. One related technology directly applies Visual Language Models (VLMs) for scene analysis and semantic matching at the two-dimensional image level, thereby achieving robot navigation control. However, Visual Language Models are typically limited to a two-dimensional feature space, resulting in insufficient positioning accuracy and making it difficult to achieve precise navigation in physical space. Summary of the Invention

[0003] This application provides a robot control method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of robot control.

[0004] The technical solution of this application embodiment is implemented as follows: This application provides a robot control method, including: The robot's pose is determined based on the images and depth information collected by the robot. For the multiple pixels included in the image, the pixel features of the multiple pixels are mapped to multiple cells included in the horizontal grid map of the robot's environment, to obtain the cells mapped with the pixel features; Based on the first correlation degree between the robot's control command text and the pixel features mapped to each cell, the target position to be reached by the robot is determined, and based on the second correlation degree between the obstacle description text of the environment and the pixel features mapped to each cell, an obstacle map of the environment is generated. Based on the pose, the target position, and the obstacle map, path planning is performed to obtain the robot's motion path; Based on the motion path, multiple waypoints to be traversed by the robot are generated, and the robot is controlled to move based on the multiple waypoints.

[0005] This application also provides a robot control device, including: The first determining module is used to determine the pose of the robot based on the images and depth information collected by the robot; The mapping module is used to map the pixel features of the multiple pixels in the image to multiple cells in the horizontal grid map of the robot's environment, thereby obtaining the cells mapped with the pixel features. The second determining module is used to determine the target position to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell, and to generate an obstacle map of the environment based on the obstacle description text of the environment and the second correlation degree between the obstacle description text of the environment and the pixel features mapped to each cell. The path planning module is used to perform path planning processing based on the pose, the target position, and the obstacle map to obtain the robot's motion path; The control module is used to generate multiple waypoints that the robot will pass through based on the motion path, and to control the robot to move based on the multiple waypoints.

[0006] This application also provides an electronic device, including: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the robot control method provided in the embodiments of this application.

[0007] This application also provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the robot control method provided in this application.

[0008] This application also provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the robot control method provided in this application.

[0009] The embodiments of this application have the following beneficial effects: The robot's pose is determined by collecting images and depth information, and the pixel features of the images are mapped to cells of a horizontal grid map of the physical environment. This accurately projects the visual semantic features, originally confined to a two-dimensional image space, into a grid space with physical scale attributes, establishing a high-precision mapping between the semantic features of the two-dimensional image space and the physical environment. Furthermore, the correlation between the control command text, obstacle description text, and the mapped cell pixel features is calculated. Based on this correlation, the target location and obstacle map are determined on the horizontal grid map. Finally, path planning is performed in conjunction with the pose, and physical waypoints are generated, achieving accurate matching and positioning of natural language commands in the real physical coordinate system. Therefore, the embodiments of this application effectively improve the spatial positioning accuracy and navigation control accuracy of the robot. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the architecture of the robot control system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3 This is a first flowchart illustrating the robot control method provided in an embodiment of this application; Figure 4 This is a second flowchart illustrating the robot control method provided in the embodiments of this application; Figure 5 This is a third flowchart illustrating the robot control method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the fourth process of the robot control method provided in the embodiments of this application.

[0011] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0015] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of a larger module or unit that includes the functionality of the module or unit.

[0016] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0017] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0018] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0019] 1) Pose is a set of structured mathematical parameters that uniquely defines the absolute or relative spatial state of a physical entity in a specific three-dimensional spatial reference coordinate system. In the embodiments of this application, the pose is determined by pose calculation (such as odometry) based on the aforementioned image data and depth information; it serves as the basis for mapping pixel features to coordinates in physical space, and as the initial starting point coordinates and initial orientation state input in subsequent path planning processing, directly determining the initial boundary conditions for motion path generation calculation.

[0020] 2) A Vision-Language Model (VLM) is a deep neural network architecture that includes a multimodal network structure, is pre-trained with parameters, and can map and align visual image data and natural language text data to a unified high-dimensional mathematical feature space. In the embodiments of this application, the visual language model is mainly responsible for performing cross-modal feature extraction operations; through its internal network nodes, it converts the input two-dimensional image into vector-form features and converts the input control commands or obstacle descriptions into corresponding text features, so that the data from the above different modalities can be uniformly calculated and compared in a high-dimensional feature space. For example, but not limited to, the visual language model may include a decoupled image encoder module and a text encoder module, and its underlying computing architecture may adopt a multi-layer Transformer self-attention network structure or introduce a multimodal deep learning entity that integrates convolutional neural networks (CNN) and text encoding networks.

[0021] 3) A horizontal grid map, also known as a bird's-eye view (BEV) map, is a two-dimensional topological representation model that projects three-dimensional physical space onto a specific two-dimensional horizontal plane along the vertical or gravitational direction and discretizes it into multiple minimal independent storage areas (i.e., cells) according to a preset physical size. In the embodiments of this application, the horizontal grid map, as the underlying spatial data structure carrying pixel features, reconstructs and aggregates data points with physical coordinates from multiple perspectives and their corresponding pixel features, establishing a high-precision geometric mapping relationship between two-dimensional image spatial semantics and the three-dimensional physical environment. This provides a unified coordinate reference system that combines scale calibration and semantic features for determining target locations and generating obstacle maps.

[0022] 4) A waypoint is a discretized spatial data node that constitutes a continuous physical movement trajectory, possessing specific spatial coordinates and associated kinematic attributes (such as attitude and velocity). In this embodiment, multiple waypoints are generated by discretizing the macroscopic motion path generated based on path planning according to a preset sampling step size. This decomposes the long-distance motion path into low-level drive commands containing specific spatial coordinates and yaw angles, which are sequentially sent to the robot's low-level motion control logic unit to directly drive the hardware actuators to complete high-precision position translation and attitude rotation actions along the time sequence. For example, but not limited to, a set of spatial state variables encapsulated using a structure or a specific protocol data packet; in this embodiment, the data packet for each waypoint precisely combines objective physical parameters such as the horizontal coordinate position, vertical coordinate position, and yaw angle guiding the robot's direction of travel in the horizontal reference plane.

[0023] This application provides a robot control method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of robot control. The following is a detailed description of the embodiments of this application based on the above explanation of the terms and concepts used.

[0024] The robot control system provided in the embodiments of this application is described below. See also... Figure 1 , Figure 1 This is a schematic diagram of the architecture of the robot control system provided in an embodiment of this application. To support an exemplary application, the robot control system 100 includes: a server 200, a network 300, and a robot 400. The robot 400 is connected to the server 200 via the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using wireless or wired links for data transmission.

[0025] Here, robot 400 collects images and depth information of its environment and sends the collected images and depth information to server 200; server 200 receives the images and depth information collected by robot 400; based on the images and depth information collected by the robot, it determines the robot's pose; for multiple pixels included in the image, the pixel features of multiple pixels are mapped to multiple cells included in the horizontal grid map of the robot's environment, resulting in cells mapped with pixel features; based on the first correlation between the robot's control command text and the pixel features mapped to each cell, the target position to be reached by the robot is determined, and based on the second correlation between the obstacle description text of the environment and the pixel features mapped to each cell, an obstacle map of the environment is generated; based on the pose, target position, and obstacle map, path planning is performed to obtain the robot's motion path; based on the motion path, multiple waypoints to be traversed by the robot are generated; multiple waypoints are returned to robot 400; robot 400 receives the multiple waypoints returned by server 200; based on the multiple waypoints, the robot is controlled to move.

[0026] The robot control method provided in this application is implemented by an electronic device. For example, it can be implemented by a robot alone, by a server alone, or by a robot and a server working together. The electronic device implementing the robot control method provided in this application can be various types of robots or servers. The server (e.g., server 200) can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Robot 400 can be various types of embodied intelligent agents, such as robots covering various scenarios, including: industrial robots (precision operation, logistics palletizing robots (efficient handling); service robots (contactless delivery, hotel reception robots (guidance and consultation); medical robots (da Vinci surgical robots (minimally invasive and precise), rehabilitation training robots (assisted rehabilitation); household robots (automatic cleaning), companion robots (interactive dialogue); and special-purpose robots (firefighting and rescue robots (operating in high-temperature environments), underwater exploration robots (deep-sea exploration)). Different types adapt to different scenario needs, with the core objective of replacing or assisting humans in completing complex, dangerous, or repetitive tasks. Robots and servers can be connected directly or indirectly via wired or wireless communication, and this application does not limit this.

[0027] In some embodiments, the terminal or server can implement the robot control method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0028] The following describes an electronic device for implementing a robot control method provided in an embodiment of this application. See also... Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 500 provided in this embodiment can be a terminal or a server. Figure 2 As shown, electronic device 500 includes at least one processor 510, memory 550, at least one network interface 520, and user interface 530. The various components in electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 540.

[0029] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0030] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0031] Memory 550 may be removable, non-removable, or a combination thereof. Memory 550 may include one or more storage devices physically located away from processor 510. Memory 550 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.

[0032] In some embodiments, memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, as illustrated below. Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as a framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks; network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including Bluetooth, Wireless Fidelity (Wi-Fi), and Universal Serial Bus (USB); presentation module 553 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with user interface 530 (e.g., a display screen, a speaker, etc.); input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.

[0033] In some embodiments, the robot control device provided in this application can be implemented in software. Figure 2 A robot control device 555 stored in memory 550 is shown. It can be software in the form of programs and plug-ins, including the following software modules: a first determination module 5551, a mapping module 5552, a second determination module 5553, a path planning module 5554, and a control module 5555. These modules are logically related and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.

[0034] The robot control method provided in the embodiments of this application is described below. As mentioned above, the robot control method provided in the embodiments of this application is implemented by an electronic device, such as a server or terminal alone, or a server and terminal working together. Therefore, the executing entity of each step will not be described again below. See Figure 3 , Figure 3 This is a first flowchart illustrating the robot control method provided in this application embodiment. The robot control method provided in this application embodiment includes: Step 101: Determine the robot's pose based on the images and depth information collected by the robot.

[0035] For step 101, "image" refers to the two-dimensional pixel matrix data acquired by the robot's vision sensors (such as cameras). "Depth information" refers to depth images or three-dimensional point cloud data that characterize the distance from each physical point in physical space to the optical center of a sensor (such as a structured light sensor, binocular or multi-view stereo vision camera, LiDAR sensor, ultrasonic sensor, and millimeter-wave radar, etc.). "Pose" refers to the robot's three-dimensional position coordinates and orientation in a specific three-dimensional reference coordinate system.

[0036] In the specific implementation process, a continuous video stream of the target environment is acquired in real time by an RGB-D camera mounted on the robot, and images and corresponding depth information are extracted from the continuous video stream. After acquiring the images and depth information, a visual simultaneous localization and mapping algorithm (such as RTAB-Map (i.e., SLAM technology)) is used to process the continuously input data stream. Specifically, a visual odometry model is used to extract feature points from the continuously acquired images and match feature associations between adjacent data frames. Combined with the three-dimensional geometric scale provided by the depth information, the continuous motion trajectory and relative displacement change of the camera in three-dimensional space are calculated. Through the above odometry calculation process, the absolute position and attitude data of the robot in the current environmental coordinate system are calculated in real time, thereby determining the robot's pose. Based on the determined pose, the calculated pose data is timestamped and spatially bound to the acquired three-dimensional point cloud, and point cloud data with pose is generated synchronously.

[0037] Step 102: For the multiple pixels in the image, map the pixel features of the multiple pixels to the multiple cells in the horizontal grid map of the robot's environment to obtain the cells mapped with pixel features.

[0038] For step 102, a pixel refers to a basic two-dimensional spatial data representation node that constitutes an image. Pixel features refer to vectorized data containing semantic association information between the image and language extracted from the pixels of an image through a pre-trained model (such as an image encoder of a visual language model). A horizontal grid map (two-dimensional horizontal grid map) refers to a two-dimensional gridded spatial representation model generated by projecting the three-dimensional physical space of the robot's environment vertically downwards onto a horizontal plane. A cell refers to the smallest independent spatial storage area in the horizontal grid map divided according to a preset physical size.

[0039] In the specific implementation process, a pre-trained visual language model image encoder is used to extract features from the image, that is, for the multiple pixels in the image, pixel features are generated for each pixel. Subsequently, based on the acquired 3D point cloud data with pose, a geometric projection correspondence between the data points in the 3D point cloud data and the horizontal grid map is established. According to this correspondence, the pixel features of multiple pixels are transformed by coordinate projection from 3D space to a 2D horizontal plane and mapped to multiple cells in the horizontal grid map. When performing the mapping operation, if multiple overlapping viewpoints due to robot movement result in pixel features from different viewpoints being projected onto the same cell in the horizontal grid map, a numerical averaging operation is performed on all pixel features mapped to the same cell, and the averaged feature result is used as the final feature data of that cell. Through the above pixel-level feature extraction and spatial projection fusion calculation process, an association mapping between the 2D pixel semantics of the image (i.e., pixel features) and the horizontal geometric space of the environment (i.e., the horizontal grid map) is established, and finally, cells mapped with pixel features are obtained, thereby realizing the construction of an environment map with semantic representation capabilities without additional data annotation.

[0040] In some embodiments, see Figure 4 In step 102, "mapping the pixel features of multiple pixels to multiple cells in the horizontal grid map of the robot's environment" can be achieved by executing the following steps 201-203: Step 201: Based on the camera intrinsic parameters and depth information of the camera that acquired the image, project each pixel into a three-dimensional space to obtain three-dimensional point cloud data, which includes data points obtained by projecting each pixel into the three-dimensional space; Step 202: For each data point, establish a correspondence between the data point and the cell containing the projected data point by projecting the data point into the cells in the horizontal grid map; Step 203: For each data point, based on the correspondence of the data points, map the pixel features of the projected data point to the cell corresponding to the data point.

[0041] Here, camera intrinsic parameters refer to the physical parameter matrix characterizing the camera's internal geometry and optical perspective projection properties. Three-dimensional space refers to a physical geometric space reference system composed of three orthogonal coordinate axes. Three-dimensional point cloud data refers to a set of discrete data nodes representing the surface morphology and spatial distribution characteristics of a physical environment in three-dimensional space. A data point refers to a specific spatial node in the three-dimensional point cloud data with a clearly defined three-dimensional coordinate position. The correspondence refers to the mathematical mapping relationship established based on geometric dimensionality reduction projection between data points in the three-dimensional physical space and cells in the two-dimensional plane.

[0042] For step 201, combining the two-dimensional pixel coordinates of the pixel on the image, the transformation matrix composed of camera intrinsic parameters, and the physical distance value provided by the depth information, the coordinate system transformation algorithm is used to transform each pixel on the image plane into a physical three-dimensional reference system (i.e., three-dimensional space), so that the generated three-dimensional point cloud data includes the data points obtained by projecting each pixel into the three-dimensional space, and each data point is given an independent and unique three-dimensional space physical coordinate attribute.

[0043] For step 202, the height axis (i.e., z-axis) value in the three-dimensional coordinates of each data point is extracted and removed, retaining its two-dimensional planar coordinate values ​​(i.e., x and y values) on the horizontal reference plane. Using these retained two-dimensional planar coordinate values, the data point is vertically projected onto a pre-constructed two-dimensional horizontal grid map to determine the cell where the data point vertically falls, thus completing the projection of the data point onto the cell and establishing the correspondence between the data point and the cell on which the data point is projected.

[0044] For step 203, based on the correspondence established in step 202, the physical association link between the data point and its source pixel, as well as the cell into which the data point falls through projection, is determined. Along this association link, the pixel features of the projected data point are extracted, and these pixel features are directly assigned to the storage space of the cell corresponding to the data point, thus achieving the mapping from pixel features to cell. Through the above operations, the spatial cross-dimensional feature fusion and transfer processing from two-dimensional image pixels to a horizontal grid map via a three-dimensional physical coordinate structure is completed.

[0045] By applying the above embodiments, two-dimensional pixels are projected into three-dimensional point cloud data using camera intrinsic parameters and depth information, and then dimensionality-reduced and projected onto a horizontal grid map. This objectively constructs a precise geometric correspondence between image pixels, three-dimensional physical nodes, and the two-dimensional horizontal grid. This correspondence serves as a data link for feature transfer, enabling high-dimensional pixel features to be accurately attached to the two-dimensional topological grid structure in physical space. This provides support for subsequent target location positioning and obstacle map generation based on correlation, offering both semantic features and physical spatial geometric accuracy.

[0046] Step 103: Based on the first correlation between the robot's control command text and the pixel features mapped to each cell, determine the target location to be reached by the robot, and based on the second correlation between the obstacle description text of the environment and the pixel features mapped to each cell, generate an obstacle map of the environment.

[0047] For step 103, the control command text refers to the natural language character sequence used to instruct the robot to perform spatial movement and navigation tasks, including single target sequences and complex long text sequences containing multiple execution levels. The first correlation refers to the mathematical similarity measure in vector space between the semantic features extracted from the encoded control command text and the pixel features attached to each cell. The target position refers to the two-dimensional grid coordinates or three-dimensional reference coordinates that the robot needs to actually reach in physical space, obtained after parsing the control command text. The obstacle description text refers to the natural language character sequence used to define the categories of physical objects in the environment that restrict passage or need to be avoided. The second correlation refers to the mathematical similarity measure in vector space between the semantic features extracted from the encoded obstacle description text and the pixel features attached to each cell. The obstacle map refers to the binary raster data structure used to represent non-navigable areas, separated and extracted from a horizontal grid map containing pixel features based on the spatial distribution of specific obstacle objects.

[0048] In the specific implementation process, a pre-trained text encoder is used to convert control command text into corresponding text features. By calculating the cosine similarity between the text features and the pixel features mapped to each cell, the first correlation between the robot's control command text and the pixel features mapped to each cell is obtained. Within the horizontal grid map, the cell corresponding to the highest first correlation is selected, and its spatial physical coordinates are extracted to determine the robot's target location. Similarly, the obstacle description text of the environment is converted into obstacle text features using a pre-trained text encoder. By calculating the cosine similarity between the obstacle text features and the pixel features mapped to each cell, the second correlation between the obstacle description text and the pixel features mapped to each cell is obtained. Cells whose second correlation satisfies a preset matching threshold are selected, and the cell values ​​that meet the condition are binarized and labeled to generate an obstacle map of the environment.

[0049] In some embodiments, in a multi-device collaborative operation scenario, multiple hardware terminal nodes are allowed to access and share the same horizontal grid map containing pixel features mapped to each cell via a communication network. Each node can independently execute the above-described matching calculation process, and based on the obstacle description text and second correlation degree of the environment it receives, independently and dynamically generate an obstacle map suitable for its current operating state. This avoids the computational overhead of repeatedly building maps while meeting the need to update the local environment state. Furthermore, since the entire process directly uses natural language features (control command text or obstacle description text) and image semantic features (i.e., pixel features) for correlation degree matching, even if the input control command text or obstacle description text of the environment contains special category data not covered in the preset training dataset, the first correlation degree and the second correlation degree can still be directly calculated based on the continuity of the semantic space. Therefore, the corresponding target location and obstacle map can be output without collecting a new dataset for new categories and retraining the model.

[0050] In some embodiments, pixel features are extracted from an image using an image encoder of a visual language model. Based on this, before "determining the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell", the following steps may also be performed: extracting the first text features of the control command text using a text encoder of a visual language model, and extracting the second text features of the obstacle description text using a text encoder; determining the first feature similarity between the first text features and the pixel features mapped to each cell, and using the first feature similarity as the first correlation degree; determining the second feature similarity between the second text features and the pixel features mapped to each cell, and using the second feature similarity as the second correlation degree.

[0051] Here, a visual language model refers to a deep neural network architecture trained to map visual and natural language text data to a unified multidimensional feature space, such as the LSeg model. An image encoder refers to a network module in a visual language model specifically designed to process images and extract multidimensional visual feature vectors. A text encoder refers to a network module in a visual language model specifically designed to process natural language character sequences and extract multidimensional text feature vectors. First text features refer to high-dimensional vector data representing the semantic information of control command text after encoding and transformation. Second text features refer to high-dimensional vector data representing the semantic information of obstacle attributes after encoding and transformation of obstacle description text. First feature similarity refers to the mathematical distance measure between the first text feature and the pixel feature in the multidimensional feature space. Second feature similarity refers to the mathematical distance measure between the second text feature and the pixel feature in the multidimensional feature space.

[0052] In the specific implementation process, pixel features are extracted from the image using an image encoder of a visual language model. Based on this, the first text features of the control command text are extracted using a text encoder of the visual language model, and the second text features of the obstacle description text are extracted using a text encoder. Specifically, the control command text is input into the text encoder of the visual language model for sequence encoding processing, outputting the first text features in high-dimensional vector form; similarly, the obstacle description text is input into the text encoder for sequence encoding processing, outputting the second text features in high-dimensional vector form. Subsequently, the first feature similarity between the first text features and the pixel features mapped to each cell is determined, and this first feature similarity is used as the first correlation degree. Specifically, in the multi-dimensional feature space, the cosine similarity vector distance calculation method is used to calculate the matching degree between the vector parameters of the first text features and the vector parameters of the pixel features mapped to each cell, obtaining the first feature similarity. Based on this, the value of the first feature similarity is directly used as the first correlation degree representing the degree of matching between the control command text and the cell. Next, the second feature similarity between the second text features and the pixel features mapped to each cell is determined, and this second feature similarity is used as the second correlation degree. Specifically, using the same vector distance calculation method, the matching degree between the vector parameters of the second text feature and the vector parameters of the pixel feature mapped to each cell is calculated to obtain the second feature similarity. The value of the second feature similarity is then used as the second correlation degree characterizing the degree of matching between the obstacle description file and the cell.

[0053] By applying the above embodiments, pixel features of images and text features of text are extracted separately using a visual language model. The calculated first feature similarity and second feature similarity are used as the correlation score, objectively aligning cross-modal natural language and visual pixel features into the same vector space. This structure allows the data processing to directly measure the matching degree between text semantics and the environmental grid using pure mathematical vector distance calculations, avoiding parsing errors caused by introducing intermediate hard-coded conversion rules, and improving the computational efficiency and accuracy when matching cross-modal features.

[0054] In some embodiments, step 103, "determining the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell", can be achieved by performing the following steps: determining the first cell from multiple cells whose first correlation degree satisfies the first correlation degree condition; determining the center position of the area where the first cell is located, and taking the center position as the target position.

[0055] Here, the first correlation condition refers to the numerical filtering rule used to determine whether the matching degree between the control instruction text and each cell meets a specific mathematical standard, such as the rule that the correlation value is the maximum value. The first cell refers to a specific two-dimensional spatial grid that meets the first correlation condition after feature comparison calculation within the scope of the horizontal grid map. The center position refers to the intersection of the geometric diagonals or the arithmetic mean coordinate point of the rectangular area covered by the first cell on the horizontal geometric plane.

[0056] In the specific implementation process, firstly, the first cell whose first correlation degree meets the first correlation degree condition is determined from multiple cells. Specifically, multiple cells within the horizontal grid map are traversed, and the calculated first correlation degree value for each cell is obtained. Cells whose first correlation degree value reaches the maximum value or exceeds a preset matching threshold are found according to a preset numerical comparison logic. These cells are extracted and marked as the first cell, thus extracting the cell in the two-dimensional grid space that best matches the semantics of the control command text as the first cell. Then, the center position of the area containing the first cell is determined and used as the target position. Specifically, the two-dimensional grid boundary coordinates of the first cell are extracted. By calculating the geometric mean of the two-dimensional grid boundary coordinates, the geometric center point coordinates of the first cell are obtained. Subsequently, the center position indicated by the geometric center point coordinates is used as the sole reference coordinate for guiding the robot's subsequent displacement, i.e., as the target position.

[0057] By applying the above embodiments, and selecting the first cell from multiple cells that meets the first correlation condition, the precise isolation and extraction of spatial regions in a complex grid map that match the semantics of control commands are objectively achieved. Furthermore, by determining the center position of the region containing the first cell and using it as the target position, the two-dimensional grid region with area attributes is reduced in dimensionality and transformed into a uniquely determined geometric coordinate point. This operation effectively eliminates the randomness and uncertainty in the coordinate values ​​of areal targets, providing a unique coordinate input with an absolute geometric reference for subsequent precise path calculations.

[0058] In some embodiments, step 103, "generating an obstacle map of the environment based on the second correlation degree between the obstacle description text of the environment and the pixel features mapped to each cell", can be achieved by performing the following steps: from the horizontal grid map, determine the second cell whose second correlation degree satisfies the second correlation degree condition, and determine the third cell whose second correlation degree does not satisfy the second correlation degree condition; set the second cell in the horizontal grid map to a first value, and set the third cell in the horizontal grid map to a second value, to obtain the obstacle map.

[0059] Here, the second correlation condition refers to the numerical threshold rule used to determine whether the matching degree between the obstacle description text and each cell reaches a specific mathematical standard. The second cell refers to a two-dimensional spatial grid in the horizontal grid map that meets the second correlation condition after comparison. The third cell refers to a two-dimensional spatial grid in the horizontal grid map that does not meet the second correlation condition after comparison. The first value refers to a discrete numerical assignment parameter used in the data structure to identify whether a specific grid is occupied by an obstacle or is not navigable, such as 1. The second value refers to a discrete numerical assignment parameter used in the data structure to identify whether a specific grid is unoccupied or freely navigable, such as 0.

[0060] In the specific implementation process, from the horizontal grid map, second cells that satisfy the second correlation condition are identified, and third cells that do not satisfy the second correlation condition are identified. Specifically, the second correlation of each grid is obtained by traversing the horizontal grid map, and compared with a preset threshold parameter. Grids with a second correlation value greater than or equal to the threshold parameter are identified and extracted as second cells, while grids with a second correlation value lower than the threshold parameter are identified and extracted as third cells, thereby achieving logical classification of obstacle attributes in spatial areas within the horizontal grid map. After completing the logical classification, the second cells in the horizontal grid map are set to the first value, and the third cells in the horizontal grid map are set to the second value, resulting in an obstacle map. Specifically, the data parameters in the storage area marked as the second cell in the horizontal grid map are uniformly replaced with the first value to represent the non-navigable occupancy status in the physical space; simultaneously, the data parameters in the storage area marked as the third cell in the horizontal grid map are uniformly replaced with the second value to represent the unobstructed passage status in the physical space. In this way, a binary data matrix consisting entirely of the first and second values ​​is generated, and then the obstacle map of the environment used for local collision avoidance calculation is output.

[0061] Applying the above embodiments, by setting the second cell that meets the second correlation condition to the first value and setting the third cell that does not meet the condition to the second value, the horizontal grid map containing high-dimensional semantic features is objectively reduced in dimensionality and converted into a binary data structure containing only two discrete states. This transformation operation accurately eliminates redundant semantic parameters unrelated to pathfinding and clearly defines the boundaries of obstacle-occupied areas in physical space. This processing method significantly reduces the storage overhead and computational complexity of the output map data, providing a well-structured and computationally efficient obstacle avoidance constraint data foundation for subsequent path planning algorithms.

[0062] In some embodiments, the following steps may also be performed to train the above-mentioned visual language model: for a target object included in the environment, a sample image including the target object is obtained; the descriptive text features of the descriptive text of the target object are extracted by the first text encoder of the first visual language model, and the sample pixel features of the sample image are extracted by the first image encoder of the first visual language model; the parameters of the first text encoder are fixed and the parameters of the first image encoder are updated based on the difference between the sample pixel features and the descriptive text features to train the first visual language model and obtain the visual language model.

[0063] Here, the target object refers to a physical entity that needs to be identified within a specific environment but is not included in the basic data categories of the pre-trained first visual language model. Therefore, the first visual language model cannot identify the target object or its recognition ability for the target object is below a preset threshold. A sample image refers to two-dimensional pixel data containing the target object. The first visual language model refers to a model that has not been fine-tuned with environmental data including the target object. The first text encoder refers to the network module in the first visual language model used to extract natural language features. Descriptive text refers to a natural language sequence used to define the semantic category of the target object. Descriptive text features refer to the high-dimensional semantic vectors output from the encoded descriptive text. The first image encoder refers to the network module in the first visual language model used to extract image features. Sample pixel features refer to the high-dimensional visual vectors output from the encoded sample images.

[0064] In the specific implementation process, for the target object included in the environment, image data including the target object is acquired as sample images, and descriptive text of the target object is also obtained. Subsequently, the descriptive text features of the target object's descriptive text are extracted through the first text encoder of the first visual language model, and the sample pixel features of the sample images are extracted through the first image encoder of the first visual language model. All network layer weight parameters within the first text encoder are frozen, and the correlation difference between the sample pixel features and the descriptive text features in the multidimensional feature space is calculated using a loss function. Thus, the gradient is updated only on the weight parameters within the first image encoder using the backpropagation algorithm. Since only a lightweight parameter update is required for a single network module, the computational power requirement of this fine-tuning process is moderate, and the calculation operation can be completed on a single graphics processing unit terminal, thereby achieving model adaptation for a specific environment.

[0065] By applying the above embodiments, and keeping the parameters of the first text encoder unchanged, while updating the parameters of the first image encoder only based on feature differences, the enormous computational cost and lengthy process of retraining all parameters of a large-scale network architecture are objectively avoided. This local parameter update strategy allows the extracted sample pixel features to be directly aligned to the fixed high-dimensional text semantic space, ensuring that the generated visual language model can quickly adapt to the environment-specific target object while fully maintaining its generalization matching ability for natural language instructions.

[0066] In some embodiments, before executing "determining the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell", the following steps may also be performed: obtaining natural language command text for the robot; parsing the natural language command text to obtain a parsing result, the parsing result including at least one of the following: entities, spatial relationships between entities, and temporal logical relationships of the robot's actions to be performed; based on the parsing result, splitting the natural language command text to obtain multiple sub-command texts, and using the multiple sub-command texts as control command texts; based on this, step 103, "determining the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell", can be achieved by executing the following steps: based on the first correlation degree between each sub-command text and the pixel features mapped to each cell, determining the sub-target location indicated by each sub-command text, the target location including multiple sub-target locations.

[0067] Here, natural language instruction text refers to an unstructured sequence of language characters containing multiple consecutive navigation intentions. Parsing results refer to the set of data features extracted after semantic reasoning of the natural language instruction text. An entity refers to an objective object in the physical environment with a clear semantic category and physical boundaries. Spatial relationships between entities refer to their relative orientation or geometric topological attributes in physical space. Temporal logical relationships refer to the order and dependencies of multiple actions occurring over time. Sub-instruction text refers to a sequence of phrases containing only a single action intention extracted from a long text. Sub-target location refers to the physical coordinate node corresponding to a single action intention.

[0068] In the specific implementation process, firstly, a large-scale language model is used to perform semantic understanding on the natural language command text to obtain the parsing results. These parsing results include at least one of the following: entities, spatial relationships between entities, and temporal logical relationships of the robot's actions to be executed. Then, based on the parsing results, the natural language command text is segmented to obtain multiple sub-command texts. These sub-command texts are then arranged according to the temporal logical relationships to obtain a text sequence composed of multiple sub-command texts, which is used as the control command text. Based on this, each sub-command text is extracted sequentially according to its order in the text sequence, and the first correlation degree between the currently extracted sub-command text and each cell is calculated. Based on the first correlation degree, target cells that meet the first correlation degree condition (e.g., selecting the cell with the highest first correlation degree) are extracted, and the sub-target position indicated by the sub-command text is determined based on the target cell. The multiple sub-command texts in the text sequence are traversed and calculated sequentially to obtain the sub-target position indicated by each sub-command text. These multiple sub-target positions constitute the target position.

[0069] For example, taking a navigation task in a commercial tour scenario, the natural language instruction text obtained for the robot is: "First go to the smart tour guide area next to the red booth, then pass between the commercial cleaning area and the smart inspection area, and finally return to the starting point." After obtaining the natural language instruction text, it is parsed to obtain the parsing result. The specific elements contained in the parsing result are as follows: The entities in the parsing result include: the red booth, the smart tour guide area, the commercial cleaning area, the smart inspection area, and the starting point; the spatial relationships between the entities in the parsing result include: "next," which indicates the physical spatial proximity between the area to be reached and the red booth; "between," which indicates that the traversal path is located in the middle area of ​​the space formed by the commercial cleaning area and the smart inspection area; the temporal logical relationship between the robot's actions in the parsing result includes: "first," "then," and "last," which indicate the execution order and dependencies of the three independent navigation actions on the timeline. Based on the above analysis results, the natural language instruction text is split into multiple sub-instruction texts. The specific splitting results are as follows: the first sub-instruction text is "Go to the smart guide exhibition area next to the red booth"; the second sub-instruction text is "Pass between the commercial cleaning exhibition area and the smart inspection exhibition area"; and the third sub-instruction text is "Return to the starting point".

[0070] By applying the above embodiments, logical parsing and segmentation of natural language command text yields multiple sub-command texts. This objectively reduces the complexity of long-sequence tasks involving multiple targets into multiple independent short-range pathfinding tasks. This multi-level decomposition processing mechanism effectively prevents the cumulative deviation generated in a single long-distance path calculation, thereby objectively reducing spatial positioning errors and improving the success rate of continuous navigation when multiple sub-target locations are involved.

[0071] Step 104: Perform path planning based on pose, target position, and obstacle map to obtain the robot's motion path.

[0072] For step 104, path planning refers to the computational process of generating spatial trajectory data that meets specific constraints within a specific spatial coordinate system, based on the starting point and ending point coordinates and combined with environmental traffic restrictions. A motion path refers to a data sequence composed of continuous or discrete spatial coordinate points and their corresponding orientation angles, used to guide the robot's underlying motion actuators to complete displacement operations in physical space.

[0073] In the specific implementation process, the current pose is obtained as the initial starting point coordinates and initial orientation state for the current planning calculation, and the target position is obtained as the planning endpoint. Simultaneously, an obstacle map containing coordinates of non-navigable areas is obtained as environmental traffic constraints. Subsequently, path planning is performed based on the pose, target position, and obstacle map. Specifically, the pose, target position, and obstacle map are input into the path planning algorithm model. Within the navigable area defined by the obstacle map, a path is calculated and generated that has the shortest spatial span from the starting position (included in the pose) to the target position, and does not spatially overlap or intersect with grids marked as obstacle attributes in the obstacle map.

[0074] In some embodiments, see Figure 5 Step 104, "Perform path planning based on pose, target position, and obstacle map to obtain the robot's motion path," can be achieved by executing the following steps 1041-1043: Step 1041, perform dilation processing on the cells representing obstacles in the obstacle map based on the robot's size to obtain a cost map; Step 1042, using the robot's position included in the pose as the starting point and the target position as the ending point, search in the cost map for multiple target free cells representing non-obstacles that connect the starting point and the ending point, where the path cost from the starting point to the ending point via multiple target free cells is the minimum; Step 1043, use the path composed of multiple target free cells as the motion path.

[0075] Here, path planning refers to the process of calculating the movement trajectory based on the current spatial state (i.e., pose) and the target state (i.e., target position) while avoiding obstacle areas. Inflation refers to the operation of expanding outwards by a specific distance based on obstacle boundaries. Cost map refers to a dataset recording the passage costs at various locations in space.

[0076] For step 1041, the cells representing obstacles in the obstacle map are inflated based on the robot's dimensions to obtain a cost map. Specifically, the robot's dimensions are acquired and converted into inflated distances. Cells representing obstacles are identified in the obstacle map, and the inflated distances are expanded outwards from these cells. The passage costs of cells within the expanded range are increased, and the passage costs of all cells are integrated to generate the cost map. This establishes a safe buffer boundary for physical obstacle avoidance.

[0077] For step 1042, starting from the robot position included in the pose and ending at the target position, search the cost map for multiple target free cells representing non-obstacles connecting the start and end points. The path cost from the start point to the end point via these multiple target free cells is minimized. Specifically, extract the robot position from the pose and set it as the start point of the path search, and the target position as the end point. Traverse the cells representing non-obstacles in the cost map and evaluate the movement cost between adjacent cells; accumulate the movement costs to calculate the overall path cost; filter multiple trajectories connecting the start and end points and compare the total cost values ​​(i.e., path costs) of different trajectories; extract the trajectories with the minimum total cost value, and identify the cells constituting the trajectories with the minimum total cost value as multiple target free cells. This ensures that the search process is free of redundancy.

[0078] For step 1043, the path consisting of multiple target free cells is used as the motion path. Specifically, multiple target free cells are connected sequentially according to their spatial adjacency topology to form a continuous spatial movement sequence, and this continuous spatial movement sequence is output as the motion path. In this way, the shortest collision-free path is planned.

[0079] Applying the above embodiments, the obstacle map is inflated based on the robot's dimensions to generate a cost map. This objectively maintains a buffer distance between the planned path and obstacles that matches the robot's physical shape, thus avoiding spatial collisions. Multiple empty target cells with the lowest path cost are searched in the cost map to generate the motion path, ensuring that the trajectory minimizes movement costs while avoiding obstacle areas, thereby improving the rationality and computational efficiency of path planning.

[0080] Step 105: Based on the motion path, generate multiple waypoints that the robot will pass through, and control the robot to move based on the multiple waypoints.

[0081] For step 105, a waypoint refers to a discrete spatial data node in the physical reference space that has specific position coordinates and a specific yaw attitude angle. Motion refers to the mechanical state changes, such as translation and attitude rotation, that occur in a hardware entity within a physical space reference frame.

[0082] In the specific implementation process, the data represented as a continuous spatial trajectory in the motion path is extracted according to a preset sampling step size to obtain a series of discrete spatial position coordinates. The desired travel orientation corresponding to each discrete spatial position coordinate is calculated by combining the trajectory curvature characteristics. The discrete spatial position coordinates and their corresponding desired travel orientations are combined and encapsulated into multiple waypoints with specific data format specifications. The data structure of each waypoint explicitly includes the horizontal x-axis coordinate, y-axis coordinate, and corresponding yaw angle value. Finally, the robot's motion is controlled based on the generated multiple waypoints. Specifically, the data sequence containing multiple waypoints is sent to the underlying motion control logic unit; after receiving the multiple waypoints, the underlying motion control logic unit parses the spatial relative displacement and attitude rotation changes between adjacent waypoints; based on the parsed relative displacement and attitude rotation changes, it generates speed control signals and pose adjustment drive commands; the adjustment drive commands are transmitted to the hardware actuator with displacement capabilities, driving the physical entity to move sequentially to the coordinate positions defined by the multiple waypoints and synchronously adjust its orientation.

[0083] In some embodiments, see Figure 6 In step 105, "generating multiple waypoints for the robot to pass through based on the motion path" can be achieved by executing the following steps 301-304: Step 301, smoothing the motion path to obtain a smooth motion path; Step 302, sampling the smooth motion path using a preset sampling step size to obtain a trajectory point sequence, which includes the position coordinates of multiple trajectory points sampled from the smooth motion path; Step 303, for each trajectory point in the trajectory point sequence, determining the yaw angle required for the robot to move from the trajectory point to the next trajectory point based on the position coordinates of the trajectory point and the position coordinates of the next trajectory point; Step 304, for each trajectory point, combining the position coordinates and yaw angle to obtain a waypoint.

[0084] Here, waypoints refer to navigation nodes that include spatial coordinates and attitude angles. Smoothing refers to the operation of eliminating sharp corners in the path to make it continuous and smooth. Yaw angle refers to the robot's orientation angle in the horizontal plane.

[0085] In step 301, the motion path is smoothed to obtain a smooth motion path. Specifically, the planned motion path is obtained, and corner nodes in the motion path are identified. A smoothing algorithm is used to adjust the connecting trajectories at the corner nodes to eliminate abrupt bends in the motion path, generating a continuous and smooth motion path. This improves the stability of subsequent motion control.

[0086] For step 302, the smooth motion path is sampled using a preset sampling step size to obtain a trajectory point sequence. The trajectory point sequence includes the position coordinates of multiple trajectory points sampled from the smooth motion path. Specifically, the preset sampling step size parameter is determined, and nodes are extracted along the extension direction of the smooth motion path according to the interval of the sampling step size. The two-dimensional spatial values ​​of the extracted nodes are obtained, and the two-dimensional spatial values ​​are confirmed as the position coordinates of the trajectory points. The trajectory points are arranged in sequence to form a trajectory point sequence. This ensures the uniformity of path sampling.

[0087] For step 303, for each trajectory point in the trajectory point sequence, based on the position coordinates of the trajectory point and the position coordinates of the next trajectory point, the yaw angle required for the robot to move from the current trajectory point to the next trajectory point is determined. Specifically, the trajectory point sequence is traversed, and for each trajectory point, the position coordinates of the current trajectory point are extracted, and the position coordinates of the adjacent next trajectory point are obtained; the direction vector from the current trajectory point to the next trajectory point is calculated, and the angle deflection value is determined based on the direction vector; the angle deflection value is determined as the yaw angle required for the robot to move from the current trajectory point to the next trajectory point.

[0088] For step 304, for each trajectory point, the position coordinates and yaw angle of the trajectory point are combined to obtain a waypoint. Specifically, the position coordinates of the trajectory point are extracted, and the yaw angle of the corresponding trajectory point is extracted; the position coordinates and yaw angle are concatenated to obtain an independent data node containing spatial position (i.e., position coordinates) and orientation attitude (i.e., yaw angle); this independent data node is used as a waypoint.

[0089] By applying the above embodiments, the motion path is smoothed and a sequence of trajectory points is generated by sampling at a preset step size, eliminating abrupt changes in the trajectory and ensuring the continuity of spatial movement. By combining the position coordinates of preceding and following trajectory points to calculate the yaw angle and combining it with the position coordinates to generate waypoints, composite navigation data containing both precise spatial coordinates and heading attitude is objectively constructed, improving the attitude control accuracy during movement transitions.

[0090] By applying the embodiments described above, the robot's pose is determined using acquired images and depth information. The pixel features of the images are then mapped to cells in a horizontal grid map of the physical environment. This accurately projects the visual semantic features, originally confined to a two-dimensional image space, into a grid space with physical scale attributes, establishing a high-precision mapping between the semantic features of the two-dimensional image space and the physical environment. Furthermore, the correlation between the control command text, obstacle description text, and the mapped cell pixel features is calculated. Based on this correlation, the target location and obstacle map are determined on the horizontal grid map. Finally, path planning is performed in conjunction with the pose, and physical waypoints are generated, achieving precise matching and positioning of natural language commands in the real physical coordinate system. Therefore, the embodiments of this application effectively improve the spatial positioning accuracy and navigation control accuracy of the robot.

[0091] The following description continues to illustrate the exemplary structure of the robot control device 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the robot control device 555 in the memory 550 may include: a first determining module 5551, used to determine the robot's pose based on the image and depth information acquired by the robot; a mapping module 5552, used to map the pixel features of the multiple pixels included in the image to multiple cells included in the horizontal grid map of the robot's environment, to obtain the cells mapped with the pixel features; a second determining module 5553, used to determine the target position to be reached by the robot based on the first correlation degree between the robot's control command text and the pixel features mapped to each cell, and to generate an obstacle map of the environment based on the obstacle description text of the environment and the pixel features mapped to each cell; a path planning module 5554, used to perform path planning processing based on the pose, the target position and the obstacle map to obtain the robot's movement path; and a control module 5555, used to generate multiple waypoints to be passed by the robot based on the movement path, and to control the robot to move based on the multiple waypoints.

[0092] In some embodiments, the mapping module 5552 is further configured to project each pixel into a three-dimensional space based on the camera intrinsic parameters of the camera that acquired the image and the depth information to obtain three-dimensional point cloud data, wherein the three-dimensional point cloud data includes data points obtained by projecting each pixel into the three-dimensional space; for each data point, a correspondence is established between the data point and the cell containing the projected data point by projecting the data point into the cell included in the horizontal grid map; for each data point, based on the correspondence of the data points, the pixel features of the pixel point of the projected data point are mapped into the cell corresponding to the data point.

[0093] In some embodiments, the pixel features are extracted from the image by an image encoder of a visual language model; the second determining module 5553 is further configured to, before determining the target position to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell, extract the first text features of the control command text by a text encoder of a visual language model, and extract the second text features of the obstacle description text by the text encoder; determine the first feature similarity between the first text features and the pixel features mapped to each cell, and use the first feature similarity as the first correlation degree; determine the second feature similarity between the second text features and the pixel features mapped to each cell, and use the second feature similarity as the second correlation degree.

[0094] In some embodiments, the second determining module 5553 is further configured to determine, from the plurality of cells, a first cell whose first correlation degree satisfies the first correlation degree condition; determine the center position of the area where the first cell is located, and use the center position as the target position.

[0095] In some embodiments, the second determining module 5553 is further configured to determine, from the horizontal grid map, a second cell whose second correlation degree satisfies the second correlation degree condition, and a third cell whose second correlation degree does not satisfy the second correlation degree condition; set the second cell in the horizontal grid map to a first value, and set the third cell in the horizontal grid map to a second value, thereby obtaining the obstacle map.

[0096] In some embodiments, the second determining module 5553 is further configured to: acquire a sample image including the target object for the target object included in the environment; extract descriptive text features of the descriptive text of the target object through a first text encoder of a first visual language model, and extract sample pixel features of the sample image through a first image encoder of the first visual language model; fix the parameters of the first text encoder unchanged, and update the parameters of the first image encoder based on the difference between the sample pixel features and the descriptive text features, so as to train the first visual language model and obtain the visual language model.

[0097] In some embodiments, the second determining module 5553 is further configured to: acquire natural language instruction text for the robot before determining the target location to be reached by the robot based on the first correlation degree between the control instruction text of the robot and the pixel features mapped to each cell; parse the natural language instruction text to obtain a parsing result, the parsing result including at least one of the following: entities, spatial relationships between entities, and temporal logical relationships of the actions to be performed by the robot; based on the parsing result, split the natural language instruction text to obtain multiple sub-instruction texts, and use the multiple sub-instruction texts as the control instruction text; the second determining module 5553 is further configured to: determine the sub-target location indicated by each sub-instruction text based on the first correlation degree between each sub-instruction text and the pixel features mapped to each cell, the target location including multiple sub-target locations.

[0098] In some embodiments, the path planning module 5554 is further configured to perform dilation processing on the cells representing obstacles in the obstacle map based on the robot's size to obtain a cost map; using the robot position included in the pose as the starting point and the target position as the ending point, search in the cost map for a plurality of target free cells representing non-obstacles connecting the starting point and the ending point, wherein the path cost from the starting point to the ending point via the plurality of target free cells is the minimum; and use the path composed of the plurality of target free cells as the motion path.

[0099] In some embodiments, the control module 5555 is further configured to smooth the motion path to obtain a smooth motion path; sample the smooth motion path using a preset sampling step size to obtain a trajectory point sequence, the trajectory point sequence including the position coordinates of multiple trajectory points sampled from the smooth motion path; for each trajectory point in the trajectory point sequence, determine the yaw angle required for the robot to move from the trajectory point to the next trajectory point based on the position coordinates of the trajectory point and the position coordinates of the next trajectory point; and for each trajectory point, combine the position coordinates and the yaw angle to obtain the waypoint.

[0100] It should be noted that the description of the device embodiments in this application is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments, so it will not be repeated here. Any technical details not covered in the robot control device provided in the embodiments of this application can be understood based on the description of the technical details in the above method embodiments.

[0101] This application also provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the robot control method provided in this application.

[0102] This application also provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the robot control method provided in this application.

[0103] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0104] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0105] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0106] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0107] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A robot control method characterized by, The method includes: The robot's pose is determined based on the images and depth information collected by the robot. For the multiple pixels included in the image, the pixel features of the multiple pixels are mapped to multiple cells included in the horizontal grid map of the robot's environment, to obtain the cells mapped with the pixel features; Based on the first correlation degree between the robot's control command text and the pixel features mapped to each cell, the target position to be reached by the robot is determined, and based on the second correlation degree between the obstacle description text of the environment and the pixel features mapped to each cell, an obstacle map of the environment is generated. Based on the pose, the target position, and the obstacle map, path planning is performed to obtain the robot's motion path; Based on the motion path, multiple waypoints to be traversed by the robot are generated, and the robot is controlled to move based on the multiple waypoints.

2. The method of claim 1, wherein, The process of mapping the pixel features of the multiple pixels to multiple cells in the horizontal grid map of the robot's environment includes: Based on the camera intrinsic parameters of the camera that acquired the image and the depth information, each pixel is projected into a three-dimensional space to obtain three-dimensional point cloud data, which includes data points obtained by projecting each pixel into the three-dimensional space. For each data point, a correspondence is established between the data point and the cell on which the data point is projected by projecting the data point onto the cell included in the horizontal grid map; For each data point, based on the correspondence between the data points, the pixel features of the projected pixel of the data point are mapped to the cell corresponding to the data point.

3. The method of claim 1, wherein, The pixel features are extracted from the image by an image encoder using a visual language model; before determining the target location to be reached by the robot based on the first correlation degree between the robot's control command text and the pixel features mapped to each cell, the method further includes: The first text features of the control instruction text are extracted using a text encoder based on a visual language model, and the second text features of the obstacle description text are extracted using the same text encoder. Determine the first feature similarity between the first text feature and the pixel feature mapped to each cell, and use the first feature similarity as the first correlation degree; Determine the second feature similarity between the second text feature and the pixel feature mapped to each cell, and use the second feature similarity as the second correlation degree.

4. The method of claim 3, wherein, The determination of the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell includes: From the plurality of cells, determine the first cell whose first correlation degree satisfies the first correlation condition; Determine the center position of the area where the first cell is located, and use the center position as the target position.

5. The method of claim 3, wherein, The second correlation between the obstacle description text based on the environment and the pixel features mapped to each cell, generating an obstacle map of the environment, includes: From the horizontal grid map, determine the second cell whose second correlation degree meets the second correlation degree condition, and determine the third cell whose second correlation degree does not meet the second correlation degree condition; The obstacle map is obtained by setting the second cell in the horizontal grid map to the first value and setting the third cell in the horizontal grid map to the second value.

6. The method of claim 3, wherein, The method further includes: For the target object included in the environment, obtain a sample image including the target object; The first text encoder of the first visual language model extracts the descriptive text features of the target object, and the first image encoder of the first visual language model extracts the sample pixel features of the sample image. The parameters of the first text encoder are kept unchanged, and the parameters of the first image encoder are updated based on the difference between the sample pixel features and the descriptive text features to train the first visual language model and obtain the visual language model.

7. The method of claim 1, wherein, Before determining the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell, the method further includes: Obtain the natural language command text for the robot; The natural language instruction text is parsed to obtain the parsing result, which includes at least one of the following: entities, spatial relationships between entities, and temporal logical relationships of the robot's actions to be executed; Based on the parsing results, the natural language instruction text is split into multiple sub-instruction texts, and the multiple sub-instruction texts are used as the control instruction text; The determination of the target location to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell includes: Based on a first correlation degree between each of the sub-instruction texts and the pixel features mapped to each of the cells, the sub-target location indicated by each of the sub-instruction texts is determined, the target location including multiple sub-target locations.

8. The method according to any one of claims 1 to 7, wherein, The path planning process based on the pose, the target position, and the obstacle map to obtain the robot's motion path includes: Based on the robot's dimensions, the cells representing obstacles in the obstacle map are inflated to obtain a cost map; Starting from the robot position included in the pose and ending at the target position, a plurality of target free cells representing non-obstacles are searched in the cost map to connect the starting point and the ending point, wherein the path cost from the starting point to the ending point via the plurality of target free cells is minimized. The path formed by the multiple target free cells is used as the motion path.

9. The method according to any one of claims 1 to 7, wherein, The process of generating multiple waypoints for the robot to traverse based on the motion path includes: The motion path is smoothed to obtain a smooth motion path; The smooth motion path is sampled using a preset sampling step size to obtain a trajectory point sequence, which includes the position coordinates of multiple trajectory points sampled from the smooth motion path; For each trajectory point in the trajectory point sequence, based on the position coordinates of the trajectory point and the position coordinates of the next trajectory point, determine the yaw angle required for the robot to move from the trajectory point to the next trajectory point. For each trajectory point, the position coordinates and yaw angle of the trajectory point are combined to obtain the waypoint.

10. A robot control device characterized by comprising: The device includes: The first determining module is used to determine the pose of the robot based on the images and depth information collected by the robot; The mapping module is used to map the pixel features of the multiple pixels in the image to multiple cells in the horizontal grid map of the robot's environment, thereby obtaining the cells mapped with the pixel features. The second determining module is used to determine the target position to be reached by the robot based on the first correlation degree between the control command text of the robot and the pixel features mapped to each cell, and to generate an obstacle map of the environment based on the obstacle description text of the environment and the second correlation degree between the obstacle description text of the environment and the pixel features mapped to each cell. The path planning module is used to perform path planning processing based on the pose, the target position, and the obstacle map to obtain the robot's motion path; The control module is used to generate multiple waypoints that the robot will pass through based on the motion path, and to control the robot to move based on the multiple waypoints.

11. An electronic device, comprising: The electronic device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the robot control method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by the processor, the robot control method according to any one of claims 1 to 9 is implemented.