Robot control method and device, electronic equipment, storage medium and program product
By collecting multi-view depth images and fusing point cloud data, the robot's pose and navigation path are updated in real time, and combined with voice command management, the problem of robot's task interruption in a dynamic environment is solved, achieving rapid, accurate and continuous execution of tasks.
Patent Information
- Application Number
- CN202510739907.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing robot systems are difficult to execute tasks quickly and accurately in dynamic and complex environments, and cannot adapt to environmental changes in real time, resulting in task interruption or incorrect execution.
By collecting multi-view depth images, fusing point cloud data of target objects, updating the robot's target pose and navigation path in real time, combining voice commands and task priority management, real-time adjustment and task continuity of the robot in complex environments.
Ensure that the robot performs tasks quickly and accurately in complex environments, achieves task continuity and accuracy, and improves the crawling success rate, navigation accuracy and voice response efficiency.
Smart Images

Figure CN120347761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and in particular to a robot control method, device, electronic equipment, storage medium and program product. Background Art
[0002] With the rapid development of artificial intelligence and automation technology, intelligent robots have been widely used in many fields such as industrial manufacturing, logistics sorting, and home services. Modern robot systems not only need to complete basic movement and obstacle avoidance functions, but also need to have complex operational capabilities such as identification, grasping, placement of target objects, and multi-task collaborative execution. Especially in a dynamically changing environment, such as when there are factors such as object displacement, occlusion, and external interference, higher requirements are placed on the robot's perception, decision-making, and execution capabilities.
[0003] At present, robot grasping and navigation technology mainly relies on depth perception equipment to obtain the spatial position information of the target object, and then performs path planning and grasping control based on this information. However, there are some common problems in such systems.
[0004] First, in terms of adaptability to dynamic environments, existing technologies are mostly designed for static scenes and lack effective perception and response mechanisms for real-time position changes of target objects. When the target object is displaced or blocked by other objects, it is difficult for the robot to update its spatial coordinates in time, resulting in grasping failure or navigation path failure.
[0005] Secondly, in terms of voice interaction and task switching, although some robot systems have integrated voice recognition modules to receive user commands, these systems generally lack effective interruption mechanisms and task priority management strategies. In the case of multi-tasking or sudden command input, the robot cannot quickly terminate the current task and turn to a new target, affecting the system's response efficiency and interactive experience.
[0006] In addition, in terms of the ability to continuously execute complex tasks, existing technologies often rely on fixed process settings when dealing with tasks involving multiple steps (such as first grabbing object A and then placing it in area B), and lack the ability to adaptively adjust to environmental changes during the process. Especially in scenarios with high requirements for multi-target recognition and spatial positioning accuracy, robots are prone to task interruptions and action conflicts due to environmental disturbances.
[0007] To sum up, how to solve the problem that existing robot systems are unable to perform tasks quickly and accurately in dynamic and complex environments, and are unable to make adaptive adjustments when the environment changes, resulting in task interruption or erroneous execution, is an important issue that needs to be urgently addressed in the field of intelligent robots. Summary of the invention
[0008] The present invention provides a robot control method, apparatus, electronic device, storage medium and program product, which are used to overcome the defects that the existing robot system cannot execute tasks quickly and accurately in a dynamic and complex environment, and cannot make adaptive adjustments when the environment changes, resulting in task interruption or incorrect execution, so as to ensure that the robot can accurately execute tasks in a complex environment. At the same time, it can also make real-time adjustments during the task execution process to ensure the continuity and accuracy of the task.
[0009] On the one hand, the present invention provides a robot control method, including: collecting multi-view depth images of a target object according to a received user instruction; determining point cloud data of the target object in the multi-view depth images based on the user instruction and the multi-view depth images; fusing the point cloud data of the target object in the multi-view depth images, and determining a target pose of the robot according to the fused point cloud data; obtaining target three-dimensional coordinates corresponding to the target object in the multi-view depth images, and mapping the target three-dimensional coordinates to a map to generate a navigation path of the robot; controlling the robot to execute a target task according to the target pose and the navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to changes in the environment where the robot is located.
[0010] Further, the user instruction is a voice instruction; correspondingly, the determining the point cloud data of the target object in the multi-view depth images based on the user instruction and the multi-view depth images includes: converting the voice instruction into a text instruction, and obtaining global point cloud data corresponding to the multi-view depth images; obtaining two-dimensional coordinates of the target object in the multi-view depth images according to the text instruction and the multi-view depth images; determining a mask of the target object in the multi-view depth images according to the two-dimensional coordinates of the target object in the multi-view depth images; and determining the point cloud data of the target object from the global point cloud data corresponding to the multi-view depth images according to the mask of the target object in the multi-view depth images, so as to obtain the point cloud data of the target object in the multi-view depth images.
[0011] Further, the fusing the point cloud data of the target object in the multi-view depth images includes: mapping the point cloud data of the target object in each view depth image in the multi-view depth images to a unified coordinate system to obtain target point cloud data of the target object in the multi-view depth images; determining weight values of each view depth image in the multi-view depth images; fusing the target point cloud data of the target object in the multi-view depth images by using weighted Kalman filtering according to the weight values of each view depth image to obtain initial fused point cloud data; and reducing the density of the initial fused point cloud data through voxel grid filtering to obtain the fused point cloud data.
[0012] Further, the obtaining of the target three-dimensional coordinates corresponding to the target object in the multi-view depth image includes: determining the target three-dimensional coordinates corresponding to the target object in the multi-view depth image according to the two-dimensional coordinates of the target object in the multi-view depth image, the internal parameters of the acquisition device of the multi-view depth image, and the depth value.
[0013] Further, the controlling the robot to execute the target task according to the target posture and navigation path of the robot includes: collecting a new multi-view depth image in real time, and obtaining the current two-dimensional coordinates of the target object in the new multi-view depth image in real time; obtaining the offset difference between the current two-dimensional coordinates and the two-dimensional coordinates; in the case where the offset difference is greater than or equal to a preset dynamic threshold, updating the target posture and navigation path of the robot, and controlling the robot to execute the target task according to the updated target posture and navigation path; in the case where the offset difference is less than the preset dynamic threshold, controlling the robot to execute the target task according to the target posture and navigation path of the robot.
[0014] Further, the user instruction is a voice instruction; when the voice instruction is received, the robot is controlled to stop executing the current task and switch to execute the target task.
[0015] In a second aspect, the present invention further provides a robot control device, including: a multi-view depth image acquisition module, configured to acquire a multi-view depth image of a target object according to a received user instruction; a target object point cloud data determination module, configured to determine point cloud data of the target object in the multi-view depth image based on the user instruction and the multi-view depth image; a target posture determination module, configured to fuse the point cloud data of the target object in the multi-view depth image and determine the target posture of the robot according to the fused point cloud data; a navigation path generation module, configured to obtain the target three-dimensional coordinates corresponding to the target object in the multi-view depth image and map the target three-dimensional coordinates to a map to generate a navigation path of the robot; a target task execution module, configured to control the robot to execute the target task according to the target posture and navigation path of the robot; wherein, the target task is determined by the user instruction, and the target posture and the navigation path will be updated in real time according to the change of the environment where the robot is located.
[0016] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the robot control method as described in any one of the above.
[0017] Fourthly, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the robot control method described in any one of the above is implemented.
[0018] Fifthly, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the robot control method described in any one of the above is implemented.
[0019] The robot control method provided by the present invention collects multi-view depth images of a target object according to a received user instruction, determines the point cloud data of the target object in the multi-view depth images based on the user instruction and the multi-view depth images, then fuses the point cloud data of the target object in the multi-view depth images, and determines the target pose of the robot according to the fused point cloud data; at the same time, obtains the target three-dimensional coordinates corresponding to the target object in the multi-view depth images, maps the target three-dimensional coordinates to a map, and generates a navigation path for the robot; furthermore, controls the robot to execute a target task according to the target pose and the navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located. This method fuses the point cloud data of the target object under multi-view depth images, coordinates the target pose determination step and the navigation path generation step, and updates the target pose and the navigation path in real time according to the change of the environment where the robot is located during the navigation process, not only ensuring that the robot quickly and accurately executes tasks in a complex environment, but also being able to make real-time adjustments during the task execution process, ensuring the continuity and accuracy of the task. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of the robot control method provided by the embodiment of the present invention.
[0022] Figure 2 It is a schematic overall flowchart of the robot control method provided by the embodiment of the present invention.
[0023] Figure 3 It is a schematic structural diagram of the robot control device provided by the embodiment of the present invention.
[0024] Figure 4 It is a schematic physical structure diagram of the electronic device provided by the embodiment of the present invention. Detailed Embodiments
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] It should be noted that with the continuous development of robot technology, the application of intelligent robots in industries, logistics, services, and other fields has gradually increased. In these applications, robots usually need to perform complex tasks such as grasping, placing, and navigation, especially in dynamic and challenging environments, such as object movement, environmental changes, or external instruction interference. However, the existing robot grasping and navigation systems face many difficulties in dealing with these complex tasks.
[0027] On the one hand, there are limitations in the existing robot grasping and navigation technologies. Currently, many robot grasping and navigation systems rely on depth perception technology for the positioning and operation of target objects. However, most of these systems can only handle static environments and lack the ability to adapt to real-time changes in object positions or environmental interference. For example, when an object is displaced or occluded, the existing technologies often cannot update the position of the target object in a timely manner, resulting in task execution failure.
[0028] On the other hand, there are problems with voice interference and multi-task processing in existing robots. Some existing robot systems have speech recognition capabilities for executing voice commands. However, these systems often lack the ability to cope with voice interference. For example, during task execution, if the user issues a new voice command, the existing system may not be able to interrupt the current task and flexibly adjust its behavior. Due to the lack of an effective task management and interruption mechanism, robots often cannot efficiently switch tasks or handle external interference.
[0029] In addition, existing robots have insufficient handling of multi-step tasks. When dealing with complex tasks involving multiple steps, existing robot systems also face great challenges in the continuous execution of spatial positioning, grasping, and placing tasks. Robots must be able to make adaptive adjustments when the position of an object changes or the environment changes, while most existing technologies lack this flexibility. Existing spatial positioning technologies usually rely on static object detection and cannot update environmental changes in a timely manner, resulting in task interruption or incorrect execution.
[0030] Considering this, the present invention proposes a new robot control method. Specifically, Figure 1 The flowchart of the robot control method provided by the embodiment of the present invention is shown.
[0031] As shown Figure 1 below, the method includes: S110, acquiring multi-view depth images of a target object according to a received user instruction; S120, determining point cloud data of the target object in the multi-view depth images based on the user instruction and the multi-view depth images; S130, fusing the point cloud data of the target object in the multi-view depth images and determining a target pose of the robot according to the fused point cloud data; S140, obtaining target three-dimensional coordinates corresponding to the target object in the multi-view depth images and mapping the target three-dimensional coordinates to a map to generate a navigation path of the robot; S150, controlling the robot to execute a target task according to the target pose and the navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to changes in the environment where the robot is located.
[0032] The following will elaborate on steps S110 - S150 and related steps in detail.
[0033] S110, acquiring multi-view depth images of a target object according to a received user instruction.
[0034] It is easy to understand that the robot will receive / monitor instructions from the user (user instructions) in real time. The user instructions can be voice instructions, text instructions, or visual recognition instructions (such as gestures), and no specific limitations are imposed here. When the user instruction is other than a text instruction, it is necessary to first convert the user instruction into a text instruction and format the text instruction so that the robot can understand the content of the text instruction.
[0035] After receiving the user instruction, start controlling multiple pre-installed acquisition devices to synchronously capture the target object from different positions and angles. These acquisition devices can be ordinary RGB cameras equipped with depth sensors or dedicated depth cameras. For example, when photographing a large object, place multiple cameras around the object, and each camera takes an image, thereby obtaining multi-view image data, that is, multi-view depth images. This method can capture information on all sides of the target object and provide a rich data source - the multi-view depth image - for subsequent analysis.
[0036] Among them, the number and installation positions of the acquisition devices can be set according to actual needs, and no specific limitations are imposed here. For example, in a specific embodiment, the acquisition devices include two (Realsense D435 and L515), one installed on the head of the robot and one installed on the chest or third view of the robot.
[0037] The target object can be obtained by parsing the user instruction. It can be a dynamic person or animal, or a static object, depending on the specific situation.
[0038] A multi-view depth image is composed of multiple single-view depth images. Compared with ordinary depth images, it provides more comprehensive three-dimensional scene information. A multi-view depth image can describe the shape, size, and positional relationship of an object from multiple angles, and can more accurately reflect the true form of the object. In particular, a multi-view depth image can avoid the occlusion problems that may occur in a single-view depth image.
[0039] Based on collecting the multi-view depth image of the object according to the received user instruction in step S110, further, step S120 is executed.
[0040] S120, based on the user instruction and the multi-view depth image, determine the point cloud data of the object in the multi-view depth image.
[0041] It is easy to understand that for the multiple single-view depth images in a multi-view depth image, since each pixel value in a single-view depth image represents the distance (depth value) from the camera to the surface of the object corresponding to the pixel, the pixel coordinates on the two-dimensional single-view depth image can be converted into three-dimensional space coordinates through the internal parameters of the camera (such as focal length, optical center position, etc.), and then the global point cloud data corresponding to the single-view depth image is generated.
[0042] At the same time, by inputting the modality-converted and formatted user instruction and the single-view depth image into a pre-trained Roborefer model, the two-dimensional coordinates of the object can be obtained as the output. Then, using the two-dimensional coordinates of the object as the input, SAM (Segment Anything Model) is called to generate a mask of the object.
[0043] Among them, the Roborefer model combines single-step spatial understanding and multi-step spatial reasoning to achieve precise control, especially in complex 3D environments.
[0044] Regarding the architecture of the Roborefer model. RoboRefer combines visual language models (VLMs) with 3D spatial perception, mainly including two key components: an RGB encoder and a depth encoder. Among them, the RGB encoder is used to process visual data (RGB images) to achieve spatial understanding; the depth encoder is used to separately extract 3D depth information to avoid interference with the RGB encoder, thereby enhancing spatial understanding without compromising image performance. The outputs of these two encoders are aligned with a large language model (LLM) through a projector to perform tasks such as question answering or point prediction to execute robot operation tasks.
[0045] Regarding the training process of the Roborefer model. The training of RoboRefer is divided into a supervised fine-tuning stage and a reinforcement fine-tuning stage. Among them, the goal of the supervised fine-tuning stage is to improve the model's understanding of spatial concepts, specifically including: Step 1: Depth alignment - Align depth information with text input to ensure the accuracy of depth perception; Step 2: Spatial understanding enhancement - Fine-tune the entire model on a preset dataset, which contains data supporting single-step spatial understanding and multi-step reasoning. The goal of the reinforcement fine-tuning stage is to enhance the model's ability to handle multi-step spatial reasoning, specifically including: Training through process reward functions (such as accuracy reward and process format reward) to improve intermediate reasoning steps and increase the accuracy of final predictions. Among them, the preset dataset can be the RefSpatial dataset, which contains 2.5 million samples and 20 million question-and-answer pairs, and also includes 31 spatial relationships, supporting multi-step reasoning tasks.
[0046] SAM is a brand-new deep learning-based general image segmentation model developed by Meta AI. Without further fine-tuning, it can perform high-quality instance segmentation or semantic segmentation on any object in any image. SAM has three core capabilities: (1) Zero-shot segmentation: It can segment any object you specify in the image without training; (2) Interactive segmentation: Users can tell the model what to segment through prompts such as points and bounding boxes; (3) Full-scene generality: It is applicable to various image types (natural images, medical images, remote sensing images, etc.).
[0047] SAM includes an image encoder, a prompt encoder, and a mask decoder. Among them, the image encoder uses VisionTransformer (such as ViT-B / 16) to encode the input image and extract high-dimensional feature representations. This part has a large computational overhead but only needs to run once. The prompt encoder encodes the user's prompt information (such as click points, bounding boxes, text descriptions) into vector form and combines it with the image features. The mask decoder generates high-quality binary or probability mask images based on the image features and prompt information to identify the pixel regions of the target object.
[0048] Subsequently, for a single-view depth image, the point cloud data of the target object in the single-view depth image is filtered from the global point cloud data corresponding to the single-view depth image using the mask of the target object.
[0049] Finally, based on the point cloud data of the target object in all single-view depth images, the point cloud data of the target object in the multi-view depth image can be obtained.
[0050] On the basis of determining the point cloud data of the target object in the multi-view depth image based on the user instruction and the multi-view depth image in step S120, further, step S130 is executed.
[0051] S130 fuses the point cloud data of the target object in the multi-view depth images and determines the target pose of the robot based on the fused point cloud data.
[0052] It is easy to understand that for the point cloud data of the target object in the obtained multi-view depth images, the point cloud data of the target object in multiple single-view depth images is fused to obtain the fused point cloud data, so as to determine the target pose of the robot based on the fused point cloud data.
[0053] Specifically, the point cloud data of the target object in multiple single-view depth images can be first converted to the same coordinate system (such as the world coordinate system), and then the corresponding weight values are determined according to the depth accuracy and visual reliability of the multiple single-view depth images. Thus, the point cloud data of the target object in the multi-view depth images is fused according to the weight values of the multiple single-view depth images, and the preliminary fused point cloud data can be obtained.
[0054] Subsequently, taking the fused point cloud data as the input of AnyGrasp, the target pose of the robot in the third-view camera coordinate system can be generated. The target pose can be a grasping pose, including the grasping point position and the grasping direction , or it can be a placing pose, which is not specifically limited here. Among them, represents the roll angle, represents the pitch angle, represents the yaw angle.
[0055] Among them, AnyGrasp is a general algorithm framework for 6-DoF Grasp Detection proposed by the team of the Institute for AI Industry Research (AIR), Tsinghua University, which is mainly used to efficiently and accurately predict the effective grasping poses that the robot can execute from 3D point cloud data. The core goal of AnyGrasp is: without prior knowledge, without object models, and without training samples, based only on single-view or full-view point cloud data, to generate high-quality and executable 6-DoF Grasp Poses for objects of any shape.
[0056] Finally, the target pose of the robot in the third-view coordinate camera system is converted to the robot coordinate system to facilitate the robot to execute the operations of the target pose.
[0057] It is worth mentioning that compared with the prior art, the fusion of the point cloud data of the target object in the multi-view depth images proposed in this embodiment is superior to the processing of the point cloud data of the target object in the single-view depth images and can further improve the accuracy.
[0058] S140. Obtain the target three-dimensional coordinates corresponding to the target object in the multi-view depth image, and map the target three-dimensional coordinates to the map to generate the navigation path of the robot.
[0059] It is easy to understand that for multiple single-view depth images in the multi-view depth image, by inputting the modality-converted and formatted user instructions and the single-view depth images into the pre-trained Roborefer model, the two-dimensional coordinates of the output target object can be obtained.
[0060] Then, according to the two-dimensional coordinates of the target object and in combination with the internal parameters of the camera (the acquisition device of the single-view depth image) (such as focal length, optical center position, etc.), the three-dimensional coordinates of the target object in the world coordinate system, that is, the target three-dimensional coordinates, can be obtained.
[0061] It should be noted that obtaining the target three-dimensional coordinates corresponding to the target object in the multi-view depth image here can refer only to the target three-dimensional coordinates corresponding to the target object in a single single-view depth image in the multi-view depth image, or can refer to the fusion result of the target three-dimensional coordinates corresponding to the target object in multiple single-view depth images. The fusion strategy can also be implemented according to the weight values of multiple single-view depth images, which will not be specifically elaborated here.
[0062] After determining the target three-dimensional coordinates corresponding to the target object, implement SLAM (Simultaneous Localization And Mapping) on the robot, map the target three-dimensional coordinates to the map, and then the navigation path of the robot can be generated. The navigation path refers to the path segment from the current position of the robot to the target object.
[0063] Specifically, the robot constructs an environmental map in real time in an unknown environment and synchronously locates its own position, and converts the target three-dimensional coordinates of the target object into global map coordinates. Based on the global environmental map and the target point (i.e., the position of the target object / the converted global map coordinates), a path planning algorithm is used to generate an executable navigation path.
[0064] It should be noted that there is no strict execution order between step S140 and steps S120 - S130. In a specific embodiment, step S140 and steps S120 - S130 are executed in cooperation.
[0065] After executing steps S110 - S140, execute step S150.
[0066] S150. Control the robot to execute the target task according to the target pose and navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located.
[0067] It is easy to understand that the robot moves to the vicinity of the target object along the navigation path and performs the target task on the target object according to the target pose. The target task here is determined by parsing the user instruction, which can be either a target object grasping task or a target object placing task, and is not specifically limited here.
[0068] It is worth mentioning that during the process of the robot moving to the vicinity of the target object along the navigation path, the robot will detect the position change of the target object in real time. If the position offset is greater than or equal to the set threshold, a new navigation path and target pose will be regenerated, and the target task will be executed according to the new navigation path and new target pose; if the position offset is less than the set threshold, the target task will continue to be executed according to the target pose and navigation path.
[0069] The robot in this embodiment can be UR5, or G1, or other types of intelligent robots, which are not specifically limited here.
[0070] In this embodiment, by receiving the user instruction, multi-view depth images of the target object are collected, and based on the user instruction and the multi-view depth images, the point cloud data of the target object in the multi-view depth images is determined. Then, the point cloud data of the target object in the multi-view depth images is fused, and the target pose of the robot is determined according to the fused point cloud data; at the same time, the target three-dimensional coordinates corresponding to the target object in the multi-view depth images are obtained, and the target three-dimensional coordinates are mapped to the map to generate the navigation path of the robot; furthermore, according to the target pose and navigation path of the robot, the robot is controlled to execute the target task; where the target task is determined by the user instruction, and the target pose and navigation path will be updated in real time according to the change of the environment where the robot is located. This method fuses the point cloud data of the target object under multi-view depth images, coordinates the target pose determination step and the navigation path generation step, and updates the target pose and navigation path in real time according to the change of the environment where the robot is located during the navigation process, which not only ensures that the robot can execute tasks quickly and accurately in a complex environment, but also can make real-time adjustments during the task execution process, ensuring the continuity and accuracy of the task.
[0071] On the basis of the above embodiment, further, taking the user instruction as a voice instruction and the target pose as a grasping pose as an example, the following will describe in detail the target pose determination step of the robot.
[0072] First, according to the received user instruction, multi-view depth images of the target object are collected.
[0073] Specifically, after receiving a voice command, the RGB-D images of the target object can be synchronously collected using the Realsense D435 installed on the robot's head and the Realsense L515 installed on the robot's chest or the third perspective, that is, a multi-perspective depth image that simultaneously contains color information (RGB) and depth information (Depth). The resolution of the multi-perspective depth image is 1280x720, and the depth accuracy is ±2mm.
[0074] Then, based on the user command and the multi-perspective depth image, the point cloud data of the target object in the multi-perspective depth image is determined, including: converting the voice command into a text command, and obtaining the global point cloud data corresponding to the multi-perspective depth image; obtaining the two-dimensional coordinates of the target object in the multi-perspective depth image according to the text command and the multi-perspective depth image; determining the mask of the target object in the multi-perspective depth image according to the two-dimensional coordinates of the target object in the multi-perspective depth image; and determining the point cloud data of the target object from the global point cloud data corresponding to the multi-perspective depth image according to the mask of the target object in the multi-perspective depth image, so as to obtain the point cloud data of the target object in the multi-perspective depth image.
[0075] Specifically, for multiple single-perspective depth images in the multi-perspective depth image, since each pixel value in the single-perspective depth image represents the distance (depth value) from the camera to the surface of the target object corresponding to the pixel, the pixel coordinates on the two-dimensional single-perspective depth image can be converted into three-dimensional space coordinates through the internal parameters of the camera (such as focal length, optical center position, etc.), and then the global point cloud data corresponding to the single-perspective depth image is generated, so as to obtain the global point cloud data corresponding to the multi-perspective depth image.
[0076] For the received voice command, first convert the voice command into a text command through the lightweight Whisper medium model, and then the text command can be translated into an English command through ChatGPT and formatted into a standard command, so that the robot can understand the content of the command.
[0077] Input the formatted standard command and the single-perspective depth image into the pre-trained Roborefer model, and the two-dimensional coordinates of the target object in the output single-perspective depth image can be obtained. Subsequently, using the two-dimensional coordinates of the target object as the input, call SAM to generate the mask of the target object in the single-perspective depth image. Thus, the mask of the target object in the multi-perspective depth image can be obtained.
[0078] According to the mask of the target object in each single-perspective depth image, the point cloud data of the target object is screened out from the global point cloud data corresponding to the corresponding single-perspective depth image. Thus, the point cloud data of the target object in all single-perspective depth images is obtained, that is, the point cloud data of the target object in the multi-perspective depth image.
[0079] Next, fuse the point cloud data of the target object in the multi-view depth image, including: mapping the point cloud data of the target object in each view depth image of the multi-view depth image to a unified coordinate system to obtain the target point cloud data of the target object in the multi-view depth image; determining the weight value of each view depth image in the multi-view depth image; according to the weight value of each view depth image, using weighted Kalman filtering to fuse the target point cloud data of the target object in the multi-view depth image to obtain the initial fused point cloud data; reducing the density of the initial fused point cloud data through voxel grid filtering to obtain the fused point cloud data.
[0080] For the point cloud data of the target object in the multi-view depth image, map the point cloud data of the target object in different single-view depth images to a unified world coordinate system to obtain the target point cloud data of the target object in the multi-view depth image.
[0081] For the multi-view depth image, the weight value of each single-view depth image can be determined according to its depth accuracy and view reliability. For example, the weight value of the single-view depth image collected by L515 is taken as 0.6, and the weight value of the single-view depth image collected by D435 is taken as 0.4. Then use weighted Kalman filtering to fuse the target point cloud data of the target object in each single-view depth image to obtain the initial fused point cloud data. Further, Voxel Grid filtering (voxel grid filtering, voxel size is 0.01m) can also be used to reduce the density of the initial fused point cloud data to reduce the amount of calculation.
[0082] Finally, determine the target pose of the robot according to the fused point cloud data.
[0083] Specifically, input the fused point cloud data into AnyGrasp, and the grasping pose in the third-view camera coordinate system can be generated, including the position of the grasping point and the grasping direction .
[0084] Further, through hand-eye calibration or robot base coordinate transformation, the grasping pose in the third-view camera coordinate system is transformed into the manipulator / robot coordinate system to facilitate the robot to perform the grasping operation of the target object.
[0085] In this embodiment, a multi-view point cloud fusion algorithm is adopted, which further improves the grasping accuracy and efficiency of the robot compared with the existing single-camera point cloud processing algorithm. At the same time, by combining the weighted Kalman filter and the Voxel Grid filter, the computational amount is reduced by about 30%, which is also better than the traditional point cloud processing algorithm. In a specific embodiment, in the experimental environment of 50 grasping tests with UR5 and G1 robots, the grasping success rate of this embodiment is increased from 80% of the prior art to 90%; in the experimental environment of Jetson OrinNX and an image resolution of 1280x720, the point cloud processing time of this embodiment is shortened from 0.8 seconds to 0.5 seconds, an increase of 37.5%.
[0086] On the basis of the above embodiments, further, the following will describe in detail the steps of generating the navigation path of the robot and the steps of real-time updating the target pose and the navigation path during navigation.
[0087] First, obtain the target three-dimensional coordinates corresponding to the target object in the multi-view depth image, including: determining the target three-dimensional coordinates corresponding to the target object in the multi-view depth image according to the two-dimensional coordinates of the target object in the multi-view depth image, the internal parameters of the acquisition device of the multi-view depth image, and the depth value.
[0088] Specifically, for multiple single-view depth images in the multi-view depth image, by inputting the modality conversion and formatted user instructions and the single-view depth image into the pre-trained Roborefer model, the two-dimensional coordinates of the output target object (such as a table) can be obtained.
[0089] According to the two-dimensional coordinates of the target object and the internal parameters of the camera (the acquisition device of the single-view depth image) (such as focal length, optical center position, etc.), the three-dimensional coordinates of the target object in the world coordinate system, that is, the target three-dimensional coordinates, can be obtained.
[0090] It should be noted that obtaining the target three-dimensional coordinates corresponding to the target object in the multi-view depth image here can refer only to the target three-dimensional coordinates corresponding to the target object in a single single-view depth image in the multi-view depth image, or to the fusion result of the target three-dimensional coordinates corresponding to the target object in multiple single-view depth images. The fusion strategy can also be implemented according to the weight values of multiple single-view depth images, which will not be specifically elaborated here.
[0091] Then, map the target three-dimensional coordinates to the map to generate the navigation path of the robot.
[0092] Specifically, after determining the target three-dimensional coordinates corresponding to the target object, SLAM is implemented on the robot, and the target three-dimensional coordinates are mapped to the map, and then the navigation path of the robot can be generated. Specifically, the robot constructs an environmental map in real time in an unknown environment and synchronously locates its own position, and converts the target three-dimensional coordinates of the target object into global map coordinates. Based on the global environmental map and the target point (i.e., the position of the target object / the converted global map coordinates), a path planning algorithm is used to generate an executable navigation path.
[0093] Furthermore, according to the target pose and navigation path of the robot, the robot is controlled to execute the target task, including: collecting new multi-view depth images in real time, and obtaining the current two-dimensional coordinates of the target object in the new multi-view depth images in real time; obtaining the offset difference between the current two-dimensional coordinates and the two-dimensional coordinates; in the case where the offset difference is greater than or equal to the preset dynamic threshold, updating the target pose and navigation path of the robot, and controlling the robot to execute the target task according to the updated target pose and navigation path; in the case where the offset difference is less than the preset dynamic threshold, controlling the robot to execute the target task according to the target pose and navigation path of the robot.
[0094] Specifically, during the process of the robot moving towards the vicinity of the target object according to the navigation path, multi-view depth images of the target object are collected in real time, and the Roborefer model is called in real time to obtain the current two-dimensional coordinates of the target object in the newly collected multi-view depth images, so as to detect the position change of the target object.
[0095] If the offset difference between the current two-dimensional coordinates of the target object and the two-dimensional coordinates of the target object obtained last time is greater than or equal to the preset dynamic threshold, the target pose and navigation path of the robot are regenerated, and the robot is controlled to execute the target task according to the regenerated target pose and navigation path.
[0096] If the offset difference between the current two-dimensional coordinates of the target object and the two-dimensional coordinates of the target object obtained last time is less than the preset dynamic threshold, there is no need to regenerate, and the robot can directly execute the target task according to the previously determined target pose and navigation path.
[0097] Among them, the preset dynamic threshold can be set according to actual needs, and no specific limitation is made here. For example, in a specific embodiment, the preset dynamic threshold takes a value of 0.1m.
[0098] In this embodiment, by introducing a preset dynamic threshold, the target pose and navigation path are updated in real time during navigation, which not only solves the problem of poor collaboration in a dynamic environment but also improves the robustness. In a specific embodiment, the navigation accuracy is increased by 20%, and the positioning error is reduced from 0.15 m to 0.12 m (experimental environment: G1 robot, 5x5 m dynamic environment). The success rate of grasping and navigation collaboration reaches 92%, supporting dynamic obstacle scenarios (such as moving objects).
[0099] On the basis of the above embodiment, further, taking the user instruction as a voice instruction as an example, the following will elaborate on the real-time voice interruption and task switching process provided by the present invention.
[0100] To achieve real-time voice control during robot movement, this embodiment uses a finite state machine (FSM) to monitor changes in voice instructions. If a new voice instruction is detected, a lightweight Whisper medium model (with a 20% reduction in the number of parameters through model pruning) is used to process the voice instructions input by the user in real time on a Jetson Orin NX, with a sampling rate of 16 kHz and a frame length of 30 ms.
[0101] Specifically, when the robot receives a voice instruction, it first converts the voice instruction into a text instruction through the lightweight Whisper medium model, and then the text instruction can be translated into an English instruction through ChatGPT and formatted into a standard instruction so that the robot can understand the content of the instruction. For example, "Grasp the red object" is converted to "Grasp redobject at position X".
[0102] For the newly received standard instruction, the robot is controlled to stop executing the current task (such as the movement of the UR5 robotic arm) through the terminal process, and the target task is switched to be executed according to the new standard instruction. That is to say, in this embodiment, a task priority queue can be designed through ROS2 Topic, and the priority rule is that "new user instructions take precedence over the current task".
[0103] In this embodiment, by combining real-time speech recognition technology (such as using the Whisper ASR model) and a task interruption mechanism, new voice instructions can be monitored during task execution. When a new voice instruction is detected, the system interrupts the current task and immediately replans the task execution according to the new voice instruction. This mechanism ensures that the robot can flexibly respond to the user's voice input during task execution and better adjust and execute tasks in a complex environment.
[0104] In some embodiments, considering that when importing fused point cloud data into AnyGrasp to calculate the grasping pose, there may be grasping poses with the grasping direction upward or parallel to the desktop, this embodiment makes some improvements.
[0105] Specifically, assuming that the target object is placed on the plate, when using SAM to determine the mask of the target object, the masks with the top three confidences are used as the mask of "target object + plate", which is also the final mask of the target object.
[0106] Assuming that the target object is not placed on the plate, after using SAM to determine the mask of the target object, the mask of the target object is inflated by one circle around it to expand the range of the retained point cloud. Specifically, with the center of the mask of the target object as the center of the circle, the largest inscribed circle of the expanded mask is taken as the final mask of the target object.
[0107] Among them, with the center of the mask as the center of the circle, the abscissa of the center is taken as the average value of the abscissas of all points in the mask of the target object, and the ordinate of the center is taken as the average value of the ordinates of all points in the mask of the target object.
[0108] The purpose of the above mask processing for the target object is to make the fused point cloud data input to AnyGrasp retain the point cloud data of the plane where the object is located.
[0109] According to the final mask of the target object in each single-view depth image, the point cloud data of the target object is screened out from the global point cloud data corresponding to the corresponding view depth image and fused to obtain the fused point cloud data.
[0110] Furthermore, the fused point cloud data is input to AnyGrasp, and the grasping postures generated by AnyGrasp are screened, and the grasping postures with the grasping direction upward or parallel to the desktop are deleted. Specifically, by converting the grasping postures to the robotic arm coordinate system through a transformation matrix, the grasping postures with too large pitch and yaw angles are deleted, and then the grasping posture with the highest confidence in the remaining grasping postures is taken as the target posture.
[0111] In some other embodiments, Figure 2 The overall flow diagram of the robot control method provided by the embodiments of the present invention is shown, in which the determination process of the grasping posture and the target tracking and dynamic adjustment process (i.e., the process of updating the grasping posture and the navigation path) are shown in detail. The specific details in the figure can be referred to the above embodiments and will not be elaborated here.
[0112] In one embodiment, in an industrial scenario, a UR5 robotic arm and a Realsense L515 camera are used. The voice command "grab the red object" is input. Whisper transcribes the command, ChatGPT4 translates it into English, Roborefer outputs the two-dimensional coordinates (0.5, 0.3) of the target object, SAM generates a mask of the target object, and AnyGrasp generates a grasping pose, and the UR5 performs the grasping. Experimental results: The success rate of 50 grasps is 90%, and the response time is 0.5 seconds.
[0113] In another embodiment, in a service scenario, the G1 robot acquires multi-view depth images of the target table through a D435 camera, and the L515 camera is used for guiding navigation. A navigation path is generated through SLAM navigation, the position of the target object is updated in real time, and the grasping pose is generated based on the multi-view depth images acquired by the D435. Experimental results: The navigation error is 0.12 m, and the grasping success rate is 92%.
[0114] Through the above technical means, the robot control method provided by the embodiments of the present invention achieves the following effects: The voice response delay is reduced to 0.5 seconds, the task switching success rate is 95%, the point cloud processing time is shortened by 37.5%, the grasping success rate is increased to 90%, the navigation accuracy is improved by 20%, and the cooperation success rate in a dynamic environment is 92%.
[0115] Corresponding to the robot control method described in the above embodiments, the present invention also provides a robot control device. Specifically, Figure 3 The structural schematic diagram of the robot control device provided by the embodiments of the present invention is shown.
[0116] As Figure 3 shown, the device includes: a multi-view depth image acquisition module 310, configured to acquire multi-view depth images of a target object according to a received user instruction; a target object point cloud data determination module 320, configured to determine point cloud data of the target object in the multi-view depth image based on the user instruction and the multi-view depth image; a target pose determination module 330, configured to fuse the point cloud data of the target object in the multi-view depth image and determine the target pose of the robot according to the fused point cloud data; a navigation path generation module 340, configured to obtain the target three-dimensional coordinates corresponding to the target object in the multi-view depth image and map the target three-dimensional coordinates to a map to generate a navigation path of the robot; a target task execution module 350, configured to control the robot to execute a target task according to the target pose and navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located.
[0117] In this embodiment, the multi-view depth image acquisition module 310 acquires multi-view depth images of the target object according to the received user instruction. The target object point cloud data determination module 320 determines the point cloud data of the target object in the multi-view depth image based on the user instruction and the multi-view depth image. Then, the target pose determination module 330 fuses the point cloud data of the target object in the multi-view depth image and determines the target pose of the robot according to the fused point cloud data. At the same time, the navigation path generation module 340 obtains the target three-dimensional coordinates corresponding to the target object in the multi-view depth image, maps the target three-dimensional coordinates to the map, and generates the navigation path of the robot. Furthermore, the target task execution module 350 controls the robot to execute the target task according to the target pose and navigation path of the robot. Among them, the target task is determined by the user instruction, and the target pose and navigation path will be updated in real time according to the change of the environment where the robot is located. This device fuses the point cloud data of the target object under the multi-view depth image, coordinates the target pose determination step and the navigation path generation step, and updates the target pose and navigation path in real time according to the change of the environment where the robot is located during the navigation process, which not only ensures that the robot can execute tasks quickly and accurately in a complex environment, but also can make real-time adjustments during the task execution process, ensuring the continuity and accuracy of the task.
[0118] It should be noted that the robot control device provided in the embodiment of the present invention can be correspondingly referred to the robot control method described in the above embodiments, and will not be elaborated here.
[0119] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the robot control method, and the method includes: acquiring multi-view depth images of the target object according to the received user instruction; determining the point cloud data of the target object in the multi-view depth image based on the user instruction and the multi-view depth image; fusing the point cloud data of the target object in the multi-view depth image and determining the target pose of the robot according to the fused point cloud data; obtaining the target three-dimensional coordinates corresponding to the target object in the multi-view depth image, mapping the target three-dimensional coordinates to the map, and generating the navigation path of the robot; controlling the robot to execute the target task according to the target pose and navigation path of the robot; among them, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located.
[0120] In addition, when the logical instructions in the memory 430 described above can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0121] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the robot control method provided by the above-mentioned various methods. The method includes: collecting multi-view depth images of a target object according to received user instructions; determining point cloud data of the target object in the multi-view depth images based on the user instructions and the multi-view depth images; fusing the point cloud data of the target object in the multi-view depth images and determining the target pose of the robot according to the fused point cloud data; obtaining the target three-dimensional coordinates corresponding to the target object in the multi-view depth images and mapping the target three-dimensional coordinates to a map to generate a navigation path for the robot; controlling the robot to execute a target task according to the target pose and navigation path of the robot; wherein, the target task is determined by the user instructions, and the target pose and the navigation path will be updated in real time according to the environment where the robot is located.
[0122] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a robot control method provided by the above-mentioned various methods. The method includes: collecting multi-view depth images of a target object according to a received user instruction; determining point cloud data of the target object in the multi-view depth images based on the user instruction and the multi-view depth images; fusing the point cloud data of the target object in the multi-view depth images, and determining a target pose of the robot according to the fused point cloud data; obtaining target three-dimensional coordinates corresponding to the target object in the multi-view depth images, and mapping the target three-dimensional coordinates to a map to generate a navigation path of the robot; controlling the robot to execute a target task according to the target pose and the navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0125] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A robot control method, characterized in that, Including: Collect multi-view depth images of a target object according to a received user instruction; Based on the user instruction and the multi-view depth images, determine the point cloud data of the target object in the multi-view depth images; Fuse the point cloud data of the target object in the multi-view depth images, and determine the target pose of the robot according to the fused point cloud data; Obtain the target three-dimensional coordinates corresponding to the target object in the multi-view depth images, and map the target three-dimensional coordinates to a map to generate a navigation path for the robot; Control the robot to execute a target task according to the target pose and navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located.
2. The robot control method according to claim 1, wherein The user instruction is a voice instruction; Correspondingly, the determining the point cloud data of the target object in the multi-view depth images based on the user instruction and the multi-view depth images includes: Convert the voice instruction into a text instruction, and obtain the global point cloud data corresponding to the multi-view depth images; According to the text instruction and the multi-view depth images, obtain the two-dimensional coordinates of the target object in the multi-view depth images; According to the two-dimensional coordinates of the target object in the multi-view depth images, determine the mask of the target object in the multi-view depth images; According to the mask of the target object in the multi-view depth images, determine the point cloud data of the target object from the global point cloud data corresponding to the multi-view depth images, and obtain the point cloud data of the target object in the multi-view depth images.
3. The robot control method according to claim 1, wherein The fusing the point cloud data of the target object in the multi-view depth images includes: Map the point cloud data of the target object in each view depth image in the multi-view depth images to a unified coordinate system to obtain the target point cloud data of the target object in the multi-view depth images; Determine the weight value of each view depth image in the multi-view depth images; According to the weight value of each view depth image, use weighted Kalman filtering to fuse the target point cloud data of the target object in the multi-view depth images to obtain initial fused point cloud data; Reduce the density of the initial fused point cloud data through voxel grid filtering to obtain the fused point cloud data.
4. The robot control method according to claim 2, characterized in that, The obtaining the target three-dimensional coordinates corresponding to the target object in the multi-view depth images includes: According to the two-dimensional coordinates of the target object in the multi-view depth images, and the internal parameters and depth values of the acquisition device of the multi-view depth images, determine the target three-dimensional coordinates corresponding to the target object in the multi-view depth images.
5. The robot control method according to claim 2, wherein, The controlling the robot to execute a target task according to the target pose and navigation path of the robot includes: Collect new multi-view depth images in real time, and obtain the current two-dimensional coordinates of the target object in the new multi-view depth images in real time; Obtain the offset difference between the current two-dimensional coordinates and the two-dimensional coordinates; In the case where the offset difference is greater than or equal to a preset dynamic threshold, update the target pose and navigation path of the robot, and control the robot to execute the target task according to the updated target pose and navigation path. When the offset difference is less than a preset dynamic threshold, control the robot to execute a target task according to the target pose and navigation path of the robot.
6. The robot control method according to claim 1, wherein The user instruction is a voice instruction; When the voice instruction is received, control the robot to stop executing the current task and switch to execute the target task.
7. A robot control device, characterized in that, It includes: A multi-view depth image acquisition module, configured to acquire multi-view depth images of a target object according to a received user instruction; A target object point cloud data determination module, configured to determine the point cloud data of the target object in the multi-view depth image based on the user instruction and the multi-view depth image; A target pose determination module, configured to fuse the point cloud data of the target object in the multi-view depth image and determine the target pose of the robot according to the fused point cloud data; A navigation path generation module, configured to obtain the target three-dimensional coordinates corresponding to the target object in the multi-view depth image and map the target three-dimensional coordinates to a map to generate a navigation path of the robot; A target task execution module, configured to control the robot to execute a target task according to the target pose and navigation path of the robot; wherein, the target task is determined by the user instruction, and the target pose and the navigation path will be updated in real time according to the change of the environment where the robot is located.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the robot control method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the robot control method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the robot control method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional point cloud-based patrol robot vision system and control method
CN108171796A
Robot control method and device and robot
CN116619374A
Robot free grabbing control method fused with target detection and robot
CN118418128A
Mechanical arm control method and system based on voice control and multi-source sensing
CN120080316A
Method and system for providing remote robotic control
US20200117212A1