Interactive robot grabbing method based on large language model
By introducing large language models and image segmentation models into the robot grasping system, combined with the crawling prediction model, the problems of insufficient semantic understanding capabilities and lack of refined grasping strategies in the existing technology are solved, and more efficient and flexible robot grasping operations are achieved.
Patent Information
- Application Number
- CN202510502298.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-06-27
AI Technical Summary
The existing natural language-driven robot crawling technology has problems such as limited semantic understanding capabilities, lack of refinement of crawling strategies, and insufficient generalization capabilities, making it difficult to achieve efficient and flexible crawling operations in complex environments.
An interactive robot crawling method based on a large language model is adopted to analyze dialogue interaction information through a large language model, combine the image segmentation model and the grab prediction model, generate task operation sequences, identify component information of the target object and grab position, and realize refined grab operations.
It improves the robot's ability to parse complex semantic instructions, improves its adaptability in refined operation tasks, and realizes an intelligent, more accurate and more generalized interactive robot crawling solution.
Smart Images

Figure CN120206528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics, and in particular to an interactive robot grasping method based on a large language model. Background Art
[0002] Interactive robot grasping refers to the robot dynamically adjusting the grasping strategy according to the information obtained through real-time interaction with the user, so as to achieve efficient and flexible grasping operations in an unstructured environment. As one of the core media for human-computer interaction, natural language enables the robot to parse and execute voice or text instructions with the help of natural language processing technology to complete specific tasks due to its universality and flexibility. The natural language-driven robot grasping technology has broad application prospects in many fields such as smart home, medical rehabilitation, industrial automation, and service robots.
[0003] Although the existing natural language-driven robot grasping technologies have made certain progress, the existing methods still have problems such as limited semantic understanding ability, lack of refinement in grasping strategies, and insufficient generalization ability. Summary of the Invention
[0004] The present invention provides an interactive robot grasping method based on a large language model, whose main purpose is to improve the robot's parsing ability for complex semantic instructions and also enhance the robot's adaptability in refined operation tasks.
[0005] In a first aspect, an embodiment of the present invention provides an interactive robot grasping method based on a large language model, including:
[0006] Based on the two-dimensional image, dialogue interaction information, and large language model of the current scene, obtain a task operation sequence, where the task operation sequence includes the target object, component information of the target object, target grasping position information, and action information;
[0007] Based on the two-dimensional image and the target grasping position information, in combination with an image segmentation model, obtain a target mask region, and based on the target mask region and the depth map of the current scene, obtain a three-dimensional point cloud map with the target grasping position;
[0008] Based on the action information, the three-dimensional point cloud map with the target grasping position, and a grasping prediction model, obtain a target grasping pose, and based on the target grasping pose, control the robot to grasp the target object.
[0009] Further, the step of obtaining the target mask region based on the two-dimensional image and the target grasping position information in combination with an image segmentation model includes:
[0010] Input the two-dimensional image and the target grasping position information into the image segmentation model to obtain the target mask region;
[0011] Dilate the target mask region to obtain the dilated target mask region, and re-use the dilated target mask region as the target mask region.
[0012] Further, the step of obtaining a three-dimensional point cloud map with target grasping positions according to the target mask region and the depth map of the current scene includes:
[0013] Generate a binary image according to the target mask region;
[0014] Convert the two-dimensional image into a three-dimensional point cloud map with target grasping positions according to the binary image and the depth map.
[0015] Further, the step of obtaining a task operation sequence according to the two-dimensional image, the dialogue interaction information, and the large language model of the current scene includes:
[0016] Obtain the two-dimensional image and the depth image of the current scene through a depth camera;
[0017] Input the two-dimensional image and the dialogue interaction information into the large language model to obtain the task operation sequence, and store the task operation sequence in JSON format.
[0018] Further, the step of obtaining the target grasping pose according to the action information, the three-dimensional point cloud map with target grasping positions, and the grasping prediction model includes:
[0019] Input the action information and the three-dimensional point cloud map with target grasping positions into the grasping prediction model to obtain multiple alternative grasping poses;
[0020] Select the target grasping pose from the multiple alternative grasping poses.
[0021] Further, the component information of the target object includes the components of the target object and the colors of the components, and the action information is one of the four actions of detection, grasping, handing over, and placing.
[0022] In a second aspect, an embodiment of the present invention provides an interactive robot grasping system based on a large language model, including a depth camera, a processor, and a robot. The depth camera is used to capture a two-dimensional image and a depth map of the current scene; the processor is used to interact with the user, obtain dialogue interaction information, and execute an interactive robot grasping method based on a large language model provided in the first aspect according to the dialogue interaction information, the two-dimensional image, and the depth map.
[0023] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned interactive robot grasping method based on a large language model are implemented.
[0024] In a fourth aspect, an embodiment of the present invention provides a computer storage medium. The computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned interactive robot grasping method based on a large language model are implemented.
[0025] For the interactive robot grasping method based on a large language model proposed by the present invention, first, the large language model is used to parse the dialogue interaction information to obtain a task operation sequence. Through the large language model, the fuzzy or implicit instructions in the dialogue interaction information can be accurately parsed, so that the true intention of the user can be accurately judged. Then, an image segmentation model is used to identify the target grasping position in the two-dimensional image. Through the two-dimensional component segmentation model and local point cloud positioning, the robot can accurately identify different functional components of the target object, ensuring the rationality and safety of grasping. Finally, a grasping prediction model is used to analyze the action information and three-dimensional point cloud to obtain the target grasping pose, and kinematic analysis is performed to control the robot to grasp the target grasping position.
[0026] Through the natural language interaction driven by the large language model in the embodiment of the present invention, combined with context reasoning, task decomposition, and component-level grasping optimization, the limitations of the prior art are broken through, and an interactive robot grasping solution with higher intelligence, higher precision, and stronger generalization ability is realized. This method not only improves the parsing ability of the robot for complex semantic instructions, but also enhances the adaptability of the robot in fine operation tasks, providing new technical support for future human-robot collaboration, service robots, and intelligent manufacturing. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flowchart of an interactive robot grasping method based on a large language model provided by an embodiment of the present invention;
[0028] Figure 2 It is an overall framework diagram of an interactive robot grasping method based on a large language model provided by an embodiment of the present invention;
[0029] Figure 3 It is a schematic diagram of the grasping effect of an interactive robot grasping method based on a large language model provided by an embodiment of the present invention;
[0030] Figure 4 It is a structural diagram of an interactive robot grasping system based on a large language model provided by an embodiment of the present invention.
[0031] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0032] The following details the implementation manners of the present application. The examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.
[0033] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0034] In the embodiments of the present application, "at least one" means one or more; "a plurality" means two or more. In the description of the present application, terms such as "first", "second", "third", etc. are only used for the purpose of distinguishing descriptions and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0035] The reference to "an implementation manner" or "some implementation manners" etc. in this specification means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more implementation manners of the present application. Thus, the terms "including", "comprising", "having" and their variants in this specification all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0036] Traditional robot grasping methods based on natural language driving mainly rely on rule matching or template matching strategies. These methods perform well when dealing with instructions with clear semantics and simple structures. However, for instructions with ambiguous semantics, implicit intentions, or strong context dependence, it is often difficult to correctly interpret the true intention of the instructions, resulting in the inability to execute the grasping operation as expected. In addition, traditional methods only focus on object-level grasping planning and lack refined analysis of object part-level features. For example, when grasping a cup, only the cup body can be grasped, and the cup handle cannot be grasped, thus limiting the application ability of the robot in tasks that require precise grasping of specific parts.
[0037] Aiming at the limited semantic parsing ability in traditional methods, it is difficult to accurately understand fuzzy or implicit instructions, resulting in the robot's inability to correctly judge the user's intention when performing complex tasks; and the grasping strategy is rigid, lacking analysis of object part-level features, and it is difficult to adapt to task scenarios that require refined grasping, such as safely grasping fragile items or tools; the generalization ability is insufficient, and it is difficult to handle new objects or unseen scenarios, restricting the adaptability of the robot in actual operations.
[0038] Figure 1 The flowchart of an interactive robot grasping method based on a large language model is provided for an embodiment of the present invention. Figure 2 The overall framework diagram of an interactive robot grasping method provided for an embodiment of the present invention. This overall system framework can include three parts: 1. Perception and Inference, 2. Point Cloud Localization, 3. Grasp Pose Detection.
[0039] As Figure 1 and Figure 2 shown, the method includes:
[0040] S110, obtaining a task operation sequence according to the two-dimensional image, dialogue interaction information, and large language model of the current scene. The task operation sequence includes the target object, the part information of the target object, the target grasping position information, and the action information.
[0041] In the embodiment of the present invention, first, the two-dimensional image of the current scene and the dialogue interaction information between the user and the robot are obtained. The two-dimensional image is obtained by the camera taking pictures of all the objects to be grasped in the current scene. As Figure 2As shown in the "Current Scene", in the embodiments of the present invention, it is necessary to grasp objects such as cups, baskets, and fruits on the shelf. Therefore, the scene is first photographed by a camera, and the obtained image is used as the two-dimensional image of the current scene; the dialogue interaction information can be text interaction information or voice interaction information. For example, the text interaction information is as follows:
[0042] {
[0043] System: I have understood the current environment. Please give instructions.
[0044] User: I need to drink some hot water.
[0045] System: I found a cup that can hold hot water in the cabinet. I can fill it up for you.
[0046] User: Confirm the execution. The hot water is very hot. Please be careful when handing it to me.
[0047] System: Okay, I will hold the handle of the cup to avoid direct contact with the high-temperature surface. The operation process is as follows:
[0048] }
[0049] Specifically, the above two-dimensional image and dialogue interaction information are input into the large language model of the fine-tuned GPT 4o. The large language model uses the two-dimensional image as environmental information. After parsing the dialogue interaction information, it outputs a task operation sequence. In the embodiments of the present invention, the task operation sequence has a fixed format, generally including the target object, the component information of the target object, the target grasping position, and the action information. Among them, the target object is the object that the robot needs to grasp. Taking the cup as an example, the component information of the target object is what components the target object specifically includes. For example, the component information of the cup is the cup body, the color of the cup body, the cup handle, and the color of the cup handle; the target grasping position information refers to the specific position where the robot grasps the target object, which can be determined according to environmental information such as dialogue interaction information and two-dimensional images. For example, in the above dialogue interaction information, it can be known that since the cup contains hot water, the target grasping position is the cup handle; the action information refers to the action of the robot on the target object. There are 4 types of this action information, namely detection, grasping, handing, and placing. Different types of action information result in different robot operations.
[0050] Among them, detection means identifying the target object in the environment through a depth camera, grasping means performing a grasping operation on the target object using the gripper at the end of the robotic arm, handing means delivering the grasped object to the user, and placing means placing the grasped object at a specified position.
[0051] In the above example, the format of the task operation sequence can be as follows:
[0052] {
[0053] Task: Submit a cup,
[0054] Step 1: Detect the target object. The component information of the target object is: cup body, white, and cup handle, white;
[0055] Step 2: Grasp the target object. The target grasping position information is: cup handle;
[0056] Step 3: Submit the target object.
[0057] }
[0058] It can be seen from this task operation sequence that the task sequence includes the target object to be grasped, the component information of the target object, the target grasping position, and actions. The formatted task operation sequence can ensure that the task operation sequence output by the large language model contains necessary information, and this fixed-format task operation sequence is easy to recognize, reducing the subsequent parsing requirements for the task operation sequence.
[0059] In the embodiment of the present invention, by utilizing the common sense information and context reasoning ability of the large language model, accurate parsing of fuzzy instructions, complex semantic expressions, and implicit intentions is realized, and an open interaction framework connecting users, large language models, and robots is constructed. And through fine-tuning the large language model, the task information obtained by parsing the instructions will be converted into a formatted task operation sequence with clear steps, which is convenient for interpreting the task operation sequence in subsequent steps.
[0060] S120, according to the two-dimensional image and the target grasping position information, combined with the image segmentation model, obtain the target mask region, and according to the target mask region and the depth map of the current scene, obtain a three-dimensional point cloud map with the target grasping position;
[0061] Input the two-dimensional image and the target grasping position information into the image segmentation model. The image segmentation model locates the target mask region in the two-dimensional image, and this target region is the region corresponding to the target grasping position in the two-dimensional image. As shown in Part 2 of Figure 2 the target grasping position information is the cup handle of the cup, and the target mask region located in the two-dimensional image is as shown in the image pointed by the red arrow in Figure 2 which is the region including the whole cup, and the cup handle of the cup is marked in green in this target mask region.
[0062] Specifically, the image segmentation model can be VL-Part, which is an open vocabulary part segmentation model and can handle object segmentation tasks more densely. Input the two-dimensional image into VL-Part, and the target mask region corresponding to the target grasping position in the two-dimensional image can be segmented.
[0063] And dilate the target mask region. By sliding a suitable dilation kernel to perform dilation operation on the generated target mask region, it can be ensured that the dilated target mask region can cover the complete edge of the target object and the surrounding background geometric information, avoiding inappropriate grasping positions generated subsequently due to the lack of background information and the interference of edge noise. Additionally, through the dilation operation, it can be ensured that the target mask region can cover the complete background information of the target object, so that the robot has a certain buffer area when grasping the target object, avoiding the problem that the target object is broken when the robot hits the target object. As Figure 2 shown, "Expansion Operation" represents the dilation operation. After performing the dilation operation on the cup handle, the dilated target mask region is obtained. The cup handle in the dilated target mask region is marked in orange. It can be seen that compared with the cup handle marked in green, the orange cup handle occupies a larger area in the figure, leaving a larger redundant space for subsequent robot operations and making it less likely to damage the cup.
[0064] Then, when constructing the scene point cloud in combination with the depth map of the current scene, according to the guidance of the target mask region, the three-dimensional point cloud map corresponding to the target grasping position is mapped, and the target grasping position is specifically marked in the three-dimensional point cloud map. Specifically, a binary image is generated using the target mask region. The binary image is Figure 2 represented by ROI in the figure. The cup handle of interest in the figure is represented by white pixel points, and other non-interested regions are represented by black pixel points. The binary image and the depth map are synthesized to obtain a three-dimensional point cloud map with the target grasping position. The three-dimensional point cloud map is Figure 2 represented by "Point Cloud" in the figure.
[0065] In the embodiment of the present invention, the target mask region is converted into a binary map for identifying the region of interest for grasping. On this basis, the depth map is converted into a single-view point cloud with camera internal parameters. The binary map is used for registration with the three-dimensional point cloud to focus on the local point cloud of the target object or component from the global point cloud, providing accurate position information and geometric details of the object for subsequent grasping pose estimation.
[0066] S130, obtain the target grasping pose according to the action information, the three-dimensional point cloud map with the target grasping position, and the grasping prediction model, and control the robot to grasp the target object according to the target grasping pose.
[0067] Input the action information and the 3D point cloud map with the target grasping position into the grasping prediction model, and the target grasping pose can be obtained. Among them, the grasping prediction model can be a 6-degree-of-freedom grasping detection network, which is used for object grasping in a cluttered scene. The 6-degree-of-freedom grasping detection network mainly consists of two modules: the grasping heat map model (abbreviated as GHM) and the non-uniform multi-grasp generator (abbreviated as NMG). GHM extracts the semantic features of the 3D point cloud map through an efficient convolutional neural network and generates four grasping heat maps as guidance for the grasping area. NMG then aggregates local points into the grasping area using the generated heat maps and detects the grasping within the local area through a lightweight point encoder.
[0068] The target grasping pose refers to the specific grasping position and angle when the robot grasps the target object, and kinematic analysis is performed according to the target grasping pose, so as to control the movement of the robot and complete the grasping of the target object.
[0069] Specifically, inputting the action information and the 3D point cloud map with the target grasping position into the 6-degree-of-freedom grasping detection network will generate multiple alternative grasping poses, such as Figure 2 shown in the 3rd part in the figure, Local Top15 represents the 15th among the 15 alternative grasping poses. Select the one with the highest score as the target grasping pose from these 15 alternative grasping poses, that is, LocalTop1 in the figure. As can be seen from the figure, the purple identification line represents the grasping posture of the robot.
[0070] Finally, perform motion planning according to the target grasping pose and control the robot to grasp the cup handle, as Figure 2 shown. The starting position of the robot is shown as ① in the figure. The robot first moves to the position where the cup is located to grasp the cup, and then picks up the cup to the position shown as ② to complete the grasping of the cup.
[0071] Figure 3 This is a schematic diagram of the grasping effect of an interactive robot grasping method based on a large language model provided by an embodiment of the present invention. As Figure 3 shown, Scene in the figure represents the two-dimensional image (RGB Image in the figure) and the depth image (DepthImage in the figure) obtained by the depth camera taking pictures of the current scene; the user inputs text instructions (DepthImage in the figure) in the computer to interact with the system (Open-ended) to obtain dialogue interaction information, and then based on the two-dimensional image, the depth image and the dialogue interaction information, perform perception and reasoning, point cloud position prediction, and grasping pose detection in sequence to obtain the target grasping pose, and control the robot to grasp the target object through the target grasping pose (6-DoF Grasp Pose).
[0072] And in Figure 3Three interaction instructions are given, which are: a simple instruction "bring me the banana", a general instruction "hand me the hammer in the optimal manner", and a complex instruction "select suitable tools and grasping points for woodworking".
[0073] These three instructions are input into the system in sequence. For example, for the simple instruction, the magenta and red markings in the figure indicate the postures for grasping the banana, achieving object-level grasping of the target object; for the general instruction, the general instruction means handing me the hammer in the optimal manner, and the optimal posture for grasping the hammer is marked in the figure, achieving part-level grasping of the object; for the complex instruction, the optimal posture for grasping the handle of the hammer is marked in the figure (part-level grasping).
[0074] From the above three instructions, it can be seen that even for complex instructions, this method can well identify the target grasping position, enabling the robot to complete the grasping.
[0075] In summary, the embodiment of the present invention provides an interactive robot grasping method based on a large language model, which can perform fine-grained analysis on the target object and convert natural language instructions of different complexities into an operation sequence with clear intent and executability; through the local point cloud positioning module, the present invention realizes refined grasping target positioning from the object level to the part level, significantly improving the accuracy and flexibility of robot grasping.
[0076] Compared with the traditional robot grasping method, the present invention has the following advantages:
[0077] 1. Enhanced natural language interaction and task understanding ability. The present invention adopts a fine-tuned large language model, enabling the robot to understand complex natural language instructions in combination with environmental information and autonomously plan grasping strategies without manually writing rules or presetting fixed instructions. It supports fuzzy instruction parsing and has common sense reasoning ability.
[0078] 2. Adaptive grasping strategy and enhanced generalization ability. Traditional methods rely on fixed feature templates or object databases and have limited generalization ability, making it difficult to handle new types of objects. The present invention utilizes the common sense reasoning ability of large language models to enable the robot to autonomously infer the functions of objects and appropriate grasping methods based on language descriptions. For example, even when faced with a new type of water cup, the robot can infer that "the handle of the water cup is an ideal grasping point" or "the glass needs to be handled gently", thus achieving more natural and stronger generalization ability operations.
[0079] 3. Enhanced human-robot collaboration ability. The present invention supports multi-round conversations and can provide real-time feedback and confirmation to the user during task execution. For example, the robot can ask "Do you want to use this cup?" or "The water temperature is high, and I will hold the handle and hand it to you", thereby improving the controllability and safety of task execution. This interaction method enables the robot to better adapt to the needs of different users and significantly enhances the human-robot collaboration experience.
[0080] Figure 4 FIG. is a structural diagram of an interactive robot grasping system based on a large language model provided by an embodiment of the present invention, as Figure 4 shown. The system includes a depth camera, a processor, and a robot. The depth camera is used to capture a two-dimensional image and a depth map of the current scene. The processor is used to interact with the user, obtain dialogue interaction information, and execute an interactive robot grasping method based on a large language model according to the dialogue interaction information, the two-dimensional image, and the depth map.
[0081] This embodiment is a system embodiment corresponding to the above method, and its specific implementation process is the same as that of the above method embodiment. For details, reference can be made to the above method embodiment, and this system embodiment will not be elaborated here.
[0082] Each module in the above interactive robot grasping system based on a large language model can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0083] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a computer storage medium and an internal memory. The computer storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the computer storage medium. The database of the computer device is used to store data generated or obtained during the execution of an interactive robot grasping method based on a large language model, such as two-dimensional images, depth images, and dialogue interaction information. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements an interactive robot grasping method based on a large language model.
[0084] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of an interactive robot grasping method based on a large language model in the above embodiment. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in an embodiment of an interactive robot grasping system based on a large language model.
[0085] In one embodiment, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of an interactive robot grasping method based on a large language model in the above embodiment. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in an embodiment of the above interactive robot grasping system based on a large language model.
[0086] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0087] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0088] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. An interactive robot crawling method based on a large language model, characterized in that: include: Obtaining a task operation sequence according to a two-dimensional image of the current scene, dialogue interaction information, and a large language model, wherein the task operation sequence includes a target object, component information of the target object, target grasping position information, and action information; According to the two-dimensional image and the target grab position information, in combination with an image segmentation model, a target mask area is obtained, and according to the target mask area and a depth map of the current scene, a three-dimensional point cloud image with the target grab position is obtained; The target grasping posture is obtained according to the action information, the three-dimensional point cloud image with the target grasping position and the grasping prediction model, and the robot is controlled to grasp the target object according to the target grasping posture.
2. The interactive robot crawling method based on a large language model according to claim 1, characterized in that: The step of obtaining a target mask area based on the two-dimensional image and the target capture position information in combination with an image segmentation model comprises: Inputting the two-dimensional image and the target capture position information into an image segmentation model to obtain a target mask area; The target mask region is expanded to obtain an expanded target mask region, and the expanded target mask region is used again as the target mask region.
3. The interactive robot crawling method based on a large language model according to claim 1, characterized in that: The step of obtaining a three-dimensional point cloud image with a target grabbing position according to the target mask area and the depth map of the current scene comprises: Generate a binary image according to the target mask area; The two-dimensional image is converted into a three-dimensional point cloud image with a target grasping position according to the binarized image and the depth map.
4. The interactive robot crawling method based on a large language model according to claim 1, characterized in that: The step of obtaining a task operation sequence according to the two-dimensional image of the current scene, the dialogue interaction information and the large language model comprises: Acquire a two-dimensional image and a depth image of the current scene through a depth camera; The two-dimensional image and the dialogue interaction information are input into the large language model to obtain the task operation sequence, and the task operation sequence is stored in a JSON format.
5. The interactive robot crawling method based on a large language model according to claim 1, characterized in that: The step of obtaining the target grasping posture according to the action information, the three-dimensional point cloud image with the target grasping position and the grasping prediction model comprises: Inputting the action information and the three-dimensional point cloud image with the target grasping position into the grasping prediction model to obtain a plurality of candidate grasping postures; The target grasping posture is selected from a plurality of candidate grasping postures.
6. The interactive robot grasping method based on a large language model according to any one of claims 1 to 5, characterized in that: The component information of the target object includes the components of the target object and the colors of the components, and the action information is one of four actions: detection, grasping, delivery and placement.
7. An interactive robot grasping system based on a large language model, characterized in that: It includes a depth camera, a processor and a robot, wherein the depth camera is used to capture a two-dimensional image and a depth map of the current scene; the processor is used to interact with a user, obtain dialogue interaction information, and execute an interactive robot grasping method based on a large language model as described in any one of claims 1 to 6 according to the dialogue interaction information, the two-dimensional image and the depth map.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the interactive robot grasping method based on a large language model as described in any one of claims 1 to 6 are implemented.
9. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of an interactive robot grasping method based on a large language model as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Intelligent control method and device for mechanical arm and electronic equipment
CN118990525A
Mechanical arm control method, device and equipment and storage medium
CN119319568A
Method and apparatus for acquiring grabbing pose of target object, and electronic device
WO2025000838A1
Object grabbing method and apparatus, robot, readable storage medium, and chip
WO2025015867A1
Cited By
Dynamic fusion method, system and equipment for multi-mode perception data of body-equipped intelligent agent
CN120495828A
Grabbing attitude generation method and system based on multi-modal large model
CN120588235A
Control instruction generation method, interaction device and industrial robot
CN120680515A
A control instruction generation method, an interaction device and an industrial robot
CN120680515B
Robot, operation method thereof, operation device, storage medium, and program product
CN120773057A