Intelligent grabbing method of mechanical arm

By combining multimodal large models and natural language parsing, an end-to-end intelligent grasping method for robotic arms was realized, which solved the problem of insufficient flexibility of traditional robotic arms, simplified the control process, and supported the automated execution of complex tasks.

CN119526384BActive Publication Date: 2025-11-18AEROSPACE SCI & IND GRP INTELLIGENT TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411508762.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-11-18
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

Traditional robotic arms lack flexibility and adaptability when faced with complex and ever-changing production demands, requiring manual intervention to adjust the program, increasing operating costs and limiting the improvement of production efficiency.

Method used

It employs a multimodal large model for image recognition and natural language parsing, combined with robotic arm motion control, to achieve an end-to-end intelligent grasping method. It can plan grasping and placement actions by inputting control commands via voice or text, and supports dynamic management of objects within the scene.

Benefits of technology

It achieves end-to-end control of the robotic arm, simplifies the usage process, expands application scenarios, lowers the control threshold, and supports the automated execution of complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119526384B_ABST
    Figure CN119526384B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of mechanical arm control, and discloses an intelligent grabbing method of a mechanical arm. The method comprises the following steps: power-on self-checking of the mechanical arm; collecting images of a task scene, and performing object recognition on the collected images by using a multimodal large model; constructing the task scene according to the recognition result; inputting a natural language control instruction for controlling the mechanical arm, analyzing the natural language control instruction by using the multimodal large model, and planning a mechanical arm motion control sequence according to the analysis result, so as to generate a structured data instruction for controlling the mechanical arm; converting the structured data instruction, generating a mechanical arm control service request, and determining a target to be grabbed and a placement position or plane according to the mechanical arm control service request; judging whether the target to be grabbed exists in the constructed task scene; and in the case that the target to be grabbed exists, outputting a grabbing task instruction and a placement task instruction to the mechanical arm, so as to control the mechanical arm to sequentially perform a grabbing action and a placement action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotic arm control technology, and in particular to an intelligent grasping method for robotic arms. Background Technology

[0002] Robotic arms play an important role in industrial automation and intelligent manufacturing, and are deeply involved in many industrial production processes. With the rapid development of artificial intelligence technology in recent years, people's demand for the automation and intelligence of robotic arms is constantly growing, in order to reduce production costs, improve production efficiency and product quality.

[0003] Traditional industrial robotic arms typically rely on pre-programmed instructions and fixed operating patterns to perform tasks, excelling in repetitive and high-precision work environments. However, these robotic arms have limited flexibility and are relatively poorly adaptable to changing production demands and complex tasks. They often require human intervention to adjust their programs to adapt to new production tasks or environmental changes, which not only increases operating costs but also limits further improvements in production efficiency. Summary of the Invention

[0004] This invention provides an intelligent grasping method for robotic arms, which can solve the problems in the prior art.

[0005] This invention provides a robotic arm intelligent grasping method, wherein the method includes:

[0006] Robotic arm performs a power-on self-test;

[0007] Collect images of the task scene and use a multimodal large model to perform object recognition on the collected images;

[0008] Construct a task scenario based on the recognition results;

[0009] Input natural language control commands for controlling the robotic arm, parse the natural language control commands using a multimodal large model, and plan the robotic arm motion control sequence based on the parsing results to generate structured data commands for controlling the robotic arm;

[0010] The structured data instructions are converted to generate a robotic arm control service request, and the target to be grasped and the placement position or plane are determined based on the robotic arm control service request.

[0011] Determine whether the target to be captured exists within the constructed task scenario;

[0012] When a target to be grasped exists, grasping task instructions and placement task instructions are output to the robotic arm to control the robotic arm to execute grasping and placement actions in sequence.

[0013] Preferably, before outputting grasping and placement instructions to the robotic arm, the method further includes:

[0014] Determine the gripping force of the robotic arm.

[0015] Preferably, control commands for controlling the robotic arm are input in a natural language manner via voice or text.

[0016] Preferably, the method further includes:

[0017] The objects in the constructed task scene are labeled and dynamically added or removed.

[0018] Preferably, the method further includes:

[0019] Manage the poses of all objects in the constructed task scene.

[0020] Preferably, managing the poses of all objects in the constructed task scene includes:

[0021] Save, display, and delete the poses of all objects in the constructed task scene, and determine the object grabbing pose and placement pose.

[0022] The above technical solution enables the grasping and placement of specific targets in a task scenario through natural language commands, thereby achieving end-to-end control of the robotic arm, greatly simplifying the usage process of the robotic arm and expanding its application scenarios. Attached Figure Description

[0023] The accompanying drawings, which form part of this specification, are provided to further illustrate embodiments of the invention and, together with the textual description, explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0024] Figure 1 A flowchart of a robotic arm intelligent grasping method according to an embodiment of the present invention is shown;

[0025] Figure 2 A schematic diagram of an example of intelligent grasping by a robotic arm according to an embodiment of the present invention is shown;

[0026] Figure 3 A schematic diagram of data flow according to an embodiment of the present invention is shown. Detailed Implementation

[0027] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0029] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following figures denote similar items; therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0030] Figure 1 A flowchart of a robotic arm intelligent grasping method according to an embodiment of the present invention is shown.

[0031] like Figure 1 As shown, this embodiment of the invention provides a robotic arm intelligent grasping method, wherein the method includes:

[0032] The robotic arm performs a power-on self-test, which ensures that the robotic arm itself is in normal working order.

[0033] Collect images of the task scene and use a multimodal large model to perform object (target) recognition on the collected images;

[0034] Construct a task scenario based on the recognition results;

[0035] The input is a natural language control command for controlling the robotic arm. This command is parsed using a multimodal large model, and the robotic arm motion control sequence is planned based on the parsing results to generate structured data commands (serialized commands) for controlling the robotic arm. The data flow process is as follows: Figure 3 As shown;

[0036] The structured data instructions are converted to generate a robotic arm control service request, and the target to be grasped and the placement position or plane are determined based on the robotic arm control service request.

[0037] Determine whether the target to be captured exists within the constructed task scenario;

[0038] When a target to be grasped exists, grasping task instructions and placement task instructions are output to the robotic arm to control the robotic arm to execute grasping and placement actions in sequence.

[0039] After that, you can continue to send the next natural language control command to control the robotic arm or end the robotic arm control task.

[0040] The above technical solution enables the grasping and placement of specific targets in a task scenario through natural language commands, thereby achieving end-to-end control of the robotic arm, greatly simplifying the usage process of the robotic arm and expanding its application scenarios.

[0041] If the target to be grabbed does not exist, the target to be grabbed and the placement position or plane can be redefined.

[0042] According to one embodiment of the present invention, before outputting grasping task instructions and placement task instructions to the robotic arm, the method further includes:

[0043] Determine the gripping force of the robotic arm.

[0044] Therefore, the target object can be grasped according to the determined gripper force, thereby achieving control of the gripper force.

[0045] According to one embodiment of the present invention, control commands for controlling a robotic arm are input in a natural language manner via voice or text (natural language control commands).

[0046] Specifically, the grasping method described in this invention can be adapted to various open-source or closed-source large models, can be flexibly deployed online and offline, and a robotic arm control system is built based on the Robot Operating System (ROS), which is widely used in the robotics field, to realize the grasping and placement of specific targets through voice and text commands.

[0047] Furthermore, in terms of functionality, this invention also designs a large-scale model prompt word template for robotic arm grasping, realizing the understanding of voice commands and text commands, task sequence planning, and standardized data output. It also designs a communication interface connecting robotic arm motion control and multimodal large-scale model, realizing the data conversion from control commands to ROS Service. For robotic arm planning and control, it designs interfaces with functions such as autonomous motion planning, target grasping with specified names, and placement at specified locations. Through the combination of these functions, it realizes end-to-end target grasping and placement functions through natural language or text commands, simplifying the traditional robotic arm control process.

[0048] According to one embodiment of the present invention, the method further includes:

[0049] The objects in the constructed task scene are labeled and dynamically added or removed.

[0050] In other words, it can establish task scenarios based on multimodal large model scene recognition results, acquire information on all targets within the task scenario, and support the dynamic addition and deletion of objects within the scenario.

[0051] According to one embodiment of the present invention, the method further includes:

[0052] Manage the poses of all objects in the constructed task scene.

[0053] Specifically, it can dynamically construct task scenarios based on perception results and manage the poses of all objects in the scenario.

[0054] According to one embodiment of the present invention, managing the poses of all objects in the constructed task scene includes:

[0055] Save, display, and delete the poses of all objects in the constructed task scene, and determine the object grabbing pose and placement pose.

[0056] For example, this invention can automatically plan the path to the target and the target grasping scheme based on a specified target name; it supports functions such as grasping with a specific posture, grasping with a specified force, placing at a specified position, and placing on a specified plane, etc. Figure 2 As shown.

[0057] As can be seen from the above embodiments, the intelligent grasping method of the robotic arm described in the above embodiments of the present invention has at least the following advantages:

[0058] (1) End-to-end robotic arm motion control combining multimodal large model

[0059] This invention provides an end-to-end robotic arm control method that combines a multimodal large model. It can be widely adapted to various open-source / closed-source large models, has online and offline deployment capabilities, and expands the control methods of the robotic arm by utilizing the large model's ability to understand and parse natural language. It realizes end-to-end control of the robotic arm through human-like natural interaction methods such as voice and text, and has broad application prospects.

[0060] (2) Automatic grasping and placement of robotic arms oriented towards natural language control

[0061] This invention provides a robotic arm control method oriented towards natural interaction. It designs control interfaces such as scene target labeling, dynamic addition and removal of targets, automatic planning and grasping of the robotic arm, and adjustment of gripper force. It realizes control modes such as grasping a specified target, grasping at a specified position, placing at a specified position, and placing on a specified plane, which greatly simplifies the robotic arm control process, lowers the threshold for robotic arm control, and provides strong support for robotic arms to perform complex tasks.

[0062] In the description of this invention, it should be understood that the orientation or positional relationship indicated by directional terms such as "front, back, up, down, left, right", "horizontal, vertical, horizontal" and "top, bottom" is generally based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing this invention and simplifying the description. Unless otherwise stated, these directional terms do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the scope of protection of this invention; the directional terms "inner" and "outer" refer to the inner and outer contours relative to the outline of each component itself.

[0063] For ease of description, spatial relative terms such as "above," "on top of," "on the upper surface of," "above," etc., are used herein to describe the spatial positional relationship of a device or feature as shown in the figures to other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation beyond the orientation of the device as described in the figures. For example, if the device in the figures were inverted, a device described as "above" or "on top of" other devices or structures would subsequently be positioned as "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both "above" and "below." The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatial relative descriptions used herein will be interpreted accordingly.

[0064] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, the above terms have no special meaning and therefore should not be construed as limiting the scope of protection of this invention.

[0065] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A robotic arm intelligent grasping method, characterized in that, The method includes: Robotic arm performs a power-on self-test; Collect images of the task scene and use a multimodal large model to perform object recognition on the collected images; Construct a task scenario based on the recognition results; Input natural language control commands for controlling the robotic arm, parse the natural language control commands using a multimodal large model, and plan the robotic arm motion control sequence based on the parsing results to generate structured data commands for controlling the robotic arm; The structured data instructions are converted to generate a robotic arm control service request, and the target to be grasped and the placement position or plane are determined based on the robotic arm control service request. Determine whether the target to be captured exists within the constructed task scenario; When a target to be grasped exists, grasping task instructions and placement task instructions are output to the robotic arm to control the robotic arm to execute grasping and placement actions in sequence.

2. The method according to claim 1, characterized in that, Before outputting grasping and placement instructions to the robotic arm, the method also includes: Determine the gripping force of the robotic arm.

3. The method according to claim 2, characterized in that, Control commands for the robotic arm can be input via voice or text in a natural language manner.

4. The method according to claim 3, characterized in that, The method also includes: The objects in the constructed task scene are labeled and dynamically added or removed.

5. The method according to claim 4, characterized in that, The method also includes: Manage the poses of all objects in the constructed task scene.

6. The method according to claim 5, characterized in that, Managing the poses of all objects in the constructed task scene includes: Save, display, and delete the poses of all objects in the constructed task scene, and determine the object grabbing pose and placement pose.

Citation Information

Patent Citations

  • Robot control method based on multi-modal large model

    CN117944052A

  • Robot planning method and system based on multi-mode large model predictive control

    CN118061186A