Mechanical arm control method based on visual language model and human feedback
Through the combination of visual language model and human feedback, the problems of semantic deviation and fault repair in robotic arm control are solved, efficient and accurate robotic arm control are achieved, and user interaction experience and system adaptability are improved.
Patent Information
- Application Number
- CN202510817411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-22
AI Technical Summary
Existing robotic arm control methods based on large language model (LLM) are prone to semantic deviations when generating fine-grained manipulation commands, and lack effective human-computer interaction mechanisms to deal with fault repair in complex environments and diverse tasks.
The visual language model is used to combine human feedback method, and scene information is obtained through the camera device, the user's natural language instructions are decomposed into multiple subtasks, and control code is generated, and a predefined API centered on the control target is introduced to achieve precise control of the robotic arm; actively communicate with the user in the event of a failure and correct the code based on the feedback.
It improves the efficiency and accuracy of robotic arm task execution, enhances the system's adaptability and user satisfaction, and can flexibly respond to diverse task scenarios in open objects and meet user expectations.
Smart Images

Figure CN120347772A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and robotics technology, and in particular to a robotic arm control method based on a visual language model and human feedback. Background Art
[0002] With the rapid development of artificial intelligence and automation technology, robotic arms are increasingly used in industrial manufacturing, logistics sorting, medical assistance and other fields. Modern robotic arms can not only perform simple repetitive actions, but also complete more diverse operations with the help of advanced perception and control algorithms. In this process, how to convert human natural language instructions into specific action sequences that the robotic arm can understand and accurately execute has become a key issue in realizing intelligent control.
[0003] In recent years, large language models (LLMs) have made significant progress in the field of natural language processing, demonstrating powerful semantic understanding and code generation capabilities. Some studies have attempted to apply LLMs to robotic arm control, directly generating robotic arm manipulation instruction codes through natural language instructions. However, existing LLM-based robotic arm control methods still face many challenges. For example, LLMs are prone to semantic deviations when generating fine-grained robotic arm manipulation commands, especially when dealing with complex environments and diverse tasks. The possible ambiguity in natural language instructions and the limitations of the robotic arm's operating range further increase the complexity of task execution. In addition, when the code generated by LLM fails to complete the task or an error occurs, there is a lack of effective mechanisms for human-computer interaction and fault repair. Therefore, how to combine the semantic understanding and code generation capabilities of LLMs with human guidance and feedback to develop a reliable, efficient, and interactive robotic arm control code generation method remains a key issue that needs to be addressed. Summary of the invention
[0004] The purpose of the present invention is to provide a robot arm control method based on visual language model and human feedback, so as to improve the intelligence level of the robot arm control system and the user interaction experience. The core is to use a large language model to understand the task semantics and decompose the user's natural language instructions into subtasks, and automatically combine the predefined application programming interface (API) in combination with scene semantics and object information to generate executable robot arm control code with stronger open task processing capabilities. In the implementation process, a human feedback mechanism is introduced. When a control failure occurs or the expected goal is not achieved, the system can actively communicate with the user, correct the code according to human feedback, and improve the robustness and adaptability of code generation.
[0005] The specific technical scheme for achieving the purpose of the present invention is as follows:
[0006] A manipulator control method based on a vision - language model and human feedback, the method comprising the following steps:
[0007] a) Obtain scene information: Use camera devices installed in multiple directions to capture objects in the task environment, and utilize the vision - language model to complete the 3D reconstruction of the task scene, obtaining information about the manipulator and the objects in the scene;
[0008] b) Generate control codes according to user instructions: Guide the vision - language model to understand the natural - language instructions expressed by the user through prompts, decompose the user instructions into multiple subtasks in the way of a chain of thought, and further generate control codes corresponding to each subtask;
[0009] c) Verify the generated control codes: Logically verify the generated control codes through the vision - language model, and based on the scene information obtained in step a), preliminarily judge whether the generated control codes are feasible for controlling the manipulator to achieve each subtask, and modify them when there are syntax errors or logical loopholes;
[0010] d) Execute the control codes: Sequentially execute the generated control codes in the order of subtasks, and judge whether the current subtask is correctly executed according to the feedback of the manipulator; If not, enter step e); If executed successfully, continue to execute the next subtask; When all subtasks are executed, call the vision - language model to analyze the position and attitude information of the objects in the scene, and judge whether the task is completed as expected; If not, enter step e); If completed smoothly, feedback the completion of the task to the user;
[0011] e) Fault repair and human - machine interaction: When the task is not completed as expected, actively feedback the reasons for not being completed as expected in the form of voice, receive and parse the user's voice feedback information, regenerate the control codes according to the user's feedback and enter step b).
[0012] Further, step a) specifically includes:
[0013] a1. Set cameras in multiple directions and collect image data of the objects in the scene from different angles;
[0014] a2. Based on the collected image data, construct the 3D point cloud of the objects in the scene to complete the 3D reconstruction of the scene;
[0015] a3. Calibrate the relative positions of the cameras and the manipulator, determine the precise position of the manipulator in the scene, and calculate the coordinates and attitudes of the objects in the scene relative to the manipulator.
[0016] Further, step b) specifically includes:
[0017] b1. Extract the natural language instructions of the user through the speech recognition system, and convert the user's speech instructions into text instructions;
[0018] b2. The text instructions will be integrated into the generated prompt, and input into the vision - language model for task decomposition and generation of control codes; the generated prompt also includes scene object information, robotic arm status information, the definition and usage method of the API, and 3 - 5 learning samples;
[0019] b3. The vision - language model, according to the generated prompt, first decomposes the user's text instructions into multiple subtasks in a chain - of - thought manner, and generates the control code for each subtask; the control code is a set of predefined API calls centered on the manipulation target generated by the vision - language model by learning the definition and usage method of the API centered on the manipulation target and the learning samples predefined in the generated prompt; the predefined API call centered on the manipulation target refers to a statement that uses the parameters inferred by the vision - language model to call the predefined API centered on the manipulation target;
[0020] b4. Combine the generated control code for each subtask with the corresponding subtask content to organize it into a complete control code.
[0021] Further, the step c) specifically includes:
[0022] c1. Integrate the complete control code in step b) into the verification prompt, where the verification prompt includes the user's instructions and the object and robotic arm information in the scene. After integration into the verification prompt, input the verification prompt into the vision - language model;
[0023] c2. The vision - language model determines whether the decomposed subtasks can achieve the user's intention according to the prompt, and estimates whether the control code for each subtask can complete the corresponding subtask; if there are problems, organize the problems into the generated prompt and return to step b) to regenerate the code; if the complete control code passes the verification, proceed to the next step.
[0024] Further, the step d) specifically includes:
[0025] d1. Restore the complete control code verified in step c) to the control code for each subtask according to the subtasks;
[0026] d2. Execute the control code of each subtask in sequence according to the subtask order. During the execution process, call the predefined API centered on the manipulation target to control the robotic arm to complete the subtask; if there is an abnormal situation where the robotic arm cannot rotate to the specified angle or move to the specified position, the predefined API centered on the manipulation target will return an exception message and stop executing the control code of each subtask; collect the exception information and the executed control code and proceed to step e).
[0027] d3. If all subtasks have been executed, use the vision-language model to analyze the position and pose information of the objects in the current scene to determine whether the user's instruction has been successfully completed; if not, determine the reason for failure and proceed to step e); otherwise, it means the task is completed.
[0028] The step of executing the control code can accurately locate the abnormal subtask and provide sufficient information for fault repair.
[0029] Furthermore, step e) specifically includes:
[0030] e1. If entering this step from step d2, organize the exception information and the executed control code into a generated prompt; if entering this step from step d3, feedback the exception to the user through voice, receive and parse the user's opinion, and then organize the user's opinion and the reason for failure into a generated prompt.
[0031] e2. Update the generated prompt and enter step b) to regenerate the control code.
[0032] Fault repair and human-computer interaction provide a high-quality fault repair strategy for robotic arm manipulation, enabling the task to be completed more in line with user expectations.
[0033] Furthermore, the predefined API centered on the manipulation target specifically includes:
[0034] Obtain the position and pose of the target to be manipulated; starting from the manipulation target, obtain the position relative to the manipulation target, which is the position where the robotic arm needs to move to; based on the pose of the manipulation target, obtain the rotation angle relative to the pose of the manipulation target, which determines the final orientation of the robotic arm; according to the obtained target position and rotation angle, control the robotic arm to complete the manipulation of the target. If it is found that the set target position and rotation angle exceed the control range of the robotic arm, or if it is found that the robotic arm does not move as expected when controlling the robotic arm, return the corresponding exception.
[0035] The design concept of the API centered on the manipulation target enables the large model to easily understand the usage of the API, and only needs to call a small number of APIs centered on the manipulation target to accurately complete the manipulation task of the target object by the robotic arm.
[0036] Furthermore, the learning sample specifically includes:
[0037] The learning sample is a piece of robotic arm control code that contains detailed annotations, decomposes the complete task into multiple subtasks, and clearly labels the subtask objectives corresponding to each code block. Such a sample guides the vision-language model to learn how to break down complex instructions into executable subtasks and generate a structured control program through the correspondence between the code and the annotations.
[0038] The beneficial effects of the present invention are as follows:
[0039] A novel robotic arm control method is proposed, which makes full use of the capabilities of the large language model in natural language task decomposition and API mapping, efficiently decomposes complex natural language instructions into multiple subtasks with clear steps, and generates control codes for each subtask by combining APIs, thereby improving the efficiency and accuracy of robotic arm task execution. An API centered on the manipulation target is designed, starting from the manipulation target, improving the fine-grainedness and accuracy of manipulation. Combining with the reasoning ability of the large language model, it can flexibly handle diverse task scenarios in an open object set, enhancing the adaptability and generalization ability of the system. A human-computer interaction framework is introduced, and the system can actively seek user feedback when encountering ambiguity or operation limitations, thereby better meeting user expectations and improving the accuracy of task execution and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the method steps of the present invention;
[0041] Figure 2 is a schematic flowchart of the three-dimensional reconstruction of the scene of the present invention;
[0042] Figure 3 is a schematic diagram of the task scenario in the embodiment;
[0043] Figure 4 is a schematic diagram of completing sub-process 2 in the embodiment;
[0044] Figure 5 is a schematic diagram of completing sub-process 3 in the embodiment;
[0045] Figure 6 is a schematic diagram of completing sub-process 4 in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings. The specific implementation examples of the present invention introduced below will help those skilled in the relevant art to further understand the present invention. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and general knowledge in the art, and the present invention has no particularly restricted content.
[0047] An example of the API centered on the manipulation target defined in advance according to the present invention is shown in Table 1.
[0048] Table 1
[0049]
[0050] Refer to Figure 1 , a manipulator control method based on a vision-language model and human feedback, specifically including:
[0051] (a) Obtain scene information: Refer to Figure 2 , collect scene images through cameras installed in two orientations, complete the three-dimensional reconstruction of the task scene through the recognition of the images by the vision model and the construction of point clouds, and then obtain the position and attitude information of the manipulator, as well as the names, positions, and states of the objects in the scene.
[0052] Step 1: Use two gemini2 cameras to capture image data of the scene and objects.
[0053] Step 2: Use GPT-4o combined with the open-vocabulary object detector OWL-ViT to identify the objects on the table and obtain the bounding boxes of the targets.
[0054] Step 3: Input the image containing the object name and bounding box into the Segment Anything model to obtain a mask, and complete the three-dimensional reconstruction of the task scene in combination with the RBG-D camera image.
[0055] Step 4: Determine the position and attitude of the manipulator according to the scene reconstruction model, and calculate the coordinates and attitude of the objects in the scene relative to the manipulator.
[0056] (b) Generate control codes according to user instructions: Recognize user instructions through speech, and in a way of 3-5 learning, guide the vision-language model to understand user instructions in a chain-of-thought manner, decompose them into multiple subtasks, and generate corresponding control codes.
[0057] Step 1: Use the Baidu speech recognition system to convert the user's speech instructions into text instructions.
[0058] Step 2: Organize the text instructions, scene object information, robotic arm status information, definitions and usage methods of predefined manipulation target-centered APIs, and 3-5 samples into the generated prompt, and input the generated prompt into the vision-language model.
[0059] Step 3: Use GPT-4o to decompose the instructions into multiple subtasks in a chain-of-thought manner according to the generated prompt, and generate a set of control codes for each subtask to call the predefined manipulation target-centered APIs shown in Table 1. Among them, the call parameter values are inferred by GPT-4o from the learning environment learning samples and the prompt according to the environmental information.
[0060] Step 4: Combine the control code of each subtask with the subtask and organize it into a complete control code.
[0061] (c) Verify the generated control code: Logically verify the generated control code through GPT-4o. According to the scene information obtained in step a), preliminarily judge whether the generated control code has the feasibility to control the robotic arm to achieve each subtask, and modify it when there are syntax errors or logical loopholes.
[0062] Step 1: Integrate the complete control code generated in step b), the user instructions, and the object and robotic arm information in the scene into the verification prompt, and input it into GPT-4o.
[0063] Step 2: Use GPT-4o to verify whether there are problems with the complete control code according to the scene information and user instructions in the verification prompt, including whether there are syntax problems, whether the decomposed subtasks can control the robotic arm to complete the user instructions, and evaluate whether the generated code can complete the corresponding subtasks. For the code with problems, guide GPT-4o to put forward modification opinions, and organize the opinions into step (b) to regenerate the code.
[0064] (d) Execute the control code: Sequentially execute the control codes of the subtasks in order. After each subtask is executed, judge whether it is successfully executed according to the feedback of the robotic arm. After all subtasks are executed, use the vision-language model to judge whether the task is successfully completed. If there are problems, feedback the problems to the user through voice. After organizing the user's opinions, enter step (b) to regenerate the control code according to the current situation.
[0065] Step 1: Sequentially execute the control codes of each subtask in the order of the subprocesses. By calling the combined APIs, use Ros to communicate with the robotic arm to complete the tasks of obtaining specific scene information or controlling the robotic arm operation through the APIs, including obtaining environmental information, obtaining the position and orientation of the target object, setting the target position or rotation angle of the robotic arm, and controlling the robotic arm to complete movement and rotation.
[0066] Step 2: During the execution of the robotic arm's actions, due to limitations in the movement range and rotation angle of the robotic arm, when the set path or rotation angle cannot be successfully completed, the corresponding abnormal situation will be returned through the API, and it will enter step (e) and stop the execution of the subsequent code.
[0067] Step 3: After all subtasks are executed, obtain the scene information through the camera, and use the vision large model GPT-4o to determine whether the current situation in the scene matches the expected result of the user's instruction. For uncompleted cases, the reasons for non-completion will be sorted out and it will enter step (e). Otherwise, it means that the user's instruction has been completed.
[0068] (e) Fault repair and human-computer interaction: Organize the abnormal situation or the reasons for uncompleted tasks into text, convert it into voice and feedback it to the user, receive the user's voice opinions, use the speech-to-text tool to organize the user's opinions, and enter step (b).
[0069] Step 1: If an exception occurs when executing the control code of the subtask, organize the exception information and the previously executed control code into the generated prompt; if an exception occurs after all subtasks are executed, feedback the exception to the user through voice, receive and parse the user's opinions, and organize the user's opinions and the reasons for failure into the generated prompt.
[0070] Step 2: Update the generated prompt and enter step b) to regenerate the control code.
[0071] Embodiment
[0072] This embodiment takes a task as an example to explain the process from natural language instructions to controlling a robotic arm.
[0073] (1) The initial state of the scene is as Figure 3 shown. Through the captured task scene image, it can be seen that there are two plates, pink and purple, in the scene, and a yellow square is placed on the pink plate.
[0074] User voice instruction: "Put the yellow square on the purple plate."
[0075] (2) Control code generation:
[0076] Recognize the user's voice instruction:
[0077]
[0078] Plan high-level execution steps:
[0079]
[0080] Call the API to implement each specific step when generating control code:
[0081]
[0082] (3)Verify the code:
[0083] Upload the generated code and the user instructions to GPT-4o-mini to verify if there are any problems with the generated code.
[0084]
[0085] (4)Execute the code step by step:
[0086] Refer to Figures 3 - 6 for the execution process.
[0087] (5)Verify if the task is completed:
[0088] Upload the scenario at the end of the execution and the user instructions to GPT-4o. Judge if the execution is completed.
Claims
1. A manipulator control method based on a vision-language model and human feedback, characterized in that, The method includes the following steps: a) Obtain scene information: Use camera devices installed in multiple directions to capture objects in the task environment, and utilize a vision-language model to complete the 3D reconstruction of the task scene, obtaining information about the robotic arm and the objects in the scene; b) Generate control code according to user instructions: Guide the vision-language model to understand the natural language instructions expressed by the user through prompts, decompose the user instructions into multiple subtasks in a chain-of-thought manner, and further generate the control code corresponding to each subtask; c) Verify the generated control code: Logically verify the generated control code through the vision-language model. Based on the scene information obtained in step a), preliminarily determine whether the generated control code has the feasibility to control the robotic arm to achieve each subtask, and modify it when there are syntax errors or logical loopholes; d) Execute the control code: Sequentially execute the generated control code in the order of subtasks, and judge whether the current subtask is correctly executed based on the feedback of the robotic arm; if not, enter step e); if executed successfully, continue to execute the next subtask; when all subtasks are executed, call the vision-language model to analyze the position and pose information of the objects in the scene, and judge whether the task is completed as expected; if not, enter step e); if successfully completed, feedback the completion of the task to the user; e) Fault repair and human-machine interaction: When the task is not completed as expected, actively feedback the reason for not being completed as expected in the form of voice, receive and parse the user voice feedback information, regenerate the control code according to the user feedback, and enter step b).
2. The manipulator control method according to claim 1, wherein, The specific steps of step a) include: a1. Set cameras in multiple directions and collect image data of the objects in the scene from different angles; a2. Based on the collected image data, construct a 3D point cloud of the objects in the scene to complete the 3D reconstruction of the scene; a3. Calibrate the relative positions of the cameras and the robotic arm, determine the exact position of the robotic arm in the scene, and calculate the coordinates and poses of the objects in the scene relative to the robotic arm.
3. The manipulator control method according to claim 1, wherein The specific steps of step b) include: b1. Extract the natural language instructions of the user through a speech recognition system and convert the user's voice instructions into text instructions; b2. The text instructions will be integrated into the generated prompts and input into the vision-language model for task decomposition and generation of control code; the generated prompts also include scene object information, robotic arm status information, the definition and usage method of APIs, and 3-5 learning samples; b3. The vision-language model, according to the generated prompts, first decomposes the user text instructions into multiple subtasks in a chain-of-thought manner and generates the control code for each subtask; the control code is a set of pre-defined API calls centered on the manipulation target generated by the vision-language model by learning the definition and usage method of the pre-defined APIs centered on the manipulation target and the learning samples in the generated prompts; the pre-defined API calls centered on the manipulation target refer to a statement that uses the parameters inferred by the vision-language model to call the pre-defined APIs centered on the manipulation target; b4. Combine the control codes of each generated subtask with the corresponding subtask content to form a complete control code.
4. The manipulator control method according to claim 1, wherein The specific steps of step c) include: c1. Integrate the complete control code of step b) into the verification prompt, which includes the user instruction and the information of the objects and the robotic arm in the scenario. After integration, input the verification prompt into the vision-language model. c2. The vision-language model determines whether the decomposed subtasks can achieve the user's intention according to the prompt, and estimates whether the control code of each subtask can complete the corresponding subtask. If there are problems, organize the problems into the generated prompt and return to step b) to regenerate the code. If the complete control code passes the verification, proceed to the next step.
5. The manipulator control method according to claim 1, wherein The specific steps of step d) include: d1. Restore the complete control code verified in step c) to the control code of each subtask according to the subtasks. d2. Execute the control codes of each subtask in sequence according to the subtask order. During the execution process, call the predefined API centered on the manipulation target to control the robotic arm to complete the subtask. If there is an abnormal situation where the robotic arm cannot rotate to the specified angle or move to the specified position, the predefined API centered on the manipulation target will return an abnormal message and stop executing the control codes of each subtask. Collect the abnormal information and the executed control codes and enter step e). d3. If all subtasks have been executed, use the vision-language model to analyze the position and pose information of the objects in the current scenario to determine whether the user instruction has been successfully completed. If not, determine the reason for failure and enter step e). Otherwise, it means the task is completed.
6. The manipulator control method according to claim 1, characterized in that The specific steps of step e) include: e1. If entering this step from step d2, organize the abnormal information and the executed control codes into the generated prompt. If entering this step from step d3, feedback the abnormality to the user through voice. After receiving and parsing the user's opinion, organize the user's opinion and the reason for failure into the generated prompt. e2. Update the generated prompt and enter step b) to regenerate the control code.
7. The manipulator control method according to claim 3, wherein The predefined API centered on the manipulation target specifically includes: Obtain the position and pose of the target to be manipulated; starting from the manipulation target, obtain the position relative to the manipulation target, which is the position where the robotic arm needs to move; based on the pose of the manipulation target, obtain the rotation angle relative to the pose of the manipulation target, which determines the final orientation of the robotic arm; according to the obtained target position and rotation angle, control the robotic arm to complete the manipulation of the target. If it is found that the set target position and rotation angle exceed the control range of the robotic arm, or if it is found that the robotic arm does not move as expected during the control of the robotic arm, return the corresponding abnormality.
8. The manipulator control method according to claim 3, wherein, The learning samples specifically include: The learning sample is a piece of robotic arm control code that contains detailed annotations, breaks down the complete task into multiple subtasks, and clearly labels the subtask objectives corresponding to each code block; such a sample guides the vision-language model to learn how to break down complex instructions into executable subtasks and generate a structured control program through the correspondence between the code and the annotations.
Citation Information
Cited By
Kitchen service robot operation method based on VLN large model
CN120902023A
Kitchen service robot operation method based on VLN large model
CN120902023B
Task planning method and system for robot
CN121061907A
Robot control method, system, device, equipment and medium
CN121290446A