Experiment task execution method and device
By combining the visual language model with the visual language action model in a dual-loop closed control architecture, the execution instability and safety issues of long-term tasks in robotic chemistry experiments are solved, and efficient and safe experimental task execution is achieved.
Patent Information
- Application Number
- CN202510939633.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing robotic chemical experiment systems have difficulty achieving efficient execution of the entire process when handling long-term tasks. They are prone to failure or errors at the connection points of steps, and have problems such as high visual recognition error rate, insufficient security and compliance.
A dual-loop closed control architecture is adopted to combine the visual language model with the visual language action model. The visual language model is used to realize semantic-level planning of experimental tasks, decompose complex tasks into controllable basic operation sequences, and fuse visual cue images, text information and experimental images through the visual language action model to accurately guide robot movements.
It significantly improves the success rate of experimental tasks, operational safety and compliance with regulations, improves the robot's recognition accuracy and operational stability in complex chemical experimental scenarios, and realizes automated execution and safety assurance of the entire process.
Smart Images

Figure CN120735015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for executing an experimental task. Background Art
[0002] While existing robotic chemistry experiment technology has made some progress, in practice, chemical experiments often involve complex, multi-step operations. Existing robotic systems, such as the Advanced Chunk Transformer (ACT), the Robotics Diffusion Transformer (RDT), and π0, struggle to efficiently execute these long-term tasks across the entire process, and are prone to failures or errors at the transition points. Therefore, an effective solution is urgently needed to address this issue. Summary of the Invention
[0003] In order to solve the above problems, the present invention provides a method and device for executing an experimental task.
[0004] The present invention provides a method for executing an experimental task, comprising: Based on the visual language model, the task information of the target experimental task is divided to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks from front to back in the execution order; Starting from the first subtask in the subtask sequence, the text information corresponding to the current subtask and the experimental image before the execution of the current subtask are processed through the visual language model to obtain a visual prompt image, and based on the text information, the experimental image and the visual prompt image processing, the visual language action model is used to guide the robot to execute the current subtask until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
[0005] According to an experimental task execution method provided by the present invention, the text information corresponding to the current subtask and the experimental image before the current subtask is executed are processed by the visual language model to obtain a visual prompt image, including: Extracting text information corresponding to the current subtask; Performing semantic analysis on the text information using the visual language model to obtain a semantic analysis result; The visual language model is used to label the experimental image based on the semantic analysis result to obtain the visual prompt image.
[0006] According to an experimental task execution method provided by the present invention, the experimental image is a multi-view experimental image; The step of labeling the experimental image based on the semantic analysis result using the visual language model to obtain the visual prompt image includes: The visual language model is used to label the front view in the multi-view experimental image based on the semantic analysis result to obtain the visual prompt image.
[0007] According to an experimental task execution method provided by the present invention, after guiding the robot to execute the current subtask based on the processing of the text information, the experimental image, and the visual prompt image, the method further includes: Acquire, through the visual language model, and based on the experimental image after the current subtask is completed, check whether the current subtask meets the requirements; If the criteria are met, the execution steps of the next subtask of the current subtask are executed; If the requirements are not met, the execution steps of the current subtask are re-executed.
[0008] According to an experimental task execution method provided by the present invention, the method guides the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image processing by using a visual language action model, including: Generate operation information based on the text information, the experimental image and the visual prompt image through the visual language action model; Based on the operation information, the robot is controlled to perform the current subtask.
[0009] According to an experimental task execution method provided by the present invention, the operation information includes at least one operation instruction, and each of the operation instructions carries an operation sequence; The controlling the robot to perform the current subtask based on the operation information includes: Starting from the operation instruction with the first operation sequence, the robot is controlled to perform the operation corresponding to the current operation instruction until the operation corresponding to the operation instruction with the last operation sequence is completed, confirming the completion of the current subtask.
[0010] The present invention also provides an experimental task execution device, comprising the following modules: a division module configured to divide the task information of the target experimental task based on the visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging the subtasks from front to back in the execution order; The execution module is configured to start from the first subtask in the subtask sequence, process the text information corresponding to the current subtask and the experimental image before the execution of the current subtask through the visual language model to obtain a visual prompt image, and guide the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image through the visual language action model until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
[0011] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any one of the above-described experimental task execution methods is implemented.
[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described experimental task execution methods.
[0013] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned experimental task execution methods.
[0014] The experimental task execution method and device provided by the present invention divides the task information of the target experimental task based on a visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks in an execution order from front to back; starting from the first subtask in the subtask sequence, the text information corresponding to the current subtask and the experimental image before the current subtask is executed are processed by the visual language model to obtain a visual prompt image, and the visual language action model is used to guide the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image processing until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed. The present invention adopts a double-loop closed control architecture, combines a visual language model with a visual language action model, realizes semantic-level planning of the experimental task through the visual language model, decomposes complex tasks into controllable basic operation sequences, and integrates visual prompt images, text information and experimental images through the visual language action model to accurately guide the robot's actions, effectively improving the success rate of the experimental task, operational safety and compliance with regulations, and has high versatility and safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 This is one of the flow charts of the experimental task execution method provided by the present invention.
[0017] Figure 2 This is the second flow chart of the experimental task execution method provided by the present invention.
[0018] Figure 3 Schematic diagram of the visual prompt image provided by the present invention.
[0019] Figure 4 It is a structural schematic diagram of the experimental task execution device provided by the present invention.
[0020] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0022] First, the relevant contents of the present invention are briefly described.
[0023] In addition to being insufficient for long-term experimental tasks, existing robotic chemical experiment technology also has the following problems: Challenges in transparent vessel recognition: Existing vision-language models (VLMs), such as VoxPoser and ReKep, rely heavily on depth sensors and object segmentation technology. However, these methods perform poorly when dealing with transparent containers commonly used in chemical experiments (such as glass beakers and test tubes). This leads to a high visual recognition error rate, which in turn affects the success rate of subsequent operations.
[0024] Shortcomings of visual prompting methods: Existing visual prompting methods, such as the Marking Open-world Keypoint Affordances (MOKA) method, can provide a certain degree of visual guidance. However, due to the lack of direct consideration of text instructions, non-compliant or unsafe grasping and operation behaviors may occur in chemical experiment scenarios with high safety requirements and strict procedural requirements.
[0025] Insufficient semantic-level feedback and safety compliance: Existing Vision-Language-Action Model (VLA) models, such as RDT and π0, perform well in executing specific actions. However, they lack global semantic understanding and closed-loop feedback mechanisms, making it difficult to proactively conduct safety checks and confirm compliance with regulations. Therefore, operational errors are prone to occur in complex or safety-critical experimental scenarios, and may even lead to safety accidents.
[0026] These issues significantly limit the widespread application of robotics in actual chemical experiment scenarios. Therefore, the present invention provides a method and device for executing experimental tasks, which can efficiently and accurately complete long-term, highly complex chemical experiment tasks while ensuring operational safety and compliance with experimental protocols.
[0027] The following combination Figure 1-Figure 5 Describe the experimental task execution method and device of the present invention.
[0028] Figure 1 This is one of the flow charts of the experimental task execution method provided by the present invention, such as Figure 1 As shown, the method includes steps 101 and 102.
[0029] Step 101: Based on the visual language model, the task information of the target experimental task is divided to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks from front to back in execution order.
[0030] Specifically, the experimental task execution method provided by the present invention is applicable to an experimental robot system, which is composed of a visual language model, a visual language action model, a robot and an imaging device arranged on the robot. The robot refers to a machine device that automatically performs work, which can be a robotic arm (such as a two-arm gripper device) or a humanoid machine that can walk independently and has two arms.
[0031] See also Figure 2 , Figure 2This is the second flow diagram of the experimental task execution method provided by the present invention: The Visual Language Model (VLM) is set in the outer loop of the experimental robot system. The VLM can serve as a triple role as a task planner (Planner), a visual prompt generator (Visual Prompt), and a task monitor (Monitor), throughout the entire experimental task execution process. The Vision-Language-Action Model (VLA) is in the inner loop of the experimental robot system. The VLA model is responsible for fusing three types of information: the language description and prompt image provided by the VLM, and the experimental image captured by the imaging device. It outputs specific action instructions to achieve high-precision control of the robot's behavior.
[0032] Furthermore, experimental robotic systems can incorporate multiple collaborative robotic arms, with multiple VLAs controlled by a single VLM, collaborating to complete multi-station or serial experimental processes. Current experimental robotic systems primarily rely on dual-arm grippers to perform various atomic manipulations. However, for tasks requiring high compliance and precise grasping (such as transferring powders or holding glass rods), these can be replaced with humanoid hands with multiple degrees of freedom and sensor feedback. This can further enhance the performance of experimental robotic systems in micromanipulation and complex manipulation tasks, while also increasing their versatility.
[0033] The target experimental task refers to the experimental task to be performed, and the experimental task can be a chemistry experimental task, a physics experimental task, etc.
[0034] In practical applications, see Figure 2 The visual language model acts as a planner for task planning: It first receives the task information (Task) of the target experimental task, which may include a description of the experimental task and equipment information. Furthermore, based on the task information, the visual language model decomposes the complex experimental process into executable atomic operations (Primitive Tasks), forming a subtask sequence (Subtask List) with a clear temporal relationship, such as "Transferring a solid," "Grasping a glass rod," and "Adding acid." The subtasks in the subtask sequence are arranged from front to back in the order of execution, such as "1.xxx" to "4.xxx."
[0035] Step 102: Starting from the first subtask in the subtask sequence, the text information corresponding to the current subtask and the experimental image before the execution of the current subtask are processed through the visual language model to obtain a visual prompt image, and based on the text information, the experimental image and the visual prompt image processing, the visual language action model is used to guide the robot to execute the current subtask until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
[0036] Specifically, the experimental image is captured by an imaging device; the imaging device can be a red, green, and blue (RGB) camera; the experimental image can be a multi-perspective experimental image, including a front view, a top view, a left view, and a right view of the current experimental scene, wherein the front view is captured by an imaging device installed directly in front of the robot, the top view can be captured by an imaging device installed at a top-down angle of the robot, the left view is captured by an imaging device installed at the left wrist of the robot, and the right view is captured by an imaging device installed at the right wrist of the robot.
[0037] The text information may be a text description of executing the subtask, such as a text instruction.
[0038] In practical applications, for each subtask in the subtask sequence, each subtask is executed in turn. When the currently executed subtask is completed, the next subtask is executed until all subtasks are executed and the target experimental task is determined to be completed.
[0039] For each subtask, the specific execution process is as follows: Using imaging equipment, capture the experimental image after the previous subtask is completed, or the experimental image before the target experimental task is executed, and use it as the experimental image before the current subtask is executed. If the current subtask is the first subtask, the experimental image before the current subtask is the experimental image before the target experimental task is executed. If the current subtask is not the first subtask, the experimental image before the current subtask is the experimental image after the previous subtask is completed.
[0040] See also Figure 2 The visual language model, acting as a visual prompt, combines the Qwen2.5-VL model with the text information (Text) and observed images corresponding to the current subtask to generate semantically understood visual prompt images (Prompted Images). These prompt images are annotated with at least one of the following: the selected area and key points of the grasped / targeted object. Furthermore, the visual language action model fuses the text and prompted images provided by the VLM with the observed images provided by the imaging device to generate action information (Action). This information then controls the robot to execute the corresponding action (Actor), achieving high-precision control of the robot's behavior.
[0041] In addition, the visual language model can generate other visual prompt information based on multimodal information such as depth maps, thermal imaging, and spectroscopic data to adapt to more complex experimental scenarios.
[0042] The experimental task execution method provided by the present invention divides the task information of the target experimental task based on a visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks in an execution order from front to back; starting from the first subtask in the subtask sequence, the text information corresponding to the current subtask and the experimental image before the execution of the current subtask are processed by the visual language model to obtain a visual prompt image, and the visual language action model is used to guide the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image processing until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed. The present invention adopts a double-loop closed control architecture, combines a visual language model with a visual language action model, realizes semantic-level planning of the experimental task through the visual language model, decomposes complex tasks into controllable basic operation sequences, and fuses visual prompt images, text information and experimental images through the visual language action model to accurately guide the robot's actions, effectively improving the success rate of the experimental task, operational safety and compliance with regulations, and has high versatility and safety.
[0043] Optionally, the step of processing the text information corresponding to the current subtask and the experimental image before the current subtask is executed by the visual language model to obtain a visual prompt image includes: Extracting text information corresponding to the current subtask; Performing semantic analysis on the text information using the visual language model to obtain a semantic analysis result; The visual language model is used to label the experimental image based on the semantic analysis result to obtain the visual prompt image.
[0044] In practical applications, the visual language model can extract the textual information corresponding to the current subtask from the subsequence task. The visual language model then acts as a visual cue generator, performing semantic understanding and analysis on the textual information corresponding to the current subtask to determine the corresponding grasped object / target, i.e., the semantic analysis result. The grasped object / target is then annotated in the experimental image, using selected regions and / or key points. This avoids center bias when handling transparent objects, thus ensuring the safety of the experimental task.
[0045] For example, see Figure 3 , Figure 3 Schematic diagram of the visual cue image provided by the present invention: the target object "test tube" is marked with a blue rectangular frame in the visual cue image, and the key point "grasping point" on the target object "test tube" is marked with a yellow dot.
[0046] The present invention generates visual prompts by combining text information with instruction perception of visual context, and uses advanced vision-language models (such as Qwen-2.5-VL) to automatically identify transparent or irregular objects that need to be operated in the experiment, and generates safe and compliant target area or key point prompts in real time in combination with language instructions, thereby improving the accuracy of object recognition and the robustness of operation, ensuring the consistency of the robot's visual perception and action execution in complex scenarios, and solving the problem of insufficient transparent vessel recognition and visual prompt methods.
[0047] Optionally, the experimental image is a multi-view experimental image; accordingly, the labeling process of the experimental image based on the semantic analysis result using the visual language model to obtain the visual prompt image includes: The visual language model is used to label the front view in the multi-view experimental image based on the semantic analysis result to obtain the visual prompt image.
[0048] In practical applications, while ensuring the effectiveness of visual cues, in order to further reduce the amount of data processing and thus improve task execution efficiency, the front view in the multi-view experimental images can be annotated to obtain visual cue images. The top view, left view, and side view in the multi-view experimental images do not need to be annotated.
[0049] Optionally, after guiding the robot to perform the current subtask based on the processing of the text information, the experimental image, and the visual prompt image, the method further includes: Acquire, through the visual language model, and based on the experimental image after the current subtask is completed, check whether the current subtask meets the requirements; If the criteria are met, the execution steps of the next subtask of the current subtask are executed; If the requirements are not met, the execution steps of the current subtask are re-executed.
[0050] In practical applications, the visual language model can also be used as a task monitor to verify the execution results of each subtask.
[0051] Specifically, see Figure 2 After the current subtask is completed (finished), the acquisition device returns the experimental images after the current subtask. The VLM then determines whether the current subtask has achieved its goal. If so, it continues to the next subtask (next task); if not, it automatically triggers a repeat (repeat). This enables closed-loop task management with error correction capabilities, breaking the limitation of only a single inner-loop feedback loop.
[0052] In this way, by establishing a complete closed-loop monitoring mechanism, VLM acts as a real-time monitor to evaluate the execution of each operation step, provide active feedback, and make timely corrections in the event of failure or unsafe conditions, ensuring operational safety and strict implementation of procedures, and solving the problem of insufficient semantic-level feedback and safety compliance.
[0053] Optionally, guiding the robot to perform the current subtask by processing the text information, the experimental image, and the visual prompt image using a visual language action model includes: Generate operation information based on the text information, the experimental image and the visual prompt image through the visual language action model; Based on the operation information, the robot is controlled to perform the current subtask.
[0054] Specifically, the operation information may include operation instructions and operation conditions, wherein the operation instructions represent the operation action, and the operation conditions represent the operation part or component, such as the operation instruction is "grab the test tube", and the operation condition is "the grabbing part is one-third of the distance from the test tube mouth".
[0055] In practical applications, the visual language action model serves as the action execution layer. It receives experimental images from four RGB cameras (front, top, and wrist), and uses them in conjunction with visual cues and text information provided by the VLM to generate action instructions and conditions. The robot then executes the current subtask according to these instructions and conditions. This enables the generation of action information, significantly improving operational stability in complex scenarios.
[0056] Optionally, the operation information includes at least one operation instruction, and each operation instruction carries an operation sequence; accordingly, controlling the robot to perform the current subtask based on the operation information includes: Starting from the operation instruction with the first operation sequence, the robot is controlled to perform the operation corresponding to the current operation instruction until the operation corresponding to the operation instruction with the last operation sequence is completed, confirming the completion of the current subtask.
[0057] In practical applications, for each subtask, the visual language action model generates an operation instruction sequence, where the operation instruction sequence includes at least one operation instruction that carries an operation order. Figure 2 The five yellow squares at the operation information indicate five operation instructions.
[0058] Furthermore, the robot performs the operations corresponding to each operation instruction in sequence from front to back to complete the current subtask.
[0059] After the robot completes each action corresponding to an instruction, the visual language action model checks whether the current subtask is complete, forming a closed loop. If the current subtask is not completed, the visual language action model will continue to control the robot to execute the current subtask. If the current subtask is completed, the visual language model will check whether the current subtask has met the requirements.
[0060] Furthermore, after the robot completes each operation corresponding to an instruction, the visual language action model can also verify whether the operation was successful. If successful, the robot performs the operation corresponding to the next instruction. If unsuccessful, the robot continues to perform the operation corresponding to the instruction. That is, starting with the instruction with the highest order, the robot is controlled to perform the operation corresponding to the current instruction. After the operation corresponding to the current instruction is successfully performed, the robot continues to perform the operation corresponding to the next instruction, until the operation corresponding to the last instruction in the order is completed, confirming the completion of the current subtask. In this way, the success rate of subtask execution can be guaranteed.
[0061] It should be noted that to improve the robustness of the visual language action model, a dynamic retry mechanism can be used to introduce failed samples during the training phase, enhancing the VLA's ability to perceive and repair execution failures. Specifically, the training process for the visual language action model involves obtaining training samples, which include positive and negative samples; training the visual language action model based on the training samples, and executing the training stop condition. Training can be stopped when the accuracy of the visual language action model reaches a set value.
[0062] In addition, based on the VLM, a small sample discrimination mechanism or a small amount of human demonstration data can be introduced to enable the system to quickly establish evaluation logic when facing unknown experimental tasks, thereby improving versatility and adaptability.
[0063] The following further describes the experimental task execution method provided by the present invention.
[0064] Task reception and parsing: The user specifies the experimental task (e.g., "complete the acid-base neutralization experiment") through natural language and inputs the task semantics (task information of the target experimental task) into the VLM.
[0065] To plan a mission: Based on the task description and the initial experimental scene image, VLM automatically decomposes the complete experimental process into a series of subtasks, such as: transferring sodium hydroxide (NaOH) solid → adding water → adding hydrogen chloride (HCl) to neutrality.
[0066] Visual cue generation: For each subtask, VLM combines the bench image and instruction semantics to automatically generate a visual cue map with grasping points, pouring points, target container boundaries, etc. Figure 3 As shown in the figure, a marking frame for grabbing the upper section of the glass rod is automatically generated to guide the VLA to perform the operation safely.
[0067] Atomic operation execution: Receive three types of input information: ① task language instructions (text information); ② observation images from four perspectives (experimental images); ③ VLM-generated prompt images (visual prompt images), and after fusion, generate control signals (operation information) to realize atomic-level action execution (such as grabbing, pouring liquid, stirring, etc.).
[0068] Execution status monitoring: After each subtask is completed, the VLM rereads the multi-view experimental images to determine whether the operation has achieved its objective. For example, in the "acid addition" task, the VLM determines whether the "neutralization" objective has been achieved by identifying whether the solution changes color. If not, the VLM automatically triggers repeated execution until the task is successful.
[0069] Task completion judgment: After all subtasks are completed, it automatically identifies whether the experiment is successful and outputs a task completion signal; there is no manual intervention or review process, and the entire process is completed by the VLM and VLA in an automatic collaborative closed loop.
[0070] Specifically, the experimental task execution method includes the following steps: user inputs the task → VLM performs semantic analysis and task planning → outputs a subtask sequence; for the subtask, VLM generates visual cues → inputs them into VLA together with language descriptions and observed images; VLA performs atomic operations; VLM analyzes the execution status image and determines whether it is successful or not → success: proceed to the next task; failure: repeat the current task; after all tasks are completed, VLM sends a "task completion" signal.
[0071] The experimental task execution method provided by the present invention completes complex experimental tasks without human intervention through an automated "planning-execution-evaluation-correction" closed-loop control system.
[0072] Comparative experiments were conducted to verify the experimental task execution method provided by the present invention. The results showed that the experimental task execution method provided by the present invention performed significantly better than the existing technology in seven types of atomic operations and five complete chemical experimental processes: In the "grasping a glass rod" task, the present invention uses visual cue images to locate the grasping point at 1 / 3 of the glass rod, avoiding the risk of slipping. The compliance rate is as high as 0.875, which is significantly better than π0 (0.200) and MOKA (0.350). The compliance rate refers to the proportion of grasping compliance. Grasping compliance means that the grasping position is correct and the grasping is successful. Grasping non-compliance includes grasping failure and / or grasping position error. In the "acid addition and neutralization" experiment, the outer loop determined that the solution color had not completely changed, automatically triggering the "repeated acid addition" operation to complete the step-by-step neutralization process, demonstrating the task adaptability guided by the VLM; Taking all the tasks into consideration, the experimental task execution method provided by the present invention has a success rate increased by 23.57% and a compliance rate increased by 0.298 compared with the existing technology, which shows that the experimental task execution method provided by the present invention has significant advantages in complex experimental scenarios. The success rate is the ratio of the number of successful tests for each task to the total number of tests.
[0073] The experimental task execution method provided by the present invention has the ability to decompose tasks into a structural form, and uses a large language model to automatically convert natural language experimental goals into structured subtasks, thereby improving the robot's task comprehension ability; it has a multimodal perception prompt mechanism: introducing instruction-related visual prompts to achieve accurate recognition and action planning of complex and transparent experimental equipment; it has semantic-level dynamic feedback control: VLM has the ability to perceive the experimental status and the degree of achievement of task goals in real time, builds a complete outer loop closed loop, and enhances the robustness of the experimental robot system; it has the ability to execute the entire process without human supervision: compared with traditional systems that require human intervention or preset program logic, the present invention uses VLM to independently evaluate and control the task process, thereby achieving automatic error correction and adaptive adjustment.
[0074] Specifically, the experimental task execution method provided by the present invention has the following advantages: Significantly improves the completion rate of long-term tasks: Traditional robotic control methods such as ACT and π0 lack a good understanding of task structure when handling complex experimental workflows, making them prone to failure at interruptions within the operational chain. This invention, by introducing a structured task decomposition and full-process scheduling mechanism based on a vision-language model (VLM), achieves multi-step disassembly and sequential execution of experimental tasks, effectively reducing failures caused by discontinuities in intermediate steps and increasing the success rate by 23.57%.
[0075] Improves the recognition and manipulation success rate of complex objects (such as transparent objects): Traditional visual cueing methods, such as MOKA, rely heavily on depth maps or masks, which often lead to errors when handling transparent objects. This invention utilizes a "command-aware visual cueing mechanism" that combines verbal commands and observed images to automatically generate a precise grasping area, addressing recognition challenges in scenes with blurred or occluded object boundaries and significantly improving grasping success and accuracy.
[0076] Achieve closed-loop assurance of safety regulations and experimental procedures: This invention uses the VLM to perform semantic-level evaluations of each subtask result (e.g., whether it changes color, whether it is full, etc.), and automatically retries if it fails to meet the standards, avoiding experimental failures or safety risks caused by execution deviations. The compliance index improved by 0.298. This mechanism is particularly effective in tasks such as titration and flame testing, ensuring reproducible and reliable experimental results.
[0077] Fully automated execution, without manual labeling, evaluation, or mid-process intervention: Other methods require manual selection of anchor points or manual setting of parameters. This invention is uniformly scheduled by VLM and automatically executed by VLA. All prompts, feedback, and corrections are completed automatically, making it suitable for large-scale deployment and unmanned laboratory environments.
[0078] The system has a clear structure and is easy to expand and migrate: it adopts a modular dual-loop structure (outer loop VLM scheduling, inner loop VLA execution), which can flexibly replace the language model or execution model to adapt to different robot hardware platforms and task requirements, and has good versatility and engineering scalability.
[0079] The experimental task execution device provided by the present invention is described below. The experimental task execution device described below and the experimental task execution method described above can be referenced to each other.
[0080] Figure 4 This is a schematic diagram of the structure of the experimental task execution device provided by the present invention. Figure 4 As shown, the device method includes: A division module 401 is configured to divide the task information of the target experimental task based on the visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging the subtasks from front to back in the execution order; The execution module 402 is configured to start from the first subtask in the subtask sequence, process the text information corresponding to the current subtask and the experimental image before the execution of the current subtask through the visual language model to obtain a visual prompt image, and guide the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image through the visual language action model until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
[0081] The experimental task execution device provided by the present invention divides the task information of the target experimental task based on a visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks in an execution order from front to back; starting from the first subtask in the subtask sequence, the text information corresponding to the current subtask and the experimental image before the execution of the current subtask are processed by the visual language model to obtain a visual prompt image, and the visual language action model is used to guide the robot to perform the current subtask based on the text information, the experimental image and the visual prompt image processing until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed. The present invention adopts a double-loop closed control architecture, combines a visual language model with a visual language action model, realizes semantic-level planning of the experimental task through the visual language model, decomposes complex tasks into controllable basic operation sequences, and integrates visual prompt images, text information and experimental images through the visual language action model to accurately guide the robot's actions, effectively improving the success rate of the experimental task, operational safety and compliance with regulations, and has high versatility and safety.
[0082] Optionally, the execution module 402 is specifically configured to: Extracting text information corresponding to the current subtask; Performing semantic analysis on the text information using the visual language model to obtain a semantic analysis result; The visual language model is used to label the experimental image based on the semantic analysis result to obtain the visual prompt image.
[0083] Optionally, the experimental image is a multi-view experimental image; The execution module 402 is specifically configured to: The visual language model is used to label the front view in the multi-view experimental image based on the semantic analysis result to obtain the visual prompt image.
[0084] Optionally, the execution module 402 is further configured to: Acquire, through the visual language model, and based on the experimental image after the current subtask is completed, check whether the current subtask meets the requirements; If the criteria are met, the execution steps of the next subtask of the current subtask are executed; If the requirements are not met, the execution steps of the current subtask are re-executed.
[0085] Optionally, the execution module 402 is specifically configured to: Generate operation information based on the text information, the experimental image and the visual prompt image through the visual language action model; Based on the operation information, the robot is controlled to perform the current subtask.
[0086] Optionally, the operation information includes at least one operation instruction, and each operation instruction carries an operation sequence; The execution module 402 is specifically configured to: Starting from the operation instruction with the first operation sequence, the robot is controlled to perform the operation corresponding to the current operation instruction until the operation corresponding to the operation instruction with the last operation sequence is completed, confirming the completion of the current subtask.
[0087] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute an experimental task execution method, which includes: dividing the task information of the target experimental task based on a visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks in an execution order from front to back; starting from the first subtask in the subtask sequence, processing the text information corresponding to the current subtask and the experimental image before the current subtask is executed using the visual language model to obtain a visual prompt image; and guiding the robot to execute the current subtask based on the text information, the experimental image, and the visual prompt image using the visual language action model until the robot completes the last subtask in the subtask sequence, thereby determining that the target experimental task is completed.
[0088] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0089] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the experimental task execution method provided by the above-mentioned methods, which includes: based on a visual language model, dividing the task information of the target experimental task to obtain a subtask sequence, and the subtask sequence is formed by multiple subtasks arranged from front to back in an execution order; starting from the first subtask in the subtask sequence, through the visual language model, processing the text information corresponding to the current subtask and the experimental image before the execution of the current subtask to obtain a visual prompt image, and through a visual language action model, based on the text information, the experimental image and the visual prompt image processing, guiding the robot to execute the current subtask until the last subtask in the subtask sequence is completed, and determining that the target experimental task is completed.
[0090] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the experimental task execution method provided by the above-mentioned methods, the method comprising: dividing the task information of the target experimental task based on a visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks from front to back in an execution order; starting from the first subtask in the subtask sequence, processing the text information corresponding to the current subtask and the experimental image before the execution of the current subtask through the visual language model to obtain a visual prompt image, and guiding the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image through a visual language action model until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0092] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for executing an experimental task, characterized in that: include: Based on the visual language model, the task information of the target experimental task is divided to obtain a subtask sequence, wherein the subtask sequence is formed by arranging multiple subtasks from front to back in the execution order; Starting from the first subtask in the subtask sequence, the text information corresponding to the current subtask and the experimental image before the execution of the current subtask are processed through the visual language model to obtain a visual prompt image, and based on the text information, the experimental image and the visual prompt image processing, the visual language action model is used to guide the robot to execute the current subtask until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
2. The experimental task execution method according to claim 1, characterized in that: The process of processing the text information corresponding to the current subtask and the experimental image before the current subtask is executed by the visual language model to obtain a visual prompt image includes: Extracting text information corresponding to the current subtask; Performing semantic analysis on the text information using the visual language model to obtain a semantic analysis result; The visual language model is used to label the experimental image based on the semantic analysis result to obtain the visual prompt image.
3. The experimental task execution method according to claim 2, characterized in that: The experimental image is a multi-view experimental image; The step of labeling the experimental image based on the semantic analysis result using the visual language model to obtain the visual prompt image includes: The visual language model is used to label the front view in the multi-view experimental image based on the semantic analysis result to obtain the visual prompt image.
4. The experimental task execution method according to claim 1, characterized in that: After guiding the robot to perform the current subtask based on the processing of the text information, the experimental image, and the visual prompt image, the method further includes: Acquire, through the visual language model, and based on the experimental image after the current subtask is completed, check whether the current subtask meets the requirements; If the criteria are met, the execution steps of the next subtask of the current subtask are executed; If the requirements are not met, the execution steps of the current subtask are re-executed.
5. The experimental task execution method according to any one of claims 1 to 4, characterized in that: The step of guiding the robot to perform the current subtask by processing the text information, the experimental image, and the visual prompt image using a visual language action model includes: Generate operation information based on the text information, the experimental image and the visual prompt image through the visual language action model; Based on the operation information, the robot is controlled to perform the current subtask.
6. The experimental task execution method according to claim 5, characterized in that: The operation information includes at least one operation instruction, and each operation instruction carries an operation sequence; The controlling the robot to perform the current subtask based on the operation information includes: Starting from the operation instruction with the first operation sequence, the robot is controlled to perform the operation corresponding to the current operation instruction until the operation corresponding to the operation instruction with the last operation sequence is completed, confirming the completion of the current subtask.
7. An experimental task execution device, characterized in that: include: a division module configured to divide the task information of the target experimental task based on the visual language model to obtain a subtask sequence, wherein the subtask sequence is formed by arranging the subtasks from front to back in the execution order; The execution module is configured to start from the first subtask in the subtask sequence, process the text information corresponding to the current subtask and the experimental image before the execution of the current subtask through the visual language model to obtain a visual prompt image, and guide the robot to execute the current subtask based on the text information, the experimental image and the visual prompt image through the visual language action model until the last subtask in the subtask sequence is completed, thereby determining that the target experimental task is completed.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the experimental task execution method as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the experimental task execution method as described in any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the experimental task execution method as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Mechanical arm action element imitation learning method based on language guidance and storage medium
CN112809689A
Robot simulation device
CN112847339A
Multi-modal data processing method for enhancing large language model
CN118070227A
Unmanned aerial vehicle visual language navigation method based on large model task analysis
CN119197530A
Intelligent mechanical arm operation method and system based on multi-mode large visual language model
CN119567268A
Cited By
Model training method and device, robot control method and device, equipment and storage medium
CN121682275A
Double-arm robot chemical operation system and method capable of achieving self-adaptive synthesis
CN121946548A
Intelligent agent system for multi-mode GUI (Graphical User Interface) interactive understanding and automatic operation of SEM (Scanning Electron Microscope)
CN121956635A