End device grasping control method and system based on visual cue guidance, and medium

CN122787993APending Publication Date: 2026-09-22SHANGHAI QIONCHE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611223542.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-13
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,发明人在研究中发现,现有的此类视觉-语言-动作模型在处理背景复杂、目标密集且外观相似的场景时,仍存在显著缺陷

Benefits of technology

通过在当前场景图像上根据目标的实时追踪位置生成带有显式标识的视觉提示图像,并将该视觉提示图像与原始场景图像、任务指令一同输入至视觉-语言-动作模型,本申请能够为模型提供一个明确、无歧义的引导信号,强制模型将注意力聚焦于唯一的待抓取目标。这从根本上解决了在目标密集、外观相似的复杂场景下,现有方法因目标不明确而导致的模型注意力分散和抓取失败问题,从而显著提高了机器人抓取任务的准确率和鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122787993A_ABST
    Figure CN122787993A_ABST
Patent Text Reader

Abstract

The application provides an end device grasping control method and system based on visual prompt guidance and a medium, comprising: acquiring a current scene image containing a target to be grasped and a grasping task instruction; acquiring a real-time tracking position of the target to be grasped; according to the real-time tracking position, explicitly identifying the target on the current scene image to generate a visual prompt image; inputting the visual prompt image, the current scene image and the task instruction into a visual-language-action model to output an action instruction for controlling an end effector; executing the action instruction and forming a closed-loop control by circulating the process. By generating an explicit visual prompt, the application forces the model to focus attention on the unique target, fundamentally solving the target confusion problem and significantly improving the accuracy and robustness of the robot grasping task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics, and more specifically, to a method, system, and medium for end-device grasping control based on visual cues. Background Technology

[0002] With the development of automation technology, autonomous picking operations by robots in retail, warehousing, and other scenarios are becoming increasingly common. A common approach to existing robotic grasping systems is open-loop control, where a vision system identifies and locates the target object, plans a fixed trajectory, and instructs the robotic arm to grasp it. However, in practical applications, errors in camera calibration, depth measurement, robotic arm movement, and the target's own placement can accumulate and cause deviations between the planned trajectory and the target's actual position. In scenarios with densely packed objects, this can easily lead to incorrect grasping, missed grasping, or grasping failure.

[0003] To address the shortcomings of open-loop control, the industry has proposed a closed-loop control method based on a vision-language-action model. This method combines high-level task instructions with real-time visual images to directly generate robot action commands and forms a control loop through continuous visual feedback, thereby dynamically correcting the actions. However, the inventors discovered that existing vision-language-action models still have significant limitations when handling scenes with complex backgrounds, densely packed targets, and similar appearances. When multiple similar objects (such as various medicine boxes closely arranged on a shelf) appear in the field of vision simultaneously, relying solely on the original scene image and high-level task instructions (such as "grab the medicine box"), the model struggles to eliminate the ambiguity of the instructions and cannot clearly determine which target should be grasped, leading to distraction and potentially resulting in target confusion and inaccurate grasping. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method, system, and medium for end-device grasping control based on visual cues.

[0005] A method for controlling the grasping of an end device based on visual cues, according to the present invention, includes the following steps: S1: Obtain the current scene image containing the target to be captured and the capture task instructions; S2: Obtain the real-time tracking position of the target to be captured; S3: Based on the real-time tracking position of the target to be captured, the target to be captured is explicitly marked on the current scene image to generate a visual cue image; S4: Input the visual cue image, the current scene image, and the grasping task instruction into the VLA model, so that the VLA model can output motion instructions for controlling the end effector; S5: Execute the action command; S6: Determine whether the crawling task is complete; and S7: If the grasping task is not completed, return to step S1 to form a closed-loop control based on real-time visual feedback.

[0006] Further, step S2 includes: obtaining the initial position of the target to be captured through a target detection algorithm, and starting a target tracking algorithm.

[0007] Furthermore, the target detection algorithm includes: Acquire scene images, identify all targets in the first frame scene image using a trained recognition model, and obtain candidate bounding boxes for each target; OCR recognition is performed on the images within each candidate bounding box to obtain the target detection bounding box of the target to be captured. Based on the target detection bounding box and the effective grasping diameter of the end device in the scene image, the grasping bounding box of the target to be grasped is obtained as bg=(xg,yg,wg,hg), where wg=wt-2δx, hg=ht-2δy, (xg,yg) are the coordinates of the upper left corner of the grasping bounding box, wg is the width of the grasping bounding box, and hg is the height of the grasping bounding box. Where wt and ht represent the width and height of the target detection bounding box, respectively, and δx and δy represent the safe shrinkage distances in the x-axis and y-axis directions preset according to the dimensions, respectively. The center coordinates of the grab bounding box are cg=(xg+wg / 2,yg+hg / 2).

[0008] Furthermore, the target tracking algorithm includes: Establish target template: Based on the captured bounding box, crop the target region from the scene image and extract the corresponding visual features to establish the target template F0, while saving the initial state State(0) = (bg, F0); Continuous target tracking: continuously acquire video streams. For the scene image It in frame t, predict the local search region Rt where the target appears based on the historical state State(t-1) saved in the previous frame. Extract image features only within the local search region Rt and match them with the historical target template. Calculate the current position of the target detection bounding box and obtain the new grasping bounding box bg(t) = (xt, yt, wt, ht). Where bg(t) = bg(t-1) + Δb(t), Δb(t) = (Δx, Δy, Δw, Δh) represents the change in the current position of the target detection bounding box relative to the previous frame; Update tracking state: Save the new grab bounding box and visual features as State(t).

[0009] Furthermore, methods for obtaining the target detection bounding box of the target to be captured through OCR recognition include: The text information is compared with the target information in the order to obtain a text matching score. The visual features are compared with the packaging feature library to obtain a packaging image matching score. The comprehensive matching score is calculated based on the text matching score and the image matching score. The candidate bounding box with the highest comprehensive matching score is selected as the target detection bounding box of the target to be captured.

[0010] Furthermore, the explicit identifier includes at least one of the following: drawing a bounding box around the target to be grabbed, drawing an arrow pointing to the target to be grabbed, masking and highlighting the target to be grabbed, color enhancement, or outline stroking.

[0011] Further, step S6 includes: The end device is a suction cup end device. The success of the grasp is determined by the signal of the adsorption state sensor installed on the suction cup end device. When the signal of the adsorption state sensor reaches a preset state, the grasp is determined to be successful.

[0012] Furthermore, it also includes: recording the initial position of the target to be captured when step S1 is executed for the first time; Furthermore, step S6 includes: analyzing the target to be grasped through a vision system to determine whether the target has left the initial position.

[0013] A visually cued end-device grasping control system according to the present invention includes: Image acquisition equipment; Robotic arm and end effector mounted on the robotic arm; The processor is configured to perform steps of the visually cued end-device grasping control method as described.

[0014] According to the present invention, a computer-readable storage medium storing a computer program is provided, wherein when the computer program is executed by a processor, the steps of the visually cued end-device grasping control method are implemented.

[0015] Compared with the prior art, the present invention has the following beneficial effects: By generating a visual cue image with explicit markings on the current scene image based on the real-time tracking position of the target, and inputting this visual cue image along with the original scene image and task instructions into the vision-language-action model, this application provides the model with a clear and unambiguous guiding signal, forcing the model to focus its attention on a single target to be grasped. This fundamentally solves the problem of model attention distraction and grasping failure caused by unclear targets in complex scenes with dense and similar targets, thereby significantly improving the accuracy and robustness of robot grasping tasks.

[0016] By integrating target tracking, visual cue generation, model decision-making, and action execution into a high-frequency, real-time closed-loop control, the robot can continuously fine-tune its actions based on visual feedback, no longer relying on fixed preset trajectories. This results in greater adaptability to camera calibration errors, robotic arm motion errors, and environmental dynamic disturbances (such as slight displacement of objects), improving operational efficiency and scene versatility. Attached Figure Description

[0017] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A schematic diagram of the architecture of the robot grasping control system provided in the embodiments of this application; Figure 2 A flowchart illustrating the robot grasping control method based on visual cues provided in this application embodiment; Figure 3 This is a schematic diagram illustrating the visual cues provided in the embodiments of this application; Figure 4 This is a timing diagram of closed-loop control signaling interaction provided in an embodiment of this application. Detailed Implementation

[0018] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0019] Example 1 This embodiment provides a basic implementation scheme for a suction cup end device gripping control method and system based on visual cues. This scheme is particularly suitable for performing precise and robust gripping operations on small, similar-looking medicine boxes in a simulated pharmacy dense shelf environment.

[0020] Reference Figure 1 , Figure 1This is a schematic diagram of the hardware and software architecture of the robot grasping control system in this application embodiment. Specifically, the system includes an industrial camera 10, a multi-degree-of-freedom robotic arm 20, a suction cup end effector 21 installed at the end of the robotic arm 20, a robot controller 30, and a server 40 as the core of the system.

[0021] As a specific implementation, the industrial camera 10 can be a global shutter color industrial camera with a resolution of 1920x1080 pixels and a frame rate of 60 frames per second. This industrial camera 10 is fixedly mounted on the wrist of the robotic arm 20 to provide an "eye-on-hand" perspective, ensuring that the robotic arm 20 can continuously obtain clear observation images as it approaches the target. The robotic arm 20 can be a six-degree-of-freedom collaborative robot, providing sufficient flexibility for movement within complex shelving spaces. The suction cup end effector 21 includes a 20mm diameter silicone suction cup connected to a vacuum generator via an air tube. It should be noted that this device also integrates a suction status sensor (specifically a digital pressure sensor) to monitor changes in air pressure inside the suction cup in real time, thereby determining whether an object has been successfully adsorbed. The robot controller 30, as the original equipment manufacturer's controller for the robotic arm 20, is responsible for receiving motion commands from the server 40 and parsing them into low-level motor control signals to drive the movement of each joint of the robotic arm 20.

[0022] Server 40, as the core computing unit of this embodiment, can be configured as an industrial computer equipped with a high-performance graphics processing unit. Server 40 runs key software modules that implement the method of this application, which can be logically divided into: a target detection module 41, a target tracking module 42, a visual cue generation module 43, and a vision-language-action decision module 44 (VLA decision module). Specifically, the target detection module 41 locates the target to be grasped in the initial image; the target tracking module 42 continuously tracks the target in subsequent consecutive image frames; the visual cue generation module 43 generates a prominent label on the image based on the tracking results; and the vision-language-action decision module 44 integrates all information to generate the final robot action command. These modules work together to achieve intelligent closed-loop control from visual perception to action execution. In this system, the information flow path is as follows: the industrial camera 10 transmits the acquired digital image stream to the server 40 in real time; after processing by various modules within the server 40, the vision-language-motion decision module 44 generates motion commands, which are then sent to the robot controller 30 via Ethernet; the robot controller 30 then drives the robotic arm 20 and the suction cup end effector 21 to perform the actions. Correspondingly, the pressure sensor signal on the suction cup end effector 21 is also fed back to the robot controller 30 and can be forwarded to the server 40 for task status determination.

[0023] The following will combine Figure 2 and Figure 3 This paper elaborates on the specific process of the robot grasping control method based on visual cues in this embodiment. Figure 2 This is a schematic diagram of the overall process of the method in this application. Figure 3 This intuitively demonstrates the generation effect of visual cues.

[0024] In a specific application scenario, suppose the task is to pick up a specific target, such as "ibuprofen capsules," from a variety of closely arranged medicine boxes on a shelf. These medicine boxes are usually similar in size (e.g., both length and width are less than 4 cm), the distance between them may be less than 2 cm, and the packaging design style is similar, so there are multiple interfering targets in the scenario.

[0025] Specifically, the method begins with step S101: acquiring images and instructions. The system first receives a grasping task instruction issued by the upper-level task management system. This instruction can be a simple text string, such as "grab ibuprofen capsules". At the same time, the industrial camera 10 installed on the wrist of the robotic arm 20 acquires images of the current scene at a frequency of 30 Hz and transmits these image frames to the server 40 in real time.

[0026] Subsequently, step S102 is executed: target detection and tracking. It should be noted that this step is fundamental for achieving continuous monitoring of a specific target. As a preferred implementation, step S102, upon initial execution, may include two phases: target detection and target tracking initiation.

[0027] The first stage is object detection. The object detection module 41 on server 40 receives the first frame of scene image and executes the object detection algorithm. In this embodiment, the object detection module 41 uses a pre-trained YOLOv8n model. This model is trained using a dataset containing thousands of medicine box images collected under different lighting, angles, and backgrounds. It can identify all medicine box targets in the current image and output the position of each medicine box. Suppose that N candidate medicine boxes are detected in the current image, then all candidate targets form a set B={b1,b2,……,bN}, where the i-th candidate bounding box is represented as bi=(xi,yi,wi,hi), where (xi,yi) represents the coordinates of the upper left corner of the bounding box; wi represents the width of the bounding box; and hi represents the height of the bounding box. Subsequently, the system performs OCR text recognition and packaging image feature matching for each candidate bounding box.

[0028] The OCR module identifies textual information such as the drug name, specifications, and manufacturer on the medicine box packaging and compares it with the target drug information in the order to obtain a text matching score S_ocr(i). Simultaneously, the packaging image feature matching module extracts visual features such as the color distribution, logo, text layout, and pattern texture of the medicine box packaging and calculates similarity with a drug packaging feature database to obtain a packaging image matching score Si_img(i). The system calculates the comprehensive matching score S(i) = αS_ocr(i) + βS_img(i), where α and β are weighting coefficients, and α + β = 1. Finally, b_target = argmaxS(i) is determined as the detection bounding box of the target medicine box.

[0029] To make subsequent suction cup gripping more stable, this invention does not directly use the target detection bounding box as a gripping reference, but generates a gripping bounding box based on the effective adsorption area of ​​the suction cup end device.

[0030] Let the effective suction diameter of the suction cup correspond to the size d in the image, then the grasping bounding box is represented as bg=(xg,yg,wg,hg). Where wg=wt-2δx, hg=ht-2δy. wt and ht represent the width and height of the target detection bounding box, respectively; δx and δy represent the preset safe retraction distance based on the suction cup's size d in the image. The center coordinates of the grasping bounding box are calculated as cg=(xg+wg / 2,yg+hg / 2). This center point serves as a unified reference position for the robotic arm's grasping trajectory planning and visual cues generation. Because the grasping bounding box avoids the edge area of ​​the medicine box, it reduces the probability of the suction cup simultaneously covering adjacent medicine boxes or adsorbing to the folded edge of the packaging, improving adsorption stability and grasping success rate.

[0031] Once the target detection module 41 completes target recognition, the system sends the captured bounding box bg to the target tracking module 42 to start the tracker.

[0032] In this embodiment, the target tracking module 42 adopts the EfficientTAM (Efficient Track Anything Model) tracking algorithm, and its specific workflow is as follows.

[0033] Step 1: Create the target template.

[0034] The target tracking module 42 crops the target region from the first frame image based on the captured bounding box, extracts the corresponding visual features, establishes the target template F0, and saves the initial state State0=(bg,F0).

[0035] Since there is no historical tracking state in the first frame, the initial state described above is directly used as the input to the EfficientTAM model. It is understood that this invention tracks the grasping bounding box after the suction cup constraint, rather than the original detection bounding box; therefore, subsequent tracking results always remain consistent with the actual grasping area of ​​the suction cup.

[0036] Step 2: Continuous target tracking.

[0037] The industrial camera continuously acquires video streams. For the current frame t, EfficientTAM receives: the current image It; and the historical state State(t-1) saved from the previous frame. The model first predicts the possible local search region Rt where the target may appear based on the historical bounding boxes.

[0038] Subsequently, image features are extracted only within the local search region Rt and matched with historical target templates to calculate the current position. The new capture bounding box is represented as bg(t) = (xt, yt, wt, ht), which can be further represented as bg(t) = bg(t-1) + Δb(t), where Δb(t) = (Δx, Δy, Δw, Δh) represents the change in the current position relative to the previous frame.

[0039] Since target matching is performed only in local search areas without re-detecting the entire image, computational complexity is significantly reduced, while the continuity of target motion between video frames can be fully utilized to improve tracking speed.

[0040] Furthermore, since the bounding box is updated continuously and recursively, the change in target position is smoother, avoiding bounding box jumps caused by occasional detection errors.

[0041] Step 3: Update the tracking status.

[0042] The target tracking module 42 saves the new capture bounding box and target features of the current frame as State(t) as the tracking input for the next frame. At the same time, the center of the capture bounding box cg(t) is sent to the visual cue generation module 43.

[0043] The visual cue generation module 43 draws visual cues in real time based on the grasping bounding box, instead of using YOLO detection boxes. Therefore, the visual cues are always generated around the actual grasping area of ​​the suction cup.

[0044] Even if the movement of the robotic arm 20 causes changes in the viewing angle, slight rotation of the medicine box, or partial occlusion, the tracking module can still continuously update the grasping bounding box, so that the visual cues can stably cover the same target medicine box.

[0045] Therefore, the present invention adopts a control strategy of "one-time detection, double confirmation, capture box generation, and continuous tracking". Target detection and target confirmation are only performed in the initialization phase, while the real-time position of the target can be obtained by simply updating the capture bounding box in the subsequent closed-loop control process.

[0046] Compared to the traditional "detection-grasp" method, this invention converts the target detection bounding box into a gripping bounding box adapted to the suction cup end effector. This gripping bounding box serves as the tracking object for EfficientTAM and as the basis for visual cues, ensuring consistency between the robotic arm's control reference position, the visual cues position, and the actual gripping position of the suction cup. This effectively reduces the deviation between the detection box and the actual gripping point, improving the robot's gripping accuracy, closed-loop control stability, and continuous operation efficiency in densely stocked pharmacy scenarios.

[0047] Next, step S104 is executed: Visual-Language-Action Model Decision. This step is crucial for achieving end-to-end intelligent control. The Visual-Language-Action Decision module 44 receives three sets of input data: 1) the visual cue image generated in step S103; 2) the original, unmodified current scene image; and 3) the initial grasping task instruction (such as the text encoding of "grab"). By using both the visual cue image and the original scene image as input, the model can obtain clear guidance about the "grab object" while retaining complete scene context information, thus making a more robust decision.

[0048] In this embodiment, the vision-language-action decision module 44 employs a vision-language-action model composed of a visual encoder, a language encoder, a multimodal fusion module, and an action prediction module. This model predicts the actions that the suction cup end effector 21 needs to perform in the next control cycle based on the visual information of the current scene, visual cues, and grasping task instructions. The vision-language-action model is trained using robot teleoperation data, enabling it to learn the correspondence between visual scenes, target visual cues, task instructions, and the robot's actual operational actions. This allows it to generate incremental robot control instructions based on real-time visual information during the actual grasping process.

[0049] Specifically, the inputs to the vision-language-action model include visual input, visual cue input, and language task input. The visual input is the original scene image captured by the camera at the current moment; the visual cue input is the image containing the target visual cue 60 generated in step S103; and the language task input is the grasping task instruction to be performed. The original scene image is used to preserve complete scene context information such as the medicine shelf, target medicine box, interference medicine box, robotic arm 20, and suction cup end effector 21. The visual cue image explicitly marks the target medicine box 50, enabling the model to clearly identify the target object for the grasping operation. The language task instruction further defines the type of task the model needs to perform.

[0050] In the visual information processing part of the model, the original scene image and the visual cue image are respectively processed by a visual encoder to extract features, obtaining visual features corresponding to the current scene and the target visual cue. The visual encoder can adopt a visual coding network based on the Transformer structure, which divides the input image into multiple image regions or image blocks, and maps each image region to a corresponding visual feature vector, thereby forming a visual feature sequence.

[0051] Specifically, for the original scene image, the visual encoder extracts scene visual features representing the medicine shelf, medicine box, robotic arm, suction cup end device and the surrounding environment; for the visual cue image, the visual encoder simultaneously extracts the visual features corresponding to the target medicine box and the visual cue 60, enabling the model to obtain the position of the target object in the current scene and its spatial relationship with the surrounding environment.

[0052] In the language information processing part of the model, the grasping task instructions are converted into corresponding language feature sequences. For example, when the task instruction is "grab the target medicine box", the language encoder converts the task instruction into corresponding language features, enabling the model to obtain the semantic constraints of the current task.

[0053] Subsequently, the visual feature sequence and the language feature sequence are input into a multimodal fusion module for joint processing. The multimodal fusion module employs a Transformer-based attention mechanism to associate visual information, visual cue information, and language task information, enabling the model to establish correspondences between the target medicine box, the location of the visual cue, the surrounding scene, and the current task.

[0054] During the training phase, robot grasping data is collected through manual teleoperation. Specifically, operators control the robotic arm 20 and the suction cup end effector 21 via teleoperation devices to perform target medicine box grasping tasks in a pharmacy retail setting or an environment with the same or similar shelf, medicine box, and robot configuration as an actual pharmacy. During manual teleoperation, the system simultaneously collects visual observation information, task instructions, and actual robot actions during task execution, thereby forming operational data for training the vision-language-motion model.

[0055] For each segment of manually operated data acquired, a visual cue is added to the corresponding visual image to indicate the target medicine box that needs to be grasped. In this embodiment, bounding boxes are used as visual cue markers. That is, based on the target medicine box actually selected during the manual remote operation, the boundary area of ​​the target medicine box is marked in the corresponding image, thereby obtaining a visual cue image corresponding to the original scene image.

[0056] Therefore, at any control time t, the corresponding training data includes at least: [D_t=(I_t,I_t^{prompt},L_t,A_t)] Where (I_t) represents the original scene image acquired at the t-th control moment; (I_t^{prompt}) represents the visual cue image obtained after visual cue annotation of the target 50 to be grasped in the original scene image; (L_t) represents the grasping task instruction corresponding to the control moment; and (A_t) represents the actual action performed by the operator through manual remote operation of the robotic arm 20 and the suction cup end device 21.

[0057] The visual cue image (I_t^{prompt}) is obtained by post-processing the manual teleoperation data. After the training data is collected, the corresponding target object is determined based on the target medicine box actually selected during the manual teleoperation, and the target object is visually cued and labeled in the scene image at the corresponding time, thereby establishing the correspondence between the target medicine box and the robot's actual operation.

[0058] In this way, the training samples at the same control moment simultaneously contain complete scene information, clear target indication information, task semantic information, and actual robot operation actions, so that the model can learn the specific operation methods that the robot should take under specific task and target indication conditions.

[0059] The above training data is input into the vision-language-action model for imitation learning training. For each training time step, the model uses the original scene image (I_t), the visual cue image (I_t^{prompt}), and the task instruction (L_t) as conditional inputs, and the actual robot action (A_t) obtained through manual teleoperation as the supervision target, enabling the model to learn the following mapping relationship: [\hat{A}*t=f*{\theta}(I_t,I_t^{prompt},L_t)] Where (f_{\theta}) represents the vision-language-action model, (\theta) represents the model parameters, and (\hat{A}_t) represents the robot action predicted by the model.

[0060] During training, the model parameters are optimized based on the difference between the model's predicted action (\hat{A}_t) and the actual manual teleoperation action (A_t). This allows the model to gradually learn how the human operator adjusts the position and posture of the robotic arm 20 and the suction cup end device 21 under different visual states, and further learns under what states the suction cup is activated for grasping.

[0061] In the motion prediction section, the fused features output by the multimodal fusion module are input into the motion prediction module, which then predicts the robot action corresponding to the current control cycle based on the fused visual, visual cues, and linguistic features. The actions are output as continuous numerical values, rather than being pre-defined as fixed robotic arm trajectories, thus enabling incremental adjustments to the robotic arm 20 based on the current visual state.

[0062] In this embodiment, the vision-language-action model outputs a 7-dimensional continuous action vector: [A_t=[\Delta x_t,\Delta y_t,\Delta z_t,\Delta r_t,\Delta p_t,\Deltay_t^{rot},s_t]] Where (\Delta x_t), (\Delta y_t), and (\Delta z_t) represent the translation increments of the suction cup end effector 21 along the X, Y, and Z directions in the robot base coordinate system, respectively, in millimeters; (\Delta r_t), (\Delta p_t), and (\Delta y_t^{rot}) represent the rotation increments of the suction cup end effector 21 around the corresponding attitude axis, respectively, in radians; and (s_t) represents the adsorption state control quantity of the suction cup end effector 21.

[0063] The adsorption state control quantity is used to control whether the suction cup initiates vacuum adsorption. In this embodiment, (s_t=0) indicates maintaining a non-adsorption state, and (s_t=1) indicates initiating adsorption.

[0064] For example, the visual-language-action model outputs the following during a certain control cycle: [A_t=[2.5,-1.0,-5.0,0,0,0.01,0]] This indicates that during the current control cycle, the suction cup end effector 21 moves 2.5 mm along the positive X-axis of the robot base coordinate system, 1.0 mm along the negative Y-axis, and 5.0 mm along the negative Z-axis, while simultaneously generating a 0.01 radian attitude rotation and keeping the suction cup in a non-adhesive state.

[0065] In actual execution, the visual cue image generated in step S103 and the current original scene image are input into the visual-language-action decision module 44. Since the target to be grasped 50 has been clearly identified by the bounding box in the visual cue image, the model can use the visual cue to determine the specific target that needs to be grasped, and at the same time use the original scene image to obtain complete environmental information around the target.

[0066] Furthermore, since step S102 continuously updates the grasping bounding box of the target 50, and step S103 continuously generates new visual cue images based on the updated grasping bounding box, the visual cues obtained by the visual-language-action model in each closed-loop control cycle correspond to the latest position of the target pillbox. The model can thus adjust its actions for the next control cycle based on the real-time positional relationship of the target pillbox relative to the suction cup end device 21, the orientation of the target pillbox, and surrounding environmental information.

[0067] For example, when there is a horizontal deviation between the target center indicated by the visual cue 60 and the projection center of the suction cup end device 21 in the current image, the model outputs the corresponding X-axis and / or Y-axis translation increments based on the manual teleoperation behavior learned during training, so that the suction cup gradually moves towards the target center; when there is an attitude deviation between the suction cup end device 21 and the surface of the target medicine box, the model outputs the corresponding attitude rotation increments, so that the suction cup gradually adjusts to a suitable attitude for adsorption; when the model determines that the suction cup has reached a suitable position for grasping based on the current visual state, it outputs an adsorption state control quantity (s_t=1), thereby controlling the suction cup to start vacuum adsorption.

[0068] Therefore, the vision-language-action model in this embodiment does not output a complete one-time robotic arm motion trajectory, but rather outputs incremental motion commands for the next control cycle based on the latest visual state obtained in each closed-loop cycle. In this way, the robotic arm 20 can continuously correct its position and posture according to the real-time changes in the target position during the grasping process, thereby improving the robot's adaptability to errors in the position of the medicine box, robotic arm motion errors, and target placement deviations.

[0069] Subsequently, step S105 is executed: Execute the action. The server 40 sends the 7-dimensional incremental motion vector generated by the vision-language-motion decision module 44 in the current control cycle to the robot controller 30. The robot controller 30 converts it into target angle commands for the six joints of the robotic arm 20 and drives the robotic arm 20 and the suction cup end effector 21 to execute this small, incremental movement.

[0070] After the action is completed, the process proceeds to step S106: determining whether the task is complete. This step is used to determine whether the entire grasping task has been successfully completed. In this embodiment, this determination is mainly based on the signal from the adsorption status sensor (i.e., pressure sensor) on the suction cup end device 21. Specifically, in the action command output by the vision-language-action model, when it determines that the suction cup is close enough to and aligned with the target surface, it sets the 7th dimension of the action vector to 1 to initiate adsorption. At this time, the robot controller 30 activates the vacuum generator according to the adsorption status control quantity, creating a negative pressure inside the suction cup. The robot controller 30 continuously monitors the reading of the pressure sensor. When the detected negative pressure value is lower than a preset threshold (e.g., -0.6 bar) and remains stable for more than 100 milliseconds, the system determines that the grasping is successful.

[0071] If the determination result of step S106 is "yes", it indicates that the grasping task is completed, and the robot controller 30 will execute the subsequent lifting and placing actions, and the entire closed-loop control process will end.

[0072] Conversely, if the judgment result is "no" (for example, the adsorption command has not been issued, or the pressure has not reached the preset threshold after the adsorption command is issued), it means that the target capture has not been completed in the current control cycle. At this time, the system returns to step S101, reacquires the current scene image, and sequentially performs target position update, visual cue generation, and visual-language-action model decision-making, thereby forming the next closed-loop control iteration.

[0073] Reference Figure 4 , Figure 4 The signaling interaction timing of the closed-loop control in this application is shown. Image acquisition device 70 (corresponding to...) Figure 1 The industrial camera 10 continuously sends image frames to the processing unit 80 (corresponding to...). Figure 1The processing unit 80 (server 40) generates action commands and sends them to the robot controller 30 after executing steps S102 to S104, which then executes the corresponding actions. It is understood that this "perception-decision-execution" cycle repeats continuously at an extremely high frequency (30 Hz in this embodiment). In each cycle, the robot fine-tunes its posture based on the latest visual information to dynamically compensate for various errors, thereby gradually approaching and ultimately precisely aligning with the target. This closed-loop control based on real-time visual feedback constitutes the core mechanism for achieving high-precision grasping in this application.

[0074] As a verification, in a comparative experiment involving 1000 grasping attempts, the grasping success rate of the method in this embodiment reached 94.7% when facing densely arranged medicine boxes, while the success rate of the traditional "target detection + fixed trajectory planning" open-loop scheme was only 81.2%. Furthermore, the average execution time of a single task using the proposed solution is 17.9 seconds, significantly improved compared to the 20.2 seconds of the traditional solution. The above experimental results fully demonstrate that the technical solution of this application has significant beneficial effects in improving grasping accuracy, robustness, and operational efficiency.

[0075] Example 2 This embodiment demonstrates a variation of the visual cues and illustrates the universality of the method for different visual cues. In this embodiment, the system hardware configuration, overall control flow, and architecture of the vision-language-action decision module are basically the same as in Embodiment 1. The main difference lies in the specific method of generating the visual cues image in step S103.

[0076] In Example 1, the visual cue 60 uses a rectangular bounding box. However, when dealing with objects that are irregularly shaped or tilted (such as an irregularly shaped cosmetic bottle), the rectangular bounding box may contain a large background area or fail to accurately reflect the actual outline of the object, thus interfering with the decision of the vision-language-action model to generate the optimal contact point.

[0077] To address this situation, this embodiment modifies the visual cue generation module 43. In step S102, in addition to outputting the bounding box of the target using the target tracking algorithm, an instance segmentation model or a segmentation head integrated into the tracker can be used to obtain a pixel-level binary mask of the target 50 to be captured. This mask is a two-dimensional array with the same size as the original image, where the pixel position corresponding to the target object is 1 and the background pixel position is 0.

[0078] Accordingly, in step S103, the visual cue generation module 43 no longer draws bounding boxes, but instead processes the visual cue image using this binary mask. One specific processing method is to overlay the area with a mask value of 1 (i.e., the target object area) onto the original scene image with a semi-transparent green highlight, for example, using a color with an RGBA value of (0, 255, 0, 128). In the visual cue image generated in this way, the target 50 to be grasped is displayed with a striking green highlight, its outline perfectly matching the actual outline of the object, while the surrounding interfering targets 51 and the background remain unchanged.

[0079] In step S104, when the vision-language-action decision module 44 receives this visual cue image based on mask highlighting, its internal convolutional feature extraction network can more accurately perceive the actual shape, pose, and graspable area of ​​the target. The model's attention mechanism is guided to the centroid or key region of the entire highlighted area, rather than a broad rectangular box. This helps the model generate more refined alignment actions. For example, when grasping a tilted bottle, the model may generate an action command with a certain rotation to keep the surface of the suction cup end device 21 parallel to the tilted surface of the bottle, thereby achieving a more reliable seal and adsorption.

[0080] It is understandable that, in addition to bounding boxes and mask highlighting, the explicit identifiers can take many other forms, such as drawing an arrow pointing to the center of the target 50, enhancing the color of the target area (e.g., increasing saturation), or outlining its contours. These different forms of visual cues 60 all serve the same core function: to provide a clear and unambiguous focus of attention for the visual-language-action model amidst complex visual information. The framework of this application has good compatibility with these specific identifier forms, allowing users to choose the most suitable cue method based on specific application scenarios and target characteristics.

[0081] Example 3 This embodiment provides a variant scheme that uses different sensor configurations and grasp completion determination logic, with the aim of enhancing the applicability of the system in scenarios that do not rely on contact sensors and reducing hardware costs.

[0082] In this embodiment, the system architecture is similar to that of Embodiment 1, but with two key modifications. First, the industrial camera 10 is replaced with an RGB-D camera. This type of camera, in addition to providing the same color image as in Embodiment 1, can simultaneously provide a registered and aligned depth map. Each pixel value in the depth map represents the distance of that point in the camera coordinate system. Second, a pressure sensor may not be installed on the suction cup end device 21, or even if installed, it may not be used as the sole criterion for successful gripping.

[0083] Accordingly, the control method's flow is also adaptively adjusted. In step S101, when acquiring the image, the system simultaneously acquires a color image and a depth map. In step S104, the input to the vision-language-action decision module 44 can be further expanded; in addition to the visual cue image, the original color image, and the task command, it can selectively receive a depth map as a fourth input. The addition of depth information allows the vision-language-action model to directly perceive the three-dimensional structure of the scene, which helps it generate more precise action commands in the depth direction. For example, it can more accurately determine the distance between the suction cup and the target surface, thereby optimizing the approach speed and timing.

[0084] It should be noted that the core difference in this embodiment lies in the implementation method of step S106, "determining whether the task is completed." Since it does not rely on pressure sensors, this embodiment employs a judgment logic based on visual analysis. Specifically, before executing the grasping task, this method can add an extra step: when step S101 is executed for the first time, after the target detection module 41 identifies the target 50 to be grasped, the system records the initial position of the target on the shelf. This initial position can be its bounding box coordinates, or it can be an estimated position in the three-dimensional world coordinate system (if combined with depth information and camera extrinsic parameters).

[0085] During closed-loop control, once the vision-language-action decision module 44 determines that alignment has been completed and issues suction and lifting action commands, the system enters the vision judgment process. After the robotic arm 20 performs the lifting action, the industrial camera 10 (i.e., the RGB-D camera) acquires a new scene image. The vision system (which can be a standalone analysis module or a function of the vision-language-action model itself) compares this new image with the recorded initial position information.

[0086] There are several specific methods for determining the target's location. One possible approach is target disappearance detection: the system re-runs the target detection algorithm in the new image, targeting the previously recorded initial location area. If the target is not detected within this area (i.e., the target has disappeared), the capture is considered successful. Another method is background subtraction: the system compares the pixel content of the target's initial location area in the current image with the pixel content of that area before capture. If a significant change has occurred (e.g., what was originally a medicine box pattern has now become the background of a shelf), the capture is considered successful.

[0087] If the visual analysis determines that the capture was successful, the process ends. If the determination fails (for example, the target is still in place), it means that the adsorption failed. At this time, the system can trigger retry logic, such as restarting the alignment process or fine-tuning the adsorption point and trying again.

[0088] Understandably, this vision-based decision-making scheme frees the entire system from reliance on sensors on the end effector, thereby reducing hardware complexity and cost, and enhancing the versatility of the solution. Furthermore, visual confirmation that an object has been picked up provides a more reliable status input for subsequent operations such as object handover and placement.

[0089] Example 4 This embodiment aims to illustrate the universality and scalability of the "visual cues-guided vision-language-action model" core control framework proposed in this application. This framework is not only applicable to suction cup end effectors, but can also be seamlessly transferred to other types of end effectors, such as mechanical grippers.

[0090] In this embodiment, the hardware configuration of the system is mostly the same as in Embodiment 1, with the main difference being the end effector. We replace the suction cup end effector 21 in Embodiment 1 with a two-finger parallel gripper, such as the Robotiq 2F-85 model. This gripper can integrate a force sensor for detecting the gripping force.

[0091] Accordingly, at the software level, especially for the vision-language-action decision module 44, adaptive adjustments are required. While its core Transformer architecture and multimodal input processing remain unchanged, its output motion space and training dataset need to be redesigned for the gripper. In this case, the output of the vision-language-action model no longer controls the translation, rotation, and suction state of the suction cup, but rather the action of the gripper. For example, its output 7-dimensional vector can be defined as follows: the first 6 dimensions are still the translation and rotation increments of the end effector in the robot's base coordinate system, while the 7th dimension represents the gripper's state, such as a value between [0, 1], where 0 represents fully closed and 1 represents fully open. Alternatively, the motion space can be designed as more advanced discrete commands, such as "align with target center," "move to pre-grip position," "close gripper," "apply 5 Newtons of gripping force," etc. To enable the model to learn to generate these new action commands, a new imitation learning dataset labeled and collected for gripper grasping tasks is needed to fine-tune or retrain the vision-language-action model.

[0092] Although the part responsible for action output and training in the visual-language-action model (which can be regarded as the "back end" of the model) has changed, the core logic responsible for perception, tracking and guiding the model's attention through visual cues (which can be regarded as the "front end" of the model) remains unchanged in this application.

[0093] In a specific grabbing task, such as grabbing a cylindrical canned beverage from a supermarket shelf, the process still follows... Figure 2The steps are as follows. The system first locates the target beverage bottle through object detection and tracking, and clearly identifies it in the real-time image stream using visual cues 60 (e.g., a bounding box or mask close to the bottle). The vision-language-action decision module 44 receives the image with visual cues and is guided to focus attention on the beverage bottle. The model first outputs a series of action commands to control the robotic arm 20 to move the open gripper to both sides of the beverage bottle and adjust its posture so that the gripper fingertips are perpendicular to the bottle surface. When the model determines that the position and posture are appropriate, it outputs a "close gripper" command (e.g., setting the 7th dimension of the output vector from 1 to 0.2). The robot controller 30 executes this command, and the gripper begins to close. Force sensors on the gripper monitor the gripping force, and when a preset stable gripping force (e.g., 5 Newtons) is reached, the system determines that the grasping is successful.

[0094] This embodiment fully demonstrates that the technical solution proposed in this application has good decoupling and scalability. Its core "visual cue guidance" mechanism, as a general attention guidance strategy, is independent of specific end effector types. When adapting to new grasping tools or tasks, the main work involves adapting the action output space of the vision-language-action model and conducting targeted training, while the front-end perception, tracking, and guidance framework remains unchanged. This characteristic greatly enhances the versatility of the method and its application potential in different industrial and commercial scenarios. For example, in addition to two-finger grippers, this method is also applicable to other end effectors such as three-finger grippers and flexible grippers.

[0095] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0096] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for controlling the grasping of an end device based on visual cues, characterized in that, Including the following steps: S1: Obtain the current scene image containing the target to be captured and the capture task instructions; S2: Obtain the real-time tracking position of the target to be captured; S3: Based on the real-time tracking position of the target to be captured, the target to be captured is explicitly marked on the current scene image to generate a visual cue image; S4: Input the visual cue image, the current scene image, and the grasping task instruction into the VLA model, so that the VLA model can output motion instructions for controlling the end effector; S5: Execute the action command; S6: Determine whether the crawling task is completed; as well as S7: If the grasping task is not completed, return to step S1 to form a closed-loop control based on real-time visual feedback.

2. The end-device grasping control method based on visual cues and guidance according to claim 1, characterized in that, Step S2 includes: obtaining the initial position of the target to be captured through a target detection algorithm, and starting a target tracking algorithm.

3. The end-device grasping control method based on visual cues and guidance according to claim 2, characterized in that, The target detection algorithm includes: Acquire scene images, identify all targets in the first frame scene image using a trained recognition model, and obtain candidate bounding boxes for each target; OCR recognition is performed on the images within each candidate bounding box to obtain the target detection bounding box of the target to be captured. Based on the target detection bounding box and the effective grasping diameter of the end device in the scene image, the grasping bounding box of the target to be grasped is obtained as bg=(xg,yg,wg,hg), where wg=wt-2δx, hg=ht-2δy, (xg,yg) are the coordinates of the upper left corner of the grasping bounding box, wg is the width of the grasping bounding box, and hg is the height of the grasping bounding box. Where wt and ht represent the width and height of the target detection bounding box, respectively, and δx and δy represent the safe shrinkage distances in the x-axis and y-axis directions preset according to the dimensions, respectively. The center coordinates of the grab bounding box are cg=(xg+wg / 2,yg+hg / 2).

4. The end-device grasping control method based on visual cues as described in claim 3, characterized in that, The target tracking algorithm includes: Establish target template: Based on the captured bounding box, crop the target region from the scene image and extract the corresponding visual features to establish the target template F0, while saving the initial state State(0) = (bg, F0); Continuous target tracking: continuously acquire video streams. For the scene image It in frame t, predict the local search region Rt where the target appears based on the historical state State(t-1) saved in the previous frame. Extract image features only within the local search region Rt and match them with the historical target template. Calculate the current position of the target detection bounding box and obtain the new grasping bounding box bg(t) = (xt, yt, wt, ht). Where bg(t) = bg(t-1) + Δb(t), Δb(t) = (Δx, Δy, Δw, Δh) represents the change in the current position of the target detection bounding box relative to the previous frame; Update tracking state: Save the new grab bounding box and visual features as State(t).

5. The end-device grasping control method based on visual cues and guidance according to claim 3, characterized in that, Methods for obtaining the target detection bounding box of the target to be captured through OCR recognition include: The text information is compared with the target information in the order to obtain a text matching score. The visual features are compared with the packaging feature library to obtain a packaging image matching score. The comprehensive matching score is calculated based on the text matching score and the image matching score. The candidate bounding box with the highest comprehensive matching score is selected as the target detection bounding box of the target to be captured.

6. The end-device grasping control method based on visual cues and guidance according to claim 1, characterized in that, The explicit identifier includes at least one of the following: drawing a bounding box around the target to be grabbed, drawing an arrow pointing to the target to be grabbed, masking and highlighting the target to be grabbed, color enhancement, or outline stroking.

7. The end-device grasping control method based on visual cues and guidance according to claim 1, characterized in that, Step S6 includes: The end device is a suction cup end device. The success of the grasp is determined by the signal of the adsorption state sensor installed on the suction cup end device. When the signal of the adsorption state sensor reaches a preset state, the grasp is determined to be successful.

8. The end-device grasping control method based on visual cues and guidance according to claim 1, characterized in that, Also includes: Record the initial position of the target to be captured when step S1 is executed for the first time; Furthermore, step S6 includes: analyzing the target to be grasped through a vision system to determine whether the target has left the initial position.

9. A visually cued end-device grasping control system, characterized in that, include: Image acquisition equipment; Robotic arm and end effector mounted on the robotic arm; The processor is configured to perform the steps of the visually cued end-device grasping control method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the visual cues-guided end-device grasping control method according to any one of claims 1 to 8.