Grape picking method and device based on visual language model
By combining a visual language model with a tracked chassis and multiple cameras, a grape harvesting method has been developed, which solves the problems of poor adaptability and insufficient occlusion recognition in existing grape harvesting devices, and achieves efficient grape harvesting results.
Patent Information
- Application Number
- CN202510898163.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-17
AI Technical Summary
Existing automated grape harvesting devices are difficult to adapt to the different growth postures of vines in different vineyards, and visual recognition solutions cannot effectively harvest grapes when the stems are obscured, resulting in low harvesting efficiency and an inability to cope with unstructured environments.
A grape-harvesting method based on a visual language model is adopted, which combines a tracked chassis, a six-axis robotic arm and multiple cameras. The visual language model identifies the distribution of grape fruits, and depth and wide-angle cameras are used to acquire image data. The position of the robotic arm is adjusted to achieve accurate harvesting and adapt to unstructured environments.
It improves the success rate and efficiency of grape harvesting, and can effectively identify and harvest grapes in unstructured environments, adapting to different harvesting scenarios.
Smart Images

Figure CN120791745A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of visual recognition, in particular to a grape picking method and device based on a visual language model. BACKGROUND
[0002] In the existing grape automatic picking device, the special mechanism capable of separating grape fruits cannot well adapt to the growth postures of plants in different vineyards in terms of size, and the current visual recognition scheme only aims at locating and separating grape fruits, that is, the current existing technology cannot pick the grape fruits once the picking point (i.e. the fruit stem part) is blocked.
[0003] Therefore, the existing picking robot application scenarios are not comprehensive, and can only blindly repeat the set actions, which not only is low in efficiency, but also cannot well solve the problems once the scene does not match the expected situation. SUMMARY
[0004] The present application aims to provide a grape picking method and device based on a visual language model, which aims to solve or improve at least one of the above technical problems.
[0005] To achieve the above-mentioned purpose, the present application provides the following solutions:
[0006] A grape picking method based on a visual language model, comprising:
[0007] obtaining target image data; the target image data comprises an RGB image and a depth image;
[0008] constructing a visual language model; the visual language model adopts a VLM fine-tuning model;
[0009] inputting the target image data into the visual language model for recognition to determine whether there is a grape fruit in the current target visual area, and if there is a grape fruit, it is determined as a picking scene, and if there is no grape fruit, it is determined as a non-picking scene;
[0010] when in the picking scene, sequentially entering a set potential grape fruit exploration action state, a found grape fruit exploration action state and a picking action state to complete the picking of the current visual area;
[0011] when in the non-picking scene, determining that the picking task in the current visual area is completed;
[0012] when all the visual areas are detected, it is determined that the picking task is completely completed.
[0013] Optionally, the processing procedure when entering the potential grape fruit exploration action state comprises:
[0014] Randomly selecting a potential fruit coordinate, performing coordinate conversion according to the potential fruit coordinate, and controlling the mechanical arm to move to the potential fruit coordinate according to the converted [R T] matrix; the [R T] matrix is composed of a rotation matrix R and a translation matrix T, and is used to describe the orientation and distance of the mechanical arm end relative to the target;
[0015] If a grape fruit is found, the mechanical arm is controlled to stop moving and enter the "discovered grape fruit exploration action state"; if no grape fruit is found, it is determined that there is no grape fruit at the current potential fruit coordinate, and the step of "inputting the target image data into the visual language model for recognition" is returned.
[0016] Optionally, the processing procedure when entering the discovered grape fruit exploration action state comprises:
[0017] Determining a target grape fruit coordinate based on the current visual area image, and performing picking point detection according to the target grape fruit coordinate; when the picking point coordinate is determined, the mechanical arm is controlled to stop moving and enter the "picking action state".
[0018] Optionally, the processing procedure when entering the picking action state comprises:
[0019] Performing coordinate conversion according to the picking point coordinate, and controlling the mechanical arm to move to the picking point coordinate according to the converted [R T] matrix, performing fruit picking by the gripper, and controlling the mechanical arm to move to the storage point coordinate to release the gripper.
[0020] The application also provides a grape picking device based on a visual language model, which applies the method as described above, and comprises:
[0021] A tracked chassis, a device warehouse and a storage warehouse are installed on the tracked chassis, and a miniPC is installed in the device warehouse;
[0022] A mechanical arm is installed on the tracked chassis, and a gripper is installed at the end of the mechanical arm;
[0023] A fixing frame is installed on the mechanical arm, and a wide-angle camera is installed at the top of the fixing frame;
[0024] A depth camera is installed on the mechanical arm, and the depth camera is located above the gripper;
[0025] The depth camera, the wide-angle camera and the mechanical arm are connected with the miniPC.
[0026] Optionally, the mechanical arm is a six-axis mechanical arm, and the gripper is a four-gripper manipulator.
[0027] The depth camera is an RGB-D sensor, and the wide-angle camera is a fisheye camera.
[0028] Optionally, the depth camera, the wide-angle camera and the mechanical arm are connected to the miniPC through USB data lines respectively.
[0029] According to the specific embodiments of the present application, the following technical effects are provided.
[0030] The track chassis moves to the picking area, the mechanical arm adjusts the position of the gripper, so as to adjust to the appropriate picking angle and coordinate, the wide-angle camera is connected with the mechanical arm through the fixing frame, so that the wide-angle camera field of view is consistent with the direction of the mechanical arm, and the visual language model is used to solve the problem of reduced recognition and picking success rate caused by occlusion in unstructured environment, thereby improving the grape picking success rate. BRIEF DESCRIPTION OF DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0032] Figure 1 is the isometric view in the embodiment;
[0033] Figure 2 is the side view in the embodiment;
[0034] Figure 3 is the structural schematic view of the mechanical arm in the embodiment;
[0035] Figure 4 is the grape picking method flowchart in the embodiment;
[0036] Figure 5 is the model operation logic schematic view in the embodiment.
[0037] 1, track chassis; 2, equipment warehouse; 3, mechanical arm; 4, storage warehouse; 5, fixing frame; 6, gripper; 7, depth camera; 8, wide-angle camera. DETAILED DESCRIPTION
[0038] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those ordinarily skilled in the art without creative effort belong to the scope of the present application.
[0039] The present application aims to provide a grape picking method and device based on a visual language model, aiming to solve or improve at least one of the above technical problems.
[0040] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0041] As shown in Figure 1 The present application provides a grape picking method based on a visual language model, which is used to adapt to different grape picking scenes. By using a wide-angle camera 8 and a depth camera 7, the distribution of grape fruits in the field of view is determined using a visual language model, and a target detection and segmentation model is reasonably used to obtain the picking point position of the grape fruits. Then, the depth camera 7 moves to the vicinity of the picking point with the mechanical arm 3 to perform secondary identification to obtain accurate picking point coordinates. The movement of the mechanical arm 3 makes the gripper 6 find a suitable picking posture to complete the separation of the grape fruit from the stem, and the grape fruit is brought back to the storage point to complete a picking. The main steps include:
[0042] Obtaining target image data; the target image data includes an RGB image and a depth image;
[0043] Constructing a visual language model; the visual language model uses a VLM fine-tuning model.
[0044] Inputting the target image data into the visual language model for identification to determine whether there are grape fruits in the current target visual area. If there are grape fruits, it is determined as a picking scene, and if there are no grape fruits, it is determined as a non-picking scene.
[0045] When in the picking scene, sequentially enter the set potential grape fruit exploration action state, the discovered grape fruit exploration action state and the picking action state to complete the picking of the current visual area; when in the non-picking scene, determine that the picking task in the current visual area is completed.
[0046] When all the visual areas are detected, it is determined that the picking task is completely completed.
[0047] In addition, a grape picking device based on a visual language model is also provided for the above-mentioned scheme, which includes a tracked chassis 1, a mechanical arm 3, a fixed frame 5 and a depth camera 7.
[0048] Wherein, the tracked chassis 1 is installed with equipment warehouse 2 and storage warehouse 4, the miniPC is installed in the equipment warehouse 2; the mechanical arm 3 is installed on the tracked chassis 1, the gripper 6 is installed at the end of the mechanical arm 3; the fixed frame 5 is installed on the mechanical arm 3, the wide-angle camera 8 is installed at the top of the fixed frame 5; the depth camera 7 is installed on the mechanical arm 3, and the depth camera 7 is above the gripper 6; wherein, the depth camera 7, the wide-angle camera 8 and the mechanical arm 3 are connected with the miniPC.
[0049] As a specific embodiment, the mechanical arm 3 adopts a six-axis mechanical arm 3, the gripper 6 adopts a four-claw mechanical hand; the depth camera 7 adopts an RGB-D sensor, and the wide-angle camera 8 adopts a fisheye camera.
[0050] As a specific embodiment, the depth camera 7, the wide-angle camera 8 and the mechanical arm 3 are connected with the miniPC through USB data lines respectively.
[0051] Based on the above technical solution, the specific embodiments are provided as follows.
[0052] The tracked chassis 1 is fixedly connected with the equipment warehouse 2, the mechanical arm 3 and the storage warehouse 4, wherein the tracked chassis 1 serves as a mobile platform; the equipment warehouse 2 stores the miniPC, which is connected with the mechanical arm 3, the gripper 6, the depth camera 7 and the wide-angle camera 8 through USB data lines; the joint motor J1-J2 (the mechanical arm 3 has six joint motors from J1 to J6 from bottom to top) connecting piece of the mechanical arm 3 is fixedly connected with the fixed frame 5, the wide-angle camera 8 is fixed on the upper end of the fixed frame 5, the joint motor J5-J6 connecting piece of the mechanical arm 3 is fixedly connected with the depth camera 7, the end of the joint motor J6 of the mechanical arm 3 is fixedly connected with the gripper 6, so as to realize the effect that the J1 rotation of the mechanical arm 3 drives the rotation of the fixed frame 5 and the wide-angle camera 8, and the end movement of the mechanical arm 3 drives the movement of the gripper 6 and the depth camera 7, so that the wide-angle camera 8 can always obtain the visual field of the whole mechanical arm 3, and the depth camera 7 can always obtain the visual field of the end of the mechanical arm 3; the storage warehouse 4 is the grape fruit storage place.
[0053] The tracked chassis 1 moves to a specified position to start the grape picking in this area; the mechanical arm 3 starts initialization calibration so that the end reaches a specified initial position; the depth camera 7 and the wide-angle camera 8 transmit the two RGB images and one depth image obtained to the miniPC in the equipment warehouse 2 through USB data lines, and the miniPC connects the cloud to call a visual language model to perform a scene judgment task.
[0054] When the visual language model judges the scene as a non-picking scene, i.e. there is no any pickable grape fruit in the field of view, the picking task in this area is completed, and the tracked chassis 1 moves into the next area.
[0055] When the visual language model judges the scene as a picking scene, it is judged that there is a grape fruit visible in the picking point in the field of view, and enters the picking action state; the miniPC in the equipment compartment 2 obtains the RGB image and depth image of the depth camera 7 through the USB data line, runs the picking point detection model, and obtains the coordinates of the picking point relative to the depth camera 7; the miniPC in the equipment compartment 2 reads the joint motor data of the mechanical arm 3 through the USB data line, obtains the coordinate relationship between the depth camera 7, the gripper 6 and the base of the mechanical arm 3, and obtains the coordinates of the picking point relative to the base of the mechanical arm 3 through coordinate transformation; the miniPC in the equipment compartment 2 sends the picking point coordinates to the mechanical arm 3 through the USB data line, controls the mechanical arm 3 to move so that the gripper 6 at the end reaches the picking point; the miniPC in the equipment compartment 2 sends instructions to control the gripper 6 to close to realize the separation and grabbing of the grape fruit at the picking point; the miniPC in the equipment compartment 2 sends instructions to control the mechanical arm 3 to move so that the gripper 6 at the end reaches the top of the storage compartment 4, places the grape fruit in the storage compartment 4, and returns to the initial position, completing a picking task.
[0056] When the visual language model judges the scene as a picking scene, it is judged that there is no grape fruit visible in the picking point in the field of view, but there is a grape fruit occluded by the picking point, and enters the autonomous exploration action state; the miniPC in the equipment compartment 2 obtains the RGB image and depth image of the depth camera 7 through the USB data line, runs the grape fruit detection model, and obtains the approximate coordinates of the grape fruit; the miniPC in the equipment compartment 2 sends instructions to control the mechanical arm 3 through the USB data line, and the end moves around the approximate coordinates of the grape fruit to find a pose that can recognize the picking point; when the picking point is recognized, the picking action state is entered, and a picking task is completed.
[0057] When the visual language model judges the scene as a picking scene, it is judged that there is no grape fruit visible in the picking point in the field of view and no grape fruit occluded by the picking point, and enters the potential grape fruit exploration action state, and the visual language model infers the coordinates of the potential grape fruit; the miniPC in the equipment compartment 2 sends instructions to control the mechanical arm 3 through the USB data line, and the end goes to the coordinates of the potential grape fruit to find the grape fruit; when the exploration finds that there is a grape fruit near the coordinate point, the autonomous exploration action state and the picking action state are entered in turn, and a picking task is completed; when no grape fruit is found after multiple exploration actions, it is judged that there is no any pickable grape fruit in this area, the picking task in this area is completed, and the tracked chassis 1 moves into the next area.
[0058] In this embodiment, the wide-angle camera 8 is a fisheye camera, the depth camera 7 is an RGB-D sensor, the gripper 6 is a four-jaw robot hand (the upper two jaws are integrated with shears, and the lower two jaws only hold), the mechanical arm 3 is a six-axis robot arm, the miniPC in the device compartment 2 uses the detection model as a neural network target detection model, and is connected to the visual language model used in the cloud as a VLM fine-tuning model.
[0059] In this embodiment, the visual language model-based grape picking state and grape fruit recognition are saved in the form of dialogue pairs. The specific examples are divided into two types:
[0060] One is to judge the state of the grape fruit and the action state of the entry, which uses the image of the wide-angle camera 8, and the input dialogue pair of the visual language model is:
[0061] {"messages": [{"role": "user", "context": ["text": "Identify the image and determine whether there is a grape fruit in the image. Return the required action state and the approximate coordinates of the grape fruit according to the specified format", "image": "Wide-angle camera image path"]}]}.
[0062] The return dialogue pair is:
[0063] {"messages": [{"role": "assistant", "context": ["status": "picking action state / autonomous exploration action state / potential grape fruit exploration action state", "2D_coordinates": [[xmin, xmax, ymin, ymax], [xmin, xmax, ymin, ymax],...]]]}.
[0064] Where xmin, xmax, ymin, ymax are the maximum and minimum pixel coordinate points of the fruit area in the image of the wide-angle camera 8.
[0065] The other is to judge the state of the grape fruit and the machine pixel coordinate point, which uses the image of the depth camera 7, and the input dialogue pair of the visual language model is:
[0066] {"messages": [{"role": "user", "context": ["text": "Identify the grape fruit in the image and return the required grape fruit state and approximate coordinates of the grape fruit according to the specified format", "image": "Depth camera image path"]}]}.
[0067] The return dialogue pair is:
[0068] {"messages":[{"role":"assistant","context":["label":"visible / occluded / potential","2D_coordinates":[[xmin,xmax,ymin,ymax],[xmin,xmax,ymin,ymax],...]]]}].
[0069] Wherein, xmin, xmax, ymin, ymax are the maximum and minimum pixel coordinate points of each fruit identified in the image of the depth camera 7, and [xmin, xmax, ymin, ymax] is taken as the identification range roi of the target detection model to obtain more accurate [xmin, xmax, ymin, ymax].
[0070] A plurality of pixel coordinates can be obtained through [xmin, xmax, ymin, ymax], and a [RT] matrix can be obtained through coordinate transformation, wherein R corresponds to a rotation matrix and T corresponds to a translation matrix, which describes the orientation and distance of the end of the mechanical arm 3 relative to the grape fruit, and can be directly used for control of the mechanical arm 3.
[0071] Therefore, the present scheme has the following advantages: (1) using RGB-D to identify and locate the grape fruit picking point, (2) using a six-axis mechanical arm 3 to autonomously plan the motion to find a suitable picking angle and coordinate, (3) using a wide-angle camera 8 connected to the mechanical arm 3 through a fixed frame 5 structure to make the wide-angle camera 8 field of view consistent with the orientation of the mechanical arm 3, (3) using a visual language model to judge the picking scene, adapting to grape picking under unstructured scene occlusion, (4) using a cloud edge scheme, deploying a lightweight model (target detection model) on the edge (miniPC in the device warehouse 2) and device end (mechanical arm 3 and part of the gripper 6), and deploying a visual language model on the cloud end, which not only relieves the running power pressure of the edge end, but also can timely adjust and improve the cloud end model.
[0072] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be mutually referred to.
[0073] The principles and implementation modes of the present application are described by applying specific examples in this paper. The above description of the embodiments is only used to help understand the core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In view of the above, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A grape picking method based on a visual language model, characterized in that: include: Obtain target image data; The target image data includes an RGB image and a depth image; Build a visual language model; The visual language model adopts the VLM fine-tuning model; Inputting the target image data into the visual language model for recognition, determining whether there are grape fruits in the current target visual area, and determining that it is a picking scene if there are grape fruits, and determining that it is a non-picking scene if there are no grape fruits; When in the picking scene, it enters the set potential grape fruit exploration action state, discovered grape fruit exploration action state and picking action state in sequence to complete the picking of the current visual area; When in a non-picking scene, it is determined that the picking task in the current visual area is completed; When detection is completed in all visual areas, the picking task is determined to be completely completed.
2. The grape picking method based on the visual language model according to claim 1, characterized in that: The processing process when entering the potential grape fruit exploration action state includes: Randomly selecting potential fruit coordinates, performing coordinate transformation according to the potential fruit coordinates, and controlling the robotic arm (3) to move to the potential fruit coordinates according to the transformed [RT] matrix; the [RT] matrix is composed of a rotation matrix R and a translation matrix T, and is used to describe the orientation and distance of the end of the robotic arm (3) relative to the target; If grape fruits are found, the robot arm (3) is controlled to stop moving and enter the "grape fruit discovery exploration action state"; if no grape fruits are found, it is determined that there are no grape fruits at the current potential fruit coordinates, and the process returns to the step of "inputting the target image data into the visual language model for recognition".
3. The grape picking method based on the visual language model according to claim 2 is characterized in that: The processing process when entering the grape fruit discovery exploration action state includes: The target grape fruit coordinates are determined based on the current visual area image, and the picking point detection is performed based on the target grape fruit coordinates. After the picking point coordinates are determined, the robot arm (3) is controlled to stop moving and enter the "picking action state".
4. The grape picking method based on the visual language model according to claim 3 is characterized in that: The processing process when entering the picking action state includes: Coordinate conversion is performed according to the coordinates of the picking point, and the robot arm (3) is controlled to move to the coordinates of the picking point according to the converted [RT] matrix, the fruit is picked by the gripper (6), and the robot arm (3) is controlled to move to the storage point coordinates to release the gripper (6).
5. A grape picking device based on a visual language model, applying the method according to any one of claims 1 to 4, characterized in that: include: A crawler chassis (1), wherein an equipment compartment (2) and a storage compartment (4) are installed on the crawler chassis (1), and a mini PC is installed in the equipment compartment (2); A mechanical arm (3), the mechanical arm (3) being mounted on the crawler chassis (1), and a clamping claw (6) being mounted at the end of the mechanical arm (3); A fixed frame (5), the fixed frame (5) is mounted on the robotic arm (3), and a wide-angle camera (8) is mounted on the top of the fixed frame (5); a depth camera (7), the depth camera (7) being mounted on the robotic arm (3), and the depth camera (7) being located above the gripper (6); Wherein, the depth camera (7), the wide-angle camera (8), and the robotic arm (3) are all connected to the miniPC.
6. The grape picking device based on the visual language model according to claim 5, characterized in that: The robotic arm (3) is a six-axis robotic arm (3), and the gripper (6) is a four-claw robotic arm; The depth camera (7) adopts an RGB-D sensor, and the wide-angle camera (8) adopts a fisheye camera.
7. The grape picking device based on the visual language model according to claim 5, characterized in that: The depth camera (7), the wide-angle camera (8), and the robotic arm (3) are respectively connected to the miniPC via USB data cables.
Citation Information
Cited By
Picking mechanical arm operation method and system based on visual feedback
CN121018601A