Tomato picking device and method based on visual and tactile combination
By combining visual and tactile methods, a depth camera and a tomato stem tactile sensor are used to accurately locate the tomato picking point, solving the problem of inaccurate picking in existing technologies. This achieves efficient and precise tomato picking and reduces the need for manual pruning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU UNIV
- Filing Date
- 2022-11-17
- Publication Date
- 2026-08-04
AI Technical Summary
Existing tomato harvesting robots cannot accurately locate the harvesting point, resulting in the need for manual trimming of the harvested tomatoes. Furthermore, the visual recognition technology suffers from occlusion issues, leading to low efficiency.
Using a combination of vision and touch, a depth camera is used to identify tomato fruits and calculate their spatial coordinates and radius. A 3D reconstruction is then performed using a tomato stalk touch sensor to determine the optimal picking point and drive an end effector to pick the fruit.
This method achieves precision and efficiency in tomato harvesting, reduces manual pruning steps, lowers the risk of damage to tomatoes, and ensures the integrity and aesthetics of the harvested tomatoes.
Smart Images

Figure CN115713761B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of tomato harvesting robots, and particularly relates to a precise positioning tomato harvesting device and method based on a combination of vision and touch. Background Technology
[0002] Currently, the research and technology level of agricultural harvesting robots in my country is not high. Traditional tomato harvesting techniques are costly, complex to operate, and inefficient. Many existing harvesting robots use visual recognition for positioning and harvesting, which not only fails to solve the problem of tomatoes being obstructed during recognition, but also cannot accurately locate the tomato harvesting point. Due to the excessively long stems, the harvested tomatoes still require manual secondary trimming after harvesting. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides a precise tomato harvesting device and method based on a combination of visual and tactile sensing. Compared to traditional harvesting robots, this device uses a tomato stem tactile sensor for further tactile recognition on top of visual recognition, precisely locating the tomato stem for harvesting. It is simple to operate, offers high harvesting accuracy, saves manpower, and ensures the integrity of the tomato fruit.
[0004] This invention uses a controller to process color images of tomatoes captured by a depth camera, identifies the tomato fruit, and calculates the spatial coordinates of the tomato's center and its diameter. A tactile sensor on the tomato stem first analyzes the tomato's firmness to determine its maturity, then performs a three-dimensional reconstruction of the stem. The controller identifies the radius of the tomato stem based on the obtained three-dimensional image and performs threshold analysis to determine the optimal harvesting point at the point of maximum radius. This drives the end effector to harvest the fruit, achieving efficient harvesting of tomatoes and other fruits. The invention features a simple and compact structure, stable operation, and high harvesting efficiency.
[0005] The present invention achieves the above-mentioned technical objectives through the following technical means.
[0006] A tomato harvesting device based on vision-touch integration for precise positioning includes a tomato harvesting robot, a depth camera, a tomato stem tactile sensor, and a controller.
[0007] The depth camera is mounted on the tomato harvesting robot to collect color images and depth information of the tomatoes in front of the end effector and transmit them to the controller.
[0008] The tomato stem tactile sensor is installed on the end effector of the tomato harvesting robot to acquire 4D light field data images and transmit them to the controller;
[0009] The controller is connected to the tomato harvesting robot, the depth camera, and the tomato stem tactile sensor, respectively.
[0010] The controller identifies tomato fruits based on color images of tomatoes, calculates the coordinates and radius of the tomatoes, determines whether the tomatoes are ripe based on 4D light field data images, reconstructs a three-dimensional image of the tomato stem, determines the optimal picking point, and drives the end effector to pick the tomatoes.
[0011] In the above scheme, the controller performs image processing on the tomato color images captured by the depth camera, identifies the tomato fruit, calculates the spatial coordinates of the tomato center and the radius of the tomato fruit, determines whether the tomato is ripe based on the obtained 4D light field data image and reconstructs the three-dimensional image of the tomato stem, and identifies the radius of the tomato fruit stem based on the three-dimensional image of the tomato stem, performs threshold analysis on the radius of the tomato fruit stem, and determines the specific location of the stem with the largest radius as the optimal picking point.
[0012] In the above scheme, the tomato stem tactile sensor includes a transparent shell, a light source, a speckle spraying device, a light field camera, and fingertip gel;
[0013] The end effector has transparent housings on both sides, and a light source and a light field camera are installed inside the transparent housings. The side of the transparent housing that is in contact with the tomato is covered with fingertip gel. The speckle spraying device is located on the transparent housing and around the fingertip gel, and is used to spray speckles onto the fingertip gel.
[0014] The light source is used to provide illumination to the light field camera, which is used to capture 4D light field data images of the fingertip gel and transmit them to the controller.
[0015] Furthermore, the controller includes a tomato recognition module, a tomato positioning and detection module, a tomato surface stress measurement module, a tomato stem three-dimensional reconstruction module, and a picking point determination module;
[0016] The tomato recognition module is used to process the color images of tomatoes captured by the depth camera and identify the tomato fruits.
[0017] The tomato positioning and detection module is used to calculate the spatial coordinates of the tomato center and the radius of the tomato fruit by combining the color image after image processing with the depth information collected by the depth camera, and to obtain the position of the tomato on the coordinates of the tomato picking robot.
[0018] The tomato surface stress measurement module is used to control the speckle spraying device of the tomato stem tactile sensor to spray a uniform black speckle onto the fingertip gel and use a light field camera to capture 4D light field data of the fingertip gel before deformation. After the end effector applies a preset force to the target tomato, it uses a light field camera to capture 4D light field data of the deformed fingertip gel. Based on the acquired 4D light field data of the deformed fingertip gel, it synthesizes a low-resolution, large-depth-of-field two-dimensional RGB image and a depth image of the speckle on the fingertip gel, calculates the deformation value, calculates the strain field of the entire tomato, and compares it with a threshold. The ripeness of the tomato fruit is determined by the hardness.
[0019] The tomato stem 3D reconstruction module is used when the controller determines that the target tomato is a ripe tomato. Then, based on the position information of the tomato on the coordinates of the tomato picking robot, the controller moves the end effector upward by a preset distance to grab the stem of the tomato. The light field camera of the tomato stem tactile sensor acquires the 4D light field data image of the tomato stem and transmits it to the controller. The controller reconstructs the 3D image of the tomato stem through the 4D light field data.
[0020] The picking point determination module is used to measure the radius of the tomato stem based on the three-dimensional image of the tomato stem, and to perform threshold analysis on the radius of the tomato stem to determine the specific location of the stem with the largest radius as the optimal picking point, thereby driving the end effector to pick the fruit.
[0021] In the above scheme, the speckle spraying device includes a built-in speckle spraying liquid storage chamber, the outlet of the chamber is provided with a spray pipe, the spray pipe is provided with a valve, and the valve is connected to a controller.
[0022] A control method for a precision positioning tomato harvesting device based on visual-touch integration includes the following steps:
[0023] Tomato recognition: The depth camera captures a color image of a tomato in front of the end effector and transmits it to the controller. The controller processes the color image of the tomato captured by the depth camera to identify the tomato fruit.
[0024] Tomato positioning and detection: The controller combines the processed color image with the depth information collected by the depth camera to calculate the spatial three-dimensional coordinates of the tomato center and the radius of the tomato fruit, thereby obtaining the position of the tomato on the coordinates of the tomato picking robot.
[0025] Tomato ripeness identification: The controller controls the speckled spraying device of the tomato stem tactile sensor to spray a uniform black speckled gel onto the fingertip and uses a light field camera to capture 4D light field data of the fingertip gel before deformation. After the end effector applies a preset force to the target tomato, it uses a light field camera to capture 4D light field data of the deformed fingertip gel. Based on the acquired 4D light field data of the deformed fingertip gel, a low-resolution, large-depth-of-field two-dimensional RGB image and a depth image of the speckled gel are synthesized. The deformation value is calculated, and then the strain field of the entire tomato is calculated and compared with a threshold. The ripeness of the tomato fruit is determined by the hardness.
[0026] 3D reconstruction of tomato stem: When the controller determines that the target tomato is a ripe tomato, it moves upward a preset distance according to the position information of the tomato on the coordinates of the tomato picking robot to control the end effector to grab the stem of the tomato. The light field camera of the tomato stem tactile sensor acquires the 4D light field data image of the tomato stem and transmits it to the controller. The controller reconstructs the 3D image of the tomato stem through the 4D light field data.
[0027] Picking point determination: The controller measures the radius of the tomato stem based on the three-dimensional image of the tomato stem, and performs threshold analysis on the radius of the tomato stem to determine the specific location of the stem with the largest radius as the optimal picking point, and drives the end effector to pick the fruit.
[0028] In the above scheme, the image processing in the tomato localization and detection step includes the following steps:
[0029] The color image covered by the Gaussian filter is smoothed, the RGB color space is converted to the HSV color space, and the red area is extracted. The red area is filled, reduced or enlarged to eliminate false alarms in circle detection. The Canny method is used for contour detection, and the circular Hough transform on the contour image is used to detect the circle in the red area image. The proportion of red pixels in the circumscribed square of the detected circle is used to determine the detection area. The image of the recognized circle is drawn on the color image, and the two-dimensional center coordinates and radius information of the recognized circle are output.
[0030] In the above scheme, the tomato location detection step is specifically as follows:
[0031] The controller obtains the depth at the center coordinates of the identification circle based on the depth information acquired by the depth camera. By mapping the viewpoint of the RGB sensor to the viewpoint of the depth sensor, it calculates a portion of the coordinates from the depth information. The depth and spatial coordinates are then displaced in the coordinate system, with the displacement equal to the radius from the center of the identification circle to the center of the circle. The actual radius is calculated based on the spatial coordinates of the tomato's center and the spatial coordinates where the displacement equals the radius. The controller then determines whether the radius of the tomato is within a preset value range. If it is within the preset value range, the controller determines that the detected circle is a tomato, draws the circle in the color image, and outputs the three-dimensional spatial coordinates of the tomato's center.
[0032] In the above scheme, the tomato maturity identification step uses the IC-GN subpixel search algorithm to calculate the three-dimensional displacement, i.e., the deformation value, and uses point-by-point local least squares fitting to calculate the strain field.
[0033] In the above scheme, in the three-dimensional reconstruction step, the controller uses a phase-shift-based sub-pixel multi-view stereo matching algorithm to estimate the depth of the light field through 4D light field data in order to reconstruct the three-dimensional image of the tomato stem.
[0034] Compared with existing technologies, the beneficial effects of this invention are as follows: Based on visual recognition, this invention employs tactile positioning technology. The controller processes color and depth images of tomatoes captured by a depth camera to identify the tomato fruit and calculate the three-dimensional coordinates of the tomato's center and the tomato's radius. The controller determines the tomato's ripeness based on the obtained 4D light field data image. Unlike visual-based ripeness color judgment, this method uses speckle pattern analysis to determine the tomato's firmness, resulting in a more accurate ripening result. The controller identifies the radius of the tomato stem based on the obtained three-dimensional image of the tomato stem and performs threshold analysis on the stem radius to determine the optimal picking point. This drives the end effector to pick the tomato, achieving stem-based harvesting, reducing the risk of damage during harvesting, ensuring the aesthetic appeal of the harvested tomatoes, solving the problem of tomato occlusion during visual recognition positioning, and reducing the need for manual pruning of harvested tomatoes with excessively long stems. The invention features a simple and compact structure, easy operation, stable operation, and precise and efficient harvesting.
[0035] Note that the description of these effects does not preclude the existence of other effects. One aspect of the invention does not necessarily have all the aforementioned effects. Effects other than those described above can be readily observed and extracted from the description, drawings, claims, etc. Attached Figure Description
[0036] Figure 1 This is a three-dimensional structural diagram of a device according to an embodiment of the present invention.
[0037] Figure 2 This is a detailed structural diagram of the end effector of an embodiment of the present invention when it is open.
[0038] Figure 3 This is a flowchart of tomato fruit detection image processing according to one embodiment of the present invention.
[0039] Figure 4 This is a flowchart of tomato fruit positioning and detection according to one embodiment of the present invention.
[0040] Figure 5This is a flowchart of a tomato surface stress measurement system according to an embodiment of the present invention.
[0041] Figure 6 This is a detailed schematic diagram of the tomato pedicel stem according to an embodiment of the present invention.
[0042] In the diagram: 1-upper arm, 2-middle arm, 3-forearm, 4-depth camera, 5-tomato stalk tactile sensor, 6-fruit bag mounting rack, 7-lighting system, 8-spot spraying device, 9-light field camera, 10-finger gel, 11-spring blade, 12-tomato stalk stem; 13-transparent shell. Detailed Implementation
[0043] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0044] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "front," "rear," "left," "right," "upper," "lower," "axial," "radial," "vertical," "horizontal," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0045] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0046] Figure 1The image shows a preferred embodiment of the vision-touch-based precision positioning tomato harvesting device, which includes a tomato harvesting robot, a depth camera 4, a tomato stem tactile sensor 5, and a controller.
[0047] The depth camera 4 is mounted on the tomato harvesting robot to capture color images of the tomatoes in front of the end effector and transmit them to the controller;
[0048] The depth camera 4 is installed on the tomato harvesting robot to collect color images and depth information of the tomatoes in front of the end effector and transmit them to the controller;
[0049] The tomato stem tactile sensor 5 is installed on the end effector of the tomato harvesting robot to acquire 4D light field data images and transmit them to the controller;
[0050] The controller is connected to the tomato picking robot, the depth camera 4, and the tomato stem tactile sensor 5, respectively.
[0051] The controller identifies tomato fruits based on color images of tomatoes, calculates the coordinates and radius of the tomatoes, determines whether the tomatoes are ripe based on 4D light field data images, reconstructs a three-dimensional image of the tomato stem, determines the optimal picking point, and drives the end effector to pick the tomatoes.
[0052] According to this embodiment, preferably, the controller performs image processing on the tomato color image acquired by the depth camera 4, identifies the tomato fruit, calculates the spatial coordinates of the tomato center and the radius of the tomato fruit, determines whether the tomato is ripe based on the obtained 4D light field data image and reconstructs the three-dimensional image of the tomato stem, identifies the radius of the tomato fruit stem based on the three-dimensional image of the tomato stem, performs threshold analysis on the radius of the tomato fruit stem, and determines the specific location of the stem at the point with the largest radius as the optimal picking point.
[0053] According to this embodiment, preferably, the tomato picking robot includes a base, a movable robotic arm, an end effector, and a fruit bag mounting frame 6. The rotating joint of the movable robotic arm is rotatably mounted on the base, the rotating joint of the end effector is mounted on the end of the robotic arm, and the fruit bag mounting frame 6 is mounted directly below the end effector.
[0054] According to this embodiment, preferably, the robotic arm includes a large arm 1, a middle arm 2, and a small arm 3 connected in sequence. The base connected to the robotic arm contains a first drive mechanism for controlling the rotational movement of the large arm 1. Second drive mechanisms are respectively built into the joints connecting the large arm 1 and the middle arm 2, and the middle arm 2 and the small arm 3, to control the rotation of the relative rotational joints. A third drive mechanism is built into the base between the small arm 3 and the end effector to control the rotational movement of the end effector.
[0055] According to this embodiment, the preferred specifications of the depth camera 4 are as follows: RGB frame resolution of 1920*1080, depth output resolution of 1024*768, ideal depth range of 0.25m-9m, and indoor use environment.
[0056] According to this embodiment, preferably, the tomato stem tactile sensor 5 includes a transparent housing 13, a light source 7, a speckle spraying device 8, a light field camera 9, and a fingertip gel 10; the end effector is provided with transparent housings 13 on both sides, the transparent housing 13 is provided with the light source 7 and the light field camera 9, and the side of the transparent housing 13 that is in contact with the tomato is provided with the fingertip gel 10; the speckle spraying device 8 is provided on the transparent housing 13 and located around the fingertip gel 10, and is used to spray speckles onto the fingertip gel 10; the light source 7 is used to provide illumination to the light field camera 9, and the light field camera 9 is used to capture 4D light field data images of the fingertip gel 10 and transmit them to the controller.
[0057] The tomato tactile sensor 5 grasps the tomato in two stages. The first grasp uses the tomato surface stress measurement module of the controller to determine the tomato's firmness. This method, unlike visual-based ripeness judgment based on color, reduces errors associated with color-based judgment, resulting in more accurate results. The second grasp uses the tomato stem, where a light field camera 9 performs 3D reconstruction and precisely locates the stem stalk 12, enabling accurate tomato harvesting. The light field camera 9 allows for pre-capturing and post-focusing, recording data from beams of light in all directions and "focusing" on any depth within the image. The controller then automatically refocuses, resulting in a clearer image. Furthermore, the light field camera 9 only requires a single image capture to obtain 3D images and 3D depth information maps with minimal lighting effects, allowing for the creation of a colorful 3D stereoscopic shape on the controller.
[0058] According to this embodiment, preferably, the controller includes at least a tomato recognition module, a tomato positioning detection module, a tomato surface stress measurement module, a tomato stem three-dimensional reconstruction module, and a picking point determination module;
[0059] The tomato recognition module is used to process the color images of tomatoes captured by the depth camera 4 and identify the tomato fruits.
[0060] The tomato positioning and detection module is used to calculate the spatial coordinates of the tomato center and the radius of the tomato fruit by combining the color image after image processing with the depth information collected by the depth camera 4, and to obtain the position of the tomato on the coordinates of the tomato picking robot.
[0061] The tomato surface stress measurement module is used to control the speckle spraying device 8 of the tomato stem tactile sensor 5 to spray uniform black speckles onto the fingertip gel 10 and use the light field camera 9 to capture the 4D light field data of the fingertip gel 10 before deformation. After the end effector applies a preset force to the target tomato, the light field camera 9 captures the 4D light field data of the deformed fingertip gel 10. Based on the acquired 4D light field data of the deformed fingertip gel 10, a two-dimensional RGB image and a depth image of speckle with low resolution and large depth of field on the fingertip gel 10 are synthesized. The deformation value is calculated, and then the strain field of the entire tomato is calculated and compared with the threshold. The ripeness of the tomato fruit is judged by the hardness.
[0062] The tomato stem 3D reconstruction module is used to control the end effector to grab the tomato stem when the controller determines that the target tomato is a ripe tomato and moves it upward a preset distance according to the position information of the tomato on the coordinates of the tomato picking robot. The light field camera 9 of the tomato stem tactile sensor 5 acquires the 4D light field data image of the tomato stem and transmits it to the controller. The controller reconstructs the 3D image of the tomato stem through the 4D light field data.
[0063] The picking point determination module is used to measure the radius of the tomato stem based on the three-dimensional image of the tomato stem, and to perform threshold analysis on the radius of the tomato stem to determine the specific location of the stem with the largest radius as the optimal picking point, thereby driving the end effector to pick the fruit.
[0064] A control method for a precision positioning tomato harvesting device based on visual-touch integration includes the following steps:
[0065] Tomato recognition: Depth camera 4 captures color images of tomatoes in front of the end effector and transmits them to the controller. The controller processes the color images of tomatoes captured by depth camera 4 to identify the tomato fruit.
[0066] Tomato positioning and detection: The controller combines the processed color image with the depth information collected by the depth camera 4 to calculate the spatial three-dimensional coordinates of the tomato center and the radius of the tomato fruit, thereby obtaining the position of the tomato on the coordinates of the tomato picking robot.
[0067] Tomato ripeness identification: The controller controls the speckled spraying device 8 of the tomato stem tactile sensor 5 to spray a uniform black speckled pattern onto the fingertip gel 10 and uses a light field camera 9 to capture the 4D light field data of the fingertip gel 10 before deformation. After the end effector applies a preset force to the target tomato, the light field camera 9 captures the 4D light field data of the deformed fingertip gel 10. Based on the acquired 4D light field data of the deformed fingertip gel 10, a low-resolution two-dimensional RGB image and a depth image of the speckled pattern on the fingertip gel 10 are synthesized. The deformation value is calculated, and then the strain field of the entire tomato is calculated and compared with the threshold. The ripeness of the tomato fruit is determined by the hardness.
[0068] 3D reconstruction of tomato stem: When the controller determines that the target tomato is a ripe tomato, it moves the end effector upward by a preset distance according to the position information of the tomato on the coordinates of the tomato picking robot to grab the stem of the tomato. The light field camera 9 of the tomato stem tactile sensor 5 acquires the 4D light field data image of the tomato stem and transmits it to the controller. The controller reconstructs the 3D image of the tomato stem through the 4D light field data.
[0069] Picking point determination: The controller measures the radius of the tomato stem based on the three-dimensional image of the tomato stem, and performs threshold analysis on the radius of the tomato stem to determine the specific location of the stem with the largest radius as the optimal picking point, and drives the end effector to pick the fruit.
[0070] Combination Figure 3 As shown, according to this embodiment, preferably, in the image processing of the tomato positioning and detection step: the depth camera 4 acquires a color image of the tomato in front of the end effector and transmits it to the controller; a Gaussian filter is used to smooth the covered color image, the RGB color space is converted to the HSV color space, and the red area is extracted; the red area is filled, eroded, and expanded to eliminate false alarms in circle detection; the Canny method is used for contour detection, and the circular Hough transform on the contour image is used to detect the circle in the red area image; the proportion of red pixels in the circumscribed square of the detected circle is used to determine the detection area; an image of the identified circle is drawn on the color image, and the center coordinates and radius information of the identified circle are output. Specifically, the steps include the following:
[0071] 1. To reduce noise and extract the red area, a Gaussian filter is used to smooth the covered color image, thus achieving noise reduction;
[0072] 2. Convert the RGB color space to the HSV color space and extract the red area;
[0073] 3. To eliminate false alarms in circle detection, fill the red area, erode the red area, and expand the red area;
[0074] 4. To detect circles in the red area image, the Canny method is used to extract the contours, and the circle Hough transform on the contour image is used to detect the circles.
[0075] 5. Define an outer bounding box and calculate the red pixels within the square. To remove false alarms, the proportion of red pixels in the outer square of the detection circle is used to determine the detection area. The proportion of red pixels is calculated and compared with the red pixel proportion threshold. Preferably, in this embodiment, the area with a red pixel proportion greater than or equal to 50% is determined as the tomato area.
[0076] 6. Draw circles within the threshold, draw the recognition circle image on the color image, output the center point coordinates and radius information of the recognition circle, and then output the recognition image. For those outside the threshold, directly output the recognition image.
[0077] Combination Figure 4 According to this embodiment, preferably, the tomato positioning detection step specifically includes:
[0078] The controller obtains the depth at the center coordinates of the identified circle based on the depth information acquired by the depth camera 4. By mapping the viewpoint of the RGB sensor to the viewpoint of the depth sensor, partial coordinates are calculated from the depth information. The depth and spatial coordinates are then displaced in the coordinate system, with the displacement equal to the radius from the center of the identified circle to that center. The actual radius is calculated based on the spatial coordinates of the tomato's center and the spatial coordinates where the displacement equals the radius. It is then determined whether the tomato's radius is within a preset value range. If it is within the preset value range, the detected circle is determined to be a tomato. A circle is drawn in the color image, and the three-dimensional spatial coordinates of the tomato's center are output. The specific steps include the following:
[0079] The tomato detection uses the identified circle and depth information acquired by depth camera 4 to calculate the spatial coordinates of the tomato's center, thereby obtaining the tomato's position in the robot's coordinate system. The specific process is as follows:
[0080] 1. The controller obtains the depth at the center coordinates of the recognition circle based on the depth information acquired by the depth camera 4. By mapping the view of the RGB sensor to the view of the depth sensor, the controller calculates the coordinates of the fused key points from the depth information.
[0081] 2. In the same way, the depth and spatial coordinates will be displaced on the coordinate system. The amount of displacement is equal to the radius from the center of the identification circle to the center of the circle. The actual diameter is calculated based on the spatial coordinates of the tomato center and the spatial coordinates where the displacement is equal to the radius.
[0082] 3. In order to remove the detected false positive tomatoes, it is determined whether the diameter of the tomato is within a preset range. Preferably, the preset range is 30-100 mm.
[0083] 4. If the diameter of the tomato is within the preset range, the detected circle is determined to be a tomato, a red circle is drawn in the color image, and the spatial coordinates of the tomato's center are output; if the diameter of the tomato is outside the preset range, it is determined to be a false positive tomato, and a white circle is drawn in the color image.
[0084] According to this embodiment, preferably, the tomato maturity identification step uses the IC-GN subpixel search algorithm to calculate the accurate three-dimensional displacement, i.e., the deformation value, and employs point-by-point local least squares fitting to calculate the strain field, specifically including the following steps:
[0085] 1. Acquire 4D light field data using a light field camera 9. Spray uniform black speckles onto the fingertip gel 10 using a speckle spraying device 8. Take a camera image of the fingertip gel 10 before deformation using the light field camera 9 at the bottom of the tomato's tactile sensor. After the end effector applies a preset force to the tomato (preferably 5N), record the deformed 4D light field data using the light field camera 9. Synthesize a low-resolution, large-depth-of-field 2D RGB image and a depth image from the camera based on the acquired deformed 4D light field data. Preferably, the speckles are high-contrast, non-repeating, and of appropriate size. The speckle size can be adjusted to suit the resolution of the light field camera 9 without reducing accuracy.
[0086] 2. Use the IC-GN subpixel search algorithm to obtain more accurate displacement values.
[0087] 3. Calculate the strain field using point-by-point local least squares fitting.
[0088] 4. Threshold judgment of tomato fruit firmness to determine whether the tomato is ripe.
[0089] The tomato stem tactile sensor 5 is small in size and can be easily attached to the end effector of the harvesting robot to obtain a high-resolution tactile image of the 3D morphology of the contact surface for accurate determination of the harvesting point.
[0090] Combination Figure 5 As shown, the data acquisition for tomato surface stress measurement is specifically as follows:
[0091] Select a reference point P0(x0,y0) on the image before deformation, with a gray value of f(x0,y0). Select a reference point P1(x1,y1) on the image after deformation, with a gray value of g(x1,y1).
[0092] In the undistorted image, a square region of size (2M+1, 2M+1) centered at P0(x0, y0) is used as the reference sub-region. In the distorted image, there must be a distorted region centered at P1(x1, y1), and the correlation between these two regions must be the highest. Then, P0 and P1 are considered corresponding matching points. The correlation function describes the magnitude of the correlation between the undistorted and distorted regions, and its mathematical formula is:
[0093]
[0094] Where: f—gray value of the reference image sub-region before deformation of the model under test; g—gray value of the target image sub-region after deformation of the model under test; These represent the average grayscale values of the pixels in the sub-region before and after deformation.
[0095] Using a first-order shape function model, there is a certain functional relationship between a point Q(x,y) in the reference sub-region and the corresponding point Q1(x1,y1) in the deformed sub-region. The shape function is:
[0096]
[0097] In the formula: Δx and Δy are the distances from point (x,y) to P0(x0,y0); u and v are the displacements of the center point of the reference sub-region in the x and y directions, respectively; u x u y and v x v y This represents the displacement gradient of the sub-region.
[0098] This allows us to calculate the displacement and displacement gradient of the center points of all target sub-regions in the x and y directions within the search area of the deformed image, thus obtaining the three-dimensional displacement of the model surface, i.e., the deformation value.
[0099] The specific method for obtaining accurate displacement values using the IC-GN subpixel search algorithm is as follows:
[0100] The relevant function model is rewritten as follows:
[0101]
[0102] In the formula: f(X) and g(X) -- the graphs before and after deformation in X = (x, y, 1) T The gray value at the coordinates, and --The average grayscale values of the sub-region before and after deformation, ζ=(dx,dy,1) T --Local coordinates of pixels within a sub-region relative to the center point
[0103] Let be the deformation parameter variable of the shape function.
[0104] The change in the deformation parameters of the shape function.
[0105] W(ζ;p) -- First-order shape function, W(ζ;Δp) is the iterative increment of the first-order shape function, specifically in the form:
[0106]
[0107] The shape function model is updated using the IC-GN method: W(ζ; p) (k+1) )=W(ζ;p (k) )·W(ζ;Δp) -1
[0108] The formula was optimized using the Gauss-Newton method to obtain W(ζ; Δp), and then a first-order Taylor expansion was performed to obtain:
[0109]
[0110] In the formula: The image gradient of the deformed sub-region. Let be the Jacobian matrix of the shape function. When the correlation function reaches an extremum, It can be deduced that:
[0111]
[0112] Where H is the Hessian matrix, and its value is:
[0113] Therefore, by performing surface fitting operations on the correlation coefficient matrix, the subpixel-level displacement can be obtained from the extreme points of the fitted surface, thereby improving the measurement accuracy of the three-dimensional displacement and deformation of the tomato model.
[0114] The specific steps for calculating the strain field by point-by-point local least squares fitting are as follows:
[0115] Since the original discrete data inevitably contains noise, point-by-point local least squares fitting is used to calculate the strain field. Because the fitting function is a two-dimensional first-order polynomial, the u and v field displacements are fitted as follows:
[0116] u(x,y)=a0+a1x+a2y
[0117] v(x,y)=b0+b1x+b2y
[0118] In the formula: x, y are the local coordinates of each data point in the displacement field; u(x, y) and v(x, y) are discrete displacement data points; a0, ..., b0 are the coefficients of the fitting polynomial to be determined.
[0119] After obtaining the polynomial coefficients a0, ..., b0 using the least squares method, the Cauchy strain components are:
[0120]
[0121]
[0122]
[0123] Where: ε x ε y ε and ε represent the strain values of the point in the x-direction, y-direction, and total strain, respectively.
[0124] According to this embodiment, preferably, in the three-dimensional reconstruction step, the controller uses a phase-shift-based sub-pixel multi-view stereo matching algorithm to estimate the depth of the light field through 4D light field data to reconstruct the three-dimensional image of the tomato stem, specifically as follows:
[0125] A phase-shift-based subpixel multi-view stereo matching algorithm is used to estimate the depth of the light field for 3D reconstruction. The core of this algorithm utilizes phase-shift theory, where a small displacement in the spatial domain is represented in the frequency domain as the product of the frequency domain expression of the original signal and the exponent of the displacement, as shown in the following formula: F{I(x+Δx)}=F{I(x)}exp 2πjΔx Therefore, the image after displacement can be represented as: I'(x)=I(x+Δx)=F -1 {F{I(x)}exp 2πjΔx The concept of phase shifting enables sub-pixel precision matching, thus addressing the short baseline issue to some extent. To facilitate matching between sub-view images, two different cost mechanisms are designed: SAD and GRAD. The final matching value C is obtained through a weighted average, which is a function of the position x and the loss number l, specifically in the following formula: C(x,l)=αC A (x,l)+(1-α)C G (x,l), where α∈[0,1] represents the SAD loss C. A and SGD loss C G The weights between them. Meanwhile, C... A It is defined in the following form: R x Let represent the rectangular region in the neighborhood of point x; τ1 is the cost cutoff value; V represents the region excluding the central viewpoint u. c Other perspectives besides the central perspective image I(u). The above formula compares the central perspective image I(u) with the other perspectives. c The loss is constructed by comparing the difference between I(u,x) and other perspectives I(u,x). Specifically, this involves continuously adjusting the value of I(u,x) from a certain perspective I(u,x). i Move a small distance around point x on the x-axis and subtract it from the central viewpoint; repeat this process until all viewpoints have been compared. The small distance mentioned is Δx in the formula, defined as: Δx(u,l)=lk(uu c ), where k represents the unit pixel of the depth / parallax layer, and Δx increases linearly with the distance between any viewpoint and the central viewpoint. Similarly, a second matching cost, SGD, can be constructed, with the following basic form:
[0126]
[0127] The Diff xThe gradient of the sub-aperture image in the x-direction is represented by β(u); β(u) controls the weights of the cost quantities in the two directions, and is represented by the relative distance between any viewpoint and the central viewpoint.
[0128]
[0129] At this point, the cost function is complete. Next, an edge-preserving filter is used to aggregate the loss of this cost function, resulting in the optimized cost. Then, an iterative optimization model is constructed to optimize the depth map.
[0130] After obtaining the depth map, input it into the Open3D library to recover the 3D point cloud of the tomato fruit stem.
[0131] Determining the precise picking point: Three-dimensional reconstruction is achieved using 4D light field data. Threshold analysis is performed on the stem of the obtained tomato fruit to measure the change in the stem radius. The specific location of the stem with a significantly increased radius is determined as the optimal picking point. The spring blade 11 is then driven to pick the fruit. The specific steps include:
[0132] The light field camera 9 of the tomato stem tactile sensor 5 acquires 4D light field data images of the tomato stem; identifies the radius of the stem in the 4D light field data image; determines the part of the stem with a significantly increased radius as the picking point, and the controller drives the spring blade 11 of the end effector to cut the stem; the end effector releases, and the fruit falls vertically into the fruit bag mounting frame 6 directly below.
[0133] This invention employs tactile positioning technology based on visual recognition. The controller processes color images of tomatoes captured by a depth camera 4 to identify the tomato fruit, calculate the spatial coordinates of the tomato's center and its radius, and then uses the obtained 3D image of the tomato stem to identify the stem's radius. Threshold analysis is performed on the stem radius to determine the optimal picking point, where the stem reaches its maximum radius. This drives the end effector to harvest the tomato, reducing the risk of damage during harvesting, ensuring the tomato's aesthetic appeal, and solving the problem of tomato occlusion during visual recognition positioning. It also reduces the need for manual trimming of harvested tomatoes with excessively long stems. The invention features a simple and compact structure, easy operation, stable performance, and precise and efficient harvesting.
[0134] It should be understood that although this specification describes various embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other implementation methods that can be understood by those skilled in the art. The series of detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and they are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.
Claims
1. A precise positioning tomato harvesting device based on visual-touch combination, characterized in that, Includes a tomato picking robot, a depth camera (4), a tomato stem tactile sensor (5), and a controller; The depth camera (4) is installed on the tomato picking robot to collect color images and depth information of the tomatoes in front of the end effector and transmit them to the controller; The tomato stem tactile sensor (5) is installed on the end effector of the tomato harvesting robot to acquire 4D light field data images and transmit them to the controller; The controller is connected to the tomato picking robot, the depth camera (4), and the tomato stem tactile sensor (5), respectively. The controller identifies tomato fruits based on color images of tomatoes, calculates the coordinates and radius of the tomatoes, determines whether the tomatoes are ripe based on 4D light field data images, reconstructs a three-dimensional image of the tomato stem, determines the optimal picking point, and drives the end effector to pick the tomatoes. The controller performs image processing on the tomato color image acquired by the depth camera (4), identifies the tomato fruit, calculates the spatial coordinates of the tomato center and the radius of the tomato fruit, and determines whether the tomato is ripe and reconstructs the three-dimensional image of the tomato stem based on the obtained 4D light field data image. It also identifies the radius of the tomato fruit stem based on the three-dimensional image of the tomato stem, performs threshold analysis on the radius of the tomato fruit stem, and determines the specific location of the stem at the maximum radius as the best picking point. The tomato stem tactile sensor (5) includes a transparent shell (13), a light source (7), a speckle spraying device (8), a light field camera (9), and fingertip gel (10). The end effector is provided with transparent housings (13) on both sides. A light source (7) and a light field camera (9) are provided inside the transparent housings (13). Finger gel (10) is provided on the side of the transparent housing (13) that is in contact with the tomato. A speckle spraying device (8) is provided on the transparent housing (13) and located around the finger gel (10) for spraying speckles onto the finger gel (10). The light source (7) is used to provide illumination to the light field camera (9), which is used to capture 4D light field data images of the fingertip gel (10) and transmit them to the controller; The tomato surface stress measurement module is used to control the speckle spraying device (8) of the tomato stem tactile sensor (5) to spray uniform black speckles on the fingertip gel (10) and use the light field camera (9) to take pictures of the 4D light field data of the fingertip gel (10) before deformation. After the end effector applies a preset force to the target tomato, it takes pictures of the 4D light field data of the deformed fingertip gel (10) through the light field camera (9). Based on the obtained 4D light field data of the deformed fingertip gel (10), it synthesizes the two-dimensional RGB image and depth image of the speckles on the fingertip gel (10), calculates the deformation value, calculates the strain field of the whole tomato, and compares it with the threshold. The ripeness of the tomato fruit is judged by the hardness.
2. The tomato harvesting device based on visual-touch integration according to claim 1, characterized in that, The controller includes a tomato recognition module, a tomato positioning and detection module, a tomato surface stress measurement module, a tomato stem three-dimensional reconstruction module, and a picking point determination module; The tomato recognition module is used to perform image processing on the tomato color images acquired by the depth camera (4) to identify tomato fruits; The tomato positioning and detection module is used to calculate the spatial coordinates of the tomato center and the radius of the tomato fruit by combining the color image after image processing with the depth information collected by the depth camera (4), and to obtain the position of the tomato on the coordinates of the tomato picking robot. The tomato stem 3D reconstruction module is used to control the end effector to grab the tomato stem when the controller determines that the target tomato is a mature tomato and moves it upward by a preset distance according to the position information of the tomato on the coordinates of the tomato picking robot. The light field camera (9) of the tomato stem tactile sensor (5) acquires the 4D light field data image of the tomato stem and transmits it to the controller. The controller reconstructs the 3D image of the tomato stem through the 4D light field data. The picking point determination module is used to measure the radius of the tomato stem based on the three-dimensional image of the tomato stem, and to perform threshold analysis on the radius of the tomato stem to determine the specific location of the stem with the largest radius as the optimal picking point, thereby driving the end effector to pick the fruit.
3. The tomato harvesting device based on visual-touch integration according to claim 1, characterized in that, The speckle spraying device (8) includes a built-in speckle spraying liquid storage chamber, the outlet of which is provided with a spray pipe and a valve, which is connected to a controller.
4. A control method for a precision positioning tomato harvesting device based on visual-touch integration according to any one of claims 1-3, characterized in that, Includes the following steps: Tomato recognition: The depth camera (4) captures a color image of the tomato in front of the end effector and transmits it to the controller. The controller performs image processing on the color image of the tomato captured by the depth camera (4) to identify the tomato fruit. Tomato positioning detection: The controller combines the color image after image processing with the depth information collected by the depth camera (4) to calculate the spatial three-dimensional coordinates of the tomato center and the radius of the tomato fruit, and obtain the position of the tomato on the coordinates of the tomato picking robot; Tomato maturity identification: The controller controls the speckled spraying device (8) of the tomato stem tactile sensor (5) to spray a uniform black speckled pattern onto the fingertip gel (10) and uses a light field camera (9) to capture the 4D light field data of the fingertip gel (10) before deformation. After the end effector applies a preset force to the target tomato, it captures the 4D light field data of the deformed fingertip gel (10) through the light field camera (9). Based on the acquired 4D light field data of the deformed fingertip gel (10), a two-dimensional RGB image and a depth image of the speckled pattern on the fingertip gel (10) are synthesized, the deformation value is calculated, the strain field of the entire tomato is calculated, and it is compared with the threshold. The ripeness of the tomato fruit is judged by the hardness. Three-dimensional reconstruction of tomato stem: When the controller determines that the target tomato is a ripe tomato, it moves the end effector upward by a preset distance according to the position information of the tomato on the coordinates of the tomato picking robot to grab the stem of the tomato. The light field camera (9) of the tomato stem tactile sensor (5) acquires the 4D light field data image of the tomato stem and transmits it to the controller. The controller reconstructs the three-dimensional image of the tomato stem through the 4D light field data. Picking point determination: The controller measures the radius of the tomato stem based on the three-dimensional image of the tomato stem, and performs threshold analysis on the radius of the tomato stem to determine the specific location of the stem with the largest radius as the optimal picking point, and drives the end effector to pick the fruit.
5. The control method for a precise positioning tomato harvesting device based on visual-touch integration according to claim 4, characterized in that, The image processing in the tomato localization detection step includes the following steps: The color image covered by the Gaussian filter is smoothed, the RGB color space is converted to the HSV color space, and the red area is extracted. The red area is filled, reduced or enlarged to eliminate false alarms in circle detection. The Canny method is used for contour detection, and the circular Hough transform on the contour image is used to detect the circle in the red area image. The proportion of red pixels in the circumscribed square of the detected circle is used to determine the detection area. The image of the recognized circle is drawn on the color image, and the two-dimensional center coordinates and radius information of the recognized circle are output.
6. The control method for a precise positioning tomato harvesting device based on visual-touch integration according to claim 4, characterized in that, The specific steps for tomato location detection are as follows: The controller obtains the depth at the center coordinates of the recognition circle based on the depth information acquired by the depth camera (4), and then... The viewpoint of the RGB sensor is mapped to the viewpoint of the depth sensor. Partial coordinates are calculated from the depth information. The depth and spatial coordinates are then displaced in the coordinate system. The displacement is equal to the radius from the center of the identified circle to the center of the circle. The actual radius is calculated based on the spatial coordinates of the tomato's center and the spatial coordinates where the displacement equals the radius. It is then determined whether the tomato's radius is within a preset value range. If it is within the preset value range, the detected circle is determined to be a tomato. A circle is drawn in the color image, and the spatial three-dimensional coordinates of the tomato's center are output.
7. The control method for a precise positioning tomato harvesting device based on visual-touch integration according to claim 4, characterized in that, The tomato maturity identification step uses the IC-GN subpixel search algorithm to calculate the three-dimensional displacement, i.e., the deformation value, and uses point-by-point local least squares fitting to calculate the strain field.
8. The control method for a precise positioning tomato harvesting device based on visual-touch combination according to claim 4, characterized in that, In the three-dimensional reconstruction step, the controller uses a phase-shift-based sub-pixel multi-view stereo matching algorithm to estimate the depth of the light field based on 4D light field data in order to reconstruct the three-dimensional image of the tomato stem.