Deep learning and point cloud processing-based string tomato detection and positioning method
By adopting deep learning and point cloud processing methods in tomato picking robots, accurate detection and positioning of string tomatoes and their rhizomes is achieved, and the problem of low picking efficiency and accuracy in the existing technology is solved, and the accuracy and efficiency of picking are improved.
Patent Information
- Application Number
- CN202411961128.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately detect and locate tomatoes composed of multiple targets in tomato picking robots, resulting in low picking efficiency and accuracy.
Using a method based on deep learning and point cloud processing, images are acquired through global and local cameras, and the global tomato rhizome detection model and local rhizome segmentation model are used, combined with SDK depth value function and European clustering algorithm to accurately detect and locate string tomatoes and their rhizomes.
It improves the accuracy and efficiency of picking tomatoes, reduces the picking error rate, is suitable for picking whole bunch of tomatoes, and enhances the robustness of robot picking.
Smart Images

Figure CN120071330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural picking robots, and particularly to a method for detecting and positioning cluster tomatoes based on deep learning and point cloud processing. Background Art
[0002] Cluster tomatoes, also known as spike tomatoes, are a new type of tomato variety that has become popular in domestic and foreign markets in recent years. They are usually harvested in clusters, and the fruits come in three sizes: large, medium, and small, weighing over 150 grams, 80 - 150 grams, and under 80 grams respectively. Cluster tomatoes are deeply loved by consumers for their excellent commercial quality and unique taste, and the cultivation area is expanding year by year.
[0003] Computer vision technology is the core of agricultural picking robots, and machine learning and deep learning algorithms also play important roles in agricultural picking robots. Accurately detecting the target tomatoes, accurately positioning the corresponding rootstocks of the target tomatoes, and stably and accurately outputting are the keys for tomato production robots to perform picking operations.
[0004] Currently, there have been some studies on the picking of similar fruits. For the Chinese patent with the application number 202311221279.1, the invention designs a device and method for tomato recognition and positioning, and completes the recognition and positioning of mature tomatoes in complex situations through the lightest Yolov5s structure of the improved Yolov5 model, meeting the accuracy and speed requirements for tomato recognition. This invention involves operations such as the improved Yolov5 model, image acquisition, image processing, and convolutional neural network processing. However, this invention only targets the case of single tomatoes and is not suitable for the picking environment of cluster tomatoes composed of multiple targets, resulting in low picking efficiency; for the Chinese patent with the application number 202210380566.6, this system uses the two - stage object detection algorithm to identify the bounding rectangle frames and center point coordinates of each tomato cluster in the RGB image, classifies single tomatoes, and determines the sticky tomatoes around all unobstructed red fruits in the current cluster of tomatoes by the distance between the center points. The unobstructed red fruit with the fewest sticky tomatoes around it is taken as the target tomato for priority picking. Finally, pose estimation is used to grasp and pick single fruits. However, there are many single - fruit pose situations and the picking point situations are relatively complex, resulting in certain limitations in the picking work. For the Chinese patent with the application number 202411129143.2, this system detects cluster tomatoes through deep learning algorithms and judges the maturity, and obtains the picking points on the fruit stalks through key - point detection to perform single - grain picking or whole - cluster picking of cluster tomatoes. However, this invention does not consider the self - constraints of the end - of - arm picking tool and is not very suitable for the picking of whole clusters of tomatoes. In summary, traditional tomato picking research often has problems such as instability, being severely affected by the environment, and poor universality and inability to be widely used. Summary of the Invention
[0005] The object of the present invention is to overcome the defects existing in the above-mentioned prior art and provide a method for detecting and positioning cluster tomatoes based on deep learning and point cloud processing, so as to achieve accurate detection and positioning of cluster tomatoes and their rhizomes.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] A method for detecting and positioning cluster tomatoes based on deep learning and point cloud processing, the method comprising:
[0008] Obtain the global image of the cluster tomatoes;
[0009] By inputting the global image into the global tomato rhizome detection model, identify the tomatoes and rhizomes, and obtain the pixel coordinates of the tomatoes;
[0010] Using the pixel coordinates of the tomatoes, calculate the depth value of the tomato center point through the SDK depth value function, and then calculate the coordinates of the tomato center point through the SDK world coordinate function to lock the target tomato;
[0011] If the global tomato rhizome detection model simultaneously identifies the tomatoes and complete rhizomes, obtain the close-up image of the target tomato, use the local rhizome segmentation model to identify the rhizomes, and calculate to obtain the rhizome inclination angle and the coordinates of the rhizome center point to locate the rhizomes;
[0012] If the global tomato rhizome detection model fails to identify the complete rhizomes, obtain the close-up image of the target tomato, collect the point cloud data of the incomplete rhizome area in the image, obtain the depth image and point cloud data and perform point cloud preprocessing; separate the incomplete rhizomes and the point cloud blocks of their surface occlusions through the Euclidean clustering algorithm to achieve the segmentation of the target rhizomes; calculate the normal vector of each point in the incomplete rhizome point cloud cluster through the PCA algorithm, regard the incomplete rhizome as a regular shape approximately, and fit the incomplete rhizome through the RANSAC model to obtain the rhizome inclination angle and the coordinates of the rhizome center point to locate the rhizomes.
[0013] Further, the training process of the global tomato rhizome detection model includes:
[0014] Obtain the global image of the cluster tomatoes under natural light, perform random horizontal or vertical flipping, cropping and scale transformation processing on the global image, and at the same time delete the pictures without tomatoes and the pictures that are too blurred;
[0015] Divide the global image into a training set, a validation set and a test set according to a ratio to form a global data set;
[0016] Perform box annotation on the global images in the global data set, and annotate the rhizome area and the tomato fruit area respectively;
[0017] By identifying the feature points of the cluster tomato fruits and rhizomes in the global image, a global tomato rhizome detection model is constructed. The global tomato rhizome detection model is trained using the loss function of object detection and the global dataset to obtain a global tomato rhizome detection model that can simultaneously identify tomatoes and rhizomes.
[0018] Further, the training process of the local rhizome segmentation model includes:
[0019] Obtain close-up images of cluster tomatoes under natural light, perform random horizontal or vertical flipping, cropping, and scale transformation on the close-up images, and at the same time delete overly blurred images;
[0020] Divide the close-up images into a training set, a validation set, and a test set in proportion to form a local dataset;
[0021] Perform box annotation on the local images in the local dataset, and only annotate the rhizome area;
[0022] By identifying the feature points of the rhizome area in the close-up image, a local rhizome segmentation model is constructed. The local rhizome segmentation model is trained using the loss function of object detection and the local dataset to obtain a local rhizome segmentation model that can identify complete and incomplete rhizomes.
[0023] Furthermore, the expression of the loss function of the object detection is:
[0024]
[0025] where y is the quality label of sample i, σ i is the predicted value of sample i, β is the loss parameter, and n is the total number of samples for calculating the loss.
[0026] Further, when the global tomato rhizome detection model simultaneously identifies tomatoes and complete rhizomes, the process of locating the rhizome includes:
[0027] Obtain a close-up image of the target tomato;
[0028] By inputting the close-up image into the local rhizome segmentation model, obtain the rhizome segmentation area and the pixel coordinates of the area anchor box;
[0029] Perform Hough line transformation on the rhizome segmentation area, detect the lines in the area and calculate the slope of the longest line to obtain the rhizome inclination angle;
[0030] Using the pixel coordinates of the area anchor box, calculate the depth value of the rhizome center point through the SDK depth value function, and then calculate the coordinates of the rhizome center point through the SDK world coordinate function to locate the rhizome.
[0031] Further, the point cloud preprocessing includes:
[0032] Downsample the point cloud of the incomplete rhizome area;
[0033] Use a statistical filter to remove the outliers in the point cloud.
[0034] Furthermore, the process of calculating the normal vector of each point in the incomplete rhizome point cloud cluster by the PCA algorithm includes:
[0035] Take a point in the incomplete rhizome point cloud cluster and establish a sphere with this point as the center. All the point clouds within the sphere form a point cloud block;
[0036] Represent the point cloud block with a matrix, and convert the point cloud block matrix into a simplified point cloud block matrix by treating each column of the point cloud block as a point;
[0037] Use the simplified point cloud block matrix to obtain the covariance matrix of the point cloud block, and decompose the covariance matrix of the point cloud block using the SVD matrix decomposition algorithm to obtain the normal vector of the selected point.
[0038] Even further, the simplified point cloud block matrix is:
[0039] p = [p x , p y , p z T ,
[0040] where p is the simplified point cloud block matrix, and p x , p y and p z represent the data sets of the point cloud data on the x, y, and z coordinates respectively.
[0041] Even further, the covariance matrix of the cloud block is:
[0042]
[0043] where p is the simplified point cloud block matrix, and p x , p y and p z represent the data sets of the point cloud data on the x, y, and z coordinates respectively, and Cov is the expanded expression of the covariance matrix of the cloud block.
[0044] Further, after obtaining the coordinates of the rhizome center point, according to the rhizome tilt angle and the coordinates of the rhizome center point, select the picking point, and determine whether there is an obstacle between the end picking tool of the picking robot and the picking point. If there is, reselect the picking point. If not, send the picking point, the coordinates of the tomato center point, the coordinates of the rhizome center point, and the rhizome tilt angle to the picking robot.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1. The present invention respectively obtains the global scene of cluster tomatoes and the images of the close-range root and stem area through the global camera and the local camera, and through secondary recognition and positioning, accurately locates the cluster tomatoes and the root and stem, reduces the picking error rate, and improves the robustness of the solution; by processing the point cloud data of the partially occluded root and stem, the present invention intuitively analyzes the pose of the target root and stem, grasps the angular state of the target root from the three-dimensional space, can also accurately locate the partially occluded root and stem, detects more comprehensively the situations that may occur in the natural scene, and locates the occluded fruits more precisely;
[0047] 2. The RANSAC model of the present invention also estimates the picking points and the picking direction of the robot for the incomplete root and stem, improving the accuracy and picking efficiency of subsequent machine picking;
[0048] 3. The present invention calculates the depth value of the center point of the root and stem through the SDK depth value function, and then calculates the coordinates of the center point of the root and stem by using the SDK world coordinate function, accurately locating the root and stem, and improving the success rate of picking cluster tomatoes;
[0049] 4. Considering the self-constraint of the end picking tool, when determining the picking point, the present invention also judges whether there is an obstacle between the end picking tool of the picking robot and the picking point, which is applicable to picking whole clusters of tomatoes and improves the picking accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is the general flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Embodiment 1
[0053] After the fruits of cluster tomatoes mature, they can be retained on the inflorescence for a long time without falling off. The inflorescence is in a fishbone shape, the fruits are closely attached to the flower spike, hard, crack-resistant, and resistant to storage and transportation. However, the closely arranged clusters also lead to the situation of fruit occlusion of the root and stem during recognition. When a picking robot picks cluster tomatoes, the existing recognition methods often have difficulty in identifying the location of the root and stem, resulting in a decrease in picking efficiency and accuracy.
[0054] This embodiment aims to disclose a method for detecting and positioning cluster tomatoes based on deep learning and point cloud processing. The method can identify and detect cluster tomatoes through the Yolov8 network model, and obtain the three-dimensional point cloud data of partially occluded cluster tomatoes through a 3D camera to achieve precise positioning. Even when the roots and stems of cluster tomatoes are occluded, the roots and stems of cluster tomatoes can be accurately located, and the positioning results are sent to the picking robot to assist the picking robot in efficient, rapid, and accurate picking. The method is as follows Figure 1 shown, and the specific steps are as follows, including:
[0055] Step S1, obtain the global image of cluster tomatoes.
[0056] In step S2, by inputting the global image into the global tomato root and stem detection model, identify tomatoes and their roots and stems, and obtain the pixel coordinates of the tomatoes; using the pixel coordinates of the tomatoes, calculate the depth value of the tomato center point through the SDK depth value function, and then calculate the coordinates of the tomato center point through the SDK world coordinate function to lock the target tomato.
[0057] The training process of the global tomato root and stem detection model includes:
[0058] Obtain the global image of cluster tomatoes under natural light, use Mosaic data augmentation to perform random horizontal or vertical flipping, cropping, and scale transformation on the global image, and at the same time delete pictures without tomatoes and pictures that are too blurred;
[0059] Divide the global image into a training set, a validation set, and a test set in proportion to form a global data set;
[0060] Perform box annotation on the global images in the global data set, and annotate the root and stem areas and the tomato fruit areas respectively;
[0061] By identifying the feature points of cluster tomato fruits and roots and stems in the global image, construct a global tomato root and stem detection model, and use the loss function of object detection and the global data set to train the global tomato root and stem detection model to obtain a global tomato root and stem detection model that can identify tomatoes and their roots and stems simultaneously.
[0062] The expression of the loss function of object detection is:
[0063]
[0064] where y is the quality label of sample i, σ i is the predicted value of sample i, β is the loss parameter, and n is the total number of samples for calculating the loss.
[0065] Step S3, determine whether the global tomato root and stem detection model has identified the complete root and stem.
[0066] Step S4: If the global tomato root and stem detection model simultaneously identifies a tomato and a complete root and stem, obtain a close-up image of the target tomato, use the local root and stem segmentation model to identify the root and stem, and calculate to obtain the root and stem tilt angle and the coordinates of the root and stem center point to locate the root and stem.
[0067] The training process of the local root and stem segmentation model includes:
[0068] Obtain a close-up close-up image of the cluster tomato under natural light;
[0069] Obtain a close-up close-up image of the cluster tomato under natural light, use Mosaic data augmentation to perform random horizontal or vertical flipping, cropping, and scale transformation on the close-up close-up image, and at the same time delete overly blurred images;
[0070] Divide the close-up close-up image into a training set, a validation set, and a test set according to a ratio to form a local data set;
[0071] Perform box annotation on the local images in the local data set, and only annotate the root and stem areas;
[0072] By identifying the feature points in the root and stem area of the close-up close-up image, construct a local root and stem segmentation model, and use the loss function of object detection and the local data set to train the local root and stem segmentation model to obtain a local root and stem segmentation model that can identify complete and incomplete root and stems.
[0073] In this embodiment, the global camera and the local camera of the picking robot are used to obtain the global image and the close-up close-up image under natural light respectively; the training sets in the two data sets are respectively annotated with the annotation tool labelme.
[0074] The process of locating the root and stem includes:
[0075] Obtain a close-up image of the target tomato;
[0076] Input the close-up image into the local root and stem segmentation model to obtain the root and stem segmentation area and the pixel coordinates of the area anchor box;
[0077] Perform Hough line transformation on the root and stem segmentation area, detect the lines in the area and calculate the slope of the longest line to obtain the root and stem tilt angle;
[0078] Use the pixel coordinates of the area anchor box, calculate the depth value of the root and stem center point through the SDK depth value function, and then calculate the coordinates of the root and stem center point through the SDK world coordinate function to locate the root and stem.
[0079] Step S5: If the global tomato root and stem detection model fails to recognize a complete root and stem, obtain a close-range image of the target tomato, collect point cloud data for the incomplete root and stem area in the image, obtain a depth image and point cloud data, and perform point cloud preprocessing; separate the point cloud blocks of the incomplete root and stem and the surface obstacles thereon through the Euclidean clustering algorithm to achieve the segmentation of the target root and stem; calculate the normal vector of each point in the point cloud cluster of the incomplete root and stem through the PCA (Principal Component Analysis) algorithm, roughly regard the incomplete root and stem as a regular shape, and fit the incomplete root and stem through the RANSAC model to obtain the inclination angle of the incomplete root and stem and the coordinates of the center point of the root and stem, and locate the root and stem.
[0080] The point cloud preprocessing includes:
[0081] Downsample the point cloud in the incomplete root and stem area;
[0082] Use a statistical filter to remove point cloud outliers.
[0083] The process of calculating the normal vector of each point in the point cloud cluster of the incomplete root and stem through the PCA algorithm includes:
[0084] Take a point in the point cloud cluster of the incomplete root and stem, establish a sphere with a radius of r centered on this point, and there are n points in the sphere, forming a point cloud block;
[0085] Represent the point cloud block with a matrix, and convert the point cloud block matrix into a simplified point cloud block matrix by regarding each column of the point cloud block as a point;
[0086] The simplified point cloud block matrix is:
[0087] p = [p x , p y , p z T ,
[0088] where p is the simplified point cloud block matrix, p x , p y and p z respectively represent the data sets of the point cloud data on the x, y, and z coordinates;
[0089] Calculate the mean value of the data set on each coordinate in the coordinate system direction, and the formula is as follows:
[0090]
[0091] where p i is the i-th data on the coordinate, n is the number of data, is the direction vector;
[0092] Decentralize each direction vector to form matrix P, and the expression is as follows:
[0093]
[0094] Among them, and respectively represent the data vectors on the x, y, and z coordinates;
[0095] Find the covariance matrix of matrix P. After matrix P is decentralized, E(P)=0. Expand the covariance matrix to obtain the expression:
[0096]
[0097] Among them, p is the point cloud block simplification matrix, p x 、p y and p z respectively represent the data sets of the point cloud data on the x, y, and z coordinates, and Cov is the expanded expression of the cloud block covariance matrix;
[0098] Use the SVD matrix decomposition algorithm to decompose the covariance matrix to form three matrices U, ∑, and V. Among them, the three elements λ 1 、λ 2 、λ 3 in the ∑ diagonal matrix are the three eigenvalues of the covariance matrix and are arranged in descending order. The corresponding three column vectors v 1 、v 2 、v 3 in the V matrix are the three eigenvectors corresponding to the three eigenvalues respectively. The eigenvector v 3 in the V matrix corresponding to the smallest eigenvalue λ 3 in the ∑ diagonal matrix is the normal vector of the extracted point.
[0099] The RANSAC model also estimates the picking points and the robot picking directions for incomplete rhizomes.
[0100] Step S6: According to the rhizome tilt angle and the rhizome center point coordinates, select the picking point, and judge whether there is an obstacle between the end picking execution tool of the picking robot and the picking point. If there is, reselect the picking point. If not, send the picking point, the tomato center point coordinates, the rhizome center point coordinates, and the rhizome tilt angle to the picking robot.
[0101] Next, take the actual usage scenario as an example:
[0102] In this example, an Intel D435i binocular depth camera is used, and the deep learning image recognition adopts the Yolov8 architecture.
[0103] The global camera and local camera of the picking robot are used to obtain the global image and the close-up close-up image under natural light respectively; in this embodiment, there are 2000 global pictures, 2000 close-up local close-ups, and a total of 4000 pictures.
[0104] Preprocess the obtained pictures and divide them into a training set, a validation set, and a test set according to a ratio of 7:2:1 to form two sets of global and local data sets. For the preprocessing stage of the images, the Mosaic data augmentation technique is adopted, including operations such as random horizontal or vertical flipping, cropping, mosaic smearing, and scale transformation, which helps to improve the diversity and robustness of the training data.
[0105] Use the annotation tool labelme to annotate the training sets in the two data sets respectively. Perform bounding box annotation on the training set in the global data set, annotating the root and stem region and the tomato fruit region respectively. Perform segmentation annotation on the training set of the local data set, only annotating the root and stem region.
[0106] When constructing the sample data set, it is necessary to classify and calibrate the sample data set. Among them, the sample data set contains the robot's global view map and local close-up map data.
[0107] Train using the two training sets respectively to construct a global tomato root and stem detection model and a local root and stem segmentation model. The global tomato root and stem detection model can receive the pictures obtained by the global camera as input, identify tomatoes and roots and stems, and output the pixel coordinates and confidence of the tomato anchor box, and the pixel coordinates and confidence of the root and stem anchor box. The local root and stem segmentation model can receive the pictures obtained by the local camera as input, segment the close-up roots and stems, and output the pixel coordinates and confidence of the root and stem segmentation region anchor box.
[0108] Use the global camera to take a global scene image, align the obtained RGB image with the depth frame, input the image into the trained global tomato and root and stem detection model, output the detected tomato pixel coordinates, and calculate the center point. Input the RGB pixel coordinates into the depth value calculation function of the camera SDK to obtain the depth value, use the world coordinate function of the depth camera SDK to return the tomato center point coordinates, and send the tomato center point coordinates to the picking robot. The robot will perform path planning functions according to the coordinate information to approach the target tomato.
[0109] Use the local camera to take a close-up image of the target tomato, input the image into the trained local root and stem segmentation model, and output the root and stem segmentation region and the anchor box pixel coordinates. Perform Hough line transformation on the segmentation region to detect lines, and calculate the slope of the longest line to obtain the root and stem tilt angle.
[0110] Input the RGB pixel coordinates into the depth value calculation function of the camera SDK to obtain the depth value, and use the world coordinate function of the depth camera SDK to return the coordinates of the center point of the root and stem.
[0111] Collect the point cloud of the partial image in the global model where only tomato fruits are detected but the complete root and stem are not detected, and obtain the corresponding depth map and point cloud data. The D435i depth camera can process the captured depth pictures and generate accurate depth information by calculating the distance from each pixel point to the camera.
[0112] Downsample the point cloud of the target area through a voxel filter to maintain the geometric structure of the point cloud and retain the density distribution of the points, thereby effectively improving the subsequent point cloud processing speed without affecting the effectiveness of the algorithm. Use a statistical filter to remove the outliers in the point cloud. By calculating the statistical mean within the neighboring point cloud, the points separated outside the mean are removed to eliminate the noise points.
[0113] Calculate the normal vector of each point in the target point cloud cluster through the PCA (Principal Component Analysis) algorithm. Regard the target root and stem as a roughly regular planar figure, and fit the target root and stem through the RANSAC model. Denote the geometric center point on the axis of the fitted plane as the picking point, and the normal vector of the points on the same plane as the picking direction.
[0114] After the point cloud preprocessing, there is a distance between the point cloud of the occluding leaves, the point cloud of the background, and the point cloud of the target root and stem. Separate the point cloud blocks of the target root and stem and its surface occluders through the Euclidean clustering algorithm using the gaps and color differences between different point cloud blocks to achieve the segmentation of the target root and stem. Obtain the inclination angle of the incomplete root and stem and the coordinates of the center point of the root and stem to locate the root and stem.
[0115] According to the inclination angle of the root and stem and the coordinates of the center point of the root and stem, select the picking point, and judge whether there is interference between the picking execution tool at the end of the picking robot and the picking point. If there is interference, return to reselect the picking point near the center point until the tool interference problem is eliminated. If there is no interference, send the information such as the center point coordinates and the inclination angle of the root and stem to the picking robot, and the robot will perform functions such as path planning and robotic arm picking path planning according to the coordinate information.
[0116] In this embodiment, key links such as image processing, neural network training, bounding box loss calculation, and depth information replacement are fully considered in the design of the system to achieve the accurate detection and positioning of the root and stem of the cluster tomato.
[0117] Embodiment 2
[0118] Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the foregoing method for detecting and positioning cluster tomatoes based on deep learning and point cloud processing.
[0119] At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the foregoing method for detecting and positioning cluster tomatoes based on deep learning and point cloud processing. Of course, in addition to the software implementation, the present invention does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.
[0120] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0121] Computer-readable media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0122] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing, characterized in that: The method comprises: Get a global image of the tomatoes on a skewer; By inputting the global image into the global tomato root detection model, tomatoes and roots are identified and tomato pixel coordinates are obtained; Use the tomato pixel coordinates to calculate the depth value of the tomato center point through the SDK depth value function, and then use the SDK world coordinate function to calculate the coordinates of the tomato center point to lock the target tomato; If the global tomato rhizome detection model recognizes both the tomato and the complete rhizome, a close-up image of the target tomato is obtained, the rhizome is identified using the local rhizome segmentation model, and the rhizome inclination angle and the coordinates of the rhizome center are calculated to locate the rhizome; If the global tomato rhizome detection model fails to identify the complete rhizome, a close-up image of the target tomato is obtained, point cloud data of the incomplete rhizome area in the image is collected, depth images and point cloud data are obtained, and point cloud preprocessing is performed; the incomplete rhizome and its surface occluders are separated by the Euclidean clustering algorithm to achieve the segmentation of the target rhizome; the normal vector of each point in the incomplete rhizome point cloud cluster is calculated by the PCA algorithm, the incomplete rhizome is roughly regarded as a regular shape, and the incomplete rhizome is fitted by the RANSAC model, the rhizome inclination angle and the coordinates of the rhizome center point are obtained, and the rhizome is located.
2. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 1, characterized in that: The training process of the global tomato root detection model includes: Obtain a global image of tomatoes on a string under natural light, randomly flip, crop, and scale the global image horizontally or vertically, and delete images without tomatoes or images that are too blurred. Divide the global image into training set, validation set and test set in proportion to form a global data set; Perform frame annotation on the global image in the global dataset, and mark the root and stem area and the tomato fruit area respectively; By identifying the feature points of tomato fruits and rhizomes in the global image, a global tomato rhizome detection model is constructed. The global tomato rhizome detection model is trained using the target detection loss function and the global dataset to obtain a global tomato rhizome detection model that can simultaneously identify tomatoes and rhizomes.
3. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 1, characterized in that: The training process of the local rhizome segmentation model includes: Acquire close-up images of tomatoes on a string under natural light, randomly flip, crop, and scale the close-up images horizontally or vertically, and delete overly blurred images; The close-up images are divided into training sets, validation sets and test sets in proportion to form a local data set; Perform frame annotation on the local images in the local dataset, and only annotate the root and stem area; By identifying the feature points of the rhizome area in close-up images, a local rhizome segmentation model is constructed. The local rhizome segmentation model is trained using the target detection loss function and local data sets to obtain a local rhizome segmentation model that can identify complete and incomplete rhizomes.
4. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 2 or claim 3, characterized in that: The expression of the loss function of the target detection is: Among them, y is the quality label of sample i, σ i is the predicted value of sample i, β is the loss parameter, and n is the total number of samples for calculating the loss.
5. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 1, characterized in that: When the global tomato root detection model recognizes both the tomato and the complete root, the root location process includes: Acquire a close-up image of the target tomato; By inputting the close-range image into the local rhizome segmentation model, the pixel coordinates of the rhizome segmentation area and the regional anchor frame are obtained; Perform Hough line transform on the rhizome segmentation area, detect the straight lines in the area and calculate the slope of the longest straight line to obtain the rhizome inclination angle; Using the pixel coordinates of the regional anchor frame, the SDK depth value function is used to calculate the depth value of the rhizome center point, and then the SDK world coordinate function is used to calculate the coordinates of the rhizome center point to locate the rhizome.
6. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 1, characterized in that: The point cloud preprocessing includes: Down-sampling of the point cloud of the incomplete rhizome area; Use statistical filters to remove outliers from point clouds.
7. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 1, characterized in that: The process of calculating the normal vector of each point in the incomplete rhizome point cloud cluster by the PCA algorithm includes: Take a point in the incomplete root point cloud group and build a sphere with this point as the center. All point clouds in the sphere form a point cloud block. The point cloud block is represented by a matrix, and each column of the point cloud block is regarded as a point, and the point cloud block matrix is converted into a point cloud block simplified matrix; The point cloud block covariance matrix is obtained by using the simplified matrix of the point cloud block, and the point cloud block covariance matrix is decomposed using the SVD matrix decomposition algorithm to obtain the normal vector of the extracted point.
8. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 7, characterized in that: The simplified matrix of the point cloud block is: p=[p x ,p y ,p z ] T , Among them, p is the simplified matrix of the point cloud block, p x 、p y and p z They represent the data sets of the point cloud data at the three coordinates of x, y, and z respectively.
9. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 7, characterized in that: The cloud block covariance matrix is: Among them, p is the simplified matrix of the point cloud block, p x 、p y and p z They represent the data sets of the point cloud data at the three coordinates of x, y, and z respectively, and Cov is the expanded expression of the cloud block covariance matrix.
10. The method for detecting and locating tomatoes on a vine based on deep learning and point cloud processing according to claim 1, characterized in that: After obtaining the coordinates of the center point of the rhizome, the picking point is selected according to the inclination angle of the rhizome and the coordinates of the center point of the rhizome, and it is determined whether there is an obstacle between the end picking execution tool of the picking robot and the picking point. If there is, the picking point is reselected. If not, the picking point, the coordinates of the center point of the tomato, the coordinates of the center point of the rhizome and the inclination angle of the rhizome are sent to the picking robot.
Citation Information
Patent Citations
A method and system for single-tomato harvesting based on target detection
CN114788455B
Device and method for identifying and positioning tomato picking
CN117079271A
Deep learning-based skewered tomato composite picking method
CN118941960A