Vegetable and fruit picking robot picking target modeling method and system based on Internet of Things
Through multi-perspective IoT visual nodes and deep image fusion technology, combined with spatial clustering and convolutional neural networks, the problems of leaf occlusion and tendril entanglement during grape picking are solved, accurate fruit recognition and picking are achieved, and the accuracy and efficiency of picking are improved.
Patent Information
- Application Number
- CN202510674268.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional grape picking methods make it difficult to effectively identify and locate the position of the fruit, especially when leaves block the view and tendrils are entangled, and it is easy to make mistakes or miss the grapes.
Multi-view IoT vision nodes are used for three-dimensional image acquisition. RGB-D multimodal time series data and depth image fusion are combined. An occlusion reconstruction model is constructed through spatial clustering and convolutional neural networks to identify the spatial structure of fruit bunches and evaluate picking points. A spatial tendril density matrix is constructed to assess the entanglement risk and generate the final picking instructions.
It achieves precise picking in the case of leaf obstruction and tendril entanglement, avoids mis-cutting or missed cutting, and improves picking accuracy and efficiency.
Smart Images

Figure CN120655820A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vision and image processing technology, and in particular to a picking target modeling method and system for a fruit and vegetable picking robot based on the Internet of Things. Background Art
[0002] With the continuous development of agricultural automation, IoT-based smart agriculture technologies are gaining widespread application in crop harvesting. In particular, accurately identifying the spatial structure of fruits and their attachments to ensure efficient and accurate harvesting has become a key research area in smart agriculture. Grapes, a fruit of high economic value, face numerous technical challenges during harvesting, particularly leaf occlusion and tendril entanglement. These challenges make it difficult for traditional automated harvesting methods to efficiently identify and locate the precise picking point.
[0003] During the actual harvesting process, grape bunches are often obscured by leaves, rendering some of the fruit invisible and hindering the robot's ability to identify the fruit's position and posture. Furthermore, due to the grape's growth habits, the bunches are often intertwined with tendrils, forming a complex spatial structure. These tendrils are tightly entangled with the bunches, making harvesting extremely difficult. In summary, leaf obstruction and tendril entanglement, in particular, affect the positioning of the cutting point, making it easy for the robot to mis-cut or miss the fruit.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for modeling picking targets of a fruit and vegetable picking robot based on the Internet of Things to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The picking target modeling method of the fruit and vegetable picking robot based on the Internet of Things includes the following steps:
[0008] Step 1: Use multi-view IoT vision nodes to collect three-dimensional images of the target fruit bunch and construct an RGB-D multimodal time series data set;
[0009] Step 2: Based on the fusion processing of the depth map and color map in the RGB-D multimodal time series data set, the preliminary three-dimensional outline of the fruit bunch is extracted, and the fruit structure domain, stalk domain, leaf extension domain, tendril extension domain and background boundary domain are distinguished by spatial clustering method to obtain domain point cloud data, and the first main axis and second reference line are identified;
[0010] Step 3: Use convolutional neural network technology (CNN) to build an occlusion reconstruction model. After training the occlusion reconstruction model using domain point cloud data, mark the connection point area of the first main axis and the second reference line in the occlusion reconstruction model, and identify whether there is an intersection with the leaf extension domain. If a first intersection area exists, collect the percentage of the area of the first intersection area to the total area of the connection point area and the leaf occlusion depth difference value to generate the corresponding cut point candidate set;
[0011] Step 4: Extract the identified tendril extension domain of each target fruit bunch and combine it with the cut point candidate set to identify the degree of entanglement between the tendril extension domain and the cut point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cut point based on the spatial tendril density matrix. j And evaluate, finally obtain the corrected actual picking and cutting point, generate the final picking instruction and execute it.
[0012] Furthermore, the step 1 includes:
[0013] S11. Within the orchard operation area, deploy multiple nodes supporting RGB-D IoT visual information collection based on the distribution characteristics of fruit trees and the hanging height of fruit bunches. Each node includes an RGB camera, a ToF or structured light depth sensor, a wireless communication device, and an embedded controller.
[0014] S12. After each node is deployed and the view angle calibration operation is performed between each node, a unified acquisition instruction is issued. Each node synchronously acquires the color image and depth image under the current view angle and records the timestamp.
[0015] S13. Pack the color images and depth images of the target fruit bunch collected within a continuous time period to generate an RGB-D multimodal time series data set containing multi-node, multi-view, and multi-frame information.
[0016] Furthermore, the step 2 includes:
[0017] S21. After pixel-level registration, denoising, and conversion of the color image and depth image in the RGB-D multimodal time series data set into a unified format, each pixel is projected into three-dimensional space to form domain point cloud data. The point cloud data is then imported into the point cloud modeling tool CloudCompare, Open3D, or PCL to generate a string of three-dimensional point cloud models.
[0018] S22. In the three-dimensional point cloud model of the fruit bunch, based on the actual production form of the fruit bunch, extract target point cloud data from the three-dimensional point cloud model, set a domain radius and a minimum number of domain points, group adjacent point clouds into clusters, obtain several point clusters to be identified, and distinguish the fruit structure domain, the fruit stalk domain, the leaf extension domain, the tendril extension domain, and the background boundary domain using a spatial clustering method;
[0019] S23. Perform PCA analysis on the fruit-particle domain and the fruit-stalk domain respectively to extract the first principal axis of the fruit-particle domain and the second baseline of the fruit-stalk domain.
[0020] Furthermore, S22 includes:
[0021] S201, the steps of fruit grain domain identification are as follows:
[0022] S2011, perform principal component analysis (PCA) on each point cluster to be identified, and calculate the ratio of the main axis length to the secondary axis length R asxis : L major is the length of the first principal axis, L minor is the secondary axis length; when 1.0≤R asxis When ≤1.4, it means that the shape of the point cluster is approximately spherical, and slightly prolate is allowed but not exceeding 40%, and the label is "first fruit label";
[0023] S2012, converting all RGB values in the point cluster into HSV color space, and calculating the average hue Hue, saturation S and average brightness V of all point clusters to be identified;
[0024] When the average hue is Hue∈[280°,340°], the fruit is purple-red; when the saturation S≥0.5, the saturation is acceptable; when the average brightness V∈[0.2,0.8], the fruit has a moderate brightness, neither too reflective nor too dim. The judgment formula is: Hue∈[280°,340°]&S≥0.5&V∈[0.2,0.8]. When the average hue, saturation S, and average brightness V meet the ranges at the same time, the label is "second fruit label";
[0025] S2013. Extract each point in each cluster and calculate its average Euclidean distance Davg and point density P with the 10 nearest points. Point density P is the ratio of cluster volume V to number of points N. When Davg ≤ 3.5 mm & P ≥ 1.8 × 10 3 points / cm 3 , the label is "Third Fruit Label";
[0026] S2014, identifying the point clusters of the "first fruit particle label", "second fruit particle label" and "third fruit particle label" obtained simultaneously in S2011-S2013, and determining them as fruit particle structure domains;
[0027] S202, the method for obtaining the fruit stalk domain is:
[0028] S2021. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis:When R asxis When ≥5, the label is “first fruit stalk label”;
[0029] S2022. After color feature extraction, Hue∈[10°,60°]&S≥0.3&V∈[0.3,0.7]. When the intervals of average hue Hue, saturation S, and average brightness V are satisfied at the same time, the label is "second fruit stalk label".
[0030] S2023, when Davg≤4.5mm&P≥1.5×103points / cm 3 , labelled as "third peduncle label";
[0031] S2024. Identify the point clusters that simultaneously obtain the "first country handle label", "second country handle label" and "third country handle label" in S2021-S2023, and determine them as the fruit handle domain.
[0032] Furthermore, S22 also includes:
[0033] S203, the method for obtaining the blade extension domain is:
[0034] S2031. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When <3, it means that the point cluster is a relatively flat structure, which conforms to the morphology of the leaf extension domain and is labeled as the “first leaf label”;
[0035] S2032. After color feature extraction, Hue∈[90°,150°]&S≥0.4&V∈[0.4,0.9]. When the intervals of average hue Hue, saturation S, and average brightness V are satisfied at the same time, the label is "second leaf label".
[0036] S2033, when Davg≤5mm&P≥1.2×103points / cm 3 , labelled as "third blade label";
[0037] S2034, identifying the point cluster of the "first leaf label", "second leaf label" and "third leaf label" obtained simultaneously in S2031-S2033, and determining it as a leaf extension domain;
[0038] S204. The method for obtaining the tendril extension domain is:
[0039] S2041. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the major axis length to the minor axis length. asxis :When R asxisWhen ≥10, it indicates that the point cluster is an elongated and curved structure, which is consistent with the morphology of the tendril extension domain and is labeled as “first tendril label”;
[0040] S2042. After color feature extraction, Hue∈[0°,50°]&S≥0.5&V∈[0.2,0.7]. When the intervals of average hue Hue, saturation S, and average brightness V are satisfied at the same time, the label is “second tendril label”.
[0041] S2043, when Davg≤2.5mm&P≥2.0×103points / cm 3 , labeled "third tendril tag";
[0042] S2044, identifying the point cluster of the "first tendril label", "second tendril label" and "third tendril label" obtained simultaneously in S2041-S2043, and determining it as the tendril extension domain;
[0043] S205: The background boundary region is obtained as follows:
[0044] S2051. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When <1.5, the point cluster has a relatively uniform structure, which is consistent with the characteristics of the background area and is labeled as the "first background label";
[0045] S2052. After color feature extraction, S≤0.3&V∈[0.1,0.6]. When the range of average hue Hue, saturation S and average brightness V is satisfied at the same time, the label is "second background label".
[0046] S2053 , identifying and simultaneously obtaining the “first background label” and “second background label” point clusters in S2051 - S2052 , and determining them as background boundary regions.
[0047] Furthermore, step three includes:
[0048] S31, extracting the boundaries of the identified fruit structure domain, fruit stalk domain, leaf extension domain, tendril extension domain, and background boundary domain of each target fruit bunch, and using Boolean operations to calculate the size of the intersection area;
[0049] S32. Construct an occlusion reconstruction model using convolutional neural network technology (CNN). After training the occlusion reconstruction model using domain point cloud data, mark the connection point area between the first principal axis and the second reference line in the occlusion reconstruction model. When a first intersection area is identified within the connection point area, it is determined that the connection point area is not obstructed by leaves and is suitable for direct shearing. Points selected every 1 cm within 10 cm of the connection point area are used as a first positioning shearing point group.
[0050] S33. When a first intersection region is identified within the connection point, the area of the first intersection region within the connection point region is calculated. When the area of the first intersection region accounts for 10%-50% of the total area of the connection point region, it indicates that there is a risk of leaf occlusion in the shear point determination. The maximum leaf occlusion depth in the intersection occlusion region and the absolute value of the depth of the connection point are further identified to obtain a leaf occlusion depth difference value.
[0051] When the blade obstruction depth difference exceeds 10 cm, it indicates that the blade obstruction is far from the connection point. Within a safe range, points within 10 cm of the central axis near the connection point and outside the blade obstruction, but not in the first intersection area, are selected as the second positioning shear point group.
[0052] If the blade obstruction depth difference does not exceed 10 cm, it indicates that the blade obstruction is close to the connection point, and there is a risk of leaf obstruction interference, which may lead to mis-cutting. The fruit bunch is marked as "First Cutting Risk Fruit Bunch";
[0053] S34. When the first intersection area is identified within the connection point, the area of the first intersection area around the connection point is calculated. When the area of the first intersection area accounts for more than 50% of the total area of the connection point area, it means that there is a risk of blade occlusion for the shear point judgment, and it is marked as the "second shear risk fruit bunch".
[0054] Furthermore, the step three further includes:
[0055] S35. When the "first shear risk fruit bunch" and the "second shear risk fruit bunch" are identified, the length of the second reference line in the fruit stem region is located. If the length of the second reference line is greater than 15 cm, points are selected at intervals of 1 cm from the second reference line within 5 cm to 10 cm from the end of the fruit stem in the direction extending from the first principal axis as a third positioning shear point group.
[0056] If the length of the second reference line is within 5 cm-15 cm, then select points at intervals of 1 cm within 5 cm from the end of the fruit stem in the direction extending from the first main axis as the fourth positioning shearing point group;
[0057] If the length of the second reference line is 5 cm, select points within 2 cm to 3 cm of the end point of the fruit stem on the second reference line at intervals of 1 cm as the fifth positioning shearing point group;
[0058] The first positioning shear point group, the second positioning shear point group, the third positioning shear point group, the fourth positioning shear point group and the fifth positioning shear point group are summarized to form a shear point candidate set.
[0059] Furthermore, the step 4 includes:
[0060] S41. Extract the identified tendril extension domain of each target fruit bunch and combine it with the cut point candidate set to identify the degree of entanglement between the tendril extension domain and the cut point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cut point based on the spatial tendril density matrix. j :
[0061] S42, preset risk threshold X, when the entanglement risk index R of the j-th shear point j If the risk threshold X is exceeded, it means that the cutting point has entanglement risk and the cutting point is removed from the cutting point candidate set;
[0062] When the winding risk index R of the j-th shear point j If the risk threshold X is not exceeded, it means that there is no entanglement risk at the shear point, and the entanglement risk index of each shear point that does not exceed the risk threshold X is calculated, and the entanglement risk index R of the j-th shear point is calculated. j The lowest value of the cutting point is marked as the corrected actual picking cutting point, and the final picking instruction is generated and executed.
[0063] Furthermore, S41 includes:
[0064] S411. In the three-dimensional point cloud model, a second intersection region of the tendril extension domain and the set of candidate shear points is calculated. The second intersection region is the portion where the tendril is wrapped. The point cloud data of the tendril extension domain is divided into U voxels in the second intersection region using a spatial partitioning algorithm. The size of each voxel is set to 1x1x1 cm. Each point cloud data point of the tendril extension domain is assigned to a corresponding voxel according to its spatial coordinates to generate a spatial tendril density matrix.
[0065] S412: The density value of each voxel is the number of point clouds within the voxel, and the density value of each voxel is normalized. The entanglement risk index R of the j-th shear point in the shear point candidate set is calculated using the following formula: j :
[0066]
[0067] Where U is the number of voxels in the tendril extension domain in the second communication area, E k is the normalized density value of the i-th voxel, indicating the number of point clouds within the k-th voxel, d k is the Euclidean distance from the kth voxel to the jth shear point.
[0068] The IoT-based fruit and vegetable picking robot picking target modeling system includes:
[0069] The image acquisition module uses a multi-view IoT vision node to collect three-dimensional images of the target fruit bunch and construct an RGB-D multimodal time series data set;
[0070] The 3D point cloud model building module is used to perform pixel-level registration, denoising, and conversion of color and depth images in the RGB-D multimodal time series data set into a unified format. Each pixel is then projected into 3D space to form domain point cloud data. The point cloud data is then imported into point cloud modeling tools such as CloudCompare, Open3D, or PCL to generate a string of 3D point cloud models.
[0071] The structure decomposition module is used to extract the preliminary three-dimensional outline of the fruit bunch based on the fusion processing of the depth map and color map in the RGB-D multimodal time series data set. It uses spatial clustering methods to distinguish the fruit structure domain, stalk domain, leaf extension domain, tendril extension domain and background boundary domain, obtain domain point cloud data, and identify the first main axis and the second reference line.
[0072] The first correction module uses convolutional neural network technology (CNN) to build an occlusion reconstruction model. After training the occlusion reconstruction model using domain point cloud data, it marks the connection point area of the first main axis and the second baseline in the occlusion reconstruction model and identifies whether there is an intersection with the leaf extension domain. If a first intersection area exists, the percentage of the area of the first intersection area to the total area of the connection point area and the difference in leaf occlusion depth are collected to generate a corresponding set of cut point candidates.
[0073] The second correction module is used to extract the identified tendril extension domain of each target fruit bunch and combine it with the cutting point candidate set, identify the degree of entanglement between the tendril extension domain and the cutting point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cutting point based on the spatial tendril density matrix. j And evaluate, finally obtain the corrected actual picking and cutting point, generate the final picking instruction and execute it;
[0074] The robot control module is used to plan the optimal grasping path based on the position information of the target fruit bunch and the shearing point of the final picking instruction. The robot control module will then guide the robotic arm to perform motion control, including translation, rotation, extension and shearing, until the fruit bunch is successfully picked.
[0075] Compared with existing technologies, the present invention offers the following advantages: by integrating RGB-D multimodal time-series data collected by multi-view IoT vision nodes with depth and color image fusion processing, it can achieve precise three-dimensional modeling of fruit bunches and accurately extract their spatial structure information. This overcomes leaf occlusion issues, ensuring that even if part of the fruit is obscured by leaves, the robot can still accurately identify and locate the fruit and its accessories.
[0076] The present invention distinguishes the fruit structure domain, fruit stalk domain, leaf extension domain, tendril extension domain and background boundary domain through a spatial clustering method, and can effectively process the complex spatial structure of grape bunches and tendrils intertwined, accurately extract the various components of the fruit bunch, and avoid the interference caused by tendril entanglement by traditional methods. Through the analysis of the tendril extension domain, this method can clearly separate the relationship between tendrils and fruit bunches, and provide clear spatial data for subsequent picking decisions. The occlusion reconstruction model constructed using a convolutional neural network (CNN) can mark the intersection area between the fruit bunch and the leaves through the reconstruction model in the face of leaf occlusion, and accurately evaluate the candidate set of cutting points based on the depth difference value of the leaf occlusion and the area of the intersection area. This process not only solves the problem of the inability to identify the fruit due to occlusion in the traditional method, but also selects the appropriate picking point by dynamically evaluating the degree of occlusion to avoid miscutting or missing cuts.
[0077] The present invention also accurately identifies the degree of influence of tendrils on picking points by constructing a spatial tendril density matrix and combining it with the calculation of the entanglement risk index. The spatial tendril density matrix evaluates the entanglement risk of each shearing point based on the actual entanglement of the tendrils and the fruit bunch, thereby avoiding picking difficulties or robotic arm jamming caused by tendril entanglement, and ensuring the smooth progress of the picking operation. Combined with the calibration of the first main axis and the second reference line, the final generated picking instructions are based on a comprehensive analysis of the fruit bunch structure, occlusion conditions, and tendril entanglement degree, ensuring that each picking is accurate and effective, and promoting improved shearing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 This is a schematic diagram of the steps of the picking target modeling method of the fruit and vegetable picking robot based on the Internet of Things of the present invention.
[0079] Figure 2 The figure is a flow chart of the target modeling system for the fruit and vegetable picking robot based on the Internet of Things of the present invention. DETAILED DESCRIPTION
[0080] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.
[0081] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0082] Example 1: Please refer to Figure 1 , the present invention provides a technical solution:
[0083] The picking target modeling method of the fruit and vegetable picking robot based on the Internet of Things includes the following steps:
[0084] Step 1: Use multi-view IoT vision nodes to collect three-dimensional images of the target fruit bunch and construct an RGB-D multimodal time series data set;
[0085] Step 2: Based on the fusion processing of the depth map and color map in the RGB-D multimodal time series data set, the preliminary three-dimensional outline of the fruit bunch is extracted, and the fruit structure domain, stalk domain, leaf extension domain, tendril extension domain and background boundary domain are distinguished by spatial clustering method to obtain domain point cloud data, and the first main axis and second reference line are identified;
[0086] Step 3: Use convolutional neural network technology (CNN) to build an occlusion reconstruction model. After training the occlusion reconstruction model using domain point cloud data, mark the connection point area of the first main axis and the second reference line in the occlusion reconstruction model, and identify whether there is an intersection with the leaf extension domain. If a first intersection area exists, collect the percentage of the area of the first intersection area to the total area of the connection point area and the leaf occlusion depth difference value to generate the corresponding cut point candidate set;
[0087] Step 4: Extract the identified tendril extension domain of each target fruit bunch and combine it with the cut point candidate set to identify the degree of entanglement between the tendril extension domain and the cut point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cut point based on the spatial tendril density matrix. j And evaluate, finally obtain the corrected actual picking and cutting point, generate the final picking instruction and execute it.
[0088] In this embodiment, the RGB-D multimodal time series data collected by the multi-view IoT visual node is combined with the fusion processing of the depth map and the color map to achieve accurate three-dimensional modeling of the fruit bunch and accurately extract the spatial structure information of the fruit bunch. This can overcome the leaf occlusion problem and ensure that even if part of the fruit is blocked by the leaves, the robot can still accurately identify and locate the fruit and its appendages. By distinguishing the fruit structure domain, the fruit stalk domain, the leaf extension domain, the tendril extension domain and the background boundary domain through the spatial clustering method, the complex spatial structure of the grape fruit bunch and the tendril intertwined can be effectively processed, and the various components of the fruit bunch can be accurately extracted, avoiding the interference caused by the tendril entanglement of the traditional method. By analyzing the tendril extension domain, this method can clearly separate the relationship between the tendril and the fruit bunch, providing clear spatial data for subsequent picking decisions. The occlusion reconstruction model constructed using a convolutional neural network (CNN) can mark the intersection area between the fruit bunch and the leaf through the reconstruction model in the face of leaf occlusion, and accurately evaluate the shear point candidate set based on the depth difference value of the leaf occlusion and the area of the intersection area. This process not only solves the problem of fruit being unable to be identified due to occlusion in traditional methods, but also selects appropriate picking points by dynamically evaluating the degree of occlusion, thus avoiding mis-cutting or missed cutting. By constructing a spatial tendril density matrix and combining it with the calculation of the entanglement risk index, the degree of influence of tendrils on picking points can be accurately identified. The spatial tendril density matrix evaluates the entanglement risk of each cutting point based on the actual entanglement of tendrils and fruit bunches, thereby avoiding picking difficulties or robotic arm jamming caused by tendril entanglement, and ensuring the smooth progress of picking operations. Combined with the calibration of the first main axis and the second reference line, the final generated picking instructions are based on a comprehensive analysis of the fruit bunch structure, occlusion and tendril entanglement, which promotes improved cutting accuracy and improves the accuracy of each picking.
[0089] Example 2: This example is an explanation of Example 1. Specifically, step 1 includes:
[0090] S11. Within the orchard operation area, multiple nodes supporting RGB-D IoT visual information collection are deployed based on the distribution characteristics of the fruit trees and the hanging height of the fruit bunches. Each node includes an RGB camera, a ToF or structured light depth sensor, a wireless communication device, and an embedded controller. By deploying multiple nodes supporting RGB-D information collection and leveraging the perspective differences between different nodes, color and depth images of the target fruit bunch can be acquired from multiple perspectives. This multi-perspective acquisition method significantly improves the recognition accuracy of the spatial structure of the fruit bunch and its appendages (such as stalks, leaves, tendrils, etc.), avoiding the blind spots associated with traditional single-perspective acquisition methods. This approach, especially for difficult-to-identify situations such as leaf occlusion and tendril entanglement, allows for more complete and detailed three-dimensional spatial data. Within the orchard operation area, multiple nodes are flexibly deployed based on the distribution characteristics of the fruit trees and the hanging height of the fruit bunches, ensuring extensive coverage of the acquisition equipment and comprehensive data collection. The rational configuration of heights and angles based on the growth patterns of the fruit trees and the distribution of the fruit bunches helps ensure data accuracy and comprehensiveness, reduces data blind spots caused by inappropriate node placement, and enhances the intelligence of the entire harvesting process.
[0091] S12: After deploying each node and performing the viewpoint calibration operation between nodes, a unified acquisition instruction is issued. Each node synchronously acquires the color image and depth image under the current viewpoint and records the timestamp. In step S12, by executing the viewpoint calibration between nodes and the unified acquisition instruction, each node synchronously acquires data and records the timestamp. This synchronization operation ensures the temporal consistency of the data collected by each node and avoids the time difference error caused by data asynchrony. It ensures that the data of different nodes can be effectively integrated during processing, improves the accuracy of data processing, and provides an accurate time reference for subsequent 3D modeling and target recognition.
[0092] S13. Packaging the color and depth images of each target fruit bunch collected over a continuous time period generates a multi-node, multi-view, and multi-frame RGB-D multimodal time series data set. This not only effectively integrates multi-source data from different nodes but also preserves time series information, facilitating dynamic tracking of target fruit bunch positional changes and improving dynamic fruit identification and precise location capabilities. This data set enables the robot's decision-making system to obtain more contextual information, enabling more intelligent picking decisions.
[0093] In this embodiment, through the collaboration and synchronous data collection of multiple nodes, not only can the recognition accuracy of the target fruit bunch be improved, but also the occlusion and complex spatial structure problems that may exist in the environment can be effectively handled, providing a solid data foundation for subsequent tasks such as fruit bunch recognition, cutting point positioning, and automated picking.
[0094] Example 3: This example is an explanation of Example 1. Specifically, step 2 includes:
[0095] S21. After pixel-level registration, denoising, and conversion of the color image and depth image in the RGB-D multimodal time series data set into a unified format, each pixel is projected into three-dimensional space to form domain point cloud data. The point cloud data is then imported into the point cloud modeling tool CloudCompare, Open3D, or PCL to generate a string of three-dimensional point cloud models.
[0096] S22. In the three-dimensional point cloud model of the fruit bunch, based on the actual production form of the fruit bunch, extract target point cloud data from the three-dimensional point cloud model, set a domain radius and a minimum number of domain points, group adjacent point clouds into clusters, obtain several point clusters to be identified, and distinguish the fruit structure domain, the fruit stalk domain, the leaf extension domain, the tendril extension domain, and the background boundary domain using a spatial clustering method;
[0097] S23, performing PCA analysis on the fruit-particle domain and the fruit-stalk domain respectively, extracting the first principal axis of the fruit-particle domain and the second baseline of the fruit-stalk domain;
[0098] For the fruit-grain structure domain, the spatial distribution of the point cloud is used to calculate the covariance of the point cloud of the fruit-grain structure domain to identify its main axis direction, and the main axis is assumed to be the "rotation axis" of the fruit bunch, marked as the first main axis; the expression of the covariance matrix C1 of the point cloud of the fruit-grain structure domain is:
[0099]
[0100] Where M1 is the total number of point clouds of the fruit grain structure domain; Λ 1,i represents the i-th point in the fruit grain structure domain point cloud, μ1 is the mean of the fruit grain structure domain point cloud, and each point Λ 1,i is a three-dimensional coordinate (x, y, z); (Λ 1,i -μ1)(Λ 1,i -μ1) T T in the equation represents transposition, (Λ 1,i -μ1) represents the column vector of the pot structure domain point cloud, transpose (Λ 1,i -μ1) T It is the row vector of this column vector that extracts the principal component of the data (i.e., the first principal axis) by performing eigenvalue decomposition on the covariance matrix, and extracts the largest eigenvector of the covariance matrix C1, which is marked as the first principal axis;
[0101] For the fruit stalk domain, the spatial distribution of the point cloud is used to calculate the covariance C2 of the point cloud in the fruit stalk domain to identify its main axis direction, and the main axis is assumed to be the "rotation axis" of the fruit stalk, marked as the second reference line; the expression of the covariance matrix C2 of the point cloud in the fruit stalk domain is:
[0102]
[0103] Where M2 is the total number of point clouds in the fruit stalk domain; Λ 2,i represents the i-th point in the fruit stalk domain point cloud, μ2 is the mean of the fruit stalk domain point cloud, and M is the number of fruit stalk domain point clouds; the principal component of the data (i.e., the second baseline) is extracted by performing eigenvalue decomposition on the covariance matrix, and the largest eigenvector of the covariance matrix C2 of the fruit stalk domain point cloud is extracted and marked as the second baseline.
[0104] In this embodiment, a spatial clustering method is used to divide the point cloud data into multiple point clusters based on the actual production form of the fruit bunch, corresponding to different fruit bunch regions, such as the kernel structure domain, the fruit stalk domain, the leaf extension domain, the tendril extension domain, and the background boundary domain. This method effectively distinguishes the spatial distribution of the various parts of the fruit bunch, ensuring clear structural division between different regions. In this way, data from different parts can be processed independently, avoiding recognition confusion caused by interlaced or overlapping regions, and improving the system's recognition accuracy for the fruit bunch. The PCA (principal component analysis) method is applied to the point cloud data in the kernel structure domain and the fruit stalk domain, respectively, to extract the main axis and reference line. This method can accurately identify the main axis direction of the fruit and the fruit stalk, marking them as the first main axis and the second reference line. By calculating the covariance matrix and performing eigenvalue decomposition, the main direction of the point cloud data can be extracted, ensuring the accurate spatial positioning of each part of the fruit bunch. This PCA-based processing method can efficiently and accurately determine the direction of the fruit and the fruit stalk, avoiding the errors in fruit structure positioning used in traditional methods.
[0105] Example 4: This example is an explanation of Example 3. Specifically, S22 includes:
[0106] S201, the steps of fruit grain domain identification are as follows:
[0107] S2011, perform principal component analysis (PCA) on each point cluster to be identified, and calculate the ratio of the main axis length to the secondary axis length R asxis : L major is the length of the first principal axis, L minor is the secondary axis length; when 1.0≤R asxis When ≤1.4, it means that the shape of the point cluster is approximately spherical, and slightly prolate is allowed but not exceeding 40%, and the label is "first fruit label";
[0108] S2012, converting all RGB values in the point cluster into HSV color space, and calculating the average hue Hue, saturation S and average brightness V of all point clusters to be identified;
[0109] When the average hue is Hue∈[280°,340°], the fruit is purple-red; when the saturation S≥0.5, the saturation is acceptable; when the average brightness V∈[0.2,0.8], the fruit has a moderate brightness, neither too reflective nor too dim. The judgment formula is: Hue∈[280°,340°]&S≥0.5&V∈[0.2,0.8]. When the average hue, saturation S, and average brightness V meet the ranges at the same time, the label is "second fruit label";
[0110] S2013. Extract each point in each cluster and calculate its average Euclidean distance Davg and point density P with the 10 nearest points. Point density P is the ratio of cluster volume V to number of points N. When Davg ≤ 3.5 mm & P ≥ 1.8 × 10 3 points / cm 3 , the label is "third fruit particle label"; "points" represents the number of points contained in the point cloud data within a given spatial volume (here per cubic centimeter). In each cubic centimeter of space, the point cloud data contains at least 1800 data points;
[0111] S2014, identifying the point clusters of the "first fruit particle label", "second fruit particle label" and "third fruit particle label" obtained simultaneously in S2011-S2013, and determining them as fruit particle structure domains;
[0112] S202, the method for obtaining the fruit stalk domain is:
[0113] S2021. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When ≥5, it means that the point cluster has obvious slender structural characteristics, which is consistent with the morphology of the fruit stalk domain and is labeled as "first fruit stalk label";
[0114] S2022. When the average hue Hue∈[10°,60°], it indicates that the area has a green hue, which is consistent with the color characteristics of the petiole area. When the saturation S≥0.3, it indicates that the saturation is qualified. When the average brightness V∈[0.3,0.7], it indicates that the brightness is low and not too reflective. The judgment formula is: after color feature extraction, Hue∈[10°,60°]&S≥0.3&V∈[0.3,0.7]. When the average hue Hue, saturation S and average brightness V are all satisfied at the same time, the label is "second fruit petiole label".
[0115] S2023, when Davg≤4.5mm&P≥1.5×103points / cm 3 , labelled as "third peduncle label";
[0116] S2024. Identify the point clusters that simultaneously obtain the "first country handle label", "second country handle label" and "third country handle label" in S2021-S2023, and determine them as the fruit handle domain.
[0117] S203, the method for obtaining the blade extension domain is:
[0118] S2031. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When <3, it means that the point cluster is a relatively flat structure, which conforms to the morphology of the leaf extension domain and is labeled as the “first leaf label”;
[0119] S2032. When the average hue Hue∈[90°,150°], it indicates that the area has a green hue, which is consistent with the color characteristics of the leaf. When the saturation S≥0.4, it indicates that the saturation is qualified. When the average brightness V∈[0.4,0.9], it indicates that the brightness is high, which is consistent with the reflective characteristics of the leaf. The judgment formula is: after color feature extraction, Hue∈[90°,150°]&S≥0.4&V∈[0.4,0.9]. When the average hue Hue, saturation S, and average brightness V are all in the range, the label is "second leaf label";
[0120] S2033, when Davg≤5mm&P≥1.2×103points / cm 3 , labelled as "third blade label";
[0121] S2034, identifying the point cluster of the "first leaf label", "second leaf label" and "third leaf label" obtained simultaneously in S2031-S2033, and determining it as a leaf extension domain;
[0122] S204. The method for obtaining the tendril extension domain is:
[0123] S2041. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the major axis length to the minor axis length. asxis :When R asxis When ≥10, it indicates that the point cluster is an elongated and curved structure, which is consistent with the morphology of the tendril extension domain and is labeled as “first tendril label”;
[0124] S2042. When the average hue Hue∈[0°,50°], it indicates that the area is yellow or light green, which meets the color characteristics of the tendril area. When the saturation S≥0.5, it indicates that the saturation is qualified. When the average brightness V∈[0.2,0.7], it indicates that the brightness is moderate and not too bright. The judgment formula is: after color feature extraction, Hue∈[0°,50°]&S≥0.5&V∈[0.2,0.7]. When the average hue Hue, saturation S and average brightness V meet the ranges of the average hue Hue, saturation S and average brightness V at the same time, the label is "second tendril label";
[0125] S2043, when Davg≤2.5mm&P≥2.0×103points / cm 3 , labeled "third tendril tag";
[0126] S2044, identifying the point cluster of the "first tendril label", "second tendril label" and "third tendril label" obtained simultaneously in S2041-S2043, and determining it as the tendril extension domain;
[0127] S205: The background boundary region is obtained as follows:
[0128] S2051. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When <1.5, the point cluster has a relatively uniform structure, which is consistent with the characteristics of the background area and is labeled as the "first background label";
[0129] S2052: When the saturation S≤0.3, it indicates that the saturation is qualified; when the average brightness V∈[0.3,0.7], it indicates that the saturation is low, which is consistent with the characteristics of the background area; when V∈[0.1,0.6], it indicates that the brightness is low, which is consistent with the ambient light characteristics of the background area;
[0130] After color feature extraction, S≤0.3&V∈[0.1,0.6], when the intervals of average hue Hue, saturation S and average brightness V are satisfied at the same time, the label is "second background label";
[0131] S2053 , identifying and simultaneously obtaining the “first background label” and “second background label” point clusters in S2051 - S2052 , and determining them as background boundary regions.
[0132] In fields like botany, computer vision, and image processing, researchers and engineers have accumulated extensive experience understanding the morphology, color, and other characteristics of different plant parts. This domain knowledge often forms the basis for setting thresholds. For example, for characteristics like leaf color and fruit shape, existing literature or expert experience can provide reference.
[0133] In this example, the present invention effectively improves the ability to accurately identify fruit, fruit bunches, and associated structures (such as stalks and leaves) in various environments by extracting detailed features for each structural domain (fruit, stalk, leaf, tendril, and background), including spatial structure analysis, color analysis, and point cloud density analysis. The extraction method for each domain is highly targeted, enabling efficient and accurate identification and classification of each component in complex environments.
[0134] Example 5: This example is an explanation of Example 4. Specifically, step 3 includes:
[0135] S31, extracting the boundaries of the identified fruit structure domain, fruit stalk domain, leaf extension domain, tendril extension domain, and background boundary domain of each target fruit bunch, and using Boolean operations to calculate the size of the intersection area;
[0136] S32. Construct an occlusion reconstruction model using convolutional neural network technology (CNN). After training the occlusion reconstruction model using domain point cloud data, mark the connection point area between the first principal axis and the second reference line in the occlusion reconstruction model. When a first intersection area is identified within the connection point area, it is determined that the connection point area is not obstructed by leaves and is suitable for direct shearing. Points selected every 1 cm within 10 cm of the connection point area are used as a first positioning shearing point group.
[0137] S33. When a first intersection region is identified within the connection point, the area of the first intersection region within the connection point region is calculated. When the area of the first intersection region accounts for 10%-50% of the total area of the connection point region, it indicates that there is a risk of leaf occlusion in the shear point determination. The maximum leaf occlusion depth in the intersection occlusion region and the absolute value of the depth of the connection point are further identified to obtain a leaf occlusion depth difference value.
[0138] When the blade obstruction depth difference exceeds 10 cm, it indicates that the blade obstruction is far from the connection point. Within a safe range, points within 10 cm of the central axis near the connection point and outside the blade obstruction, but not in the first intersection area, are selected as the second positioning shear point group.
[0139] If the blade obstruction depth difference does not exceed 10 cm, it indicates that the blade obstruction is close to the connection point, and there is a risk of leaf obstruction interference, which may lead to mis-cutting. The fruit bunch is marked as "First Cutting Risk Fruit Bunch";
[0140] S34. When the first intersection area is identified within the connection point, the area of the first intersection area around the connection point is calculated. When the area of the first intersection area accounts for more than 50% of the total area of the connection point area, it means that there is a risk of blade occlusion for the shear point judgment, and it is marked as the "second shear risk fruit bunch".
[0141] In this embodiment, by calculating the areas and relative positions of different regions (such as the first intersection region and the connection point region), it is possible to determine whether there is a clipping risk when there is significant leaf occlusion. In particular, when the difference in occlusion depth exceeds a certain threshold, the system can promptly mark the cluster as a "first clipping risk cluster," preventing clipping in the dangerous area.
[0142] By calculating the intersection area of the connection point area and analyzing occlusion at different depths, the most appropriate cutting points (such as the first and second positioning cutting point groups) are selected. This method helps optimize the cutting path, minimize fruit damage, and improve cutting efficiency. When occlusion significantly affects cutting judgment, the system can mark it as a "first cutting risk fruit bunch" or "second cutting risk fruit bunch." This helps operators identify potential problems in advance and avoid errors.
[0143] Example 6: This example is an explanation of Example 5. Specifically, step 3 further includes:
[0144] S35. When the "first shear risk fruit bunch" and the "second shear risk fruit bunch" are identified, the length of the second reference line in the fruit stem region is located. If the length of the second reference line is greater than 15 cm, points are selected at intervals of 1 cm from the second reference line within 5 cm to 10 cm from the end of the fruit stem in the direction extending from the first principal axis as a third positioning shear point group.
[0145] If the length of the second reference line is within 5 cm-15 cm, then select points at intervals of 1 cm within 5 cm from the end of the fruit stem in the direction extending from the first main axis as the fourth positioning shearing point group;
[0146] If the length of the second reference line is 5 cm, select points within 2 cm to 3 cm of the end point of the fruit stem on the second reference line at intervals of 1 cm as the fifth positioning shearing point group;
[0147] The first positioning shear point group, the second positioning shear point group, the third positioning shear point group, the fourth positioning shear point group and the fifth positioning shear point group are summarized to form a shear point candidate set.
[0148] In this embodiment, the cutting point is adjusted based on the length of the fruit stem: by analyzing the length of the second reference line, the system can adjust the position of the cutting point according to the different lengths of the fruit stem. This approach ensures that fruit bunches with different types of fruit stems can receive the appropriate cutting point selection, avoiding miscutting or improper cutting on shorter or longer fruit stems. If the second reference line is longer than 15cm, the point group 5-10cm from the end of the fruit stem is selected to ensure that the cutting point is appropriately close to the base of the fruit stem.
[0149] For fruit stalks between 5cm and 15cm in length, select a 5cm area near the end for cutting, ensuring that the cutting point is close to the fruit stalk;
[0150] For particularly short fruit stalks (less than 5 cm), choose to cut in an area 2-3 cm close to the end of the fruit stalk to avoid inaccurate operations caused by cutting points that are too far away.
[0151] By selecting different shear point groups (e.g., the first, second, third, fourth, and fifth positioning shear point groups), candidate shear points can be provided in different areas of the fruit bunch. This multi-level setup allows the system to flexibly respond to changes in fruit bunches and stems, ensuring more accurate and adaptable shearing operations.
[0152] Example 7: This example is an explanation of Example 6. Specifically, step 4 includes:
[0153] S41. Extract the identified tendril extension domain of each target fruit bunch and combine it with the cut point candidate set to identify the degree of entanglement between the tendril extension domain and the cut point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cut point based on the spatial tendril density matrix. j :
[0154] S411. In the three-dimensional point cloud model, the second intersection area of the tendril extension domain and the shear point candidate set is calculated. The second intersection area is the part where the tendril is entangled. The point cloud data of the tendril extension domain is divided into U voxels in the second intersection area through a spatial partitioning algorithm. The size of each voxel is set to 1x1x1cm, and each point cloud data of the tendril extension domain is assigned to the corresponding voxel according to its spatial coordinates to generate a spatial tendril density matrix. Step S411 uses the second intersection area of the tendril extension domain and the shear point candidate set to identify the area where the tendril is entangled. These areas are divided into multiple voxels using the three-dimensional point cloud data, and the size of each voxel is 1x1x1cm. This method can be refined to the microscopic level and accurately capture the tendril entanglement. The detailed spatial division improves the accuracy of the model and helps to identify subtle tendril entanglements in actual picking to avoid omissions.
[0155] S412: The density value of each voxel is the number of point clouds within the voxel, and the density value of each voxel is normalized. The entanglement risk index R of the j-th shear point in the shear point candidate set is calculated using the following formula: j :
[0156]
[0157] Where U is the number of voxels in the tendril extension domain in the second communication area, E kis the normalized density value of the i-th voxel, indicating the number of point clouds within the k-th voxel, d k It is the Euclidean distance from the kth voxel to the jth shear point, that is, the straight-line distance between the voxel and the shear point. The density value of each voxel is the number of point clouds within the voxel, and then it is normalized so that the density value of each voxel is between 0 and 1. Through these density values, the density of tendrils within each voxel can be evaluated, and then the entanglement risk index of the shear point can be calculated. The entanglement risk index of the jth shear point can be calculated through a formula (which involves the number of voxels and distance calculation). This formula takes into account the spatial distribution of tendrils and the distance between shear points, and can effectively identify shear points with higher entanglement risks.
[0158] S42, preset risk threshold X, when the entanglement risk index R of the j-th shear point j If the risk threshold X is exceeded, it means that the cutting point has entanglement risk and the cutting point is removed from the cutting point candidate set;
[0159] When the winding risk index R of the j-th shear point j If the risk threshold X is not exceeded, it means that there is no entanglement risk at the shear point, and the entanglement risk index of each shear point that does not exceed the risk threshold X is calculated, and the entanglement risk index R of the j-th shear point is calculated. j The cutting point with the lowest value in is marked as the corrected actual picking cutting point, and the final picking instruction is generated and executed.
[0160] In this embodiment, through the calculation of the spatial tendril density matrix and the entanglement risk index, the system can accurately assess the risk of the cutting point to avoid miscutting or damaging the fruit due to tendril entanglement. By eliminating the cutting points with a higher risk of entanglement, the system can effectively prevent operational errors caused by tendril interference and ensure the safety of the picking process. By intelligently screening the cutting points, unnecessary adjustments and erroneous operations are avoided, and the overall picking efficiency is improved. When the entanglement risk index does not exceed the risk threshold X, the system will record all the cutting points that meet the requirements and select the cutting point with the lowest risk index as the actual picking cutting point. Through this method, the safest and most suitable picking point can be selected to ensure the accuracy and efficiency of picking.
[0161] Example 8, please refer to Figure 2 ,The picking target modeling system of the fruit and vegetable picking robot based on the Internet of Things includes:
[0162] The image acquisition module uses a multi-view IoT vision node to collect three-dimensional images of the target fruit bunch and construct an RGB-D multimodal time series data set;
[0163] The 3D point cloud model building module is used to perform pixel-level registration, denoising, and conversion of color and depth images in the RGB-D multimodal time series data set into a unified format. Each pixel is then projected into 3D space to form domain point cloud data. The point cloud data is then imported into point cloud modeling tools such as CloudCompare, Open3D, or PCL to generate a string of 3D point cloud models.
[0164] The structure decomposition module is used to extract the preliminary three-dimensional outline of the fruit bunch based on the fusion processing of the depth map and color map in the RGB-D multimodal time series data set. It uses spatial clustering methods to distinguish the fruit structure domain, stalk domain, leaf extension domain, tendril extension domain and background boundary domain, obtain domain point cloud data, and identify the first main axis and the second reference line.
[0165] The first correction module uses convolutional neural network technology (CNN) to build an occlusion reconstruction model. After training the occlusion reconstruction model using domain point cloud data, it marks the connection point area of the first main axis and the second baseline in the occlusion reconstruction model and identifies whether there is an intersection with the leaf extension domain. If a first intersection area exists, the percentage of the area of the first intersection area to the total area of the connection point area and the difference in leaf occlusion depth are collected to generate a corresponding set of cut point candidates.
[0166] The second correction module is used to extract the identified tendril extension domain of each target fruit bunch and combine it with the cutting point candidate set, identify the degree of entanglement between the tendril extension domain and the cutting point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cutting point based on the spatial tendril density matrix. j And evaluate, finally obtain the corrected actual picking and cutting point, generate the final picking instruction and execute it;
[0167] The robot control module is used to plan the optimal grasping path based on the position information of the target fruit bunch and the shearing point of the final picking instruction. The robot control module will then guide the robotic arm to perform motion control, including translation, rotation, extension and shearing, until the fruit bunch is successfully picked.
[0168] In this embodiment, the system integrates advanced sensing, image processing, machine learning, and robotic control technologies to create an efficient, precise, and intelligent fruit and vegetable picking system. Its key advantages lie in its high-precision bunch identification and positioning, occlusion management, entanglement risk assessment, path optimization, and automated picking. This system facilitates intelligent and automated harvesting in large-scale agricultural production, reducing labor costs, improving production efficiency, and ensuring harvest quality.
[0169] It should be noted that all calculation formulas in this application document utilize, including but not limited to, regression analysis within machine learning algorithms to deeply analyze the collected parameters and identify their natural trends and interrelationships. Professional software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Model performance is then objectively evaluated through methods such as cross-validation, combined with continuous feedback and optimization to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their validity and accuracy, and ensuring that the calculation process complies with the constraints of natural laws rather than being based on artificially set rules.
[0170] The technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0171] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0172] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
[0173] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A picking target modeling method for a fruit and vegetable picking robot based on the Internet of Things, characterized in that: The following steps are involved: Step 1: Use multi-view IoT vision nodes to collect three-dimensional images of the target fruit bunch and construct an RGB-D multimodal time series data set; Step 2: Based on the fusion processing of the depth map and color map in the RGB-D multimodal time series data set, the preliminary three-dimensional outline of the fruit bunch is extracted, and the fruit structure domain, stalk domain, leaf extension domain, tendril extension domain and background boundary domain are distinguished by spatial clustering method to obtain domain point cloud data, and the first main axis and second reference line are identified; Step 3: Use convolutional neural network technology (CNN) to build an occlusion reconstruction model. After training the occlusion reconstruction model using domain point cloud data, mark the connection point area of the first main axis and the second reference line in the occlusion reconstruction model, and identify whether there is an intersection with the leaf extension domain. If a first intersection area exists, collect the percentage of the area of the first intersection area to the total area of the connection point area and the leaf occlusion depth difference value to generate the corresponding cut point candidate set; Step 4: Extract the identified tendril extension domain of each target fruit bunch and combine it with the cut point candidate set to identify the degree of entanglement between the tendril extension domain and the cut point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cut point based on the spatial tendril density matrix. j And evaluate, finally obtain the corrected actual picking and cutting point, generate the final picking instruction and execute it.
2. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 1, characterized in that: The step one comprises: S11. Within the orchard operation area, deploy multiple nodes supporting RGB-D IoT visual information collection based on the distribution characteristics of fruit trees and the hanging height of fruit bunches. Each node includes an RGB camera, a ToF or structured light depth sensor, a wireless communication device, and an embedded controller. S12. After each node is deployed and the view angle calibration operation is performed between each node, a unified acquisition instruction is issued. Each node synchronously acquires the color image and depth image under the current view angle and records the timestamp. S13. Pack the color images and depth images of the target fruit bunch collected within a continuous time period to generate an RGB-D multimodal time series data set containing multi-node, multi-view, and multi-frame information.
3. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 2, characterized in that: The second step includes: S21. After pixel-level registration, denoising, and conversion of the color image and depth image in the RGB-D multimodal time series data set into a unified format, each pixel is projected into three-dimensional space to form domain point cloud data. The point cloud data is then imported into the point cloud modeling tool CloudCompare, Open3D, or PCL to generate a string of three-dimensional point cloud models. S22. In the three-dimensional point cloud model of the fruit bunch, based on the actual production form of the fruit bunch, extract target point cloud data from the three-dimensional point cloud model, set a domain radius and a minimum number of domain points, group adjacent point clouds into clusters, obtain several point clusters to be identified, and distinguish the fruit structure domain, the fruit stalk domain, the leaf extension domain, the tendril extension domain, and the background boundary domain using a spatial clustering method; S23. Perform PCA analysis on the fruit-particle domain and the fruit-stalk domain respectively to extract the first principal axis of the fruit-particle domain and the second baseline of the fruit-stalk domain.
4. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 3, wherein: S22 includes: S201, the steps of fruit grain domain identification are as follows: S2011, perform principal component analysis (PCA) on each point cluster to be identified, and calculate the ratio of the main axis length to the secondary axis length R asxis : L major is the length of the first principal axis, L minor is the secondary axis length; when 1.0≤R asxis When ≤1.4, the label is "first fruit label"; S2012, converting all RGB values in the point cluster into HSV color space, and calculating the average hue Hue, saturation S and average brightness V of all point clusters to be identified; When the average hue is Hue∈[280°,340°], the fruit is purple-red; when the saturation S is ≥0.5, the saturation is acceptable; when the average brightness V is in the range of [0.2,0.8], the fruit has a moderate brightness, neither too reflective nor too dim. The judgment formula is: Hue∈[280°,340°]&S≥0.5&V∈[0.2,0.8]. When the ranges of average hue, saturation S, and average brightness V are all met, the label is "second fruit label"; S2013. Extract each point in each cluster and calculate its average Euclidean distance Davg and point density P with the 10 nearest points. Point density P is the ratio of cluster volume V to number of points N. When Davg ≤ 3.5 mm & P ≥ 1.8 × 10 3 points / cm 3 , the label is "Third Fruit Label"; S2014, identifying the point clusters of the "first fruit particle label", "second fruit particle label" and "third fruit particle label" obtained simultaneously in S2011-S2013, and determining them as fruit particle structure domains; S202, the method for obtaining the fruit stalk domain is: S2021. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When ≥5, the label is "first fruit stalk label"; S2022. After color feature extraction, Hue∈[10°,60°]&S≥0.3&V∈[0.3,0.7]. When the ranges of average hue Hue, saturation S, and average brightness V are satisfied at the same time, the label is "second fruit stalk label". S2023, when Davg≤4.5mm&P≥1.5×103points / cm 3 , labelled as "third stalk label"; S2024. Identify the point clusters that simultaneously obtain the "first country handle label", "second country handle label" and "third country handle label" in S2021-S2023, and determine them as the fruit handle domain.
5. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 3, characterized in that: The S22 also includes: S203, the method for obtaining the blade extension domain is: S2031. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When <3, it indicates that the point cluster is a relatively flat structure, which conforms to the morphology of the leaf extension domain and is labeled as "first leaf label"; S2032. After color feature extraction, Hue∈[90°,150°]&S≥0.4&V∈[0.4,0.9]. When the ranges of average hue Hue, saturation S, and average brightness V are satisfied at the same time, the label is "second leaf label". S2033, when Davg≤5mm&P≥1.2×103points / cm 3 , labelled "third leaf label"; S2034, identifying the point cluster of the "first leaf label", "second leaf label" and "third leaf label" obtained simultaneously in S2031-S2033, and determining it as the leaf extension domain; S204. The method for obtaining the tendril extension domain is: S2041. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the major axis length to the minor axis length. asxis :When R asxis When ≥10, it indicates that the point cluster is an elongated and curved structure, which is consistent with the morphology of the tendril extension domain and is labeled as "first tendril label"; S2042. After color feature extraction, Hue∈[0°,50°]&S≥0.5&V∈[0.2,0.7]. When the ranges of average hue Hue, saturation S, and average brightness V are satisfied at the same time, the label is "second tendril label". S2043, when Davg≤2.5mm&P≥2.0×103points / cm 3 , labeled "third tendril tag"; S2044, identifying the point cluster of the "first tendril label", "second tendril label" and "third tendril label" obtained simultaneously in S2041-S2043, and determining it as the tendril extension domain; S205: The background boundary region is obtained as follows: S2051. Perform principal component analysis (PCA) on each point cluster to be identified and calculate the ratio R of the main axis length to the secondary axis length. asxis :When R asxis When <1.5, the point cluster has a relatively uniform structure, which is consistent with the characteristics of the background area and is labeled as "first background label"; S2052. After color feature extraction, S≤0.3&V∈[0.1,0.6]. When the range of average hue Hue, saturation S and average brightness V is satisfied at the same time, the label is "second background label". S2053: Identify and simultaneously obtain the "first background label" and "second background label" point clusters in S2051-S2052, and determine them as background boundary regions.
6. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 4, characterized in that: Step three includes: S31, extracting the boundaries of the identified fruit structure domain, fruit stalk domain, leaf extension domain, tendril extension domain, and background boundary domain of each target fruit bunch, and using Boolean operations to calculate the size of the intersection area; S32. Construct an occlusion reconstruction model using convolutional neural network technology (CNN). After training the occlusion reconstruction model using domain point cloud data, mark the connection point area between the first principal axis and the second reference line in the occlusion reconstruction model. When a first intersection area is identified within the connection point area, it is determined that the connection point area is not obstructed by leaves and is suitable for direct shearing. Points selected every 1 cm within 10 cm of the connection point area are used as a first positioning shearing point group. S33. When a first intersection region is identified within the connection point, the area of the first intersection region within the connection point region is calculated. When the area of the first intersection region accounts for 10%-50% of the total area of the connection point region, it indicates that there is a risk of leaf occlusion in the shear point determination. The maximum leaf occlusion depth in the intersection occlusion region and the absolute value of the depth of the connection point are further identified to obtain a leaf occlusion depth difference value. When the blade obstruction depth difference exceeds 10 cm, it indicates that the blade obstruction is far from the connection point. Within a safe range, points within 10 cm of the central axis near the connection point and outside the blade obstruction, but not in the first intersection area, are selected as the second positioning shear point group. If the blade obstruction depth difference does not exceed 10 cm, it indicates that the blade obstruction is close to the connection point, and there is a risk of blade obstruction interference, which may lead to miscutting. The fruit bunch is marked as "First Cutting Risk Fruit Bunch"; S34. When the first intersection area is identified within the connection point, the area of the first intersection area around the connection point is calculated. When the area of the first intersection area accounts for more than 50% of the total area of the connection point area, it indicates that there is a risk of blade occlusion for the shear point judgment, and it is marked as the "second shear risk fruit cluster".
7. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 5, characterized in that: The step three also includes: S35. After identifying the "first shear risk fruit bunch" and the "second shear risk fruit bunch," locate the length of the second reference line in the fruit stem region. If the length of the second reference line is greater than 15 cm, select points within 5 cm to 10 cm of the second reference line in a direction extending from the end of the fruit stem toward the first principal axis, with intervals of 1 cm between the two points, as the third located shear point group. If the length of the second reference line is within 5 cm-15 cm, then select points at intervals of 1 cm within 5 cm from the end of the fruit stem in the direction extending from the first main axis as the fourth positioning shearing point group; If the length of the second reference line is 5 cm, select points within 2 cm to 3 cm of the end point of the fruit stem on the second reference line at intervals of 1 cm as the fifth positioning shearing point group; The first positioning shear point group, the second positioning shear point group, the third positioning shear point group, the fourth positioning shear point group and the fifth positioning shear point group are summarized to form a shear point candidate set.
8. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 6, wherein: The fourth step includes: S41. Extract the identified tendril extension domain of each target fruit bunch and combine it with the cut point candidate set to identify the degree of entanglement between the tendril extension domain and the cut point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cut point based on the spatial tendril density matrix. j : S42, preset risk threshold X, when the entanglement risk index R of the j-th shear point j If the risk threshold X is exceeded, it means that the cutting point has entanglement risk and the cutting point is removed from the cutting point candidate set; When the winding risk index R of the j-th shear point j If the risk threshold X is not exceeded, it means that there is no entanglement risk at the shear point, and the entanglement risk index of each shear point that does not exceed the risk threshold X is calculated, and the entanglement risk index R of the j-th shear point is calculated. j The cutting point with the lowest value in is marked as the corrected actual picking cutting point, and the final picking instruction is generated and executed.
9. The method for modeling a picking target for a fruit and vegetable picking robot based on the Internet of Things according to claim 8, characterized in that: S41 includes: S411. In the three-dimensional point cloud model, a second intersection region of the tendril extension domain and the set of candidate shear points is calculated. The second intersection region is the portion where the tendril is wrapped. The point cloud data of the tendril extension domain is divided into U voxels in the second intersection region using a spatial partitioning algorithm. The size of each voxel is set to 1x1x1 cm. Each point cloud data point of the tendril extension domain is assigned to a corresponding voxel according to its spatial coordinates to generate a spatial tendril density matrix. S412: The density value of each voxel is the number of point clouds within the voxel, and the density value of each voxel is normalized. The entanglement risk index R of the j-th shear point in the shear point candidate set is calculated using the following formula: j : Where U is the number of voxels in the tendril extension domain in the second communication area, E k is the normalized density value of the i-th voxel, indicating the number of point clouds within the k-th voxel, d k is the Euclidean distance from the kth voxel to the jth shear point.
10. A fruit and vegetable picking robot picking target modeling system based on the Internet of Things, applied to the fruit and vegetable picking robot picking target modeling method based on the Internet of Things according to any one of claims 1 to 9, characterized in that: include: The image acquisition module uses a multi-view IoT vision node to collect three-dimensional images of the target fruit bunch and construct an RGB-D multimodal time series data set; The 3D point cloud model building module is used to perform pixel-level registration, denoising, and conversion of color and depth images in the RGB-D multimodal time series data set into a unified format. Each pixel is then projected into 3D space to form domain point cloud data. The point cloud data is then imported into point cloud modeling tools such as CloudCompare, Open3D, or PCL to generate a string of 3D point cloud models. The structure decomposition module is used to extract the preliminary three-dimensional outline of the fruit bunch based on the fusion processing of the depth map and color map in the RGB-D multimodal time series data set. It uses spatial clustering methods to distinguish the fruit structure domain, stalk domain, leaf extension domain, tendril extension domain and background boundary domain, obtain domain point cloud data, and identify the first main axis and the second reference line. The first correction module uses convolutional neural network technology (CNN) to build an occlusion reconstruction model. After training the occlusion reconstruction model using domain point cloud data, it marks the connection point area of the first main axis and the second baseline in the occlusion reconstruction model and identifies whether there is an intersection with the leaf extension domain. If a first intersection area exists, the percentage of the area of the first intersection area to the total area of the connection point area and the difference in leaf occlusion depth are collected to generate a corresponding set of cut point candidates. The second correction module is used to extract the identified tendril extension domain of each target fruit bunch and combine it with the cutting point candidate set, identify the degree of entanglement between the tendril extension domain and the cutting point candidate set, construct a spatial tendril density matrix, and construct the entanglement risk index R of the j-th cutting point based on the spatial tendril density matrix. j And evaluate, finally obtain the corrected actual picking and cutting point, generate the final picking instruction and execute it; The robot control module is used to guide the robotic arm to perform motion control based on the position information of the target fruit bunch and the shearing point of the final picking instruction, including translation, rotation, extension and shearing of the robotic arm until the fruit bunch is successfully picked.
Citation Information
Cited By
Miniaturized unmanned aerial vehicle mounting type high branch and leaf intelligent sampling method and system, and storage medium
CN121353943A