A dual-arm cherry tomato harvesting robot and its harvesting method

CN122539404APending Publication Date: 2026-08-11ZHONGKAI UNIV OF AGRI & ENG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,圣女果果实个体小、成簇(串状)生长、成熟期不一致且果皮脆弱易损的特性,使其规模化高效采摘面临严峻挑战

Benefits of technology

本发明通过维纳滤波抑制启停运动模糊、光流法剔除动态阴影干扰、贝塞尔曲线补全断线、曲率自适应权重优化航向,能够在光照波动、叶片阴影遮挡、垄道线断续等场景下稳定导航,有效提升复杂光照下导航线识别准确率与航向控制精度,减少偏垄、脱线故障。同时,本发明融合CoT模块的语义分割提升了细长结果枝的检出率,减少果梗漏检;椭球先验谱聚类解决了密集串生果实的粘连分割难题,提升单果分割准确率;偏振辅助的立体匹配让果梗区域深度精度提升,结合主茎几何约束实现采摘点的精确定位,满足精准剪切的需求,避免串生场景的干扰。此外,本发明兼顾“色泽外观”与“弹性模量(硬度)”双维度的成熟度分级,有效排除果皮反光、转色期果实的误判,精准筛选出硬度适中、适合储运的适采果实,降低过熟软果的采摘损耗,提升商品果一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122539404A_ABST
    Figure CN122539404A_ABST
Patent Text Reader

Abstract

This invention provides a dual-arm cherry tomato harvesting robot, relating to the field of intelligent harvesting, including a wire-controlled chassis (10), dual robotic arms (20), an end effector (30), a detection and positioning system, and a navigation system. The wire-controlled chassis (10) includes a chassis frame (11), a shock absorption system, an Ackermann steering system, and a rear-drive transmission power system. The dual robotic arms (20) are set on the end face of the chassis frame (11), including a fixed bracket (201) and two sets of robotic arm structures. An end effector (30) is set at the end of the robotic arm structure. The end effector includes a motor connector (31), a rotary motor (32), a base plate (33), a stationary blade (34), an insulating base (35), a moving blade (36), and a transmission servo motor (37). This harvesting robot can automatically detect and position cherry tomatoes, thereby achieving precise clamping, cutting, and collection of cherry tomatoes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent harvesting technology, specifically to a dual-arm cherry tomato harvesting robot and its harvesting method. Background Technology

[0002] Cherry tomatoes, a popular nightshade vegetable, are not only unique in flavor but also rich in vitamins, lycopene, and various antioxidants, offering significant nutritional value for human health, such as boosting immunity and reducing oxidative stress. However, the small size, clustered growth, uneven ripening, and fragile skin of cherry tomatoes pose significant challenges to large-scale, efficient harvesting. In recent years, with the migration of rural labor to cities and the continuous rise in agricultural labor costs, the shortage of harvesting labor during the harvest season has become increasingly prominent. This makes timely and precise harvesting impossible, leading to overripe and rotten fruit, decreased quality, and yield losses.

[0003] Traditional manual harvesting methods rely mainly on workers' visual identification and manual operation, which has significant drawbacks: First, it requires highly experienced harvesters who need to accurately judge the maturity (color, firmness) of each small fruit and whether it is easy to separate from the stem, which is greatly affected by human factors and inconsistent standards; Second, in long-term, high-intensity repetitive work, workers are prone to visual fatigue and decreased attention, which leads to a significant reduction in the accuracy of locating small fruits and judging their maturity, seriously affecting harvesting efficiency and fruit integrity rate; Third, in densely packed environments, manual judgment and operation are prone to misharvesting (unripe fruits), missing harvesting (hidden fruits), and mechanical damage to fruits caused by squeezing and scratching during the harvesting process, directly affecting commercial value and economic benefits.

[0004] In addition, in order to cope with the labor shortage, some automated picking equipment or single-arm robots used in existing technologies do not perform well in dealing with the complex scenarios of cherry tomatoes: they generally have problems with low recognition accuracy and inaccurate positioning of small fruits in dense, clustered, and shaded environments, making it difficult to achieve millimeter-level precision picking; at the same time, single picking strategies or end effector designs often lead to improper control of gripping force, which can easily damage the delicate fruits or break the fruit stems, and the rates of missed picking and mispicking remain high. Summary of the Invention

[0005] To address the problems existing in the prior art, the present invention aims to provide a dual-arm cherry tomato harvesting robot that can automatically detect and locate cherry tomatoes, thereby achieving precise gripping, cutting, and collection of cherry tomatoes.

[0006] Another objective of this invention is to provide a harvesting method for the above-mentioned harvesting robot, which has high positioning accuracy, good harvesting consistency, and high harvesting efficiency.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] A dual-arm cherry tomato harvesting robot includes a wire-controlled chassis, dual robotic arms, an end effector, a detection and positioning system, and a navigation system. The wire-controlled chassis includes a chassis frame, a shock absorption system, an Ackermann steering system, and a rear-drive power system. The navigation system is integrated inside the chassis frame and includes a vision system and a central control device. The dual robotic arms are mounted on the end face of the chassis frame and include a fixed bracket and two sets of robotic arm structures. The fixed bracket is located on the end face of the chassis frame, and the two sets of robotic arm structures are symmetrically arranged on it. An end effector is located at the end of the robotic arm structure (i.e., the end away from the chassis frame). The end effector includes a motor connector, a rotary motor, a base plate, a stationary blade, an insulating base, a moving blade, and a transmission servo motor. The detection and positioning system includes a binocular camera and a central control device. The binocular camera is mounted on the robotic arm structure and located directly above the end effector. The central control device is located inside the chassis frame.

[0009] Based on further optimization of the above scheme, the chassis frame is a frame structure; the shock absorption system includes a shock absorption movable beam and a shock absorption spring, the shock absorption movable beam is hinged to the chassis frame via a pin, and the shock absorption spring is located between the shock absorption movable beam and the chassis frame; the Ackermann steering system is located at the front of the bottom surface of the chassis frame, including a steering joint, a steering tie rod and a front wheel axle, and uses a hydraulic or electric push rod to drive the steering tie rod to achieve steering; the rear wheel drive power system is located at the rear of the bottom surface of the chassis frame, including a rear wheel drive motor and a differential transmission, driving the rear wheels to rotate and achieve differential driving.

[0010] Based on further optimization of the above scheme, the robotic arm structure includes an X-axis main guide rail, two X-axis auxiliary guide rails, a Y-axis guide rail, a vertical guide rail, a cantilever, and a fruit collection mechanism. The two X-axis auxiliary guide rails are arranged on both sides of the corresponding X-axis main guide rail, and the X-axis auxiliary guide rails and the X-axis main guide rail are arranged in parallel. Both the X-axis auxiliary guide rails and the X-axis main guide rail are located on the end face of the fixed bracket. The Y-axis guide rail is slidably arranged on the end face of the X-axis main guide rail and the two X-axis auxiliary guide rails and is driven by an X-axis drive motor arranged on the X-axis main guide rail. The vertical guide rail is slidably arranged on the end face of the Y-axis guide rail and is driven by a Y-axis drive motor arranged on the Y-axis guide rail. A cantilever is slidably arranged on the vertical guide rail, and an end effector is arranged at the end of the cantilever. A vertical lifting motor for controlling the lifting and lowering of the cantilever is arranged on the vertical guide rail. A fruit collection mechanism is arranged at the end of the cantilever and is located below the end effector.

[0011] Based on further optimization of the above scheme, the motor connecting seat is fixedly installed at the end of the cantilever and a rotary motor is installed inside it; the output shaft of the rotary motor is connected to the base plate through the end connecting flange (for driving the end actuator to rotate as a whole); a stationary blade and an insulating base are installed on the base plate, and the moving blade is rotatably connected to the stationary blade through a rotating shaft; the transmission servo is fixedly installed below the base plate, its output shaft passes through the base plate and is fixedly sleeved with the drive gear on the upper side of the base plate, and a sector gear is fixedly installed at the end of the moving blade through an anti-loosening nut and the sector gear meshes with the drive gear (thereby driving the moving blade to rotate and realize the shearing action).

[0012] Based on further optimization of the above scheme, an upper shell and a lower shell are respectively provided on the upper and lower end faces of the base plate for mechanism protection.

[0013] Based on further optimization of the above scheme, the inner side of the front end of the moving blade is provided with clamping foam, which, together with the foam stop block fixed on the moving blade, completes the flexible clamping of the fruit stem while the moving blade closes to cut the fruit stem.

[0014] A method for harvesting cherry tomatoes using a dual-arm robotic cherry tomato harvesting robot, comprising: Step S1, Navigation Path Generation: Generate a navigation path based on the navigation system, drive the wire-controlled chassis to control the picking robot to reach the picking location, and trigger the picking operation; Step S2, Picking point identification: Trigger the detection and positioning system to acquire images, detect fruit maturity and obtain picking points, and convert the picking points to the robot arm base coordinate system; Step S3, Harvesting Control: The robotic arm executes the motion plan and drives the end effector to perform clamping and shearing harvesting; Step S4, Fruit Collection and Mechanism Reset: After the fruit is picked, the end effector transports the fruit to the fruit collection mechanism and resets the end effector and robotic arm structure. Step S5, Cyclic Operation: Based on the fruit maturity detection results within the field of view, if there are still mature fruits, repeat steps S2 to S4; if there are no mature fruits, trigger the chassis to move and repeat steps S1 to S5 until the end of the ridge is reached.

[0015] Based on further optimization of the above scheme, in step S1, generating the navigation path based on the navigation system specifically involves: Step S11, Trigger Image Acquisition: After the chassis speed returns to zero and vibration decays, the navigation system (monocular camera) is triggered to acquire images, and Wiener filtering is used to restore slight motion blur. ; In the formula: (u, v) Represents frequency domain spatial coordinates; Represents the frequency domain estimation of the restored image; G(u, v)The frequency domain represents the degraded image; H(u, v) Indicates the degenerate transfer function; H (u, v) Represents the complex conjugate of the degenerate transfer function; S n (u, v) Represents the noise power spectral density. S f (u, v) Represents the power spectral density of the original sharp image; Represents the regularization coefficient; Output clear, unblurred original navigation image; Meanwhile, as the chassis travels along the ridge, the monocular camera continuously acquires images at a fixed frame rate and directly outputs the original navigation images; Step S12, Image Preprocessing: First, convert the original navigation image to grayscale: ; In the formula: These represent the pixel values ​​for the red, green, and blue channels, respectively. Then, perform non-linear grayscale correction on the grayscale image: ; In the formula: V in (x, y) Indicates the grayscale value of the input pixel; Indicates the correction index; Finally, iterate through every pixel in the entire image, calculating the local Gaussian mean and local threshold sequentially to complete the binary classification: With the current pixel (x, y) Select size as the center. w × w Generate a Gaussian weight kernel from the square neighborhood: ; In the formula: Indicates the Gaussian standard deviation; Obtain Gaussian weighted local mean Mean(x, y) With local threshold Thy(x, y) : ; ; In the formula: C py Indicates an offset constant; like V out (x, y) < Thy(x, y) If it is determined to be a navigation line (foreground), the pixel value is assigned to 0 (black); if Vout (x, y) ≥ Thy(x, y) If the value is 255 (white), then it is determined to be the background and the pixel value is assigned to 255; connect the black connected components to generate the initial navigation line binary image; Step S13: Based on the motion consistency characteristics of the optical flow field, distinguish between static navigation lines and dynamic shadow artifacts: First, using two consecutive grayscale images as input, employ... Farneback Dense optical flow algorithm calculates the optical flow vector for each pixel in an image. u g ,v g This process obtains the dense optical flow field of the entire image. Then, all black connected components in the binary image are masked out, and the remaining region is used as a static background sample. The median of the optical flow vectors of all pixels within the static background sample region is taken to obtain the global motion vector. v global ,like If the output is positive, directly output the original binary image; otherwise, perform 8-connected component detection on the initial navigation line binary image to obtain all independent black connected regions. B 1 ,B 2 ,…,B n And record the set of pixel coordinates for each region; finally, for each connected component B i Extract the optical flow vectors of all pixels within the region and obtain the region's average motion vector. v Bi And calculate the motion consistency score: ; Preset motion difference judgment threshold Thm ,like S motion-i > Thm If the value is 0, it is determined to be a dynamic shadow; all pixels within that connected component are set to white and removed from the binary image. S motion-i ≤ Thm If so, it is determined to be a real navigation line and the area is retained; Step S14: First, based on the dimensions and height of the binary image... h ez The navigation line is divided into three independent sampling regions based on spatial distance: distant ROI ∈ [0, ... h ez / 3), mid-range ROI∈[ h ez / 3,2 h ez / 3), near-field ROI∈[2 h ez / 3, hez -1); Then, for each ROI region, the lateral centroid coordinates of the navigation line foreground within that region are calculated using the geometric moment method: foreground centroid x near Mid-field center x mid Visionary core x far Then, the curvature of the navigation line is calculated using the centroids of the three regions, and the weights of the distant and near views are updated adaptively. Finally, the navigation target point is calculated using weighted averages to generate the heading deviation. Step S15, Line Break Completion: If the number of foreground pixels in a certain area is less than 10, the centroid value of the corresponding area in the previous frame is temporarily used for preservation, and subsequent Bézier curve line break completion is triggered; if there are no valid pixels for 3 consecutive frames, it is marked as a navigation line break; the navigation line break is based on the historical tangent direction and uses Bézier curve completion: ; In the formula: t bes Indicates the parameters of the normalized curve; P bes0 Indicates the starting control point; P bes3 Indicates the endpoint control point; P bes1 , P bes2 These represent the tangent control points at the start and end points, respectively. Output the completed continuous navigation lines and the corrected heading deviation; Step S16: The visual pixel deviation is converted into the physical error of vehicle dynamics through "pixel-physical scale conversion". Then, a model predictive controller (MPC) is constructed based on the linear two-degree-of-freedom error dynamics model to output the target front wheel steering angle command of the Ackerman steering mechanism to achieve steering control.

[0016] Based on further optimization of the above scheme, step S2 specifically includes: Step S21, Image Acquisition and Preprocessing: Synchronous RGB images at three polarization angles (0°, 45°, and 90°) are acquired using the binocular camera of the detection and positioning system. The images at these three polarization angles are then converted to grayscale and subjected to Gamma correction. The grayscale value sequence of each pixel at different polarization angles is calculated, and the degree of polarization of each pixel is obtained based on the grayscale response at different polarization angles. ; In the formula: DOLP(p) Represents pixels p degree of polarization; I max (p) , Imin (p) Representing pixels p Maximum and minimum gray values ​​at three polarization angles; Finally, we obtained a three-angle polarization grayscale image, a pixel-level polarization map, and a color original image after Gamma correction; Step S22, Semantic Segmentation: Based on the improved DeepLabV3 model, pixel-level semantic segmentation of fruit-fruiting branch-main stem is achieved, and a mature fruit mask is output. M fruit , resultant branch mask M branch Main stem cover M stem ; The improved DeepLabV3 model uses a lightweight MobileNetV3-large as the backbone feature extraction network for the encoder, and introduces context Transformer modules at the output positions of its low-level and high-level features; the decoder adopts a multi-scale feature fusion strategy, fusing features from different levels enhanced by the CoT module; the model uses a stochastic gradient descent (SGD) optimizer combined with a cosine annealing learning rate decay strategy for iterative training, and the loss function is cross-entropy loss; Step S23, Point Cloud Completion: First, construct a multi-constraint matching cost that integrates grayscale, structure, and polarization, then perform cost aggregation and disparity optimization. ; In the formula: These represent the corresponding weight coefficients; This represents the total matching cost of pixel p under disparity d; This represents the cost of the absolute difference in grayscale values; Indicates the cost of the Census transformation; Then, extract the cloud of fruit stalk skeleton points { S i The corresponding effective parallax is d i For any invalid point within the mask p Find the adjacent projection points on its skeleton. S k , S k+1 Perform interpolation completion: ; Output high-precision disparity map and original 3D point cloud in camera coordinate system P raw ; Step S24, Clustering and Segmentation: For any point pi Calculate its to k The average distance between the nearest neighbors global mean Standard deviation is ,like If so, it will be discarded as a noise point; Then, a spectral clustering algorithm incorporating prior knowledge of cherry tomato ellipsoidal shapes is used to obtain a set of independent fruit point cloud clusters. C fruit} and the main stem point cloud cluster set { C stem}; Step S25, Maturity Grading: Based on the set of independent fruit point cloud clusters { C fruit The point cloud of a single fruit is projected onto the image plane, and a semantic mask is used to obtain the precise pixel region (ROI) of the fruit. For each pixel within the ROI, a reflected light intensity equation is fitted based on the gray values ​​of three polarization angles to solve for the diffuse reflection and specular reflection components. The diffuse reflection component is used to reconstruct the pure natural color image of the fruit, eliminating highlights on the peel and ambient reflections. ; In the formula: I d Indicates the intensity of diffuse reflection. I s Indicates the intensity of light reflected from the specular surface; Indicates the principal polarization direction angle of specular reflection; Obtain pixel grayscale values ​​at three polarization angles I(0 0 )、I(45 0 )、I(90 0 ) ; Then, the mean polarization degree of a single fruit region is extracted. DOLP mean Polarization degree variance DOLP std Mean of specular diffuse reflectance R s_mean Three features were used, and these features were input into an offline-calibrated second-order polynomial regression model to obtain the fruit's elastic modulus. E tan : ; In the formula: These represent the regression coefficients calibrated offline; like E tan If the pressure is >0.8 MPa, it is considered an unripe, hard fruit; if it is ≤0.3 MPa... E tanFruits with a pressure ≤0.8MPa are considered suitable for harvesting. E tan If the pressure is less than 0.3 MPa, it is considered an overripe, soft fruit. Finally, the diffuse RGB image is converted to the Lab color space, and the color is extracted. a Regional mean of channel (red-green hue) As the core feature of color maturity; constructing feature vectors Input the pre-trained LightGBM lightweight classifier, the classifier outputs three probabilities: "unripe", "suitable for picking", and "overripe", and take the class with the highest probability and a confidence of not less than 0.8 as the final grade of a single fruit; Step S26: Obtain the picking point: Based on the fruit point cloud clusters and corresponding main stem point cloud clusters after maturity grading in Step S25, calculate the centroid and covariance matrix of the bounding box point cloud: ; Perform eigenvalue decomposition on the covariance matrix Cov to obtain three eigenvalues ​​sorted in descending order, and three corresponding orthogonal eigenvectors. ezj 1. ezj 2. ezj 3; The three eigenvectors form an orthogonal local coordinate system, that is, the three axes of the smallest bounding box; Project all centered points onto the three principal axes respectively, obtain the minimum and maximum projection values ​​on each axis, and determine the bounding box boundary: for any point p i In the j The projection values ​​on each principal axis are: Traverse all points to obtain the projected extrema of the three principal axes: ; [ min j , max j That is, the bounding box in the first... j The range on each main axis; Calculate the major axis (e.g.) ezj 1 is the geometric center of the two end faces (where the major axis is 1): ; Obtain the center of the upper surface of the fruit enclosure box P a : ; In the formula: This represents the centroid of the main stem dot cloud; Based on the long axis of the main stem ezj s1Obtain the center of the upper end face of the main stem: ; Then, calculate the direction vector from the centroid of the main stem to the centroid of the fruit: ; Distance from the center of the top surface to the boundary of the bounding box d u-b for: ; The midpoint of the upper border of the main stem enclosure box closest to the fruiting branch is: ; Finally, the picking location was obtained: ; Will P pick Convert to picking point coordinates in camera coordinate system P cam .

[0017] Based on further optimization of the above scheme, in step S24, the set of independent fruit point cloud clusters is obtained { C fruit} and the main stem point cloud cluster set { C stem Specifically: First, for each point p i Extract its k p The standard ellipsoid equation is fitted to the neighborhood point cloud using the RANSAC robust fitting algorithm, with the nearest neighbor points as the reference points. ; In the formula: ( x 0 ,y 0 ,z 0) represents the coordinates of the ellipsoid center; a ty 、b ty 、c ty These represent the semi-major axes of the ellipsoid. Reconstruct the ellipsoidal prior affinity matrix: ; In the formula: e ij This represents the average of the fitting residuals between two points; Indicates spatial distance bandwidth; This represents the bandwidth of the ellipsoid fitting residual; Obtain the construction degree matrix: Simultaneously, it is constructed as a normalized Laplace matrix: L = D -W ; Laplace matrix L Eigenvalue decomposition is performed to obtain n eigenvalues ​​sorted from smallest to largest, along with their corresponding eigenvectors. Then, the number of clusters is automatically determined using the eigenvalue gap method or silhouette coefficient method, and the eigenvectors corresponding to the k smallest eigenvalues ​​are selected and concatenated column-wise to form an n×k dimensional feature matrix. ; feature matrix Each row is treated as a k-dimensional sample point. K-means clustering is performed on all row vectors, resulting in the number of clusters. k Based on the clustering results, the original 3D point cloud was divided into... k There are three independent point cloud clusters, each corresponding to a single cherry tomato fruit. After removing tiny clusters with too few points, output the set of independent fruit point cloud clusters. C fruit}; Set a clustering distance threshold dth; traverse the point cloud, and if the spatial distance between two points is less than dth, they are grouped into the same cluster; finally, m independent stem point cloud clusters are obtained, and clusters with too few points and small noise points are simultaneously removed, outputting a set of independent stem point cloud clusters { C stem}

[0018] Based on further optimization of the above scheme, step S3 specifically includes: Step S31: Input the coordinates of the picking point P cam Based on the offline calibrated full-stroke error table, the systematic error of the current position is obtained through cubic spline interpolation (taking the X-axis as an example, and the Y-axis and Z-axis are similar; the basic pulse equivalent of the Z-axis is 128 pulses / mm, that is, replacing 320 in the correction pulse number with 128): ; In the formula: M i , M i+1 These represent the second derivative values ​​at the endpoints of the segment; error i , error i+1 They represent the calibration points respectively. x i , x i+1 The measured error; Through systematic errors error(x) Corrected pulse count: ; In the formula: K x Indicates the pulse correction coefficient; When a reverse motor direction is detected, a gap compensation pulse is superimposed: ; In the formula: This represents the measured backlash value; This represents the reciprocal of the pulse equivalent. Direction reversal criterion: , v(t) express t The speed of the motor at all times; Final output compensation target pulse number: ; In the formula: flag res This indicates the inversion flag; it is 1 when inverting and 0 otherwise. Step S32: Employ modal orthogonal decoupling and multimodal input shaping to eliminate mechanical vibration from the command end. Modal decoupling: ; In the formula: m i Indicates the first i Equivalent mass of the first mode, g i Indicates the first i First-order modal equivalent stiffness, c i Indicates the first i First-order mode equivalent damping; Let represent the vibration amplitude, modal velocity, and modal acceleration of the i-th mode, respectively. f i Indicates the first i Modal generalized excitation force For each mode, a zero-vibration (ZV) input shaper is designed: For the i-th vibration mode, the shaper consists of two impulse pulses, triggered at times t1 and t2 respectively, with amplitude weights A1 and A2. ; ; In the formula: Indicates the first i First-order modal damping ratio; T i Indicates the first i First-order modal natural period Output the decoupled and shaped speed planning instructions for each axis; Step S33: First, perform FOC vector current decoupling and current loop closed-loop control: First, perform three-phase current sampling and Clarke transformation (3s→2s stationary frame transformation): for the three-phase stator current of the motor I a 、I b 、I c After sampling, the current is converted into a two-phase stationary coordinate system current through Clarke transformation, thus eliminating three-phase coupling. ; Then perform the Park transformation: combine the rotor electrical angle feedback from the magnetic encoder. The static coordinate system current is transformed into a dq coordinate system that rotates synchronously with the rotor (converting AC quantity into DC quantity): ; In the formula: I d , I q These represent direct-axis current (excitation component) and quadrature-axis current (torque component), respectively. Independent PID closed-loop regulation is applied to the d-axis and q-axis currents respectively, converting the current error into voltage commands for the corresponding axes: ; In the formula: Indicates the reference value of the d-axis current; This represents the q-axis current reference value obtained by superimposing the feedforward compensation amount of the DOB disturbance observer onto the q-axis current command; These represent the proportional gain, integral gain, and derivative gain along the d-axis, respectively. These represent the proportional gain, integral gain, and derivative gain along the q-axis, respectively. The voltage command in the rotating coordinate system is inversely transformed back to the stationary coordinate system, and then the PWM drive signal for the three-phase inverter is generated through space vector pulse width modulation (SVPWM). ; Then, the load disturbance torque is estimated in real time using the DOB disturbance observer: based on the nominal mechanical model of the motor, the difference between the control input and the actual position feedback is used to estimate the external load disturbance torque such as shear impact in real time. Motor nominal model: , J G Indicates the moment of inertia of the motor. B G Indicates the damping coefficient; Estimated disturbance: ; In the formula: Indicates the motor angle; u(s)Indicates the control quantity; Indicates the filter time constant; Then, the estimated disturbance torque is converted into quadrature-axis current compensation, which is added to the original current command to compensate in advance before the location of the disturbance, thereby achieving rapid impact suppression. ; In the formula: This indicates the original quadrature-axis current command output by the speed loop; This represents the external load disturbance torque estimated by DOB; Indicates feedforward gain; This indicates the final quadrature-axis current command after superposition compensation (sent to the current loop PID control). Finally, the system adopts a three-level cascaded closed-loop structure of position loop → velocity loop → current loop, with the response bandwidth increasing step by step from the outside to the inside; after the movement is completed, the main controller is notified asynchronously through the CAN bus to trigger the next picking action, thereby improving the system's collaborative efficiency. Step S34: Use token ring distributed scheduling for dual robotic arms.

[0019] Based on further optimization of the above scheme, step S34 specifically involves: First, calculating the fruit density on both sides and generating a dynamic stepping distance. ; In the formula: D base Indicates the reference step distance; These represent the fruit densities on the left and right sides, respectively. Indicates the standard density threshold; Then, token-based scheduling is used to prioritize assigning tasks to the idle side; Finally, the safe distance is adjusted in real time: ; In the formula: d safe Indicates the minimum safe distance; Trigger collision avoidance.

[0020] The following are the technical effects of the present invention: This invention employs Wiener filtering to suppress motion ambiguity during start-stop operations, optical flow to eliminate dynamic shadow interference, Bézier curve completion for broken lines, and curvature adaptive weighting to optimize heading. This enables stable navigation under conditions of fluctuating light, leaf shading, and discontinuous row lines, effectively improving the accuracy of navigation line recognition and heading control under complex lighting conditions, and reducing row deviation and deviation faults. Simultaneously, the semantic segmentation of the CoT module improves the detection rate of slender fruiting branches, reducing missed detections of fruit stalks; ellipsoidal prior spectrum clustering solves the problem of segmenting densely clustered fruits, improving the accuracy of single-fruit segmentation; polarization-assisted stereo matching improves the depth accuracy of the fruit stalk region, and combined with the geometric constraints of the main stem, achieves precise positioning of the picking point, meeting the needs of precise cutting and avoiding interference from clustered fruit. Furthermore, this invention considers both "color and appearance" and "elastic modulus (hardness)" for maturity grading, effectively eliminating misjudgments of fruits with reflective peels and those in the color-changing stage, accurately selecting fruits with moderate hardness suitable for storage and transportation, reducing harvesting losses of overripe soft fruits, and improving the consistency of marketable fruit.

[0021] In terms of the design of the harvesting robot, this invention avoids direct contact with the fruit pulp by using the flexible foam clamping and precise cutting of the end effector, which greatly reduces scratches and squeezing damage to the delicate fruit peel; it uses a DOB disturbance observer to quickly suppress shearing impact, and with input shaping and vibration suppression, it reduces secondary damage to the fruit caused by mechanical vibration; through the collaborative operation of the dual gantry arms and dynamic step scheduling, it effectively improves the harvesting efficiency per unit time and ensures that the left and right rows are covered without any omissions. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the overall structure of the harvesting robot in an embodiment of the present invention.

[0023] Figure 2 This is a schematic diagram of the wire-controlled chassis of the harvesting robot in an embodiment of the present invention.

[0024] Figure 3 This is a schematic diagram of the robotic arm structure of the harvesting robot in an embodiment of the present invention.

[0025] Figure 4 This is a schematic diagram of the end effector of the harvesting robot in an embodiment of the present invention.

[0026] Figure 5 This is a schematic diagram of the picking point detection and positioning method in an embodiment of the present invention.

[0027] Among them, 10. Wire-controlled chassis; 11. Chassis frame; 121. Shock-absorbing movable beam; 122. Shock-absorbing spring; 131. Steering joint; 132. Steering tie rod; 133. Front axle; 141. Rear wheel drive motor; 142. Differential transmission; 20. Dual robotic arms; 201. Fixed bracket; 21. X-axis main guide rail; 22. X-axis auxiliary guide rail; 23. Y-axis guide rail; 24. Vertical guide rail; 25. Cantilever; 26. Fruit collection mechanism; 30. End effector; 31. Motor connector; 32. Rotary motor; 33. Base plate; 331. Upper shell; 332. Lower shell; 34. Stationary blade; 35. Insulating base; 36. Moving blade; 361. Sector gear; 362. Clamping foam; 363. Foam stop; 37. Transmission servo; 371. Drive gear. Detailed Implementation

[0028] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0029] Example 1: A dual-arm cherry tomato harvesting robot includes a wire-controlled chassis 10, dual robotic arms 20, an end effector 30, a detection and positioning system, and a navigation system. The wire-controlled chassis 10 includes a chassis frame 11, a shock absorption system, an Ackermann steering system, and a rear-drive power system. The chassis frame 11 has a frame structure. The shock absorption system includes a shock-absorbing movable beam 121 and a shock-absorbing spring 122. The shock-absorbing movable beam 121 is hinged to the chassis frame 11 via a pin, and the shock-absorbing spring 122 is disposed between the shock-absorbing movable beam 121 and the chassis frame 11 (e.g., ...). Figure 2 (As shown); The Ackermann steering system is located at the front of the underside of the chassis frame 11, including a steering joint 131, a steering tie rod 132, and a front axle 133. Steering is achieved by hydraulically or electrically driving the steering tie rod. The rear-wheel drive system is located at the rear of the underside of the chassis frame 11, including a rear-wheel drive motor 141 and a differential transmission 142, driving the rear wheels to achieve differential driving (e.g., ...). Figure 2 (As shown).

[0030] The navigation system is integrated inside the chassis frame 11, including the vision system and central control equipment.

[0031] The dual robotic arms 20 are mounted on the end face of the chassis frame 11, including a fixed bracket 201 and two sets of robotic arm structures. The fixed bracket 201 is mounted on the end face of the chassis frame 11, and two sets of robotic arm structures are symmetrically arranged on it. The robotic arm structure includes an X-axis main guide rail 21, two X-axis auxiliary guide rails 22, a Y-axis guide rail 23, a vertical guide rail 24, a cantilever 25, and a fruit collection mechanism 26. The two X-axis auxiliary guide rails 22 are located on both sides of the corresponding X-axis main guide rail 21, and the X-axis auxiliary guide rails 22 and the X-axis main guide rail 21 are arranged in parallel. Both the X-axis auxiliary guide rails 22 and the X-axis main guide rail 21 are mounted on the fixed bracket 201. The Y-axis guide rail 23 is slidably mounted on the end face of the X-axis main guide rail 21 and the two X-axis auxiliary guide rails 22 and is driven by the X-axis drive motor mounted on the X-axis main guide rail 21. The vertical guide rail 24 is slidably mounted on the end face of the Y-axis guide rail 23 and is driven by the Y-axis drive motor mounted on the Y-axis guide rail 23. A cantilever 25 is slidably mounted on the vertical guide rail 24, and an end effector 30 is mounted at the end of the cantilever 25. A vertical lifting motor for controlling the lifting and lowering of the cantilever 25 is mounted on the vertical guide rail 24. A fruit collection mechanism 26 is mounted at the end of the cantilever 25 and is located below the end effector 30 (e.g., ...). Figure 3 (As shown).

[0032] Reference Figure 4 As shown: The end effector 30 includes a motor connector 31, a rotary motor 32, a base plate 33, a stationary blade 34, an insulating base 35, a moving blade 36, and a transmission servo motor 37. The motor connector 31 is fixedly installed at the end of the cantilever 25 and houses the rotary motor 32. The output shaft of the rotary motor 32 is connected to the base plate 33 via an end connecting flange (for driving the overall rotation of the end effector). The stationary blade 34 and the insulating base 35 are mounted on the base plate 33, and the moving blade 36 is rotatably connected to the stationary blade 34 via a rotating shaft. The transmission servo motor 37 is fixedly installed below the base plate 33, and its output shaft passes through the base plate 33 and is fixedly sleeved with a drive gear 371 on the upper side of the base plate 33. A sector gear 361 is fixedly installed at the end of the moving blade 36 via a lock nut, and the sector gear 361 meshes with the drive gear 371 (thus driving the moving blade 36 to rotate and achieve the shearing action). An upper shell 331 and a lower shell 332 are respectively provided on the upper and lower end faces of the base plate 33 for mechanism protection. The inner side of the front end of the moving blade 36 is provided with clamping foam 362, which works in conjunction with the foam stop block 363 fixed on the moving blade 36 to flexibly clamp the fruit stem while the moving blade 36 closes and cuts the fruit stem.

[0033] The detection and positioning system includes a binocular camera and a central control device. The binocular camera is mounted on the robotic arm structure and located directly above the end effector 30. The central control device is located inside the chassis frame 11.

[0034] Example 2: As another preferred embodiment of the present invention, a harvesting method for a dual-arm cherry tomato harvesting robot, employing the harvesting robot as described in Example 1, includes: Step S1, Navigation Path Generation: Generate a navigation path based on the navigation system, specifically as follows: Step S11: Trigger image acquisition: After the chassis speed returns to zero and vibration decays, trigger the navigation system (monocular camera) to acquire images, and use Wiener filtering to restore slight motion blur (eliminating image blur caused by intermittent start-stop). ; In the formula: (u, v) Represents the frequency domain spatial coordinates (determined by the width and height of the input image, and automatically generated by calculating through a two-dimensional discrete Fourier transform). Represents the frequency domain estimation of the restored image; G(u, v) The frequency domain of the degraded image is represented (obtained by performing a two-dimensional fast Fourier transform on the input blurred image). H(u, v) The degradation transfer function (characterizing the degree to which different spatial frequency components are attenuated / distorted by the fuzzy system, and is the Fourier transform of the spatial point spread function) represents the degradation transfer function. H ​ Represents the complex conjugate of the degenerate transfer function; S n ​ This represents the noise power spectral density (which describes the distribution of noise energy at different frequencies and is obtained through the Wiener-Khinchin theorem). S f ​ The power spectral density represents the original sharp image (describing the frequency energy distribution of the real scene itself, obtained through the Wiener-Khinchin theorem); This represents the regularization coefficient (typically 0.01 to 0.05). Output clear, unblurred original navigation image; Meanwhile, as the chassis travels along the ridge, the monocular camera continuously acquires images at a fixed frame rate (such as 10-15fps) and directly outputs the original navigation images; Step S12, Image Preprocessing: First, convert the original navigation image to grayscale: ; In the formula: These represent the pixel values ​​for the red, green, and blue channels, respectively. Then, perform non-linear grayscale correction on the grayscale image: ; In the formula: V in ​ Indicates the grayscale value of the input pixel; This indicates the correction index (usually 2). Finally, iterate through every pixel in the entire image, calculating the local Gaussian mean and local threshold sequentially to complete the binary classification: With the current pixel ​ Select size as the center. w × w The square neighborhood (general) w (Take 31), generate Gaussian weight kernel: ; In the formula: Gaussian standard deviation (general) ); Obtain Gaussian weighted local mean ​ With local threshold ​ : ; ; In the formula: C py This represents an offset constant (typically 5). like V out ​ < ​ If it is determined to be a navigation line (foreground), the pixel value is assigned to 0 (black); if V out ​ ≥ ​ If the value is 255 (white), then it is determined to be the background and the pixel value is assigned to 255; connect the black connected components to generate the initial navigation line binary image; Step S13: Based on the motion consistency characteristics of the optical flow field, distinguish between static navigation lines and dynamic shadow artifacts (remove interference): First, using two consecutive grayscale images as input, employ... ​ Dense optical flow algorithm calculates the optical flow vector for each pixel in an image. u g ,v g This process obtains the dense optical flow field of the entire image. Then, all black connected components in the binary image are masked out, and the remaining region is used as a static background sample. The median of the optical flow vectors of all pixels within the static background sample region is taken to obtain the global motion vector. v global ,like If the chassis is approximately stationary, the original binary image is directly output (to avoid the denominator approaching 0, which could lead to abnormal scores). Otherwise, an 8-connected component detection is performed on the initial navigation line binary image to obtain all independent black connected regions. B 1 ,B 2 ,…,B nAnd record the set of pixel coordinates for each region; finally, for each connected component B i Extract the optical flow vectors of all pixels within the region and obtain the region's average motion vector. v Bi And calculate the motion consistency score: ; Preset motion difference judgment threshold ​ (Usually 0.3), if S motion-i > ​ If the value is 0, it is determined to be a dynamic shadow; all pixels within that connected component are set to white and removed from the binary image. S motion-i ≤ ​ If so, it is determined to be a real navigation line and the area is retained; Step S14: First, based on the dimensions and height of the binary image... h ez The navigation line is divided into three independent sampling regions based on spatial distance: distant ROI ∈ [0, ... h ez / 3) (corresponding to the navigation line in the distance of the field of view, used for pre-aiming and prediction of curves), mid-range ROI∈[ h ez / 3,2 h ez / 3)(corresponding to the navigation line in the middle of the field of view, used for transition between near and far views), near ROI ∈ [2 h ez / 3, h ez -1) (corresponding to the navigation line in the foreground of the field of view, used for precise alignment correction); then, for each ROI region, the lateral centroid coordinates of the foreground of the navigation line in that region are calculated using the geometric moment method: foreground centroid x near Mid-field center x mid Visionary core x far Then, the curvature of the navigation line is calculated using the centroids of the three regions, and the weights for distant and near views are adaptively updated, specifically as follows: First, obtain three centroids: foreground, middle ground, and background. The foreground centroid ( x far ,v far ), v far = h ez / 6 (Vertical midpoint of distant area); Mid-field centroid point ( x mid,v mid ), v mid = h ez / 2 (Midfield vertical midpoint); Foreground centroid ( x near , v near ), v near =5 h ez / 6 (vertical midpoint of the foreground area); Then, the curvature of the navigation line in the mid-ground is obtained using the three centroids: ; The weights of foreground and background views are dynamically adjusted based on the curvature of the navigation line. ; ; In the formula: This indicates the baseline weight for the straight-line perspective (typically 0.4). This represents the curvature sensitivity coefficient (typically 50). This indicates the mid-range weight (a fixed value, typically 0.3). These represent the weights of the foreground and background, respectively. Finally, the navigation target point is calculated using weighted averages to generate the heading deviation, specifically: ; ; In the formula: x center This represents the horizontal center reference value of the image.

[0035] Step S15, Discontinuity Line Completion: If the number of foreground pixels in a certain area is less than 10 (no valid navigation lines), the centroid value of the corresponding area in the previous frame is temporarily used for preservation, and subsequent Bézier curve discontinuity line completion is triggered; if there are no valid pixels for 3 consecutive frames, it is marked as a navigation line discontinuity; navigation line discontinuity is based on the historical tangent direction and Bézier curve completion is used (to ensure navigation continuity): ; In the formula: t bes Indicates the parameters of the normalized curve; P bes0 Indicates the starting control point (i.e., the near end of the broken line); P bes3 Indicates the endpoint control point (i.e., the far end of the broken line); Pbes1 , P bes2 These represent the tangent control points at the start and end points, respectively (obtained through the slope of historical navigation lines). Output the completed continuous navigation lines and the corrected heading deviation; Step S16: Visual pixel deviation is converted into physical error quantities in vehicle dynamics through "pixel-physical scale conversion." Then, a model predictive controller (MPC) is constructed based on a linear two-degree-of-freedom error dynamics model to output the target front wheel steering angle command for the Ackerman steering mechanism, thereby achieving steering control. Specifically: First, the visual pixel deviation is converted into a physical error quantity through "pixel-physical scale conversion": ; In the formula: k pix This represents the horizontal pixel equivalent (obtained by dividing the actual horizontal distance of the bottom surface corresponding to the image width by the image pixel width). These are the decomposed components of the heading deviation; D near , D far These represent the distances from the ground in front of the vehicle to the foreground and background areas, respectively. Then, construct the path tracking error dynamic state-space model: The state vector is: ,in, Indicates the centroid sideslip angle. Indicates yaw rate; control input is: ,in, Indicates the front wheel steering angle; The continuous-time error dynamic equation is: ; In the formula: v xs Indicates driving speed; C f 、C r These represent the lateral stiffness of the front and rear wheels, respectively. a m , b m These represent the distances from the center of mass to the front and rear axles, respectively. m Indicates the overall vehicle weight; I z This represents the yaw moment of inertia of a vehicle. State-space equations in matrix form: ; Then, the continuous model was converted into a discrete model using the Euler forward discretization method, with a sampling period of... ​ : ; In the formula: I represents a 4th-order identity matrix; Assuming the prediction time domain is n p Control time domain is n c ( n c ≤ n p ), based on the state of k at the current time. X zt (k) Using the initial value, the future is obtained recursively. n p Step state sequence: ; Control increment sequence As optimization variables, the predicted state matrix takes the form of: ; In the formula: They represent respectively by A d , B d , D d Constructed prediction coefficient matrix; Finally, construct the optimization objective function: ; In the formula: These represent the weighting coefficients for lateral deviation and heading deviation, respectively. This represents the weighting coefficient that controls the increment; Constraints: (in, , It is typically 0.52 rad; , (Typically 0.05 rad / control cycle).

[0036] The drive-by-wire chassis controls the picking robot to reach the picking location and triggers the picking operation. Step S2, Picking Point Recognition: Trigger the detection and positioning system to acquire images, detect fruit ripeness, and obtain the picking point, then convert the picking point to the robotic arm's base coordinate system; specifically: Step S21, Image Acquisition and Preprocessing: Synchronous RGB images at three polarization angles (0°, 45°, and 90°) are acquired using the binocular camera of the detection and positioning system. Grayscale conversion and Gamma correction are performed on the images at each polarization angle (grayscale conversion and Gamma correction are the same as in step S12, with a correction index of 2). The grayscale value sequence of each pixel at different polarization angles is calculated, and the degree of polarization of each pixel is obtained based on the grayscale response at different polarization angles. ; In the formula: ​ Represents pixels p degree of polarization; I max (p) , I min (p) Representing pixels p Maximum and minimum gray values ​​at three polarization angles; Finally, we obtained a three-angle polarization grayscale image, a pixel-level polarization map, and a color original image after Gamma correction; Step S22, Semantic Segmentation: Based on the improved DeepLabV3 model, pixel-level semantic segmentation of fruit-fruiting branch-main stem is achieved, and a mature fruit mask is output. M fruit , resultant branch mask M branch Main stem cover M stem ; The improved DeepLabV3 model's encoder uses a lightweight MobileNetV3-large as its backbone feature extraction network, and introduces a Contextual Transformer (CoT) module at its low-level and high-level feature output positions. CoT's unique dynamic contextual attention mechanism is used to mine the deep semantic association between cherry tomato clusters and branches, effectively enhancing the model's feature capture ability for slender branches and solving the problem of missed or false detections of small branches. The decoder employs a multi-scale feature fusion strategy, fusing features from different levels enhanced by the CoT module (gradually restoring pixel-level details of the image, ensuring the accuracy of segmentation boundaries, and providing reliable support for subsequent deep localization). The model uses a Stochastic Gradient Descent (SGD) optimizer combined with a cosine annealing learning rate decay strategy for iterative training, and the loss function is cross-entropy loss. Step S23, Point Cloud Completion: First, construct a multi-constraint matching cost that integrates grayscale, structure, and polarization, then perform cost aggregation and disparity optimization. ; In the formula: These represent the corresponding weighting coefficients (generally) (The costs corresponding to the weighting coefficients are all dimensionless after Min-Max normalization). This represents the total matching cost of pixel p under disparity d; This represents the cost of the absolute difference in grayscale values; Indicates the cost of the Census transformation; Then, extract the cloud of fruit stalk skeleton points { S i The corresponding effective parallax is d i For any invalid point within the mask p Find the adjacent projection points on its skeleton. S k , S k+1 Perform interpolation completion: ; Output high-precision disparity map and original 3D point cloud in camera coordinate system P raw ; Step S24, Clustering and Segmentation: For any point p i Calculate its to k The average distance between the nearest neighbors global mean Standard deviation is ,like If so, it will be discarded as a noise point; Then, a spectral clustering algorithm incorporating prior knowledge of cherry tomato ellipsoidal shapes is used to obtain a set of independent fruit point cloud clusters. C fruit} and the main stem point cloud cluster set { C stem}; Obtain the set of independent fruit point cloud clusters { C fruit} and the main stem point cloud cluster set { C stem Specifically: First, for each point p i Extract its k p We use 50 nearest neighbor points (typically) and employ the RANSAC robust fitting algorithm to fit the standard ellipsoid equation to the neighborhood point cloud: ; In the formula: ( x 0 ,y 0 ,z 0) represents the coordinates of the ellipsoid center;a ty 、b ty 、c ty These represent the semi-major axes of the ellipsoid. Reconstruct the ellipsoidal prior affinity matrix: ; In the formula: e ij This represents the average of the fitting residuals between two points; Indicates the spatial distance bandwidth (typically 5mm); This represents the bandwidth of the ellipsoid fitting residual (typically 0.02). Obtain the construction degree matrix: (an n×n diagonal matrix, where the diagonal elements are the sum of all affinities in the corresponding row); and also constructed as a normalized Laplace matrix: L = D - W ; Laplace matrix L Eigenvalue decomposition is performed to obtain n eigenvalues ​​sorted from smallest to largest, along with their corresponding eigenvectors. Then, the number of clusters is automatically determined using the eigenvalue gap method or silhouette coefficient method, and the eigenvectors corresponding to the k smallest eigenvalues ​​are selected and concatenated column-wise to form an n×k dimensional feature matrix. ; feature matrix Each row is treated as a k-dimensional sample point. K-means clustering is performed on all row vectors, resulting in the number of clusters. k Based on the clustering results, the original 3D point cloud was divided into... k Each cluster represents an independent point cloud and corresponds to a single cherry tomato fruit. Small clusters with too few points (e.g., <30 points, considered residual noise) are removed, and the set of independent fruit point cloud clusters is output. C fruit}; Set a clustering distance threshold dth (typically 20mm, greater than the stem diameter and less than the distance between adjacent stems); traverse the point cloud, and if the spatial distance between two points is less than dth, they are grouped into the same cluster; finally, m independent stem point cloud clusters are obtained, and clusters with too few points and small noise are simultaneously removed, outputting a set of independent stem point cloud clusters { C stem}

[0037] Step S25, Maturity Grading: Based on the set of independent fruit point cloud clusters { C fruitThe point cloud of a single fruit is projected onto the image plane, and a semantic mask is used to obtain the precise pixel region (ROI) of the fruit. For each pixel within the ROI, a reflected light intensity equation is fitted based on the gray values ​​of three polarization angles to solve for the diffuse reflection and specular reflection components. The diffuse reflection component is used to reconstruct the pure natural color image of the fruit, eliminating highlights on the peel and ambient reflections. ; In the formula: I d This indicates the intensity of diffuse reflected light (the color information of the fruit itself, and its non-polarized characteristics). I s Indicates the intensity of specularly reflected light (spectral reflection and linear polarization characteristics of ambient light); Indicates the principal polarization direction angle of specular reflection; Obtain pixel grayscale values ​​at three polarization angles I(0 0 )、I(45 0 )、I(90 0 ) ; Then, the mean polarization degree of a single fruit region is extracted. ​ mean Polarization degree variance ​ std Mean of specular diffuse reflectance R s_mean Three features were used, and these features were input into an offline-calibrated second-order polynomial regression model to obtain the fruit's elastic modulus. E tan : ; In the formula: These represent the regression coefficients calibrated offline (obtained from the texture analyzer puncture experiment); like E tan If the pressure is >0.8 MPa, it is considered an unripe, hard fruit; if it is ≤0.3 MPa... E tan Fruits with a pressure ≤0.8MPa are considered suitable for harvesting. E tan If the pressure is less than 0.3 MPa, it is considered an overripe, soft fruit. Finally, the diffuse RGB image is converted to the Lab color space, and the color is extracted. a Regional mean of channel (red-green hue) As the core feature of color maturity; constructing feature vectors Input the pre-trained LightGBM lightweight classifier, the classifier outputs three probabilities: "unripe", "suitable for picking", and "overripe", and take the class with the highest probability and a confidence of not less than 0.8 as the final grade of a single fruit; Step S26: Obtain the picking point: Based on the fruit point cloud clusters and corresponding main stem point cloud clusters after maturity grading in Step S25, calculate the centroid and covariance matrix of the bounding box point cloud: ; Perform eigenvalue decomposition on the covariance matrix Cov to obtain three eigenvalues ​​sorted in descending order, and three corresponding orthogonal eigenvectors. ​ 1 (First principal direction, the direction in which the point cloud extends the longest, corresponding to the major axis) ​ 2 (Second primary direction, secondary extension direction of point cloud) ​ 3 (Third principal direction, the direction in which the point cloud extends the shortest); the three eigenvectors form an orthogonal local coordinate system, that is, the three axes of the smallest bounding box; Project all centered points onto the three principal axes respectively, obtain the minimum and maximum projection values ​​on each axis, and determine the bounding box boundary: for any point p i In the j The projection values ​​on each principal axis are: Traverse all points to obtain the projected extrema of the three principal axes: ; [ ​ j ​ j That is, the bounding box in the first... j The range on each main axis; Calculate the major axis (e.g.) ​ 1 is the geometric center of the two end faces (where the major axis is 1): ; Obtain the center of the upper surface of the fruit enclosure box P a : ; In the formula: This represents the centroid of the main stem dot cloud; Based on the long axis of the main stem ​ s1 Obtain the center of the upper end face of the main stem: ; Then, calculate the direction vector from the centroid of the main stem to the centroid of the fruit: ; Distance from the center of the top surface to the boundary of the bounding box du-b for: ; The midpoint of the upper border of the main stem enclosure box closest to the fruiting branch is: ; Finally, the picking location was obtained: ; Will P pick Convert to picking point coordinates in camera coordinate system P cam .

[0038] Step S3, Harvesting Control: The robotic arm executes the motion plan and drives the end effector to perform clamping and harvesting; specifically: Step S31: Input the coordinates of the picking point P cam Based on the offline calibrated full-stroke error table, the systematic error of the current position is obtained through cubic spline interpolation (taking the X-axis as an example, and the Y-axis and Z-axis are similar; the basic pulse equivalent of the Z-axis is 128 pulses / mm, that is, replacing 320 in the correction pulse number with 128): ; In the formula: M i , M i+1 These represent the second derivative values ​​at the endpoints of the segment (obtained by solving the spline boundary conditions (natural spline or clamped spline) simultaneously); ​ i , ​ i+1 They represent the calibration points respectively. x i , x i+1 The measured error; Through systematic errors ​ Corrected pulse count: ; In the formula: K x Indicates the pulse correction coefficient; When a reverse motor direction is detected, a gap compensation pulse is superimposed: ; In the formula: This represents the measured backlash value; This represents the reciprocal of the pulse equivalent. Direction reversal criterion: , ​ express t The speed of the motor at all times; Final output compensation target pulse number: ; In the formula: ​ res This indicates the inversion flag; it is 1 when inverting and 0 otherwise. Step S32: Employ modal orthogonal decoupling and multimodal input shaping to eliminate mechanical vibration from the command end (avoiding modal coupling resonance in the gantry three-axis structure): Modal decoupling: ; In the formula: m i Indicates the first i Equivalent mass of the first mode, g i Indicates the first i First-order modal equivalent stiffness, c i Indicates the first i First-order mode equivalent damping; Let represent the vibration amplitude, modal velocity, and modal acceleration of the i-th mode, respectively. f i Indicates the first i Modal generalized excitation force For each mode, a zero-vibration (ZV) input shaper is designed: For the i-th vibration mode, the shaper consists of two impulse pulses, triggered at times t1 and t2 respectively, with amplitude weights A1 and A2. ; ; In the formula: Indicates the first i First-order modal damping ratio (obtained by hammer impact method and modal testing); T i Indicates the first i First-order modal natural period Output the decoupled and shaped speed planning instructions for each axis; Step S33: First, perform FOC vector current decoupling and current loop closed-loop control: First, perform three-phase current sampling and Clarke transformation (3s→2s stationary frame transformation): for the three-phase stator current of the motor I a 、I b 、I c After sampling, the current is converted into a two-phase stationary coordinate system current through Clarke transformation, thus eliminating three-phase coupling. ; Then perform the Park transformation: combine the rotor electrical angle feedback from the magnetic encoder. The static coordinate system current is transformed into a dq coordinate system that rotates synchronously with the rotor (converting AC quantity into DC quantity): ; In the formula: I d , I q These represent direct-axis current (excitation component) and quadrature-axis current (torque component), respectively. Independent PID closed-loop regulation is applied to the d-axis and q-axis currents respectively, converting the current error into voltage commands for the corresponding axes: ; In the formula: Indicates the reference value of the d-axis current; This represents the q-axis current reference value obtained by superimposing the feedforward compensation amount of the DOB disturbance observer onto the q-axis current command; These represent the proportional gain, integral gain, and derivative gain along the d-axis, respectively. These represent the proportional gain, integral gain, and derivative gain along the q-axis, respectively. The voltage command in the rotating coordinate system is inversely transformed back to the stationary coordinate system, and then the PWM drive signal for the three-phase inverter is generated through space vector pulse width modulation (SVPWM). ; Then, the load disturbance torque is estimated in real time using the DOB disturbance observer: based on the nominal mechanical model of the motor, the difference between the control input and the actual position feedback is used to estimate the external load disturbance torque such as shear impact in real time. Motor nominal model: , J G Indicates the moment of inertia of the motor. B G Indicates the damping coefficient; Estimated disturbance: ; In the formula: Indicates the motor angle (satisfying) ); ​ Indicates the control quantity; This represents the filtering time constant (typically 0.5ms). Then, the estimated disturbance torque is converted into quadrature-axis current compensation, which is added to the original current command to compensate in advance before the location of the disturbance, thereby achieving rapid impact suppression. ; In the formula: This indicates the original quadrature-axis current command output by the speed loop; This represents the external load disturbance torque estimated by DOB; This represents the feedforward gain (characterized by the electromagnetic torque generated per ampere of quadrature-axis current, obtained from the reciprocal of the motor torque coefficient). This indicates the final quadrature-axis current command after superposition compensation (sent to the current loop PID control). Finally, the system adopts a three-stage cascaded closed-loop structure of position loop → speed loop → current loop, with the response bandwidth increasing step by step from the outside to the inside (position loop: the input is the target position command, the feedback is the actual position of the magnetic encoder, and the output speed command is adjusted by PID; speed loop: the input is the speed command output by the position loop, the feedback is the actual speed obtained by encoder differential, and the output original quadrature-axis current command is adjusted by PID). Current loop: the innermost loop with the highest response bandwidth, ensuring rapid current tracking of commands); after the movement is completed, the main controller is asynchronously notified via the CAN bus to trigger the next picking action, improving system coordination efficiency; Step S34: Employ token ring distributed scheduling for the dual robotic arms. Specifically: First, calculate the fruit density on both sides and generate the dynamic step distance. ; In the formula: D base This indicates the baseline step distance (typically 300mm). These represent the fruit density (number of mature fruits / field of view area) on the left and right sides, respectively. Indicates the standard density threshold; Then, token-based scheduling is used to prioritize assigning tasks to the idle side; Finally, the safe distance is adjusted in real time: ; In the formula: d safe This indicates the minimum safe distance (usually 50mm). Trigger collision avoidance.

[0039] Step S4, Fruit Collection and Mechanism Reset: After the fruit is picked, the end effector transports the fruit to the fruit collection mechanism and resets the end effector and robotic arm structure. Step S5, Cyclic Operation: Based on the fruit maturity detection results within the field of view, if there are still mature fruits, repeat steps S2 to S4; if there are no mature fruits, trigger the chassis to move and repeat steps S1 to S5 until the end of the ridge is reached.

[0040] Example 3: As another preferred embodiment of the present invention, based on embodiment 2, the harvesting of the end effector is specifically as follows: when the end effector reaches the target harvesting point, the rotary motor first drives the end connecting flange to rotate, adjusting the end posture to adapt to the fruit stem growth angle; subsequently, the transmission servo motor drives the active gear to drive the driven gear to rotate, causing the moving blade to close to the stationary blade; during the closing process, the clamping foam located at the front end of the moving blade, together with the foam stop block, first flexibly clamps the fruit stem to prevent the fruit from falling, and then the moving blade cuts the fruit stem, realizing the "clamping and cutting in one" operation; after the cutting is completed, the collection drive motor of the fruit collection mechanism drives the collection connector to open the entrance of the fruit collection mechanism, and the robotic arm finely adjusts its posture to release the fruit, allowing it to slide into the collection box along the flexible pipe, completing the non-destructive harvesting.

Claims

1. A dual-arm cherry tomato harvesting robot, characterized in that: The system includes a drive-by-wire chassis, dual robotic arms, an end effector, a detection and positioning system, and a navigation system. The drive-by-wire chassis includes a chassis frame, a shock absorption system, an Ackermann steering system, and a rear-drive powertrain. The navigation system is integrated inside the chassis frame and includes a vision system and a central control unit. The dual robotic arms are located on the end face of the chassis frame and include a fixed bracket and two sets of robotic arm structures. The fixed bracket is located on the end face of the chassis frame, and the two sets of robotic arm structures are symmetrically arranged on it. The end effector is located at the end of each robotic arm structure. The end effector includes a motor connector, a rotary motor, a base plate, a stationary blade, an insulating base, a moving blade, and a transmission servo motor. The detection and positioning system includes a binocular camera and a central control unit. The binocular camera is mounted on the robotic arm structure and located directly above the end effector. The central control unit is located inside the chassis frame.

2. The dual-arm cherry tomato harvesting robot according to claim 1, characterized in that: The chassis frame is a frame structure; the shock absorption system includes a shock absorption movable beam and shock absorption springs. The shock absorption movable beam is hinged to the chassis frame via a pin, and the shock absorption springs are located between the shock absorption movable beam and the chassis frame; the Ackermann steering system is located at the front of the chassis frame's underside and includes a steering joint, steering tie rods, and a front axle. Steering is achieved by using hydraulic or electric push rods to drive the steering tie rods; the rear-wheel drive power system is located at the rear of the chassis frame's underside and includes a rear-wheel drive motor and a differential transmission.

3. A dual-arm cherry tomato harvesting robot according to claim 1 or 2, characterized in that: The robotic arm structure includes an X-axis main guide rail, two X-axis auxiliary guide rails, a Y-axis guide rail, a vertical guide rail, a cantilever, and a fruit collection mechanism. The two X-axis auxiliary guide rails are arranged on both sides of the corresponding X-axis main guide rail, and the X-axis auxiliary guide rails and the X-axis main guide rail are arranged in parallel. Both the X-axis auxiliary guide rails and the X-axis main guide rail are located on the end face of the fixed support. The Y-axis guide rail is slidably arranged on the end face of the X-axis main guide rail and the two X-axis auxiliary guide rails and is driven by an X-axis drive motor arranged on the X-axis main guide rail. The vertical guide rail is slidably arranged on the end face of the Y-axis guide rail and is driven by a Y-axis drive motor arranged on the Y-axis guide rail. A cantilever is slidably arranged on the vertical guide rail, and an end effector is arranged at the end of the cantilever. A vertical lifting motor for controlling the lifting and lowering of the cantilever is arranged on the vertical guide rail. The fruit collection mechanism is arranged at the end of the cantilever and is located below the end effector.

4. A dual-arm cherry tomato harvesting robot according to claim 2 or 3, characterized in that: The motor connector is fixedly installed at the end of the cantilever and a rotary motor is installed inside it; the output shaft of the rotary motor is connected to the base plate through the end connecting flange; a stationary blade and an insulating base are installed on the base plate, and the moving blade is rotatably connected to the stationary blade through a rotating shaft; the transmission servo is fixedly installed below the base plate, its output shaft passes through the base plate and is fixedly sleeved with the drive gear on the upper side of the base plate, and a sector gear is fixedly installed at the end of the moving blade through an anti-loosening nut and the sector gear meshes with the drive gear.

5. A dual-arm cherry tomato harvesting robot according to claim 3 or 4, characterized in that: The bottom plate is provided with an upper shell and a lower shell on its upper and lower end faces, respectively.

6. A dual-arm cherry tomato harvesting robot according to claim 3 or 4, characterized in that: The inner side of the front end of the moving blade is provided with clamping foam, which, together with the foam block fixed on the moving blade, completes the flexible clamping of the fruit stem while the moving blade closes to cut the fruit stem.

7. The harvesting method of a dual-arm cherry tomato harvesting robot according to claim 4, characterized in that: include: Step S1, Navigation Path Generation: Generate a navigation path based on the navigation system, drive the wire-controlled chassis to control the picking robot to reach the picking location, and trigger the picking operation; Step S2, Picking point identification: Trigger the detection and positioning system to acquire images, detect fruit maturity and obtain picking points, and convert the picking points to the robot arm base coordinate system; Step S3, Harvesting Control: The robotic arm executes the motion plan and drives the end effector to perform clamping and shearing harvesting; Step S4, Fruit Collection and Mechanism Reset: After the fruit is picked, the end effector transports the fruit to the fruit collection mechanism and resets the end effector and robotic arm structure. Step S5, Cyclic Operation: Based on the fruit maturity detection results within the field of view, if there are still mature fruits, repeat steps S2 to S4; if there are no mature fruits, trigger the chassis to move and repeat steps S1 to S5 until the end of the ridge is reached.

8. The harvesting method of a dual-arm cherry tomato harvesting robot according to claim 7, characterized in that: In step S1, generating the navigation path based on the navigation system specifically involves: Step S11: Trigger Image Acquisition: After the chassis speed returns to zero and vibration decays, the navigation system is triggered to acquire images, and Wiener filtering is used to restore slight motion blur. ; In the formula: (u,v) Represents frequency domain spatial coordinates; Represents the frequency domain estimation of the restored image; G(u,v) The frequency domain represents the degraded image; H(u,v) Indicates the degenerate transfer function; H (u,v) Represents the complex conjugate of the degenerate transfer function; S n (u,v) Represents the noise power spectral density. S f (u,v) Represents the power spectral density of the original sharp image; Represents the regularization coefficient; Output clear, unblurred original navigation image; Meanwhile, as the chassis travels along the ridge, the monocular camera continuously acquires images at a fixed frame rate and directly outputs the original navigation images; Step S12, Image Preprocessing: First, convert the original navigation image to grayscale: ; In the formula: These represent the pixel values ​​for the red, green, and blue channels, respectively. Then, perform non-linear grayscale correction on the grayscale image: ; In the formula: V in (x,y) Indicates the grayscale value of the input pixel; Indicates the correction index; Finally, iterate through every pixel in the entire image, calculating the local Gaussian mean and local threshold sequentially to complete the binary classification: With the current pixel (x,y) Select size as the center. w × w Generate a Gaussian weight kernel from the square neighborhood: ; In the formula: Indicates the Gaussian standard deviation; Obtain Gaussian weighted local mean Mean(x,y) With local threshold Thy(x,y) : ; ; In the formula: C py Indicates an offset constant; like V out (x,y) < Thy(x,y) If it is, then it is determined that the navigation line and pixel are assigned a value of 0; if V out (x,y) ≥ Thy(x,y) If the value is 255, then it is determined to be the background and the pixel value is assigned to 255; connect the black connected components to generate the initial navigation line binary image; Step S13: Based on the motion consistency characteristics of the optical flow field, distinguish between static navigation lines and dynamic shadow artifacts; Step S14: First, based on the dimensions and height of the binary image... h ez The navigation line is divided into three independent sampling regions based on spatial distance: distant ROI ∈ [0, ... h ez / 3), mid-range ROI∈[ h ez / 3,2 h ez / 3), near-field ROI∈[2 h ez / 3, h ez -1); Then, for each ROI region, the lateral centroid coordinates of the navigation line foreground within that region are calculated using the geometric moment method: foreground centroid x near Mid-field center x mid Visionary core x far Then, the curvature of the navigation line is calculated using the centroids of the three regions, and the weights of the distant and near views are updated adaptively. Finally, the navigation target point is calculated using weighted averages to generate the heading deviation. Step S15, Line Completion: If the number of foreground pixels in a certain area is less than 10, the centroid value of the corresponding area in the previous frame is temporarily used to maintain it, and the subsequent Bézier curve line completion is triggered. If there are no valid pixels for 3 consecutive frames, it is marked as a navigation line break; the navigation line break is based on the historical tangent direction and is completed using Bézier curves; the output is the completed continuous navigation line and the corrected heading deviation; Step S16: The visual pixel deviation is converted into the physical error of vehicle dynamics through "pixel-physical scale conversion". Then, a model predictive controller is constructed based on the linear two-degree-of-freedom error dynamics model to output the target front wheel steering angle command of the Ackerman steering mechanism to achieve steering control.

9. The harvesting method of a dual-arm cherry tomato harvesting robot according to claim 7, characterized in that: Step S2 specifically involves: Step S21, Image Acquisition and Preprocessing: Synchronous RGB images at three polarization angles (0°, 45°, and 90°) are acquired using the binocular camera of the detection and positioning system. The images at these three polarization angles are then converted to grayscale and subjected to Gamma correction. The grayscale value sequence of each pixel at different polarization angles is calculated, and the degree of polarization of each pixel is obtained based on the grayscale response at different polarization angles. ; In the formula: DOLP(p) Represents pixels p degree of polarization; I max (p) , I min (p) Representing pixels p Maximum and minimum gray values ​​at three polarization angles; Finally, we obtained a three-angle polarization grayscale image, a pixel-level polarization map, and a color original image after Gamma correction; Step S22, Semantic Segmentation: Based on the improved DeepLabV3 model, pixel-level semantic segmentation of fruit-fruiting branch-main stem is achieved, and a mature fruit mask is output. M fruit , resultant branch mask M branch Main stem cover M stem ; Step S23, Point Cloud Completion: First, construct a multi-constraint matching cost that integrates grayscale, structure, and polarization, then perform cost aggregation and disparity optimization. ; In the formula: These represent the corresponding weight coefficients; This represents the total matching cost of pixel p under disparity d; This represents the cost of the absolute difference in grayscale values; Indicates the cost of the Census transformation; Then, extract the cloud of fruit stalk skeleton points { S i }, corresponding effective parallax is d i For any invalid point within the mask p Find the adjacent projection points on its skeleton. S k , S k+1 Perform interpolation completion: ; Output high-precision disparity map and original 3D point cloud in camera coordinate system P raw ; Step S24, Clustering and Segmentation: For any point p i Calculate its to k The average distance between the nearest neighbors global mean Standard deviation is ,like If so, it will be discarded as a noise point; Then, a spectral clustering algorithm incorporating prior knowledge of cherry tomato ellipsoidal shapes is used to obtain a set of independent fruit point cloud clusters. C fruit } and the main stem point cloud cluster set { C stem }; Step S25, Maturity Grading: Based on the set of independent fruit point cloud clusters { C fruit The point cloud of a single fruit is projected onto the image plane, and a semantic mask is used to obtain the precise pixel region (ROI) of the fruit. For each pixel within the ROI, a reflected light intensity equation is fitted based on the gray values ​​of three polarization angles to solve for the diffuse reflection and specular reflection components. The diffuse reflection component is used to reconstruct the pure natural color image of the fruit, eliminating highlights on the peel and ambient reflections. ; In the formula: I d Indicates the intensity of diffuse reflection. I s Indicates the intensity of light reflected from the specular surface; Indicates the principal polarization direction angle of specular reflection; Obtain pixel grayscale values ​​at three polarization angles I(0 0 )、I(45 0 )、I(90 0 ) ; Then, the mean polarization degree of a single fruit region is extracted. DOLP mean Polarization degree variance DOLP std Mean of specular diffuse reflectance R s_mean Three features were used, and these features were input into an offline-calibrated second-order polynomial regression model to obtain the fruit's elastic modulus. E tan : ; In the formula: These represent the regression coefficients calibrated offline; like E tan If the pressure is >0.8 MPa, it is considered an unripe, hard fruit; if it is ≤0.3 MPa... E tan Fruits with a pressure ≤0.8MPa are considered suitable for harvesting. E tan If the pressure is less than 0.3 MPa, it is considered an overripe, soft fruit. Finally, the diffuse RGB image is converted to the Lab color space, and the color is extracted. a regional mean of the channel As a core characteristic of color maturity; Constructing feature vectors Input the pre-trained LightGBM lightweight classifier, the classifier outputs three probabilities: "unripe", "suitable for picking", and "overripe", and take the class with the highest probability and a confidence of not less than 0.8 as the final grade of a single fruit; Step S26: Obtain the picking point: Based on the fruit point cloud clusters and corresponding main stem point cloud clusters after maturity grading in Step S25, calculate the centroid and covariance matrix of the bounding box point cloud: Perform eigenvalue decomposition on the covariance matrix Cov to obtain three eigenvalues ​​sorted in descending order, and three corresponding orthogonal eigenvectors. ezj 1. ezj 2. ezj 3; The three eigenvectors form an orthogonal local coordinate system, that is, the three axes of the smallest bounding box; Project all the centered points onto the three principal axes respectively, and obtain the minimum and maximum projection values ​​on each axis to determine the boundary of the bounding box; Obtain the center of the upper surface of the fruit enclosure box P a and the midpoint of the upper border of the main stem enclosure box near the fruiting branch. P b ; Final picking location: ; Will P pick Convert to picking point coordinates in camera coordinate system P cam .