A smart chili pepper harvesting system and harvesting method

By combining visual positioning with a six-degree-of-freedom robotic arm, the system achieves precise harvesting of chili peppers, solving the problem of limited harvesting accuracy and range in existing technologies, improving harvesting efficiency and reducing costs.

CN119302125BActive Publication Date: 2026-04-03SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing chili-harvesting robots and devices have limitations in harvesting accuracy and range, making it impossible to harvest chilies efficiently and precisely.

Method used

Using a visual positioning device and a six-degree-of-freedom robotic arm, combined with a binocular camera and a motor screw drive mechanism, the system can accurately identify and harvest the chili peppers, and complete the harvesting action through a shearing drive mechanism.

Benefits of technology

This technology enables precise harvesting of chili peppers, reducing breakage rates and harvesting costs while improving harvesting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119302125B_ABST
    Figure CN119302125B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent chili pepper harvesting system and method. The intelligent chili pepper harvesting system includes a visual positioning device for identifying the coordinate position of chili peppers, a chili pepper harvesting device for harvesting chili peppers based on their coordinates, and a control device. The chili pepper harvesting device includes a base, a robotic arm mounted on the base, a harvester at the end of the robotic arm, and a harvesting drive mechanism for driving the robotic arm to move along the X and Y axes. The harvester includes a support, a fixed blade, a moving blade, and a shearing drive mechanism for driving the moving blade to swing in coordination with the fixed blade to perform cutting and clamping actions. The fixed blade is fixed to the support; the moving blade is rotatably connected to the support and located above the fixed blade. The visual positioning device includes a binocular camera mounted on the base of the robotic arm. This intelligent chili pepper harvesting system enables precise harvesting of chili peppers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of chili harvesting, specifically relating to an intelligent chili harvesting device and harvesting method. Background Technology

[0002] The harvesting stage is the most time-consuming and labor-intensive phase in chili pepper cultivation, and the quality of the harvest directly affects the yield, storage, and processing of the peppers. Therefore, reducing harvesting costs, saving human resources, and improving harvesting efficiency are crucial.

[0003] Among them, the invention patent application with publication number CN117918119A discloses a "chili pepper harvesting machine," which includes a support frame, a walking device, a robotic arm device, and a drive device. The walking device is located on the support frame; the robotic arm device includes a robotic arm and a gripper structure located at the end of the robotic arm. The robotic arm includes a rigid component and a flexible component. The rigid component is rotatably mounted on the support frame, and the flexible component is connected between the rigid component and the gripper structure; the drive device is connected to the gripper structure and the flexible component to drive the gripper structure to open or close, and to cause the flexible component to bend. The robotic arm of the aforementioned chili pepper harvesting machine includes both a rigid and a flexible component, allowing the flexible component to bend under the drive of the drive device, while simultaneously driving the gripper structure to open or close, thus harvesting chilies as needed. This avoids the excessive freedom caused by an entirely flexible component, which could lead to the robotic arm twisting into a knot when rotating.

[0004] Patent application CN116472868A discloses a "wheeled automatic chili-harvesting robot," which includes a mobile chassis, a body, and other components. The body is fixed to the top of the mobile chassis. A gripping device and a collection box are mounted on the top of the body. A laser radar and a display are fixed to the front of the body. A depth camera is mounted on the gripping device. Steering wheels and drive wheels are fixed to the lower end of the mobile chassis. The steering wheels are connected to a steering device, and the drive wheels are connected to a drive device. The gripping device, display, laser radar, depth camera, steering device, and drive device are all connected to a control system. The control system enables the mobile chassis and gripping device to work collaboratively, thereby achieving autonomous gripping, precise harvesting, and automatic collection of chilies in complex environments without affecting the quality of the chilies.

[0005] However, the chili harvesting machine mentioned above has limited harvesting accuracy and range due to its structural limitations; while the wheeled chili automatic harvesting robot mentioned above has a small working space because it is not equipped with guide rails. Summary of the Invention

[0006] In order to overcome the shortcomings of the existing technology, the present invention provides an intelligent chili harvesting system that can achieve precise harvesting of chilies.

[0007] The second objective of this invention is to provide an intelligent chili harvesting method for an intelligent chili harvesting system.

[0008] The technical solution of the present invention to solve the above-mentioned technical problems is:

[0009] A smart chili pepper harvesting system includes a visual positioning device for identifying the coordinate position of chili peppers, a chili pepper harvesting device for harvesting chili peppers based on their position coordinates, and a control device. The chili pepper harvesting device includes a base, a robotic arm mounted on the base, a harvester at the end of the robotic arm, and a harvesting drive mechanism for driving the robotic arm to move along the X and Y axes. The harvester includes a support, a fixed blade, a movable blade, and a shearing drive mechanism for driving the movable blade to swing in coordination with the fixed blade to perform a cutting action. The fixed blade is fixed to the support; the movable blade is rotatably connected to the support and located above the fixed blade. The visual positioning device includes a binocular camera mounted on the base of the robotic arm.

[0010] Preferably, the binocular camera is used to acquire the first relative position coordinates of the chili pepper stem relative to the binocular camera, and transmits the first relative position coordinates to the control device. The control device calculates the third relative position coordinates of the chili pepper stem relative to the base of the robotic arm based on the first relative position coordinates and the second relative position coordinates of the binocular camera and the base of the robotic arm.

[0011] Preferably, the robotic arm is a six-degree-of-freedom robotic arm.

[0012] Preferably, both the X-axis drive mechanism and the Y-axis drive mechanism adopt a drive method that combines a motor and a lead screw transmission mechanism.

[0013] Preferably, both the X-axis drive mechanism and the Y-axis drive mechanism further include a guide mechanism, which adopts a combination of slide rail and slider, or a combination of guide rod and guide sleeve.

[0014] Preferably, the shearing drive mechanism includes an electric push rod; the electric push rod and the moving blade are connected by a connecting rod, one end of the connecting rod is rotatably connected to the telescopic rod of the electric push rod, and the other end is rotatably connected to the moving blade.

[0015] A smart chili pepper harvesting method includes the following steps:

[0016] Step S1: Perform hand-eye calibration to obtain the rotation transformation matrix between the image coordinate system of the binocular camera and the coordinate system of the robotic arm;

[0017] Step S2: The binocular camera identifies the chili peppers, processes the acquired chili pepper images, and obtains the image coordinates of the chili pepper stem in the chili pepper image;

[0018] Step S3: Based on the rotation transformation matrix obtained in step S1 and the image coordinates of the chili stem obtained in step S2, calculate the three-dimensional coordinates of the chili stem in the robotic arm coordinate system.

[0019] Step S4: Based on the three-dimensional coordinates of the chili stem obtained in step S3, the robotic arm drives the harvester to harvest the chili.

[0020] Step S5: Repeat steps S2-S4 until the binocular camera no longer detects the chili peppers, at which point the harvesting work is complete.

[0021] Preferably, in step S2, the process of processing the chili pepper image includes:

[0022] Step 201: Train the YOLOv10 model using the chili pepper images acquired by the binocular camera to construct a YOLOv10 training model in ONNX format;

[0023] Step 202: Preprocess the collected chili pepper images, feed the preprocessed chili pepper images into the YOLOv10 training model to perform forward inference, and parse the output from the inferred vector to obtain the best inference result.

[0024] Step 203: Combine the obtained best inference results, execute the GrabCut algorithm on the chili pepper image to generate an effective foreground mask; continuously solve the partial derivatives of the connected regions of the foreground mask from bottom to top, find the maximum value of the partial derivative and determine its direction to obtain the approximate two-dimensional coordinates of the chili pepper stem, and perform offset correction on the approximate two-dimensional coordinates to determine the final two-dimensional coordinates of the chili pepper stem.

[0025] Step 204: Load the archived intrinsic and extrinsic parameter matrices of the stereo camera; after inputting the chili pepper images from the left and right eyes of the stereo camera, perform monocular distortion correction on the chili pepper images using the intrinsic and extrinsic parameter matrices, followed by stereo epipolar correction to ensure that the same object in the chili pepper images of the left and right eyes is aligned in the same row; calculate the disparity map of the corrected chili pepper image relative to the left eye of the stereo camera in the current field of view using a global stereo matching algorithm; fill the holes in the disparity map using a multi-level mean filtering method; after filling, obtain the depth map matrix of the current chili pepper image relative to the left eye of the stereo camera by using the relationship between disparity and depth;

[0026] Step 205: Project the obtained two-dimensional coordinates of the chili stem onto the depth map matrix to obtain the three-dimensional coordinates of the chili stem in the depth map. Calculate the three-dimensional coordinates of the chili stem in the robotic arm coordinate system using the rotation transformation matrix obtained in step S1.

[0027] Preferably, in step S201, more than 1200 color pepper images, including various weather conditions, lighting conditions and target poses, are used to train the YOLOv10 model to obtain an ONNX format YOLOv10 training model.

[0028] Preferably, in step S202, the preprocessing process is as follows: the chili pepper image acquired by the binocular camera is scaled using an image downscaling algorithm; then the scaled chili pepper image is normalized.

[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0030] 1. The intelligent chili harvesting system of the present invention can quickly identify the location coordinates of chili peppers (stems), thereby achieving precise harvesting of chili peppers and reducing the breakage rate of chili pepper harvesting.

[0031] 2. The intelligent chili harvesting system of the present invention can effectively improve chili harvesting efficiency and reduce chili harvesting costs, and has a good market application prospect. Attached Figure Description

[0032] Figure 1 and Figure 2 These are schematic diagrams of the intelligent chili harvesting system of the present invention from two different perspectives.

[0033] Figure 3 This is a schematic diagram of the harvester.

[0034] Figure 4 This is a flowchart of the intelligent chili pepper harvesting method of the present invention.

[0035] Figure 5 This is a logic block diagram of an image processing program.

[0036] Figure 6 A logic diagram for performing forward reasoning.

[0037] Figure 7 A logic block diagram for performing foreground and background binarization separation.

[0038] Figure 8 Image of a chili pepper.

[0039] Figure 9 This is a foreground mask.

[0040] Figure 10A logic block diagram for performing binocular stereo depth matching.

[0041] Figure 11 This is a disparity map. Detailed Implementation

[0042] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0043] See Figures 1-3 The intelligent chili pepper harvesting system of the present invention includes a mobile vehicle 1, a visual positioning device for identifying the coordinate position of chili peppers mounted on the mobile vehicle 1, a chili pepper harvesting device for harvesting chili peppers according to their position coordinates, and a control device. The chili pepper harvesting device includes a base, a robotic arm 3 mounted on the base, a harvester 4 at the end of the robotic arm 3, and a harvesting drive mechanism 2 for driving the robotic arm 3 to move along the X and Y axes. The harvester 4 includes a bracket 401, a fixed blade 406 mounted on the bracket 401, a movable blade 405, and a mechanism for driving the movable blade 405 to swing in coordination with the fixed blade 406 to cut the chili peppers. The cutting drive mechanism includes a fixed blade 406 fixed on a support 401; a movable blade 405 rotatably connected to the support 401 and located above the fixed blade 406; the visual positioning device includes a binocular camera mounted on the base of the robotic arm 3; the binocular camera is used to acquire the first relative position coordinates of the chili stem relative to the binocular camera and transmit the first relative position coordinates to the control device; the control device calculates the third relative position coordinates of the chili stem relative to the base of the robotic arm 3 based on the first relative position coordinates and the second relative position coordinates of the binocular camera and the base of the robotic arm 3.

[0044] In this embodiment, the binocular camera is mounted directly in front of the base of the robotic arm 3 for identifying chili peppers.

[0045] See Figures 1-3 The shearing drive mechanism includes an electric push rod 402; the electric push rod 402 and the moving blade 405 are connected by a connecting rod 404, one end of the connecting rod 404 is rotatably connected to the telescopic rod of the electric push rod 402, and the other end is rotatably connected to the moving blade 405; the electric push rod 402 drives the telescopic rod 403 to extend or retract, thereby driving the moving blade 405 to rotate around its rotation fulcrum to complete the shearing action.

[0046] In this embodiment, the lower ends of the movable blade 405 and the fixed blade 406 are provided with elastic clamping blocks 407. When the movable blade 405 and the fixed blade 406 work together to cut off the chili stem, the elastic clamping blocks 407 at the lower ends of the movable blade 405 and the fixed blade 406 clamp the chili stem, thereby achieving integrated cutting and clamping.

[0047] See Figures 1-3 The robotic arm 3 is a six-degree-of-freedom robotic arm; both the X-axis drive mechanism and the Y-axis drive mechanism adopt a drive method combining a motor and a lead screw transmission mechanism; in addition, in order to ensure motion accuracy, both the X-axis drive mechanism and the Y-axis drive mechanism also include a guide mechanism, which adopts a combination of slide rail and slider, or a combination of guide rod and guide sleeve.

[0048] See Figure 4 The intelligent chili harvesting method of the present invention includes the following steps:

[0049] Step S1: Perform hand-eye calibration to obtain the rotation transformation matrix between the image coordinate system of the binocular camera and the coordinate system of the robotic arm;

[0050] Step S2: The binocular camera identifies the chili peppers, processes the acquired chili pepper images, and obtains the image coordinates of the chili pepper stem in the chili pepper image;

[0051] Step S3: Based on the rotation transformation matrix obtained in step S1 and the image coordinates of the chili stem obtained in step S2, calculate the three-dimensional coordinates of the chili stem in the robotic arm coordinate system.

[0052] Step S4: Based on the three-dimensional coordinates of the chili stem obtained in step S3, the robotic arm drives the harvester to harvest the chili.

[0053] Step S5: Repeat steps S2-S4 until the binocular camera no longer detects the chili peppers, at which point the harvesting work is complete.

[0054] The intelligent chili pepper harvesting method of this invention uses a program written in a mix of Python and C++, with OpenCV 4.9.0 as the backend computer algorithm library, the ONNX target recognition and detection model from YOLOv10 as the reasoning basis, ROS-noetic as the communication framework, and StereolabsZEDSDK as the camera driver and frontend library.

[0055] In step S1, before starting work, hand-eye calibration of the intelligent chili harvesting system of the present invention is required; in this embodiment, the calibration method of "eye outside hand" is adopted: assuming that the transformation matrix of the camera coordinate system relative to the robotic arm coordinate system is... Take n images with a camera. For each image, we have:

[0056]

[0057] In the formula: It is an unknown quantity; It can be determined by the robot arm's own pose parameters; This can be determined by camera calibration;

[0058] For each image, let:

[0059]

[0060] Solve the matrix equation:

[0061] A i X = XB i , i∈[1,n-1];

[0062] The rotation transformation matrix can then be calculated. Using the rotation transformation matrix obtained above, the three-dimensional coordinates of the chili pepper in the robotic arm coordinate system can be calculated, which is the position that the harvester at the end of the robotic arm should reach.

[0063] See Figure 5 In step S2, the image processing flow mainly consists of six parts: program initialization, forward inference, foreground and background binarization separation, stereo matching, point cloud reading, and message sending / visualization. These six main parts are connected sequentially and interpolated to select and filter at each level. Their respective code implementations are executed either sequentially or in parallel.

[0064] In the "program initialization" process, necessary dependent modules, such as OpenCV, ONNXRuntime, and ROS, are loaded first. Next, the deep learning model trained from YOLOv10 is loaded into memory for subsequent forward inference. Then, the ROS node is started, preparing for the publication of ROS topic messages. Finally, the startup parameters for the stereo camera are set, and the stereo camera is started. Each step has a self-checking and error handling process, using Python object references combined with try...except blocks for checking. If an error occurs at any step, it is determined that the program has not been correctly initialized, and the program immediately exits; if it has been correctly initialized, the core program loop begins.

[0065] This invention uses the YOLOv10 training model based on the PyTorch deep learning framework. It uses more than 1,200 images of chili peppers, including color images of various weather conditions, lighting conditions, and target poses, to train the model. The training is accelerated using NVIDIA GPU, and finally a YOLOv10 training model file in ONNX format is obtained, which will be used for the subsequent inference process.

[0066] See Figure 6 The first few steps of the "forward inference" process are handled by the Ultralytics module. To ensure stability and compatibility, the program scales the original 1280×720 color chili pepper image from the stereo camera to 720×720 using a local mean image downscaling algorithm, and then normalizes it. Let the size transformation matrix be C. Next, the chili pepper image is fed into an ONNX model trained on YOLOv10 for forward inference, and its output is parsed from the inferred vector. Since the chili pepper image has already been scaled and normalized, it is necessary to reverse engineer it to obtain the mapping of these output data onto the source chili pepper image, that is, to map the obtained original inference vector to C. -1 The calculation is performed. Finally, the Ultralytics module yields a value in the format (c... x c y The tuple contains (c, w, h), where each data point represents the center coordinate (c, w, h) of the detection bounding box of the chili pepper target on the stereo camera's view. x c y The coordinates (w, h) and width (w, h) of each bounding box are used. Since a single chili pepper image may contain multiple available data points, it is necessary to verify the availability of these data, i.e., verify the reasonableness of the coordinate dimensions of each bounding box, and its center coordinates (c...). x c y The area (w, h) must not exceed the width and height of the chili pepper image, and must not extend beyond the chili pepper image area. Assume the chili pepper image size is (p... x p y If the condition is met, then the following conditions should be satisfied:

[0067] 0≤c x ≤p x And 0≤c y ≤p y ;

[0068] 0≤c x +w≤p x And 0≤c y +w≤p y ;

[0069] Next, nonmaximum suppression and optimal target selection are manually performed on the dataset (accuracy). This filters out detection boxes located at the edges of regions with excessive overlap, retaining the detection boxes with the highest confidence in the central region. The center coordinates and ROI (regional area of ​​interest) of the detection boxes in the central region are defined as follows: The default confidence level is thres = 0.65.

[0070] After the above post-processing, the optimal inference result can be obtained. This optimal inference result and its related numerical values ​​are retained. bestThe bounding box is loaded into the computer memory for subsequent foreground and background binarization separation.

[0071] See Figure 7 The process of "performing foreground and background binarization separation" requires the use of OpenCV 4.9.0 as the image processing algorithm library in some parts, and the use of a self-developed algorithm for precise two-dimensional localization of chili peppers (stems) in others. Specifically:

[0072] First, initialize the parameters required for the GrabCut algorithm, such as the number of iterations k and the offset correction parameter ω, and input the images from the stereo camera. Next, input the detection bounding box of the target object from which to extract the foreground; this detection bounding box is obtained from the previous "perform forward inference" step. The effect is as follows Figure 8 As shown.

[0073] The GrabCut algorithm is executed iteratively for k iterations to generate a valid foreground image. The GrabCut algorithm is a composite iterative algorithm that iterates and calls the following operations: Gaussian Mixture Model, Foreground Estimate, and Background Color Distribution. A Markov Random Field is constructed using pixel labels to build the foreground and background. GraphCutOptimization is then applied to achieve the final segmentation effect.

[0074] After obtaining the foreground image, the image matrix can be further processed to obtain an effective foreground mask. The foreground is then processed to be pure white in the grayscale image, and the background is processed to be pure black in the grayscale image. The effect is as follows: Figure 9 As shown.

[0075] Subsequently, using the distance L between the two endpoints (e1, e2) in the same row of the white part of the mask matrix as the independent variable, starting from the bottom of the connected region... bottom h starts moving towards the top. top h, calculate the distance length d on the line segment. L Regarding line segment d h The maximum value of the absolute value of the derivative:

[0076]

[0077] The row containing the chili pepper shoot is the row with the largest absolute value of its derivative, located at the top of the connected mask, and with decreasing distances from the top. In this case, i should satisfy:

[0078]

[0079] By taking the perpendicular bisectors of the left and right endpoints of the mask in that row, the approximate coordinates of the chili pepper stem can be obtained. Compensation and correction are then performed based on the pose in the image to determine the final two-dimensional coordinates (t) of the chili pepper stem in the source image. x , t y ), where:

[0080]

[0081] The coordinate values ​​and the binarized image of the foreground mask are retained for use in subsequent processes.

[0082] See Figure 10 The "perform binocular depth matching" process mainly relies on the Global Stereo Matching (SGBM) algorithm in OpenCV 4.9.0 and previously archived binocular camera calibration parameters, specifically:

[0083] First, load the archived intrinsic parameter matrix M of the stereo camera. CamL M CamR V DistL V DiskR The intrinsic and extrinsic parameter matrices R, T, E, F, and Q are not directly read from a file but should be pre-loaded during the overall program initialization. After inputting the left and right eye images from the binocular camera, the aforementioned intrinsic and extrinsic parameter matrices are used to perform monocular distortion correction on the images, followed by binocular epipolar stereo correction, ensuring that the same object in the left and right eye images is roughly aligned in the same line, facilitating subsequent matching. Next, the necessary parameters for the global stereo matching algorithm are initialized, and the previously corrected binocular images are input along with the parameters. The disparity map of the current binocular camera's field of view relative to the left eye of the binocular camera is calculated, as shown below. Figure 11 As shown.

[0084] Due to uneven lighting or differences in object surface materials, disparity maps may contain many low-reliability holes. To fill these holes, the program uses a self-developed multi-level mean filtering method. Assume the hole region Ω = {(x, y) | D(x, y) = 0}, where D(x, y) = 0 indicates missing depth information. Define a local window W. r (x,y)={(x′,y′)||x′-x|≤r,||y'-y||≤r}, where r is the window radius.

[0085] For each scale r, calculate the local mean:

[0086]

[0087] After iterating from 1 to R multiple times using different scales for accumulation or weighted averaging, the final result of multi-level filtering is:

[0088]

[0089] Predict the hole parallax and fill it in evenly with the parallax of surrounding objects; this is the process of:

[0090]

[0091] After reliable filling is completed, the parallax-depth conversion formula is derived using the geometric relationships of parallel binocular vision. Where f is the focal length, B is the baseline distance, and D is the target parallax; the depth map matrix of the current image relative to the left eye is obtained; then, after data processing such as mean normalization, a deep copy is made to retain a visual grayscale depth map image for subsequent visualization / analysis.

[0092] Clearly, the size of the depth map matrix should be similar to the size of the left and right monocular images input from the stereo camera. If the size of the depth map (w...) d h d (p) does not meet the left / right monocular image size requirement x ±5px, p y If the image size is within ±5px, it is considered unqualified, and a new depth map or original image from a stereo camera should be obtained.

[0093] The process of "obtaining high-dimensional point mapping" is relatively simple, involving the lookup and retrieval of matrix data. Through simple matrix operations and pointer / subscript operator manipulation, the precise coordinates (t) of the chili pepper stalk in the left image of the stereo camera, which were previously calculated, can be retrieved. x , t y Projected onto the depth map matrix, and 3D coordinates (t) in millimeters are obtained. rx , t ry , t rz Subsequently, the rotation transformation matrix of the binocular camera-robotic arm base was obtained using hand-eye calibration. The three-dimensional coordinates (t) rx , t ry , t rz Transformed into the coordinates of the chili pepper stem in the robotic arm's coordinate system (t) bx , t by , t bz Clearly, the depth of this coordinate point must meet a certain requirement; otherwise, it is an invalid coordinate point.

[0094] 0≤t bz ≤2×10 4 ;

[0095] After verifying the effectiveness of all the above data analysis and processing procedures, the target's 3D point data can be packaged into a topic message object in ROS and broadcast using the ROS-noetic framework at a frequency of 2–3 Hz. Simultaneously, the stereo camera source image marking the chili pepper target, the foreground mask grayscale image with foreground and background binarized separation, and the depth map are refreshed and displayed on the computer monitor. Important parameters are refreshed and displayed in the console for user observation or developer debugging. After completing the above process, a new loop automatically begins, starting the analysis and recognition of the target in the next frame of the chili pepper image.

[0096] Unless the user manually confirms exiting the program or the program process is not killed, the program will continue to run in a loop, constantly reading images and performing analysis and calculations.

[0097] The stereo camera used in this invention is a Stereolabs ZED stereo camera, which, by default, can acquire multiple color 8-bit RGB format stereo images in various formats such as 3840×1080 30fps, 2560×720 30fps, and VGA 100fps. Taking 2560×720 30fps as an example, after cropping, it can obtain color images of 1280×720 30fps on the left and right respectively. In addition, the camera has a baseline length of 120mm, a depth map of 1280×720 30fps 32bit, an F / 2.0 aperture, and functions such as automatic white balance and automatic exposure adjustment. Compared with depth cameras such as RGB-D and TOF, this camera is more suitable for field operations or outdoor scenes.

[0098] The above are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above content. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A smart chili pepper harvesting method, characterized in that, A smart chili harvesting system is used to harvest chilies. The system includes a visual positioning device for identifying the coordinates of the chilies, a chili harvesting device for harvesting the chilies based on their coordinates, and a control device. The chili harvesting device includes a base, a robotic arm mounted on the base, a harvester at the end of the robotic arm, and a harvesting drive mechanism for driving the robotic arm to move along the X and Y axes. The harvester includes a support, a fixed blade, a moving blade, and a shearing drive mechanism for driving the moving blade to swing and cooperate with the fixed blade to complete a cutting action. The fixed blade is fixed to the support; the moving blade is rotatably connected to the support and located above the fixed blade. The visual positioning device includes a binocular camera mounted on the base of the robotic arm. The specific intelligent chili harvesting method includes the following steps: Step S1: Perform hand-eye calibration to obtain the rotation transformation matrix between the image coordinate system of the binocular camera and the coordinate system of the robotic arm; Step S2: The binocular camera identifies the chili peppers and processes the acquired chili pepper images to obtain the image coordinates of the chili pepper stem within the image. The process of processing the chili pepper images includes: Step 201: Train the YOLOv10 model using the chili pepper images acquired by the binocular camera to construct a YOLOv10 training model in ONNX format; Step 202: Preprocess the acquired chili pepper images, feed the preprocessed chili pepper images into the YOLOv10 training model to perform forward inference, and parse the output from the inferred vectors to obtain the best inference result; the steps for performing forward inference are as follows: The original 1280×720 color chili pepper image from the stereo camera was scaled down to 720×720 using a local mean image downscaling algorithm and then normalized. The size transformation matrix is ​​as follows: ; Pepper images are fed into an ONNX model trained from YOLOv10 for forward inference, and the output is parsed from the inferred vector. Backpropagation is then performed to obtain the mapping of these output data onto the source pepper images, i.e., the original inference vectors are compared with... Perform calculations; obtain a format using the Ultralytics module. The tuple contains data points representing the center coordinates of the detection bounding box of the chili pepper target on the stereo camera's image. With width and height ; Verify the rationality of the coordinate dimensions of each detection box, and the center coordinates of the detection box. The dimensions must not exceed the width and height of the chili pepper image. The image must not extend beyond the chili pepper image area; assuming the chili pepper image size is... Then it should satisfy: and ; and ; Next, nonmaximum suppression and optimal target selection are manually performed on the dataset. This filters out edge-based detection boxes within regions with multiple overlapping detection boxes, retaining the detection boxes with the highest confidence in the central region. The center coordinates and ROI (Region of Interest) of the detection boxes in the central region are defined as follows: Default confidence level ; After the above processing, the optimal reasoning result can be obtained. This optimal reasoning result and its related numerical values ​​are retained. Stored in computer memory for subsequent foreground and background binarization separation; Step 203: Combine the obtained best inference results, execute the GrabCut algorithm on the chili pepper image to generate an effective foreground mask; continuously solve the partial derivatives of the connected regions of the foreground mask from bottom to top, find the maximum value of the partial derivative and determine its direction to obtain the approximate two-dimensional coordinates of the chili pepper stem, and perform offset correction on the approximate two-dimensional coordinates to determine the final two-dimensional coordinates of the chili pepper stem. Step 204: Load the archived intrinsic and extrinsic parameter matrices of the stereo camera; after inputting the chili pepper images from the left and right eyes of the stereo camera, perform monocular distortion correction on the chili pepper images using the intrinsic and extrinsic parameter matrices, followed by stereo epipolar correction to ensure that the same object in the chili pepper images of the left and right eyes is aligned in the same row; calculate the disparity map of the corrected chili pepper image relative to the left eye of the stereo camera in the current field of view using a global stereo matching algorithm; fill the holes in the disparity map using a multi-level mean filtering method; after filling, obtain the depth map matrix of the current chili pepper image relative to the left eye of the stereo camera by using the relationship between disparity and depth; Step 205: Project the obtained two-dimensional coordinates of the chili stem onto the depth map matrix to obtain the three-dimensional coordinates of the chili stem in the depth map. Calculate the three-dimensional coordinates of the chili stem in the robotic arm coordinate system using the rotation transformation matrix obtained in step S1. Step S3: Based on the rotation transformation matrix obtained in step S1 and the image coordinates of the chili stem obtained in step S2, calculate the three-dimensional coordinates of the chili stem in the robotic arm coordinate system. Step S4: Based on the three-dimensional coordinates of the chili stem obtained in step S3, the robotic arm drives the harvester to harvest the chili. Step S5: Repeat steps S2-S4 until the binocular camera no longer detects the chili peppers, at which point the harvesting work is complete.

2. The intelligent chili pepper harvesting method according to claim 1, characterized in that, The binocular camera is used to acquire the first relative position coordinates of the chili pepper stem relative to the binocular camera, and transmits the first relative position coordinates to the control device. The control device calculates the third relative position coordinates of the chili pepper stem relative to the base of the robotic arm based on the first relative position coordinates and the second relative position coordinates of the binocular camera and the base of the robotic arm.

3. The intelligent chili pepper harvesting method according to claim 1, characterized in that, The robotic arm is a six-degree-of-freedom robotic arm.

4. The intelligent chili pepper harvesting method according to claim 1, characterized in that, Both the X-axis drive mechanism and the Y-axis drive mechanism adopt a drive method that combines a motor and a lead screw transmission mechanism.

5. The intelligent chili harvesting method according to claim 4, characterized in that, Both the X-axis drive mechanism and the Y-axis drive mechanism further include a guide mechanism, which adopts a combination of slide rail and slider, or a combination of guide rod and guide sleeve.

6. The intelligent chili pepper harvesting method according to claim 4, characterized in that, The shearing drive mechanism includes an electric push rod; the electric push rod and the moving blade are connected by a connecting rod, one end of the connecting rod is rotatably connected to the telescopic rod of the electric push rod, and the other end is rotatably connected to the moving blade.

7. The intelligent chili pepper harvesting method according to claim 1, characterized in that, In step S201, more than 1,200 color pepper images, including various weather conditions, lighting conditions, and target poses, are used to train the YOLOv10 model, resulting in an ONNX format YOLOv10 training model.

8. The intelligent chili pepper harvesting method according to claim 7, characterized in that, In step S202, the preprocessing process is as follows: the chili pepper image acquired by the binocular camera is scaled using an image downscaling algorithm; then the scaled chili pepper image is normalized.

Citation Information

Patent Citations

  • Wheel type pepper automatic picking robot

    CN116472868A

  • Pepper picking machine

    CN117918119A

  • Harvesting robot control system and control method based on visual servo

    CN111673755A

  • Pepper harvesting and managing integrated robot based on machine vision

    CN118542149A