Root vegetable leaf cutting method and system based on deep learning
By using deep learning technology to accurately locate the cutting point of root vegetables and combining it with multi-threaded processing, the problems of flesh damage and inconsistent leaf length in existing mechanical leaf cutting processes have been solved, achieving safe and efficient leaf cutting operations for root vegetables.
Patent Information
- Application Number
- CN202511715091.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-10
AI Technical Summary
Existing harvesting machinery for root vegetables is prone to damaging the fleshy parts during the leaf-cutting process, especially in the leaf-cutting operation of root vegetables such as white radish, making it difficult to ensure the consistency and safety of the cut leaf length.
A deep learning-based leaf-cutting method for root and stem vegetables is adopted. Through multi-threaded collaborative processing of image acquisition, fleshy body detection model and leaf-cutting decision model, the cutting point is accurately located and leaf-cutting operation is performed. This includes image preprocessing, real-time image data processing, prediction thread interpolation and decision thread identification of cutting position.
It significantly improves the safety and precision of leaf cutting operations for root vegetables, avoids damage to the flesh, ensures consistent leaf length, and enhances the level of automation.
Smart Images

Figure CN121505599A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of root vegetable harvesting machines, in particular to a root vegetable leaf cutting method and system based on deep learning. BACKGROUND
[0002] After the harvesting machine completes the pulling operation of the root vegetable, the excess leaves need to be cut off, and the leaves of a specific length need to be retained according to the user's requirements, and the length of the retained leaves needs to be consistent. However, due to the different depths of root fruits buried in the soil, and the differences in soil viscosity and water content, the pulling of root vegetables requires different forces and difficulties. These factors cause the height of the pulled root fruits below the conveying chain to have a large difference, and it is difficult to achieve the operation of consistent leaf cutting length. The existing root vegetable harvesting machines, such as carrot harvesters, usually use two conveying chains to determine the cutting length of the leaves during the leaf cutting process. The specific operation method is as follows: after successfully pulling the vegetable, an upper inclined conveying chain clamps the carrot leaves and conveys the carrot upward and backward; another horizontal conveying chain located below and intersecting with the upper conveying chain also clamps the carrot leaves and conveys them backward. During the conveying process, the upper inclined chain continuously pulls the carrot upward, and the lower horizontal chain pulls the carrot downward at the same time, until the neck of the carrot reaches the lower edge of the horizontal chain. At this time, due to the blocking action of the chain, the carrot stops rising, so that all the carrots are pulled to the same height to ensure the consistent cutting length of the carrot leaves.
[0003] Although the above mechanical method has a certain leaf cutting effect on some root vegetables such as carrots, when cutting the leaves of white carrots, the flesh of the carrot may contact the lower end of the horizontal chain during the upward pulling process. In a more serious case, the flesh may even pass over the lower end of the horizontal chain and be clamped by the chain. In either case, the flesh of the carrot will be damaged during the cutting process. SUMMARY
[0004] The present application aims to at least solve the technical problem in the prior art that the flesh of the carrot will be damaged during the cutting process. The present application particularly innovatively provides a root vegetable leaf cutting method and system based on deep learning.
[0005] In order to achieve the above-mentioned purpose of the present application, the present application provides a root vegetable leaf cutting method based on deep learning, which comprises the following steps: S1, collecting image data of the root vegetable, and positioning the cutting point based on the image data using a flesh detection model; S2, preprocessing the image data, and training a leaf cutting decision model using the preprocessed image data; S3, a photographing thread using the leaf cutting decision model collects real-time images of the root vegetable in the harvesting process, locates the cutting point of the real-time image using the fleshy body detection model, and obtains a positioning result; based on the positioning result and a timestamp, an fruit object is packaged and the fruit object is pressed into a real object queue; S4, a prediction thread using the leaf cutting decision model performs position prediction and interpolation on the real object queue to obtain a smooth prediction object queue; S5, a decision thread using the leaf cutting decision model identifies the cutting position in the prediction object queue, and performs leaf cutting operation on the root part of the root vegetable based on the cutting position.
[0006] In another aspect, the present application also provides a root vegetable leaf cutting system, which is used to implement the deep learning-based root vegetable leaf cutting method; the system comprises: A transport unit is obliquely arranged to continuously transport the root vegetable; An image acquisition module is arranged on the transport unit to acquire image data of the root vegetable on the transport unit; A decision module is connected with the image acquisition module to execute a photographing thread, a prediction thread and a decision thread of the leaf cutting decision model according to the image data, to obtain real-time image data through the photographing thread, to locate the cutting point using the fleshy body detection model and to package the fruit object into the real object queue; to perform motion trajectory prediction and interpolation processing on the real object queue through the prediction thread to generate a smooth prediction object queue; and to identify the cutting position based on the prediction object queue through the decision thread and to output a control instruction; A lifting and cutting assembly is arranged at the bottom of the transport unit to perform leaf cutting according to the control instruction; A pick-up unit is arranged below the lifting and cutting assembly to pick up the cut fruit.
[0007] The application has the beneficial effects that: the application accurately locates the leaf cutting point through a deep learning meat body detection model, and effectively solves the problem of meat body damage caused by the height difference of rhizomes in the traditional mechanical leaf cutting process through the multi-thread cooperative processing of the leaf cutting decision model (the photographing thread collects images in real time, the prediction thread interpolates and smooths the position, and the decision thread dynamically identifies the cutting position). Specifically, the meat body detection model accurately identifies the meat body boundary and locates the cutting point through a convolutional neural network, avoiding direct contact or clamping damage of chain mechanical pulling on white radish meat body; the leaf cutting decision model improves the robustness of image recognition under different light and soil conditions through real-time image preprocessing, visual feature clustering optimization and data enhancement training, and ensures the consistency of leaf cutting length; the prediction thread compensates for the time delay of image acquisition and cutting action through position interpolation and smoothing of the queue trajectory, so that the cutting position prediction is more consistent with the actual motion trajectory; the decision thread further reduces the risk of cutting error through the height cooperative judgment of the neighbor cutting object while ensuring the cutting efficiency. The scheme significantly improves the safety, accuracy and automation level of rhizome vegetable leaf cutting operation, especially solves the problem of meat body protection of white radish and other easily damaged rhizome vegetables in the leaf cutting process.
[0008] Additional aspects and advantages of the application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and / or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and / or additional aspects and advantages of the application will become apparent and be readily appreciated from the description of the embodiments, which follows, when considered in connection with the following drawings, in which: Figure 1 is a flow chart of a rhizome vegetable leaf cutting method based on deep learning in embodiment 1 of the application; Figure 2 is a structural schematic diagram of a rhizome vegetable leaf cutting system based on deep learning in embodiment 2 of the application; Figure 3 is a structural schematic diagram of a rhizome vegetable leaf cutting system based on deep learning in embodiment 2 of the application; Figure 4 is a front view of a rhizome vegetable leaf cutting system based on deep learning in embodiment 2 of the application; Figure 5 is a structural schematic diagram of a lifting and cutting assembly of a rhizome vegetable leaf cutting system based on deep learning in embodiment 2 of the application; Figure 6 is a bottom view of a lifting and cutting assembly of a rhizome vegetable leaf cutting system based on deep learning in embodiment 2 of the application.
[0010] In the diagram: 1. Transport unit, 2. Image acquisition module, 3. Lifting and shearing assembly, 301. Rotating base, 302. Square lifting rod, 303. Shearing blade, 304. Connector, 305. Mounting bracket, 306. Rotary bearing, 307. Lifting rack, 308. Lifting driver, 309. Rotary driver, 310. Slider, 4. Pick-up and drop-off unit. Detailed Implementation
[0011] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0012] Example 1 like Figure 1 As shown, a deep learning-based method for cutting leaves from root vegetables is described, the method comprising: S1. Collect image data of root vegetables, and locate the cutting point based on the image data using a fleshy body detection model; S2. Preprocess the image data and use the preprocessed image data to train the leaf cutting decision model; S3. Use the leaf-cutting decision model's photo-taking thread to collect real-time images of root vegetables during the harvesting process, use the fleshy body detection model to locate the cut points of the real-time images, and obtain the location results; based on the location results and timestamps, encapsulate the fruit objects and push the fruit objects into the real object queue. In step S3, it is necessary to explain in detail that the image-capturing thread continuously captures dynamic images of the transported root vegetables using a high frame rate camera, with each frame accompanied by precise timestamp information. The fleshy body detection model performs convolution operations on the real-time images, identifies the boundary contours of the fleshy bodies through a multi-layer feature extraction network, and accurately locates the coordinates of the tangent points using a non-maximum suppression algorithm. The location results and corresponding timestamps are encapsulated into a fruit object data structure containing location, time, and morphological features. This data structure is pushed into the real object queue in chronological order through a queue management mechanism. The queue uses a circular buffer for dynamic storage, ensuring that the latest 100 frames of fruit object data are traceable. When the queue length reaches a preset threshold, the system automatically triggers the prediction thread to perform trajectory interpolation processing. At the same time, the decision module continuously monitors the timeliness of the objects at the head of the queue. If there are outdated objects that have not been updated for more than 0.5 seconds, the queue cleanup mechanism is activated to remove them, ensuring the real-time performance and accuracy of the predicted object queue.
[0013] S4. Use the prediction thread of the leaf-cutting decision model to predict and interpolate the positions of the real object queue to obtain a smooth predicted object queue. S5. Use the decision thread of the leaf-cutting decision model to identify the cutting position in the prediction object queue, and perform leaf-cutting operation on the root and stem parts of root vegetables based on the cutting position.
[0014] In this embodiment, a deep learning-based leaf-cutting method for root vegetables works as follows: First, image data of the root vegetables during transportation is acquired using a camera. Then, a fleshy body detection model is used to process the image data. This model employs a pre-trained convolutional neural network architecture, accurately identifying fleshy body boundaries and locating cutting points through multi-layer feature extraction and existing boundary regression algorithms. After obtaining the location results, each root vegetable is encapsulated into a fruit object based on timestamp information and pushed into a real object queue to ensure data orderliness and real-time performance. Subsequently, the prediction thread of the leaf-cutting decision model performs motion trajectory prediction and interpolation processing based on the real object queue. A linear interpolation algorithm is used to compensate for the time delay between image acquisition and the cutting action, generating a smooth prediction object queue. The decision thread then identifies the cutting position from this queue and dynamically adjusts the start-up timing and height parameters of the lifting and lowering cutting components by calculating the time it takes for the cutting object to move to the target position. Specifically, this includes detecting the horizontal distance of neighboring cutting objects; if it is less than a preset threshold, a higher cutting height is used collaboratively to avoid miscutting and improve efficiency. Finally, the control unit drives the lifting and shearing assembly to perform precise leaf-cutting operations according to the instructions output by the decision thread, while the receiving and conveying unit receives the cut fruit, realizing fully automated processing.
[0015] As an optional embodiment of the present invention, optionally, locating the cutting point based on the fleshy body detection model using image data in step S1 includes: S101. Preprocess the image data, including color space conversion or contrast sharpening, Gaussian filtering for noise reduction, and size normalization. In step S101, it should be noted that color space conversion enhances the color difference between the fleshy body and the background by converting the original RGB image to HSV or Lab color space, which facilitates subsequent feature extraction; contrast sharpening uses the Laplacian operator to enhance image edge details and improve the clarity of the fleshy body boundary; Gaussian filtering denoising suppresses image noise through a 3×3 convolution kernel while retaining key feature information; size normalization scaling all images to a uniform resolution of 256×256 pixels ensures the consistency of input data dimensions.
[0016] S102. Input the preprocessed image data into the pre-trained fleshy body detection model. The fleshy body detection model adopts a convolutional neural network architecture, identifies fleshy body regions through the feature extraction layer, and outputs bounding box coordinates. In step S102, it should be noted that the fleshy body detection model in this embodiment adopts the YOLOv5 architecture. This architecture extracts features through the backbone network, reduces computation and improves feature fusion capabilities by utilizing the Cross-Stage Local Network (CSPNet), achieves multi-scale feature fusion through the Path Aggregation Network (PANet) in the neck network, and finally outputs detection results containing bounding box coordinates and class probabilities in the head network. Specifically, the input image is first sliced by the Focus module, dividing the original image into four independent feature maps and then stitching them together to achieve downsampling without losing information. Subsequently, deep feature extraction is performed through multiple BottleneckCSP modules, each containing a residual connection to alleviate the gradient vanishing problem. In the feature fusion stage, the SPP module expands the receptive field through max pooling, while PANet enhances the expressive power of features at different levels through bidirectional fusion paths from top to bottom and bottom to top. Finally, the output bounding box coordinates are processed by non-maximum suppression (NMS) to obtain accurate fleshy body region localization results.
[0017] The training method for the fleshy body detection model is as follows: First, an image dataset containing different types of root vegetables (long white radish, round white radish, red-skinned radish, purple-skinned radish, yellow radish, carrot, garlic, single-clove garlic, jicama, sweet potato, beet, and other root crops) is collected. This dataset needs to cover various lighting conditions, shooting angles, and growth stages to ensure data diversity. The dataset is then labeled, marking the bounding boxes of the fleshy body regions, and divided into training, validation, and test sets.
[0018] A transfer learning strategy was employed, loading the weights of a YOLOv5 model pre-trained on the COCO dataset as initial parameters. Fine-tuning was then performed on a root vegetable dataset using backpropagation. During training, a stochastic gradient descent (SGD) optimizer was used, with an initial learning rate of 0.001, a momentum parameter of 0.937, and a weight decay coefficient of 0.0005. To improve model robustness, Mosaic data augmentation was used during training, stitching four randomly scaled and cropped images together into a single image for training. Simultaneously, the brightness, contrast, saturation, and Gaussian noise were randomly adjusted. Model performance was evaluated on a validation set after every 10 epochs, using mean accuracy (mAP) as the evaluation metric. If the validation set mAP did not improve for three consecutive epochs, the learning rate was decayed to 0.1 times its original value. The final trained fleshy body detection model achieved an mAP accuracy of 96.2% on the test set, and maintained a detection accuracy of over 90% even in low-light (<50 lux) and occlusion (occlusion area >30%) scenarios.
[0019] S103. Use the vertices of the bounding box coordinates as tangent points; In step S103, it should be noted that using the vertices of the bounding box coordinates as tangent points is based on the precise localization of the succulent region using the succulent detection model. Since the bounding box accurately defines the succulent region, its vertex positions typically correspond to the key points where the succulent connects to the rootstock, which are the tangent points that need to be precisely located during leaf cutting. This method ensures that leaf cutting is performed in the correct position, avoiding unnecessary damage to the succulent, while guaranteeing the accuracy and consistency of leaf cutting. In practical applications, this localization method is adaptable to different varieties and growth stages of root vegetables, exhibiting high versatility and practicality.
[0020] S104. Perform real-time calibration of the fleshy body detection model.
[0021] In step S104, it should be noted that the purpose of real-time calibration is to ensure that the fleshy body detection model can maintain high-precision positioning under different environmental conditions (such as changes in light, soil pollution interference, and stains on the surface of vegetables). The specific calibration process includes: first, continuously acquiring images of root vegetables in the current environment through the image acquisition module and extracting key features (such as texture and edge gradient); then, comparing the current features with the standard feature library of the pre-trained model and calculating the feature offset; if the offset exceeds a preset threshold (such as 5%), a dynamic calibration mechanism is triggered—updating the model parameters through online incremental learning, specifically using the mini-batch gradient descent method, using the current batch of images as input, and only adjusting the weights of the last layer of the classifier in the model to avoid destroying the converged feature extraction layer; at the same time, a teacher-student model architecture is introduced, using the original high-precision model as the teacher model, generating soft labels to supervise the update process of the student model (the currently running model) to prevent overfitting due to insufficient data; after calibration, the model performance is verified on the test set, and if the mean accuracy (mAP) recovers to above 95%, calibration is stopped, otherwise iterative optimization continues. This real-time calibration mechanism significantly improves the robustness of the model in complex agricultural scenarios, ensuring that the tangent point positioning accuracy always meets production requirements.
[0022] As an optional embodiment of the present invention, training the leaf-cutting decision model in step S2 may include: S201. Label the image data and use it as initial training data to train the initial detection model; In step S201, it should be noted that the annotation content must cover key information such as the cut point location, fleshy body boundary, and overall outline of the root and stem of root vegetables. This annotated data will serve as initial training data for training the initial detection model. The annotation process must ensure the accuracy and consistency of the data, typically employing a combination of manual and semi-automatic annotation to improve efficiency and accuracy. Manual annotation is performed by professionals based on image features to ensure the reliability of the annotation results; semi-automatic annotation utilizes image processing algorithms (such as edge detection, region segmentation, etc.) to assist in generating preliminary annotation results, which are then corrected and confirmed manually. The initial detection model can employ a classic convolutional neural network architecture (such as VGG, ResNet, etc.), using backpropagation algorithms and optimizers (such as stochastic gradient descent, Adam, etc.) to iteratively optimize the model parameters, enabling the model to initially identify key features in the image and output prediction results. During training, the initial training data needs to be divided into a training set, a validation set, and a test set, used for model training, parameter tuning, and performance evaluation, respectively. By continuously adjusting the model structure and hyperparameters (such as learning rate and batch size), the model's performance on the validation set is optimized, and finally, the model's generalization ability is verified on the test set.
[0023] S202. Extract visual feature vectors from the initial training data using the initial detection model; In step S202, it should be noted that the visual feature vector is a numerical representation of key features in an image (such as cut point location, fleshy body boundary, root outline, etc.), typically generated through the feature extraction layer of a convolutional neural network. Specifically, the initial detection model (such as VGG or ResNet) extracts multi-scale features from low-level texture to high-level semantics layer by layer through multiple convolution, pooling, and activation operations. For example, shallow convolutional layers may capture edge and color information, while deep convolutional layers can identify more complex shapes and structures. After passing through global average pooling or fully connected layers, these features are compressed into a fixed-dimensional vector (such as 512-dimensional or 1024-dimensional), i.e., the visual feature vector. This vector not only retains the key information of the image but also reduces computational complexity through dimensionality reduction, facilitating subsequent classification or regression tasks. In the training of the leaf-cutting decision model, the visual feature vector will be used as input to train the position interpolation model of the prediction thread and the shearing position recognition model of the decision thread, thereby improving the model's adaptability and positioning accuracy for different root vegetables.
[0024] S203. Cluster the visual feature vectors in all the initial training data to obtain feature clusters; In step S203, it should be noted that clustering is an unsupervised learning method. Its purpose is to divide samples in a dataset into several clusters, such that samples within the same cluster have high similarity, while samples in different clusters have low similarity. In this embodiment, the K-means clustering algorithm is used to cluster visual feature vectors. First, the number of clusters K needs to be determined, which can be determined by methods such as the elbow rule or silhouette coefficient. After determining the value of K, K visual feature vectors are randomly selected as initial cluster centers. Then, the distance between each visual feature vector and each cluster center is calculated, and the vector is assigned to the cluster containing the nearest cluster center. Next, the cluster center of each cluster (i.e., the mean of all visual feature vectors in that cluster) is recalculated, and the sample division and cluster center update are performed again until the cluster centers no longer change or the preset number of iterations is reached. Through clustering, root vegetable image data with similar features can be grouped into one category to obtain feature clusters. These feature clusters can reflect the distribution patterns of visual features of different root vegetables. For example, different varieties of root vegetables may differ in shape, texture, etc. Clustering can reflect these differences, enabling the model to better learn the characteristics of different varieties of vegetables, thereby improving the accuracy of leaf cutting decisions.
[0025] S204. Set a distance threshold; when a visual feature vector exceeds the distance threshold, remove that visual feature vector; obtain the training dataset. In step S204, it should be noted that the main purpose of setting the distance threshold is to remove outliers or isolated points. These data points may have inaccurate feature extraction due to factors such as noise, occlusion, and abnormal lighting during image acquisition, thus affecting the training effect and generalization ability of the model. Specifically, for each visual feature vector, the distance between it and the center of its feature cluster is calculated (usually using Euclidean distance or cosine similarity as a metric). If the distance exceeds the preset distance threshold, the visual feature vector is considered an outlier and is removed from the initial training data. After removing outliers, the remaining visual feature vectors constitute the training dataset. This dataset has higher data quality and consistency, and can better reflect the visual feature distribution patterns of different root vegetables. In this way, the training efficiency and accuracy of the leaf-cutting decision model can be improved, making it more robust and adaptable in practical applications.
[0026] S205. Based on the training dataset, use image interpolation algorithms to reduce the distance between training datasets; In step S205, it should be noted that the image interpolation algorithm in this embodiment is as follows: First, a pre-trained diffusion model is fine-tuned using training data; from the clustering results of the training data, a pair of images with a distance exceeding a certain predefined threshold is selected; or the distance between images is defined by the vector distance between visual features output by the detection model, and a pair of images with a distance exceeding the predefined threshold is still selected; a new visual feature vector is calculated using the interpolation algorithm; this new vector is used as a guide to generate multiple new images using the diffusion model; finally, valid newly generated images are manually or automatically selected, and object detection boxes are labeled.
[0027] In step S205, it should be further explained that the core of the above image interpolation algorithm lies in filling in the feature gaps in the dataset caused by differences in variety or growth stage by generating new images between the original images. Specifically, the diffusion model fine-tuning stage adopts a pre-trained StableDiffusion architecture, freezing its encoder-decoder backbone network and only adjusting the parameters of the attention mechanism in the intermediate layers to adapt it to the texture features of root vegetables. In the image pair selection stage, in addition to the cluster center distance threshold based on Euclidean distance (e.g., set to 0.8), a Mahalanobis distance metric in the feature space is also introduced to eliminate the dimensional differences between different feature dimensions. During interpolation calculation, spherical linear interpolation (Slerp) based on manifold learning is used to generate new features with smooth transitions on the unit hypersphere of the feature vector, avoiding semantic distortion caused by direct linear interpolation. In the diffusion generation stage, vegetable variety labels and growth stage labels are injected by controlling the conditional encoder to ensure the semantic consistency of the generated images. The new image selection process employs a dual-metric evaluation: firstly, it calculates the prediction confidence of the generated images using the initial detection model (≥0.95); secondly, it uses the Structural Similarity Index (SSIM) to measure visual similarity with the original images (≥0.85). The final qualified images are then corrected for cutpoint positions using a semi-automatic annotation tool, achieving a 40% improvement in annotation efficiency compared to purely manual annotation, with an annotation consistency (Kappa coefficient) exceeding 0.92. This interpolation process increases the distribution density of the training dataset by three times, significantly improving the model's localization accuracy for rare varieties (such as purple radishes) and special growth morphologies (such as forked carrots), resulting in an 8.7 percentage point improvement in mAP on the test set.
[0028] S206. Use data augmentation techniques to expand the training dataset and generate diverse training samples. In step S206, it should be noted that data augmentation expands the scale and diversity of the training dataset by performing a series of transformations on the original training dataset to generate new data samples that are similar to but not exactly the same as the original data. In this embodiment, the data augmentation methods used include, but are not limited to, the following: First, geometric transformations, such as random rotation, flipping (horizontal and vertical flipping), scaling, and cropping. Random rotation can rotate the image within a certain angle range (e.g., -30 degrees to 30 degrees) to simulate vegetable images under different shooting angles; flipping operations can increase the symmetry of the image; scaling and cropping can change the size and local content of the image, allowing the model to adapt to the characteristics of vegetables of different sizes and positions. Second, color transformations, such as adjusting the brightness, contrast, and saturation of the image. By randomly changing these parameters, vegetable images under different lighting conditions can be simulated, enhancing the model's robustness to changes in lighting. For example, reducing brightness can simulate a cloudy day or a dimly lit environment, while increasing brightness simulates a sunny day or a brightly lit environment. Third, adding noise, such as Gaussian noise or salt-and-pepper noise, to the image. Adding noise can simulate interference factors that may occur during image acquisition, such as sensor noise and transmission noise, enabling the model to accurately identify and locate cut points even when faced with noisy images. Fourthly, sample mixing involves combining multiple images to generate new image samples. Alpha mixing can be used to fuse two or more images in a certain proportion, creating new images with various feature combinations and further enriching the diversity of training data. Through these data augmentation methods, the generated large number of diverse training samples can effectively improve the generalization ability of the leaf cutting decision model, enabling it to better cope with various complex scenarios and different varieties of root vegetables in practical applications.
[0029] S207. The leaf-cutting decision model is trained using the expanded training dataset. The leaf-cutting decision model adopts a deep neural network structure and optimizes the model parameters through multiple rounds of iteration.
[0030] In step S207, it should be noted that the leaf-cutting decision model employs a deep neural network structure, which consists of multiple hidden layers, including convolutional layers, pooling layers, and fully connected layers, enabling it to automatically learn complex features and patterns in images. During training, the expanded training dataset is first divided into a training set, a validation set, and a test set, used for model training, parameter tuning, and performance evaluation, respectively. During training, the backpropagation algorithm and optimizers (such as stochastic gradient descent, Adam, etc.) are used to iteratively optimize the model parameters. Specifically, in each iteration, the model receives a batch of training images as input, calculates the prediction results through forward propagation, and calculates the loss function (such as cross-entropy loss, mean squared error loss, etc.) with the true labels to quantify the difference between the prediction results and the true values. Next, the gradient of the loss function with respect to each parameter is calculated using the backpropagation algorithm, and the optimizer updates the model parameters based on the gradient, gradually reducing the loss function. During iteration, the model's performance is monitored using the validation set. When the loss function on the validation set fails to decrease for several consecutive epochs or reaches the preset number of iterations, training is stopped to prevent overfitting. Finally, the model's generalization ability is evaluated on the test set, requiring the model to achieve a mean accuracy (mAP) of over 90% on the test set to ensure the accuracy and robustness of the leaf cutting decision. Furthermore, to further improve model performance, a transfer learning strategy can be employed, using model parameters pre-trained on a large-scale image dataset (such as ImageNet) to initialize some layers of the leaf cutting decision model, accelerating model convergence and improving feature extraction capabilities.
[0031] As an optional embodiment of the present invention, the expression of the leaf-cutting decision model is optionally: in, This represents the optimal parameter set of the shearing leaf decision model. This represents the total number of training samples. Represents the loss function. This represents a leaf-cutting decision model, taking augmented image data as input and predicting the shearing location as output. Indicates the data augmentation result of the first... training samples, express, Indicates the first The actual cut position label corresponding to each sample. Indicates the weighting coefficient. This represents the L2 regularization term. This indicates a data augmentation operation. This represents an image interpolation algorithm. This represents the data cleaning function. This represents a feature extraction network. This represents the original image after preprocessing. Represents the first clustering of visual feature vectors. Cluster center, This represents the upper limit threshold for determining whether a feature vector is an outlier.
[0032] Here is another method for training the leaf-cutting decision model in step S2; specifically including: A201. Label the image data to obtain initial training data; The annotation method in step A201 is the same as that in step S201 above, so it will not be described in detail.
[0033] A202. The initial training data is processed using cleaning algorithms, image interpolation methods, and data augmentation methods to obtain the processed training dataset. The training dataset is then divided into training data and test data. A generation detection model is trained using the training data. In step A202, it should be noted that the cleaning algorithm is mainly used to remove noisy data and outliers from the initial training data. These outliers may be introduced due to image acquisition equipment malfunctions, environmental interference, or human factors, and can seriously affect the training effect of the model. The cleaning algorithm filters image data by setting a series of rules and thresholds. For example, a sharpness threshold can be set to remove images with a sharpness below the threshold; or a color range threshold can be set to remove images with abnormal colors (outside the normal color range of vegetables). The image interpolation method serves the same purpose as described in step S205 above, aiming to fill in feature gaps in the dataset caused by differences in varieties or growth stages, and to increase the diversity and richness of the dataset by generating new images that fall between the original images. The data augmentation method is also the same as the method described in step S206 above. By performing geometric transformations, color transformations, adding noise, and mixing samples on the original images, a large number of new data samples that are similar to but not exactly the same as the original data are generated, further expanding the scale of the training dataset. The cleaned, interpolated, and augmented training dataset is divided into training data and test data. The training data is used to train the first-generation detection model, while the test data is used for subsequent performance evaluation of the model. When training the first-generation detection model using the training data, a deep neural network structure similar to that in step S207 above is adopted. The model parameters are optimized through multiple rounds of iteration, enabling the model to automatically learn complex features and patterns in the image and achieve accurate prediction of the vegetable cutting position.
[0034] A203. Calculate the detection performance of the first-generation detection model using test data to obtain the first-generation performance index; In step A203, it's important to note that the first-generation performance index is a crucial metric for evaluating the performance of a first-generation detection model, reflecting its performance on test data. Specifically, various evaluation metrics can be used to calculate the first-generation performance index, such as accuracy, recall, F1 score, and mean AP. Accuracy represents the proportion of correctly predicted samples out of the total number of samples, directly reflecting the model's predictive accuracy. Recall represents the proportion of correctly predicted positive samples out of the actual number of positive samples, reflecting the model's ability to identify positive samples. The F1 score is the harmonic mean of accuracy and recall, comprehensively considering both the model's accuracy and recall capabilities. Mean AP is the average precision at different recall rates, providing a more comprehensive evaluation of the model's performance at different thresholds. By calculating these evaluation metrics, the first-generation performance index can be obtained, which comprehensively reflects the overall performance of the first-generation detection model on test data. If the first-generation performance index fails to meet the preset performance requirements, it indicates that the model's performance is poor under the current training data and parameter settings, requiring further optimization and improvement.
[0035] A204. Use cleaning algorithms, image interpolation methods, and data augmentation methods to perform secondary processing on the initial training data to obtain a new training dataset; use the new training dataset to train the second-generation detection model. In step A204, it should be noted that if the initial performance index does not meet the preset standard, secondary cleaning, interpolation, and enhancement processing of the initial training data are required. The secondary cleaning stage adds a dynamic threshold adjustment mechanism to the original rules, such as automatically correcting the sharpness threshold based on the local variance of the image, or dynamically determining the range of color anomalies through cluster analysis. The image interpolation stage introduces a Generative Adversarial Network (GAN) to assist in further optimizing texture details on the basic feature vector generated by spherical linear interpolation, enabling the new image to possess more realistic vegetable surface features while maintaining semantic consistency. The data augmentation strategy adds an elastic deformation operation, enhancing the model's adaptability to deformed samples by simulating non-rigid deformations (such as bending and twisting) during vegetable growth. The dataset after secondary processing uses stratified sampling to construct training and validation sets, ensuring the uniform distribution of rare variety samples across the data subsets. The second-generation detection model adopts an improved YOLOv8 architecture, embedding an attention mechanism module in the neck network, enabling the model to dynamically focus on the boundary region between vegetable leaves and stems. During training, a course learning strategy is introduced, initially using simple scenario samples for rapid convergence, and gradually increasing the weight of complex samples in later stages. Knowledge distillation is also employed, using the first-generation model as a teacher network to guide the second-generation model in learning more robust feature representations through soft label propagation. In the validation phase, five-fold cross-validation is used, ultimately requiring the second-generation model to achieve at least a 15% improvement in mAP on the independent test set compared to the first generation, and a recall rate of over 92% for detecting special shapes such as forked carrots.
[0036] A205. Calculate the detection performance of the second-generation detection model using the test data to obtain the second-generation performance index; the method for calculating the detection performance of the second-generation detection model is the same as that for the first-generation detection model, and will not be repeated here.
[0037] A206. Compare the performance index of the first generation and the performance index of the second generation to obtain the comparison result. When the comparison result reaches the performance threshold set by the grass fruit, repeat step S204. Otherwise, stop training and use the second generation detection model as the leaf-cutting decision model.
[0038] In step A206, it should be noted that the performance threshold is set based on the minimum accuracy requirements of the model in the actual application scenario. For example, in a high-speed sorting line, the mAP should be no less than 85% and the processing time for a single image should be less than 200ms. If the comparison results show that the second-generation performance index still does not meet the standard, the system will automatically trigger the iterative optimization process in step A204. At this time, a more complex model structure, such as a Transformer-based detection head, will be introduced, and a self-supervised pre-training strategy will be used to learn features on unlabeled vegetable image data. When the second-generation performance index exceeds the set threshold, the system will terminate the training process and freeze the model parameters, while generating a complete delivery package containing the model structure, weight file, and deployment code. Before final deployment, the stability of the model during continuous 72-hour operation must be verified through stress testing to ensure that it reaches a 99.9% availability standard in industrial application scenarios.
[0039] As an optional embodiment of the present invention, optionally, in step S4, the prediction thread of the leaf-cutting decision model is used to predict and interpolate the positions of the real object queue to obtain a smooth predicted object queue, including: S401. Extract the front-end cropped object from the real object queue. If no front-end cropped object is extracted, it means that no new real-time image data has been captured in the real object queue. Then proceed to step S404. In step S401, it should be noted that the real object queue acts as a dynamic data stream. When the queue management module detects new image data flowing in, it immediately triggers the front-end object extraction thread. This thread uses existing frame difference algorithms to locate significant change regions in the image and performs matching verification using a vegetable morphology feature library. If no target matching the contour features of root vegetables is detected for three consecutive frames, the queue is determined to be in an empty state. This indicates that no new real-time image data has been captured in the real object queue. At this time, the system automatically enters a low-power standby mode, maintaining only basic detection and sending a status query command to the image acquisition device every 5 seconds. When valid image data is detected again, the system immediately wakes up the core processing thread and resumes the real-time leaf-cutting decision service.
[0040] S402. If the front-end cropping object is extracted, the cropping object in the new real-time image data is horizontally aligned and matched with the original cropping object in the real object queue to obtain the aligned prediction object queue. In step S402, it should be noted that horizontal alignment is achieved through feature point registration technology, specifically using geometric feature points on the vegetable stem (such as bifurcation points and nodule centers) as the registration reference. The system uses the SIFT algorithm to extract locally invariant feature points of the vegetable stem in the old and new images, and performs coarse feature point matching using the FLANN matcher; subsequently, the RANSAC algorithm is used to remove mismatched point pairs. During the alignment process, it is robust to slight rotations (within ±5°) and scale changes (within ±10%) caused by conveyor belt vibration. The matching stage uses bidirectional optical flow tracing for verification. Using the key area of the vegetable (the top 20% of the stem-leaf junction) within the detection box of the previous frame as a reference, a dense optical flow field is calculated in the corresponding area of the new image. When the angle between the average motion vector direction and the conveyor belt direction is less than 15° and the displacement is within a preset pixel range, it is determined to be a consecutive frame of the same vegetable. The final generated aligned prediction object queue contains the spatiotemporal trajectory index of each vegetable instance, and its position confidence threshold is set to 0.85.
[0041] S403. If the match is successful, the position of the cut object in the new real-time image data is updated to the predicted position of the original cut object; if the match fails, the realism of the cut object in the real-time image data is reduced, and an updated predicted object queue is obtained. In step S403, it should be noted that upon successful matching, the system uses a Kalman filter to fuse and update the old and new position data. The process noise covariance is set to 0.01, and the measurement noise covariance is set to 0.05 to achieve a smooth transition in position prediction. During the update process, the new position data is integrated into the original predicted position using a weighted average algorithm. The weight coefficient is dynamically adjusted based on the object's real-time movement speed: when the conveyor belt speed is below 0.5 m / s, the new data weight is 0.8; when the speed is above 1.0 m / s, the weight is reduced to 0.6 to suppress the jitter caused by high-speed movement. The updated predicted position is immediately written to the shared memory area for the leaf-cutting execution module to call and trigger the position verification thread. The Euclidean distance threshold (default set to 5 pixels) is used to verify the reasonableness of the update result. If the position deviation is less than the threshold for three consecutive updates, the system automatically increases the object's confidence to 0.9 and reduces the subsequent image processing frame rate to 10 fps to save computational resources. Meanwhile, the matching history recorder increments the successful match count for the object. When the cumulative number of successful matches reaches 10, it is marked as a stable tracking object, and subsequent matching processes skip the coarse matching stage and directly enter optical flow verification, accelerating the processing flow. If a match fails (e.g., due to occlusion or sudden changes in illumination), the system activates an anomaly handling mechanism: first, it lowers the authenticity score of the shearing object to 0.6, and starts a re-detection thread to attempt re-matching within the next 3 frames; if it still fails, the object is marked as an outlier, removed from the prediction queue, and an alarm log is triggered. The updated prediction object queue is output through a double buffering mechanism to ensure data consistency and real-time performance, while generating a structured data packet containing location coordinates, confidence level, and timestamp for accurate execution by the downstream shearing control system.
[0042] S404. Remove cut objects with a truth value lower than the truth value threshold from the updated prediction object queue and re-sort them to obtain a smooth prediction object queue.
[0043] In step S404, it should be noted that the realism threshold in this embodiment is set to 0.75, and a dynamic adjustment mechanism is adopted: when the conveyor belt speed exceeds 1.2m / s, the threshold is increased to 0.85 to cope with the influence of motion blur; when the light intensity is below 100 lux, it is reduced to 0.7 to adapt to low visibility environments. The rejection operation adopts a two-stage verification: first, based on spatiotemporal continuity analysis, the standard deviation of the target's position within 5 consecutive frames is calculated. If it exceeds the preset displacement tolerance (default 15 pixels), it is marked as an abnormal trajectory; second, through feature consistency detection, the cosine similarity between the target's stem and leaf morphology features in the current frame and historical features is compared. If it is lower than 0.6, forced rejection is triggered. The re-sorting adopts a priority queue mechanism, and the sorting weight is composed of the formula W = 0.4 × position confidence + 0.3 × target size proportion + 0.2 × outline integrity + 0.1 × color saturation. For targets that temporarily disappear due to occlusion and then reappear, the system activates the trajectory backtracking module: a local search is initiated within a 50-pixel radius of the target's disappearance location. If the target is recaptured within 30 milliseconds and the feature matching degree is above 0.8, the original target ID is inherited and the missing trajectory is interpolated; otherwise, a new target ID is generated and the tracking parameters are initialized. The final output smooth prediction object queue must meet industrial-grade real-time requirements: single-frame processing latency does not exceed 33ms (corresponding to a 30fps video stream), and the queue length is dynamically compressed to within 10 objects. Each object contains three-dimensional spatial coordinates (x, y, scale), motion vectors (vx, vy), and a quality evaluation score (0-1). The system automatically performs queue integrity verification every 100 frames processed, ensuring output stability by calculating the effective target ratio (>85%) and the position jitter coefficient (<0.05).
[0044] As an optional embodiment of the present invention, optionally, identifying the shearing position in the prediction object queue using the decision thread of the leaf-cutting decision model in step S5 includes: S501. Find the first cut object that has not yet been cut from the smoothed prediction object queue. In step S501, it should be noted that the system determines the cutting status by detecting the processed status flag (initialized to unprocessed) of each object in the queue. The search process uses an efficient existing two-pointer scanning algorithm: the main pointer traverses sequentially from the head of the queue, and the auxiliary pointer records the position of the object where the cutting operation was last successfully performed. When the object pointed to by the main pointer is unprocessed and its authenticity score is higher than the current dynamic threshold (default 0.75), the object is immediately locked as the target to be cut, and its spatial coordinates and contour feature vector are loaded into the cutting execution buffer. If no matching object is found after scanning to the tail of the queue (e.g., the queue is empty or all objects are processed), the system triggers the queue refresh mechanism: pausing the current processing thread for 1 millisecond, waiting for new image data to be injected, and then re-executing the scanning process. In the case of auxiliary pointer failure due to system restart or abnormal interruption, the system starts the position backtracking module: based on the displacement data recorded by the conveyor encoder, combined with the object's historical trajectory, the current position is predicted, and local feature matching is started within a 20-pixel radius of the predicted position. After successful matching, the auxiliary pointer is updated and the object status is reset. Once the target is locked, the system will immediately update its status to "processing" and generate a task instruction package containing the target ID and the estimated processing time.
[0045] S502, Adjust the start time of the lifting shearing component based on the time it takes for the first shearing object to move to the shearing position; In step S502, it should be noted that the time adjustment mechanism is based on the real-time speed data (in m / s) fed back by the conveyor belt encoder and the predicted position coordinates (x, y) of the object. The start-up delay is calculated using the kinematic formula Δt=(dp) / v, where d is the fixed coordinate value of the shearing position (default x=1280 pixels), p is the x-coordinate of the current object's center point, and v is the instantaneous speed of the conveyor belt (collected by a Hall sensor at a sampling frequency of 1kHz). The system employs a dual closed-loop control strategy: the main loop dynamically compensates for the mechanical response delay using a PID controller (proportional coefficient Kp=0.8, integral time Ti=0.2s); the secondary loop uses a photoelectric position sensor to calibrate the actual position of the object in real time. When the detected position deviation exceeds ±3 pixels, a recalculation process is immediately triggered. The start-up time setting accuracy reaches ±1ms, ensuring that the shearing blade is triggered synchronously when the object's stem-leaf junction enters the cutting area. If the object's trajectory changes abruptly due to vibration (position standard deviation > 5 pixels), the system automatically switches to safety mode, suspends shearing, and starts a backup visual positioning module to re-collect data.
[0046] S503. If the position of the object to be cut is after the cut position, then the object to be cut is selected as the first object to be cut; if no object to be cut is selected, then return to step S501. In step S503, it should be noted that the system determines spatial relationships by comparing the center coordinates of the first uncut object in the predicted object queue with the preset cutting position coordinates (default x = 1280 pixels). When the object's center x-coordinate is greater than or equal to the cutting position, an electronic tag containing the object ID, contour feature vector, and timestamp is immediately generated, and the cutting preparation process is initiated. If all objects are located in front of the cutting position (x < 1280 pixels), the system maintains the current conveyor belt operation state and re-executes the position detection loop every 50ms.
[0047] S504. In the selected cutting objects, if the horizontal distance between the second cutting object and the first cutting object is less than the preset distance threshold, it is determined that the second cutting object is the neighbor of the second cutting object. When the two are cut at the same time, the higher cutting height of the two is selected. In step S504, it should be noted that the preset distance threshold is dynamically adjusted according to the type of vegetable. For example, carrots are set to 50 pixels. This threshold is corrected in real time based on the conveyor belt speed: when the speed exceeds 1.0m / s, the threshold is reduced by 20% to cope with the risk of motion blur.
[0048] When determining whether an object is a neighbor, the system first uses an existing spatial indexing structure (such as an R-tree) to quickly filter out candidate objects whose horizontal distance from the first cut object is less than a threshold. Then, it uses the existing IOU (Intersection over Union) algorithm for precise matching. When the bounding box overlap area of two objects exceeds 30%, they are determined to be neighbors. For detected neighbor objects, the system extracts the height value h from their 3D spatial coordinates (x, y, h) and compares it with the height value of the first cut object, selecting the larger value as the joint cut height reference. This height value is dynamically compensated: when the conveyor belt vibration amplitude exceeds ±2mm, the system automatically increases the height margin by 5% to ensure complete cut. If no neighbor object matching the criteria is detected within 30ms, the system defaults to single-object cut mode and uses the height value of the first cut object as the final execution parameter.
[0049] S505, Repeat steps S501 to S504.
[0050] Example 2 like Figure 2 , 3 As shown in Figure 4, a leaf-cutting system for root vegetables is provided, which implements a deep learning-based leaf-cutting method for root vegetables; the system includes: Transport unit 1 is tilted and used for continuous transport of root vegetables; like Figure 2As shown, the transport unit 1 in this embodiment consists of two cooperating conveyor belts. Root vegetables are placed in the V-shaped groove formed between the two conveyor belts, and the vegetables are transported forward by the continuous movement of the conveyor belts. The surface of the conveyor belts is designed with an anti-slip texture to increase the friction between the vegetables and the conveyor belts, preventing the vegetables from slipping or rolling during transport and ensuring the stability of the transport. Transport unit 1 is also equipped with a speed adjustment device, which can flexibly adjust the running speed of the conveyor belts according to actual production needs to meet the leaf cutting requirements of different types of root vegetables.
[0051] Image acquisition module 2, installed on transport unit 1, is used to acquire image data of root vegetables on transport unit 1; like Figure 2 and 4 As shown, the image acquisition module 2 is a camera, fixed to the bottom of the transport unit 1 with screws, used to acquire images of root vegetables on the transport unit 1. The camera has high resolution, capable of clearly capturing the detailed features of the vegetables, and its shooting angle has been adjusted to obtain omnidirectional image information of the vegetables. Simultaneously, the image acquisition module 2 also has a data caching function, which can temporarily store the acquired image data when network transmission fails, and resume transmission after the network is restored, avoiding data loss.
[0052] The decision-making module, connected to the image acquisition module 2, is used to execute the image capture thread, prediction thread, and decision-making thread of the leaf-cutting decision model based on image data. The image capture thread acquires image data in real time, uses the fleshy body detection model to locate the cutting point, encapsulates it into a fruit object, and pushes it into the real object queue. The prediction thread performs motion trajectory prediction and interpolation processing on the real object queue to generate a smooth predicted object queue. The decision-making thread identifies the cutting position based on the predicted object queue and outputs control commands. In this embodiment, the decision-making module is located within the control unit, which is the leaf-cutting decision model in Embodiment 1. The decision-making module employs a high-performance processor with powerful computing capabilities, enabling it to quickly process image data transmitted from the image acquisition module. Its image-taking thread acquires image data in real time at a set frequency, ensuring that the acquired images accurately reflect the current state of the vegetables. The fleshy body detection model, trained on a large amount of data, has high positioning accuracy, accurately locating the cutting point and encapsulating the positioning results into a real object queue. The prediction thread combines the historical movement information and current state of the vegetables to predict and interpolate the movement trajectory of the real object queue, generating a smooth predicted object queue. Based on the predicted object queue, the decision-making thread comprehensively considers factors such as the vegetable's position, size, and shape to accurately identify the cutting position and output control commands to control the leaf-cutting execution unit to perform precise leaf-cutting operations.
[0053] The lifting and shearing assembly 3 is located at the bottom of the transport unit 1 and is used to cut the blades according to control commands; the lifting and shearing assembly 3 includes two sets of lifting and shearing units arranged opposite to each other. In this embodiment, the two sets of lifting and shearing units are used to control two shearing blades 303 respectively, and control the two shearing blades 303 to rotate in opposite directions to achieve the shearing function.
[0054] The lifting and shearing unit includes: The rotating base 301 is movably mounted on the transport unit 1 and has a square through hole inside; like Figure 5 As shown, the rotating base 301 is rotatably mounted on the transport unit 1 via bearings.
[0055] A square lifting rod 302 is mounted on the transport unit 1, and its bottom passes through the square through hole of the rotating base 301 via a slider 310; The square through hole of the rotating base 301 is matched with the square lifting rod 302. When the rotating base 301 starts to rotate, the square lifting rod 302 rotates synchronously.
[0056] The shearing blade 303 is mounted on the bottom of the square lifting rod 302 via a connector 304; the specific structure of the shearing blade 303 can be designed as needed. Mounting bracket 305 is connected to the top of square lifting rod 302 via rotating bearing 306; the function of mounting bracket 305 is to connect lifting rack 307 to square lifting rod 302, so that square lifting rod 302 rises and falls synchronously with lifting rack 307.
[0057] The lifting rack 307 is fixedly connected to the mounting bracket 305 at its top; the lifting rack 307 is a commercially available product.
[0058] The lifting driver 308 is mounted on the transport unit 1 and connected to the bottom gear of the lifting rack 307, and is used to drive the lifting rack 307 to move up and down; the lifting driver 308 is a servo motor with a reduction gear; the reduction gear is connected to the gear of the lifting rack 307. A rotary driver 309, mounted on the transport unit 1, is geared to the rotary base 301 and is used to drive the rotary base 301 to rotate on the transport unit 1. The rotary driver 309 is a servo motor with a reduction gear. The reduction gear is connected to the rotary base 301.
[0059] like Figures 2 to 6As shown, when the lifting and shearing assembly 3 is in use, the lifting driver 308 receives a height control command sent by the decision module, drives the lifting rack 307 to move vertically, and drives the mounting bracket 305 and the square lifting rod 302 to rise and fall axially, thereby achieving precise height positioning of the shearing blade 303. The lifting speed is adjustable according to the type of vegetable. At the same time, the rotary driver 309 drives the rotating base 301 to rotate according to the angle command (adjustable from 0-180°) issued by the decision module. Through the geometric constraints of the square through hole and the square lifting rod 302, the shearing blade 303 is driven to rotate around the vertical axis to the predetermined azimuth angle. When the shearing action is performed, the two sets of lifting and shearing units work synchronously: first, according to the joint shearing height calculated by the decision thread, the blade is raised to the target position to wait; when the target vegetable stem and leaves enter the cutting area, the system uses the two rotary drivers 309 to make the blade complete the downward cutting action at a linear speed of 0.8m / s.
[0060] The receiving and delivering unit 4 is located below the lifting and shearing assembly 3 and is used to receive and deliver the cut fruit. like Figure 2 As shown, the receiving unit 4 in this embodiment consists of two cooperating conveyor belts. After the root vegetables are cut, they fall into the V-shaped groove formed between the two conveyor belts. The continuous movement of the conveyor belts drives the root vegetables forward.
[0061] The control unit is connected to the transport unit 1, the image acquisition module 2, the decision module, the lifting and shearing assembly 3, and the receiving and delivering unit 4. The control unit is used to receive control commands output by the decision thread and dynamically adjust the shearing timing and shearing height of the lifting and shearing assembly 3 based on the control commands.
[0062] In this embodiment, the control unit is located in the system operation cabinet and uses a programmable logic controller (PLC) as the core control element, possessing powerful logic processing capabilities and stability. It internally stores pre-written control programs, enabling it to respond quickly to control commands output by the decision module. The control unit is connected via multiple signal lines to the speed adjustment device of the transport unit 1, the image acquisition module 2, the decision module, the lifting driver 308 and rotary driver 309 of the lifting shearing assembly 3, and the receiving unit 4. Upon receiving control commands from the decision thread regarding the shearing timing and height of the lifting shearing assembly 3, the control unit first parses the commands to identify the specific shearing time requirements and height parameters. Then, based on the parsing results, the control unit sends a signal to the speed adjustment device of the transport unit 1 to ensure the vegetables reach the shearing position at the appropriate time; simultaneously, it sends a height control signal to the lifting driver 308 of the lifting shearing assembly 3 to drive the shearing blade 303 to the specified shearing height; and, as needed, it sends angle commands to the rotary driver 309 to adjust the azimuth angle of the shearing blade 303. During the shearing process, the control unit monitors the operating status of each component in real time. If any abnormality occurs, such as abnormal conveyor belt speed or incomplete shearing blade lifting, an alarm signal is immediately issued, and corresponding protective measures are taken, such as pausing the shearing operation and adjusting equipment parameters, to ensure the safe and stable operation of the entire leaf-cutting system. In addition, the control unit also has data recording and storage functions, capable of recording relevant data for each leaf-cutting operation, such as vegetable type, shearing time, shearing height, and shearing effect, for subsequent data analysis and equipment optimization.
[0063] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A deep learning-based method for cutting leaves from root vegetables, characterized in that, The method includes: S1. Collect image data of root vegetables, and locate the cutting point based on the image data using a fleshy body detection model; S2. Preprocess the image data and train the leaf-cutting decision model using the preprocessed image data; S3. Use the image capture thread of the leaf cutting decision model to collect real-time images of root vegetables during the harvesting process, use the fleshy body detection model to locate the cut point of the real-time image, and obtain the location result; based on the location result and timestamp, encapsulate the fruit object, and push the fruit object into the real object queue; S4. Use the prediction thread of the leaf-cutting decision model to perform position prediction and interpolation on the real object queue to obtain a smooth predicted object queue. S5. Using the decision thread of the leaf-cutting decision model, identify the cutting position in the prediction object queue, and perform leaf-cutting operation on the root and stem parts of root vegetables based on the cutting position.
2. The deep learning-based leaf-cutting method for root and stem vegetables as described in claim 1, characterized in that, In step S1, locating the cut point based on the image data using the fleshy body detection model includes: S101. Preprocess the image data, including Gaussian filtering for noise reduction and size normalization. S102. The preprocessed image data is input into a pre-trained fleshy body detection model. The fleshy body detection model adopts a convolutional neural network architecture, identifies fleshy body regions through a feature extraction layer, and outputs bounding box coordinates. S103. Use the vertices of the bounding box coordinates as the tangent points; S104. Perform real-time calibration on the fleshy body detection model.
3. The deep learning-based leaf-cutting method for root and stem vegetables as described in claim 1, characterized in that, Training the leaf-cutting decision model in step S2 includes: S201. The image data is labeled and used as initial training data to train the initial detection model; S202. Extract the visual feature vector from the initial training data using the initial detection model; S203. Cluster the visual feature vectors in all the initial training data to obtain feature clusters; S204. Set a distance threshold, and when a visual feature vector exceeds the distance threshold, remove the visual feature vector; obtain the training dataset; S205. Based on the training dataset, use the image interpolation algorithm to reduce the distance between the training datasets; S206. Expand the training dataset using data augmentation techniques to generate diverse training samples; S207. The leaf-cutting decision model is trained using the expanded training dataset. The leaf-cutting decision model adopts a deep neural network structure and optimizes the model parameters through multiple rounds of iteration.
4. The deep learning-based leaf-cutting method for root and stem vegetables as described in claim 1, 2, or 3, characterized in that, The expression for the leaf-cutting decision model is: in, This represents the optimal parameter set of the shearing leaf decision model. This represents the total number of training samples. Represents the loss function. This represents a leaf-cutting decision model, taking augmented image data as input and predicting the shearing location as output. Indicates the data augmentation result of the first... training samples, express, Indicates the first The actual cut position label corresponding to each sample. Indicates the weighting coefficient. This represents the L2 regularization term. This indicates a data augmentation operation. This represents an image interpolation algorithm. This represents the data cleaning function. This represents a feature extraction network. This represents the original image after preprocessing. Represents the first clustering of visual feature vectors. Cluster center, This represents the upper limit threshold for determining whether a feature vector is an outlier.
5. The deep learning-based leaf-cutting method for root and stem vegetables as described in claim 1, characterized in that, Training the leaf-cutting decision model in step S2 includes: A201. Label the image data to obtain initial training data; A202. The initial training data is processed using cleaning algorithms, image interpolation methods, and data augmentation methods to obtain a processed training dataset. The training dataset is then divided into training data and test data. A generation detection model is trained using the training data. A203. Calculate the detection performance of the first-generation detection model using the test data to obtain the first-generation performance index; A204. The initial training data is processed a second time using cleaning algorithms, image interpolation methods, and data augmentation methods to obtain a new training dataset; the second-generation detection model is trained using the new training dataset. A205. Calculate the detection performance of the second-generation detection model using the test data to obtain the second-generation performance index; A206. Compare the first-generation performance index and the second-generation performance index to obtain the comparison result. When the comparison result reaches the performance threshold set by the grass fruit, repeat step S204. Otherwise, stop training and use the second-generation detection model as the leaf-cutting decision model.
6. The deep learning-based leaf-cutting method for root and stem vegetables as described in claim 1, characterized in that, In step S4, the prediction thread of the leaf-cutting decision model is used to predict and interpolate the positions of the real object queue to obtain a smooth predicted object queue, including: S401. Extract the front-end cropped object from the real object queue. If no front-end cropped object is extracted, it means that no new real-time image data has been captured in the real object queue. Then proceed to step S404. S402. If the front-end cropping object is extracted, the cropping object in the new real-time image data is horizontally aligned and matched with the original cropping object in the real object queue to obtain the aligned prediction object queue. S403. If the match is successful, the position of the cut object in the new real-time image data is updated to the predicted position of the original cut object; if the match fails, the realism of the cut object in the real-time image data is reduced, and an updated predicted object queue is obtained. S404. Remove cut objects with a realism lower than the realism threshold from the updated prediction object queue and re-sort them to obtain a smooth prediction object queue.
7. The deep learning-based leaf-cutting method for root and stem vegetables as described in claim 1, characterized in that, In step S5, identifying the shearing position in the prediction object queue using the decision thread of the shearing decision model includes: S501. Find the first cut object that has not yet been cut from the smoothed prediction object queue; S502. Adjust the start time of the lifting and shearing assembly based on the time it takes for the first shearing object to move to the shearing position; S503. If the position of the object to be cut is after the cut position, then the object to be cut is selected as the first object to be cut; if no object to be cut is selected, then return to step S501. S504. In the selected cutting objects, if the horizontal distance between the second cutting object and the first cutting object is less than the preset distance threshold, it is determined that the second cutting object is the neighbor of the second cutting object. When the two are cut at the same time, the higher cutting height of the two is selected. S505, Repeat steps S501 to S504.
8. A leaf-cutting system for root and stem vegetables, characterized in that, The system is used to implement the deep learning-based leaf-cutting method for root and stem vegetables as described in any one of claims 1 to 6; the system includes: The transport unit (1) is tilted and used for continuous transport of root vegetables; Image acquisition module (2) is installed on the transport unit (1) and is used to acquire image data of root vegetables on the transport unit (1); The decision module is connected to the image acquisition module (2) and is used to execute the image capture thread, prediction thread and decision thread of the leaf cutting decision model according to the image data. The image capture thread acquires image data in real time, uses the fleshy body detection model to locate the cutting point and encapsulates the fruit object into the real object queue; the prediction thread performs motion trajectory prediction and interpolation processing on the real object queue to generate a smooth prediction object queue; the decision thread identifies the cutting position based on the prediction object queue and outputs control commands. The lifting and shearing assembly (3) is located at the bottom of the transport unit (1) and is used to cut the leaf according to the control command; The receiving and delivering unit (4) is located below the lifting and shearing assembly (3) and is used to receive and deliver the cut fruit.
9. The root vegetable leaf-cutting system as described in claim 8, characterized in that, The lifting and shearing assembly (3) includes two sets of lifting and shearing units arranged opposite to each other; the lifting and shearing unit includes: A rotating base (301) is movably mounted on the transport unit (1) and has a square through hole inside; A square lifting rod (302) is provided on the transport unit (1), and its bottom passes through the square through hole of the rotating base (301) via a slider (310); The shearing blade (303) is disposed at the bottom of the square lifting rod (302) via a connector (304); The mounting bracket (305) is connected to the top of the square lifting rod (302) via a rotary bearing (306); The lifting rack (307) is fixedly connected to the mounting bracket (305) at its top; A lifting drive (308) is disposed on the transport unit (1) and connected to the bottom gear of the lifting rack (307) for driving the lifting rack (307) to move up and down; A rotary driver (309) is disposed on the transport unit (1) and gear-connected to the rotary base (301) for driving the rotary base (301) to rotate on the transport unit (1).
10. The root vegetable leaf-cutting system as described in claim 8, characterized in that, The system also includes a control unit, which is connected to the transport unit (1), image acquisition module (2), decision module, lifting and shearing assembly (3) and pick-up and drop-off unit (4); The control unit is used to receive control commands output by the decision thread and dynamically adjust the shearing timing and shearing height of the lifting shearing component (3) based on the control commands.
Citation Information
Cited By
Leafy vegetable adaptive processing control method and system based on multi-view visual fusion
CN122613743A