Lightweight estimation method and system for fruit yield of result type horticultural crops

By building a lightweight fruit detection model and using an RGB-D camera, combined with YOLO11 and FasterNet modules for object detection and segmentation, the problem of difficult to balance the accuracy, efficiency and cost of fruit yield estimation methods in the prior art is solved, and efficient and accurate fruit yield detection is achieved.

CN120071333AActive Publication Date: 2025-05-30JIANGSU UNIV

Patent Information

Application Number
CN202510130774.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The existing fruit yield estimation methods are difficult to balance between accuracy, efficiency and cost. The RGB image method has occlusion problems, while the three-dimensional point cloud method has complex calculations and high equipment costs.

Method used

A lightweight estimation method and system for fruit yield of fruits of horticultural crops is proposed. By constructing a lightweight fruit detection model, image and depth data are acquired in combination with RGB-D cameras, and the target detection and segmentation are performed. The YOLO11 model and FasterNet module are used for target detection and segmentation, and yield estimation is performed with area tracking counting and regression model.

Benefits of technology

It realizes accurate identification, classification and segmentation of fruits in complex environments, improves the accuracy and efficiency of fruit yield detection, and reduces equipment costs and calculation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071333A_ABST
    Figure CN120071333A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight estimation method and system for fruit yield of fruit-bearing horticultural crops. The method comprises the steps that firstly, a lightweight fruit detection model is constructed, then continuous multi-frame RGB images and depth images are acquired through an RGB-D camera, the RGB images are input into the lightweight fruit detection model, and fruit target frame information and semantic segmentation information output by the lightweight fruit detection model are acquired; fruits beyond a set camera-fruit distance threshold are filtered out according to camera-fruit distance information in the step fruit target frame information depth image, and then fruits within a set distance in each frame of RGB image are obtained; tracking and counting fruits within a set distance in each frame of RGB image through a region tracking and counting method; the fruit area and the fruit weight are obtained; and finally, displaying statistical information of the number of the fruits, the area of the fruits and the weight of the fruits in a visual interface to monitor the yield in real time. According to the invention, accurate identification, classification and segmentation of fruits in a complex environment can be realized, and accurate yield detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of precision agriculture technology, and in particular, to a lightweight estimation method and system for the fruit yield of fruit-bearing horticultural crops. Background Art

[0002] Estimating the fruit yield of fruit-bearing horticultural crops (such as fruits, apples, oranges, etc.) is an important research topic in modern agriculture, which is of great significance for optimizing planting management, improving agricultural production efficiency, and realizing precision agriculture. At present, traditional fruit yield estimation is mainly completed by manual visual estimation. Although this method is simple and intuitive, it is inefficient and greatly affected by human subjective factors, making it difficult to be popularized and applied in large-scale planting scenarios. With the development of computer vision and image processing technologies, researchers have proposed fruit yield estimation methods based on RGB images. These methods detect and count fruits by using a camera to capture fruit images and applying machine learning or deep learning algorithms. However, RGB images can only provide two-dimensional information of fruits and cannot accurately reflect the three-dimensional spatial distribution and occlusion relationship of fruits, resulting in limited estimation accuracy. To overcome the limitations of two-dimensional images, some studies use 3D point cloud data to model and estimate fruits. Although this method can provide more accurate spatial information, it usually requires expensive lidar equipment or complex multi-view image stitching techniques, with high costs and large computational complexity, and is not suitable for scenarios with high real-time requirements.

[0003] Current fruit yield estimation methods are difficult to achieve a balance among accuracy, efficiency, and cost. The RGB image method has occlusion problems, while the three-dimensional point cloud method is computationally complex and has high equipment costs. In addition, there is still much room for improvement in the lightweight of equipment, simplification of algorithms, and low-cost implementation in the prior art. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention proposes a lightweight estimation method and system for the fruit yield of fruit-bearing horticultural crops, which can accurately identify, classify, and segment fruits in complex environments and achieve accurate yield detection.

[0005] The above object is achieved by the following technical solutions:

[0006] The present invention first provides a lightweight estimation system for the fruit yield of fruit-bearing horticultural crops, including the following steps:

[0007] S1. Construct a lightweight fruit detection model: Based on the YOLO11 model as the basic framework, introduce the FasterNet module into the residual feature learning C3k2 unit of the YOLO11 model, change the Conv of the YOLO11 model to DSCOnv, and set the regression loss of the YOLO11 model as the weighted loss of the DFL distribution focal loss and the GIoU regression loss; Train the dataset labeled with three different maturity labels to obtain a lightweight fruit detection model;

[0008] S2. Obtain multiple consecutive RGB images and depth images through an RGB-D camera. The multiple consecutive RGB images and depth images are obtained by the image acquisition device for image acquisition of the fruit of the result-type horticultural crops in the current area to be estimated for yield, and the fruit is the current object to be detected;

[0009] S3. Align the RGB image and the depth image obtained in step S2;

[0010] S4. Input the RGB images obtained in step S3 into the lightweight fruit detection model constructed in step S1 in sequence, and obtain the fruit target box information and semantic segmentation information output by the lightweight fruit detection model;

[0011] S5. Filter out the fruits outside the set camera-fruit distance threshold according to the fruit target box information obtained in step S4 and the camera-fruit distance information in the depth image obtained in step S3, and then obtain the fruits within the set distance in each frame of the RGB image;

[0012] S6. Track and count the fruits within the set distance in each frame of the RGB image obtained in step S5 by the region tracking and counting method, and obtain the total number of fruits in multiple consecutive RGB images;

[0013] S7. Obtain the fruit area according to the fruit semantic segmentation information obtained in step S4 and the camera-fruit distance information in the depth image, and obtain the fruit weight according to the established fruit area-weight regression model;

[0014] S8. Display the quantity of fruits obtained in step S6, the fruit area obtained in step S7, and the fruit weight statistical information in a visualization interface for real-time yield monitoring.

[0015] Furthermore, the lightweight fruit detection model includes an image input module, a feature extraction module, a feature fusion module, and an identification module; The feature extraction module includes multiple cascaded first neural network units, the feature fusion module includes multiple cascaded second neural network units, and the identification module includes a weighted loss function of the VFL zoom loss function, the DFL distribution focal loss function, and the GIoU regression loss function.

[0016] Further, the regional tracking and counting method described in step S6 specifically includes the following sub-steps:

[0017] S6-1. Select the regions of consecutive multiple frames of RGB images and depth images symmetric about the center line as random regions;

[0018] S6-2. Calculate the number of fruits in the random region obtained in step S6-1: Take the size of the random region as the input of the particle swarm optimization algorithm and bring it into the Bytetrack tracking algorithm. First, detect the video: Use a lightweight fruit detection model to detect the fruits in the video to obtain the target detection boxes and corresponding features; then predict the target: Use the Kalman filter algorithm to predict the position and state of the target in the next frame of the video; finally, match the target: Use the Hungarian algorithm to perform optimal matching on several targets between two consecutive frames of the video to obtain the trajectory of the target in the video. The target trajectories that fail to match will be temporarily saved and continue to participate in the prediction and matching of subsequent frames until the target fails to match for consecutive multiple frames and is regarded as a disappeared fruit, and then the trajectory is deleted; The Bytetrack tracking algorithm outputs the id number of the fruits in the random region, that is, the number of fruits in the random region;

[0019] S6-3. Take the root mean square error between the number of fruits in the random region obtained in step S6-2 and the true value as the fitness value of the particle swarm optimization algorithm. The size of the random region corresponding to the smallest convergence of the fitness value is the optimal region size;

[0020] S6-4. Take the optimal region size obtained in step S6-3 as the region size used in yield estimation, and bring it into the Bytetrack tracking algorithm to obtain the number of fruits in the optimal region.

[0021] Further, the fruit area-weight regression model described in step S7 measures the area and weight of 35 large, medium, and small tomatoes in three weight ranges of over 200g, 100 - 200g, and 0 - 100g, and establishes a weight-fruit area regression relationship according to the position of the tomatoes in the image (upper / lower or middle region).

[0022] Further, the camera-fruit distance threshold described in step S5 is set according to the crop planting row spacing of the fruits.

[0023] Further, the dataset labeled with three different maturity labels in step S1 is obtained through the following method:

[0024] Obtain the original dataset containing fruits and leaves;

[0025] Determine the seed pixels in the original dataset, and segment the fruit images from the original dataset according to the seed pixels;

[0026] Use the fruit image as the object and the original dataset as the background for image synthesis;

[0027] Then, use the region - based information segmentation method to segment the background and fruits in the image, and expand the segmented fruit images: Manually select the seed pixels in the original dataset, retrieve the pixel points near the seed pixels, aggregate the similar regions, and traverse all the pixel points in the image through the region - based information segmentation method to complete the image segmentation. After that, use the segmented fruit images as the "objects" and other original datasets as the "background" for image synthesis. In this way, the number of fruit samples in the image is increased, and a dataset with the number of fruits meeting the preset quantity requirement is obtained;

[0028] After obtaining the dataset, use the LabelImg annotation software to add bounding boxes to the fruits with different maturities in the dataset to obtain the annotated dataset. Then, divide the annotated dataset into a training set, a test set, and a validation set according to a preset ratio.

[0029] The present invention also provides a lightweight estimation system for the fruit yield of result - type horticultural crops. The system includes a fruit recognition and localization processor and a program or instruction that can run on the fruit recognition and localization processor. When the program or instruction is executed by the fruit recognition and localization processor, it implements the lightweight estimation method for the fruit yield of result - type horticultural crops as described above.

[0030] Further, the system further includes an image acquisition device and a visualization device; the image acquisition device and the visualization device are respectively connected to the fruit recognition and localization processor;

[0031] The image acquisition device is used to collect the video stream of the fruits planted in the area to be estimated for yield in real - time and send the video stream to the fruit recognition and localization processor for the fruit recognition and localization processor to obtain a continuous multi - frame image based on the video stream;

[0032] The visualization device is used to display the information output by the fruit recognition and localization processor in real - time.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a lightweight estimation method for the fruit yield of result - type horticultural crops as described above.

[0034] The present invention also provides a non - transitory computer - readable storage medium. A computer program is stored on the non - transitory computer - readable storage medium. When the computer program is executed by a processor, it implements a lightweight estimation method for the fruit yield of result - type horticultural crops as described above.

[0035] The beneficial effects of the present invention compared with the prior art are as follows:

[0036] The present invention obtains the number of fruits based on the fruit target box information and RGB-D information; obtains the weight of the fruits based on the fruit segmentation information, RGB-D information, and regression model; displays the yield (quantity and weight) of the fruits in a visualization interface for real-time monitoring; the lightweight fruit detection model is trained based on a dataset labeled with three different maturity tags. The lightweight fruit detection model is based on the YOLO11 model framework, introduces the FasterNet module in the residual feature learning C3k2 unit of the YOLO11 model, changes the Conv of the YOLO11 model to DSCOnv, and sets the regression loss of the YOLO11 model as the weighted loss of the DFL distribution focal loss and the GIoU regression loss. Thus, the present invention can achieve accurate identification, classification, and segmentation of fruits in complex environments and realize accurate yield detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a schematic flowchart of a lightweight estimation method for the fruit yield of fruit-bearing horticultural crops provided by the present invention;

[0038] Figure 2 is a schematic structural diagram of the lightweight fruit detection model provided by the present invention;

[0039] Figure 3 is a schematic structural diagram of the C3k2-F unit provided by the present invention;

[0040] Figure 4 is a schematic structural diagram of a lightweight estimation system for the fruit yield of fruit-bearing horticultural crops provided by the present invention;

[0041] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0043] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings. Figure 1It is one of the flow schematic diagrams of a lightweight estimation method for the fruit yield of result-type horticultural crops provided by the present invention. The execution subject of each step in this method can be a fruit yield detection device, which can be implemented by software and / or hardware, and can be integrated in an electronic device. The electronic device can be a terminal device (such as a smart phone, a personal computer, a learning machine, etc.), or a server (such as a local server or a cloud server, or a server cluster, etc.), or a processor, or a chip, etc. As Figure 1 shown, this method may include the following steps:

[0044] S1. Build a lightweight fruit detection model: Based on the YOLO11 model as the basic framework, introduce the FasterNet module into the residual feature learning C3k2 unit of the YOLO11 model, change the Conv of the YOLO11 model to DSCOnv, and set the regression loss of the YOLO11 model as the weighted loss of the DFL distribution focal loss and the GIoU regression loss; Train the dataset labeled with three different maturity labels to obtain a lightweight fruit detection model, and the specific description is as follows:

[0045] The YOLO (You Only Look Once) model is an object recognition and localization algorithm based on a deep neural network. YOLO has the characteristics of fast specific running speed and being suitable for real-time operation, etc.

[0046] FasterNet (Faster Network) is an efficient network architecture designed to improve the inference speed by simplifying the model design and optimizing the calculation while maintaining excellent performance. Different from traditional large models, FasterNet adopts a hierarchical design and an efficient operator layout, reducing the number of parameters and the amount of calculation while ensuring the model's expressive ability and generalization performance. Its core feature lies in significant improvement in model efficiency through structural innovations such as lightweight convolutional modules and improved data flow paths. In addition, FasterNet introduces a task-specific adaptive adjustment mechanism that can dynamically optimize the calculation allocation, enabling the model to run efficiently on various hardware platforms. This design gives FasterNet significant advantages in terms of inference efficiency and resource utilization, and is suitable for scenarios with high requirements for speed and performance.

[0047] VFL (VariFocal Loss, variable focal loss) is a dynamically scaled binary cross-entropy loss. It adopts an asymmetric training example weighting method. VFL reduces the weight of negative samples that solve the class imbalance problem during training, while increasing the weight of positive samples that generate high quality. This focuses the training on high-quality positive samples and better solves the problem of class imbalance.

[0048] DFL (Distribution Focal Loss) adjusts the weights of samples by introducing a focus factor and a distribution parameter. The focus factor is mainly used to alleviate the class imbalance problem, assigning higher weights to rare classes and making the model pay more attention to the classification accuracy of rare classes. The distribution parameter, on the other hand, is used to control the shape of the sample distribution and has a further weighting effect when dealing with difficult samples.

[0049] GIoU (Generalized Intersection over Union) is an improved object detection loss function designed to address the deficiencies of IoU in cases where there is no overlap or partial overlap between bounding boxes. GIoU introduces the concept of a closure region on the basis of IoU. By calculating the minimum closure region of the predicted box and the target box, it enhances the discrimination ability of the loss function for non-fully overlapping boxes. Compared with IoU, GIoU can more effectively optimize the regression of bounding boxes and improve the performance of object detectors in complex scenarios.

[0050] Specifically, in the embodiments of the present invention, based on the YOLO11 model in the prior art as the basic framework, first, a FasterNet lightweight module is added to its residual feature learning C3k2 unit, effectively reducing the parameters of the model without reducing the model accuracy; second, the VFL classification loss function better solves the problem of the imbalance in the number of fruit samples and enhances the reliability of the network model. Finally, using the GIoU regression loss + DFL distribution focal loss greatly solves the problem of difficult recognition and positioning caused by dense fruit growth and mutual occlusion.

[0051] In this embodiment, the training and testing tasks of the model can be carried out on a workstation equipped with an 11th Gen Intel i9 CPU and an NVIDIA GTX3080Ti GPU, where the Pytorch version used is 1.10.2 for the construction of the network model. In terms of parameter settings: during the model training process, a weight file trained on the coco dataset is pre-loaded; the image resolution size of the input fruit dataset is 640×640, the number of training epochs is 200, and the batch size is 8. The model weights are updated once per iteration, and the performance of the model is tested using the test set data. The model weights with the highest recognition accuracy are saved as the weights of the trained lightweight fruit detection model.

[0052] In this embodiment, the dataset labeled with three different maturity labels is obtained through the following method:

[0053] Obtain the original dataset containing fruits and leaves;

[0054] Manually determine the seed pixels in the original dataset, and segment the fruit images from the original dataset according to the seed pixels;

[0055] Take the fruit image as the object and the original dataset as the background for image synthesis to obtain a dataset with the number of fruits in the image meeting the preset quantity requirement.

[0056] It should be noted that since the sample data of fruits with different maturities in the original dataset collected by the image acquisition device may be unbalanced, in this embodiment, the region information segmentation method can be used to segment the background and fruits in the image, and the segmented fruit images are expanded to ensure the balance of the sample data of fruits with different maturities.

[0057] Specifically, randomly select a pixel in the image as the seed pixel, retrieve the pixel points near the seed pixel, aggregate the similar regions, and traverse all the pixel points in the image through the region information segmentation method, that is, the image segmentation is completed. Then, take the segmented fruit image as the "object" and the other original datasets as the "background" for image synthesis, so as to increase the number of fruit samples in the image and obtain a dataset with the number of fruits meeting the preset quantity requirement.

[0058] After obtaining the dataset, the LabelImg annotation software can be used to add marking frames to the fruits with different maturities in the dataset to obtain the annotated dataset. According to the preset ratio, the annotated dataset is divided into a training set, a test set, and a validation set for training, and then a trained lightweight fruit detection model can be obtained.

[0059] The embodiment of the present invention solves the problem of unbalanced sample ratios of fruits with different maturities well by performing pixel-level data augmentation on the original dataset, can improve the robustness of the lightweight fruit detection model, and can increase the applicable range of the lightweight fruit detection model.

[0060] In some embodiments, the lightweight fruit detection model includes an image input module, a feature extraction module, a feature fusion module, and an identification module. The feature extraction module includes multiple cascaded first neural network units. The feature fusion module includes multiple cascaded second neural network units. The identification module includes a weighted loss function of a VFL zoom loss function, a DFL distribution focal loss function, and a GIoU regression loss function;

[0061] Among them, the first neural network unit includes any one of a convolutional DSConv unit, a residual feature learning C3k2-F unit based on the FasterNet lightweight module, a feature fusion SPPF unit, and a C2PSA unit based on the PSA self-attention mechanism module;

[0062] The second neural network unit includes any one of a convolutional DSConv unit, a residual feature learning C3k2-F unit based on the FasterNet lightweight module, an image splicing Concat unit, and an image upsampling Upsample unit.

[0063] In this embodiment, referring to Figure 2 , the lightweight fruit detection model includes an input image input module 201, a Backbone feature extraction module 202, a Neck feature fusion module 203, and a Prediction recognition module 204.

[0064] The input image input module 201 is used to input an RGB image with a resolution size of 640×640. The Backbone feature extraction module 202 includes a plurality of cascaded first neural network units. The Neck feature fusion module 203 is constructed based on the PANet network and includes a plurality of cascaded second neural network units. The input of the current neural network unit is the output of the previous neural network unit, or the output of the previous neural network unit and the previous N neural network units, where N is a positive integer greater than 1. The Prediction recognition module 204 includes a weighted loss function of a VFL zoom loss function, a DFL distribution focal loss function, and a GIoU regression loss function.

[0065] The residual feature learning C3k2-F unit is used to learn residual features. The convolutional DSConv unit is used to perform convolution, normalization processing, and activation function calculation on the input image. The feature fusion SPPF unit is used to fuse the input images of different sizes, enriching the expression ability of the output image and facilitating the detection of targets with large size differences in different growth cycles. The feature fusion C2PSA unit is used to fuse the input image through comprehensive processing of channel and spatial features, improving the model's attention ability to the target area and global feature expression ability, and facilitating the detection of targets or the segmentation of fine structures in complex scenes. The image upsampling Upsample unit is used to upsample the input image, and the image splicing Concat unit is used to perform Concat function calculation on the input image.

[0066] In this embodiment, in order to improve the recognition accuracy of the lightweight fruit detection model for targets, reduce the model parameters to a certain extent, and improve the real-time detection speed of the model, a residual feature learning C3k2-F unit based on the FasterNet module is proposed, so as to increase the feature extraction ability of the backbone network and provide a reliable front-end information source for the feature fusion stage. The residual feature learning C3k2-F unit based on the FasterNet module is constructed based on the FasterNet lightweight module, and the FasterNet lightweight module is used to reduce the parameters of the model. Specifically, referring to Figure 3 , the C3k2-F unit is composed of a DSConvBNSiLU network layer 301, a Split network layer 302, n FasterNet Block network layers 303, a Concat network layer 304, a ConvBNSiLU network layer 305, and a C3k-F network layer 306. In this embodiment, the number of the FasterNet Block network layer 303 and the C3k-F network layer 306 can be flexibly adjusted according to the actual situation, and no limitation is imposed thereon.

[0067] Optionally, in this embodiment, the feature extraction module 202 may be cascaded by at least three DSConv units, at least three C3k2-F units, at least one SPFF unit, and at least one C2PSA unit, and any two C3k2-F units are not adjacent, and the C2PSA unit is the last unit in the feature extraction module 202.

[0068] It should be noted that in the embodiment of the present invention, the number and cascading manner of the DSConv unit, the C3k2-F unit, the SPFF unit, and the C2PSA unit in the Backbone feature extraction module 202 can be determined according to prior knowledge, and the number and cascading manner of the DSConv unit, the C3k2-F unit, the SPFF unit, and the C2PSA unit in the Backbone feature extraction module 202 are not specifically limited in the embodiment of the present invention.

[0069] Optionally, the Neck feature fusion module 203 may be cascaded by at least three DSConv units, at least three C3k2-F units, multiple Concat units, and an Upsample unit, and any two C3k2-F units are not adjacent.

[0070] In the embodiments of the present invention, the number and cascading manner of the DSConv unit, C3k2-F unit, concat unit, and Upsample unit in the Neck feature fusion module 203 can be determined according to prior knowledge, and the embodiments of the present invention do not specifically limit the number and cascading manner of the DSConv unit, C3k2-F unit, Concat unit, and Upsample unit in the Neck feature fusion module 203.

[0071] For the convenience of understanding the lightweight fruit detection model in the embodiments of the present invention, the following uses an example to illustrate the lightweight fruit detection model in the embodiments of the present invention. The connection relationships of each neural network unit in the lightweight fruit detection model are shown in Table 1 and Figure 2 as shown, and the model parameters of the lightweight fruit detection model are shown in Table 1.

[0072] Among them, the note "n = 1" of the C3k2-F unit 210 in Table 1 indicates that the number of FasterNetBlock modules or C3k-F modules in the internal structure of the C3k2-F unit is 1; the note "n = 2" of the C3k2-F unit 214 in Table 1 indicates that the number of FasterNet Block modules or C3k-F modules in the internal structure of the C3k2-F unit is 2.

[0073] The note "210, 217" of the Concat unit 218 in Table 1 indicates that the inputs of the Concat unit 218 are the outputs of the C3k2-F unit 210 and the Upsample unit 217; the note "213, 220" of the Concat unit 221 in Table 1 indicates that the inputs of the Concat unit 221 are the outputs of the C3k2-F unit 213 and the Upsample unit 220.

[0074] Table 1 Structural relationship and model parameter table of the lightweight fruit detection model

[0075]

[0076]

[0077] Specifically, in this embodiment, the step of sequentially inputting each frame of the image into the lightweight fruit detection model to obtain the coordinates of the bounding box of the fruit output by the lightweight fruit detection model includes:

[0078] Inputting the image into the feature extraction module through the image input module, and obtaining a first feature map output by the first C3k2-F unit, a second feature map output by the second C3k2-F unit, and a third feature map output by the C2PSA unit in the feature extraction module, where the second C3k2-F unit is the subordinate unit of the first C3k2-F unit;

[0079] Input the first feature map, the second feature map, and the third feature map into the first Concat unit, the second Concat unit, and the Upsample unit in the feature fusion module in sequence, and obtain the first feature matrix output by the first C3k2-F unit in the feature fusion module, the second feature matrix output by the second C3k2-F unit, and the third feature matrix output by the third C3k2-F unit. The second Concat unit is the subordinate unit of the first Concat unit, the third C3k2-F unit is the subordinate unit of the second C3k2-F unit, and the second C3k2-F unit is the subordinate unit of the first C3k2-F unit;

[0080] Output the first feature matrix, the second feature matrix, and the third feature matrix to the recognition module, and identify the fruit through the weighted loss function of the VFL zoom loss function, the DFL distribution focal loss function, and the GIoU regression loss function in the recognition module, and segment the pixel size in the fruit through the annotation box containing spatial coordinate information.

[0081] In this embodiment, after the input image input module 201 inputs the current frame image into the Backbone feature extraction module 202, the Backbone feature extraction module 202 can extract fruit features from the current frame image. The C3k2-F unit 210 (i.e., the first C3k2-F unit) in the Backbone feature extraction module 202 can output the first feature map to the Concat unit 218 in the feature fusion module 203; the C3k2-F unit 212 (i.e., the second C3k2-F unit) in the feature extraction module 202 can output the second feature map to the Concat unit 221 in the feature fusion module 203; the C2PSA unit 216 in the feature extraction module 202 can output the third feature map to the Upsample unit 222 and the Concat unit 227 in the feature fusion module 203.

[0082] The C3k2-F unit 217 (i.e., the first C3k2-F unit) in the Neck feature fusion module 203 can output a first feature matrix to the Loss function module 205 in the Prediction recognition module 204. The C3k2-F unit 225 (i.e., the second C3k2-F unit) in the Neck feature fusion module 203 can output a second feature matrix to the Loss function module 205. The C3k2-F unit 228 (i.e., the third C3k2-F unit) in the Neck feature fusion module 203 can output a third feature matrix to the Loss function module 205. Among them, the size of the first feature matrix is 80×80×45, the size of the second feature matrix is 40×40×45, and the size of the second feature matrix is 20×20×45.

[0083] The Loss function module 205 can identify the target fruit based on the first feature matrix, the second feature matrix, and the third feature matrix, using the GIoU+DFL regression loss function 229 and the VFL classification loss function 230, and can label the recognized fruit in the current frame image through a bounding box, and then output the spatial coordinate information of the labeled fruit.

[0084] S2. Obtain a continuous multi-frame RGB image and a depth image through an RGB-D camera. In this embodiment, the continuous multi-frame images are obtained in the form of a video stream, which can avoid the problem of fruit information loss caused by missed shots. Specifically, the image acquisition device can move within the area to be measured for yield according to a preset route, for real-time acquisition of the video stream of the fruits planted in the area to be measured for yield, so as to obtain each frame of the image containing fruits from the video stream. Among them, the area to be measured for yield is determined according to the actual situation. For example, it can be an open-field area for open-air planting, or a greenhouse area. This embodiment of the present invention does not make specific limitations on this. The preset route is the route for identifying the fruits planted in the area to be measured for yield, and the preset route can be determined according to prior knowledge and / or actual situation. This embodiment of the present invention does not make specific limitations on the preset route.

[0085] S3. Align the RGB image and the depth image obtained in step S2;

[0086] S4. Input the RGB image obtained in step S3 into the lightweight fruit detection model constructed in step S1 in sequence, and obtain the fruit target box information and semantic segmentation information output by the lightweight fruit detection model;

[0087] S5. Filter out the fruits outside the set camera-fruit distance threshold according to the fruit target box information obtained in step S4 and the camera-fruit distance information in the depth image obtained in step S3, so as to obtain the fruits within the set distance in each frame of the RGB image;

[0088] S6. Track and count the fruits within a set distance in each frame of RGB image obtained in step S5 through the regional tracking and counting method, and obtain the total number of fruits in multiple consecutive frames of RGB images; in this embodiment, it specifically includes the following sub-steps:

[0089] S6-1. Select the regions of multiple consecutive frames of RGB images and depth images that are symmetric about the center line as random regions;

[0090] S6-2. Calculate the number of fruits in the random region obtained in step S6-1: Take the size of the random region as the input of the particle swarm optimization algorithm and bring it into the Bytetrack tracking algorithm. First, detect the video: Use a lightweight fruit detection model to detect the fruits in the video to obtain the target detection boxes and corresponding features; then predict the target: Use the Kalman filter algorithm to predict the position and state of the target in the next frame of the video; finally, match the target: Use the Hungarian algorithm to perform optimal matching on several targets between two consecutive frames of the video to obtain the trajectory of the target in the video. The target trajectories that fail to match will be temporarily saved and continue to participate in the prediction and matching of subsequent frames until the target fails to match for multiple consecutive frames and is regarded as a disappeared fruit, and then the trajectory is deleted; The Bytetrack tracking algorithm outputs the id number of the fruits in the random region, that is, the number of fruits in the random region;

[0091] S6-3. Take the root mean square error between the number of fruits in the random region obtained in step S6-2 and the true value as the fitness value of the particle swarm optimization algorithm. The size of the random region corresponding to the smallest convergence of the fitness value is the optimal region size;

[0092] S6-4. Take the optimal region size obtained in step S6-3 as the region size used in yield estimation, and bring it into the Bytetrack tracking algorithm to obtain the number of fruits in the optimal region.

[0093] S7. Obtain the fruit area based on the fruit semantic segmentation information obtained in step S4 and the camera-fruit distance information in the depth image, and obtain the fruit weight according to the established fruit area-weight regression model; specifically, obtain the fruit weight based on the fruit segmentation information, RGB-D information, and regression model. By traversing the pixel points within the target bounding box, obtain the RGB values of these pixel points. Since the RGB values of the segmentation masks of fruits with three different maturities are different, determine which maturity level the segmented fruit belongs to by judging the RGB values of each pixel point within each target bounding box, and then accumulate and count these pixel points to finally obtain the pixel areas of fruits with different maturities. The actual size of the fruit is obtained by converting the pixel area of the target fruit segmented by the improved YOLO11 algorithm and the measured distance from the camera to the fruit through a formula. The fruit area-weight regression model classifies fruits into three weight levels: large, medium, and small according to their weights, and divides each fruit into three regions: upper, middle, and lower. Linear regression relationships are established between the areas of the upper, middle, and lower regions of the fruit and the three weight levels of large, medium, and small respectively.

[0094] S8. Display the fruit quantity obtained in step S6, the fruit area obtained in step S7, and the fruit weight statistical information in the visualization interface for real-time yield monitoring. This visualization interface is mainly divided into three areas: The function selection area allows users to start and pause camera detection; the detection result display area provides a visualization interface for the algorithm of this study; the yield information display area shows the quantities of fruits with three different maturities, the quantities of fruits in three different weight ranges, and the total quantity of all fruits. Specifically, click the start detection button in the GUI interface, then the improved YOLOv8 algorithm can be called according to the pre-defined button controls for object detection and semantic segmentation. After the detection is completed, the type, distance, and segmented pixel area of the object are output, and the actual size of the fruit is calculated by combining the camera internal parameters. Then, establish the relationship between the size and the weight to obtain the weight of the target fruit. Obtain the results of object detection and semantic segmentation from the camera image. The type and confidence level are displayed above the object, and four values are displayed below the object, representing the central pixel coordinates (x, y) of the target fruit, the measured distance in cm, and the weight in g respectively. The detection result display area shows the quantities of fruits with different maturities, the quantities of fruits in different weight ranges, and the total quantity obtained through statistics. Among them, a regional counting function is designed within the camera image. Draw a line in the middle three-fifths area of the camera image, and set a motion tracking counting function. Once an object is detected within this area, the type of the object will be judged and accumulated, and at the same time, a weight range is set to judge the weight of the object and count it. Finally, the results of each accumulated count are displayed in real time behind the corresponding text on the interface. Click the stop detection button in the GUI interface, then the camera will pause to obtain the image. Click start detection, and the camera will start to obtain the image again for detection and counting.

[0095] In addition, the present invention also provides a lightweight estimation system for the fruit yield of fruit-bearing horticultural crops, as Figure 4 shown, the system includes a fruit recognition and positioning processor 701 and a program or instruction that can run on the fruit recognition and positioning processor.

[0096] It should be noted that for the specific process of the program or instruction running on the fruit recognition and positioning processor being executed by the fruit recognition and positioning processor to perform the above-mentioned fruit recognition and positioning method, reference can be made to the content of the above-mentioned embodiments, and details will not be repeated in the embodiments of the present invention.

[0097] Optionally, the above-mentioned fruit recognition and positioning processor can be an NVIDIA Jeston Xavier NX development board with a power of 15W. The above-mentioned fruit recognition and positioning processor can execute the above-mentioned fruit recognition and positioning method based on the Pytorch 1.10.2 framework.

[0098] Based on the content of the above-mentioned embodiments, the lightweight estimation system for the fruit yield of fruit-bearing horticultural crops of the present invention further includes: a power supply 702, an image acquisition device 703, and a display device 704.

[0099] The power supply 702 is connected to the fruit recognition and positioning processor 701 and the display device 704, so as to provide power for the fruit recognition and positioning processor 701 and the display device 704.

[0100] The image acquisition device 703 is used to collect the video stream of the fruits planted in the area to be detected for yield in real time, and send the video stream to the fruit recognition and positioning processor, so that the fruit recognition and positioning processor can obtain the current frame image based on the video stream;

[0101] The display device 704 is used to receive and display the quantity and weight information of the fruits sent by the fruit recognition and positioning processor 701.

[0102] The yield detection system of the embodiments of the present invention improves the robustness of fruit recognition at different maturities, can improve the precision rate during fruit recognition at different maturities, and can achieve accurate yield detection.

[0103] Figure 5 It is a schematic structural diagram of an electronic device provided by the present invention. As shown in the figure, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the lightweight estimation method for fruit yield based on RGB-D information, and the method includes:

[0104] Obtain a series of consecutive frames of images, which are obtained by an image acquisition device for image acquisition of fruits within the current production to be detected, and the fruits are the current objects to be detected;

[0105] Input each frame of the images into a lightweight fruit detection model in sequence to obtain the target box information and segmentation information of the fruits output by the lightweight fruit detection model;

[0106] Obtain the quantity of fruits according to the fruit target box information and RGB-D information;

[0107] Obtain the weight of fruits according to the fruit segmentation information, RGB-D information and a regression model;

[0108] Display the yield (quantity and weight) of the fruits in a visualization interface for real-time monitoring;

[0109] Among them, the lightweight fruit detection model is trained according to a dataset labeled with three different maturity labels. The lightweight fruit detection model takes the YOLO11 model as the basic framework, introduces the FasterNet module in the residual feature learning C3k2 unit of the YOLO11 model, changes the Conv of the YOLO11 model to DSCOnv, and sets the regression loss of the YOLO11 model as the weighted loss of the DFL distribution focal loss and the GIoU regression loss for construction.

[0110] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs and other various media that can store program codes.

[0111] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute a lightweight estimation method for the yield of fruit of horticultural crops of the result class based on an RGB-D camera provided by the above-mentioned various methods. The method includes:

[0112] Obtain a series of consecutive frames of images, where the series of consecutive frames of images are obtained by an image acquisition device collecting images of fruits within the current production volume to be detected, and the fruits are the current objects to be detected;

[0113] Input each frame of the images into a lightweight fruit detection model in sequence to obtain the target box information and segmentation information of the fruits output by the lightweight fruit detection model;

[0114] Obtain the quantity of fruits based on the fruit target box information and RGB-D information;

[0115] Obtain the weight of fruits based on the fruit segmentation information, RGB-D information, and a regression model;

[0116] Display the production volume (quantity and weight) of the fruits in a visualization interface for real-time monitoring;

[0117] Among them, the lightweight fruit detection model is trained based on a dataset labeled with three different maturity labels. The lightweight fruit detection model takes the YOLO11 model as the basic framework, introduces the FasterNet module in the residual feature learning C3k2 unit of the YOLO11 model, changes the Conv of the YOLO11 model to DSCOnv, and sets the regression loss of the YOLO11 model as a weighted loss of the DFL distribution focal loss and the GIoU regression loss for construction.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0119] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A lightweight method for estimating the fruit yield of fruit-bearing horticultural crops, characterized in that: The method comprises the following steps: S1. Build a lightweight fruit detection model: Use the YOLO11 model as the basic framework and introduce the FasterNet module into the residual feature learning C3k2 unit of the YOLO11 model, change the Conv of the YOLO11 model to DSCOnv, and set the regression loss of the YOLO11 model to the weighted loss of the DFL distribution focus loss and the GIoU regression loss; A lightweight fruit detection model is obtained by training a dataset labeled with three different maturity labels; S2. Acquire a continuous multi-frame RGB image and a depth image by using an RGB-D camera, wherein the continuous multi-frame RGB image and the depth image are obtained by performing image acquisition on the fruit of a fruit-bearing horticultural crop in the current yield estimation area by an image acquisition device, wherein the fruit is the current object to be detected; S3. Align the RGB image and the depth image obtained in step S2; S4. Input the RGB image obtained in step S3 into the lightweight fruit detection model constructed in step S1 in sequence, and obtain the fruit target frame information and semantic segmentation information output by the lightweight fruit detection model; S5. Filter out fruits outside the set camera-fruit distance threshold according to the fruit target frame information obtained in step S4 and the camera-fruit distance information in the depth image obtained in step S3, and then obtain fruits within the set distance in the RGB image of each frame; S6. Track and count the fruits within a set distance in each frame of the RGB image obtained in step S5 by a regional tracking and counting method to obtain the total number of fruits in multiple consecutive RGB images; S7. Obtain the fruit area according to the fruit semantic segmentation information obtained in step S4 and the camera-fruit distance information in the depth image, and obtain the fruit weight according to the established fruit area-weight regression model; S8. Display the statistical information of the number of fruits obtained in step S6, the fruit area and the fruit weight obtained in step S7 in a visual interface to monitor the yield in real time.

2. A lightweight method for estimating fruit yield of fruit-bearing horticultural crops according to claim 1, characterized in that: The lightweight fruit detection model includes an image input module, a feature extraction module, a feature fusion module and a recognition module; the feature extraction module includes multiple cascaded first neural network units, the feature fusion module includes multiple cascaded second neural network units, and the recognition module includes a weighted loss function of VFL zoom loss function, DFL distribution focus loss function and GIoU regression loss function.

3. The lightweight method for estimating the fruit yield of fruit-bearing horticultural crops according to claim 1, characterized in that: The area tracking and counting method described in step S6 specifically includes the following sub-steps: S6-1. Selecting a region of a continuous plurality of frames of RGB images and depth images symmetrical to the center line as a random region; S6-2. Calculate the number of fruits in the random area obtained in step S6-1: use the size of the random area as the input of the particle swarm optimization algorithm and bring it into the Bytetrack tracking algorithm. First, detect the video: use the lightweight fruit detection model to detect the fruits in the video and obtain the target detection box and corresponding features; then predict the target: use the Kalman filter algorithm to predict the position and state of the target in the next frame of the video; finally, match the target: use the Hungarian algorithm to optimally match several targets between the two frames before and after the video to obtain the trajectory of the target in the video. The target trajectory that fails to match will be temporarily saved and continue to participate in subsequent frame prediction matching until the target fails to match for multiple consecutive frames and is regarded as a disappeared fruit, and then the trajectory is deleted; the Bytetrack tracking algorithm outputs the number of fruit ids in the random area, that is, the number of fruits in the random area; S6-3. The root mean square error between the number of fruits in the random region obtained in step S6-2 and the true value is used as the fitness value of the particle swarm optimization algorithm. The size of the random region corresponding to the minimum convergence of the fitness value is the optimal region size; S6-4. The optimal area size obtained in step S6-3 is used as the area size used for yield estimation, and is brought into the Bytetrack tracking algorithm to obtain the number of fruits in the optimal area.

4. A lightweight method for estimating fruit yield of fruit-bearing horticultural crops according to claim 1, characterized in that: The fruit area-weight regression model in step S7 measures the area and weight of multiple tomatoes of different weight sizes, wherein the area measurement is completed by collecting the RGB image of the tomato with an RGB-D camera, and then the weight-fruit area regression relationship is established according to the position of the tomato in the image (above / below or in the middle area).

5. The lightweight method for estimating fruit yield of fruit-bearing horticultural crops according to claim 1, characterized in that: The camera-fruit distance threshold described in step S5 is set according to the crop planting row spacing of the fruit.

6. A lightweight method for estimating fruit yield of fruit-bearing horticultural crops according to claim 1, characterized in that: The data set annotated with three different maturity labels in step S1 is obtained by the following method: Get the original data set containing fruits and leaves; Manually determining seed pixels in the original data set, and segmenting a fruit image from the original data set according to the seed pixels; Performing image synthesis using the fruit image as an object and the original data set as a background; Then, the background and fruit in the image are segmented by the region information segmentation method, and the segmented fruit image is expanded: a pixel in the image is randomly selected as a seed pixel, and the pixels near the seed pixel are retrieved, similar regions are aggregated, and all the pixels in the image are traversed by the region information segmentation method to complete the image segmentation; then, the segmented fruit image is used as the "object" and the other original data sets are used as the "background" to synthesize the image, thus increasing the number of fruit samples in the image and obtaining a data set with the number of fruits reaching the preset number requirement; After obtaining the data set, use LabelImg annotation software to add marking boxes to the fruits of different maturity in the data set to obtain the annotated data set. According to the preset ratio, the annotated data set is divided into training set, test set and validation set.

7. A lightweight estimation system for fruit yield of fruit-bearing horticultural crops, characterized in that: The system includes a fruit identification and positioning processor and a program or instruction that can be run on the fruit identification and positioning processor. When the program or instruction is executed by the fruit identification and positioning processor, it implements the lightweight estimation method for fruit yield of fruit-bearing horticultural crops as described in one of claims 1-6.

8. A lightweight fruit yield estimation system for fruit-bearing horticultural crops according to claim 7, characterized in that: The system also includes an image acquisition device and a visualization device; the image acquisition device and the visualization device are respectively connected to the fruit recognition and positioning processor; The image acquisition device is used to collect video streams of fruits planted in the area to be estimated in real time, and send the video streams to the fruit recognition and positioning processor, so that the fruit recognition and positioning processor can obtain continuous multiple frames of images based on the video streams; The visualization device is used for real-time display according to the information output by the fruit identification and positioning processor.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for lightweight estimation of fruit yield of fruit-bearing horticultural crops as described in any one of claims 1 to 6 is implemented.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements a lightweight method for estimating the fruit yield of fruit-bearing horticultural crops as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Fruit detection and yield estimation method and system based on machine vision

    CN114663814A

  • Field fruit counting method and system based on video target tracking

    CN117036238A

  • Deep learning-based camellia oleifera yield rapid prediction method

    CN117994701A

Cited By

  • Pig breeding monitoring and behavior intervention method and system based on improved YOLOv11

    CN120954099A

  • Pig breeding monitoring and behavior intervention method and system based on improved YOLOv11

    CN120954099B

  • DEIM-FC-based yield estimation method for potato pickup machine

    CN121121735A