Method for identifying banana peduncle and banana cluster based on lightweight convolutional neural network

By replacing the backbone network of Mask RCNN with Mobile Net and making improvements, combined with transfer learning, the problems of large size and inaccurate recognition of deep learning models were solved, enabling fast and accurate recognition of banana shafts and banana bunches, which is suitable for use by harvesting robots.

CN115690776BActive Publication Date: 2025-12-19BEIJING FORESTRY UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211344118.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-12-19
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

Existing deep learning models are bulky and require high computational resources, making them unsuitable for deployment on mobile devices. Furthermore, existing banana recognition methods fail to effectively address the issues of recognizing irregular objects and overlapping targets, leading to malfunctions in the harvesting robot.

Method used

The backbone network of Mask RCNN is replaced by a lightweight convolutional neural network, Mobile Net. The image feature extraction FPN, mask prediction branch, and dilated convolution are improved. Combined with instance segmentation methods, the model is trained on a banana sample dataset through transfer learning to generate a lightweight model.

Benefits of technology

It enables fast and accurate identification of banana stems and bunches on mobile devices, reducing computing resource requirements, improving identification efficiency and accuracy, and reducing harvesting errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690776B_ABST
    Figure CN115690776B_ABST
Patent Text Reader

Abstract

The application discloses a method for identifying banana peduncle and banana cluster based on a lightweight convolutional neural network, comprising the following steps: constructing a banana sample image dataset containing banana peduncle and banana cluster, and performing data enhancement; replacing the backbone network ResNet of a MaskRCNN network model with a Mobile-Net lightweight convolutional neural network, and performing improvement and enhancement in three aspects of image feature extraction FPN, mask prediction branch and hollow convolution to obtain a Mask-Mobile model; transferring the learning results of the Mask-Mobile model in a COCO dataset to the banana sample image dataset through a transfer learning method, and obtaining a weight file generated by the Mask-Mobile model after retraining; and deploying the Mask-Mobile model and the weight file to a mobile terminal for identifying banana peduncle and banana cluster. The improved identification model is obtained by simultaneously performing three kinds of improvement and enhancement on the MaskRCNN network, and is suitable for accurately and quickly identifying banana peduncle and banana cluster in a banana picking site of a picking robot.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of banana intelligent recognition, in particular to a method for recognizing banana peduncle and banana cluster based on a light convolutional neural network. BACKGROUND

[0002] As one of the four fruits (litchi, pineapple, coconut and banana) in tropical and subtropical regions, banana is one of the most important crops in tropical and subtropical developing countries. In recent years, the demand for bananas in China has been increasing, and the banana yield ranks second in the world, only next to India. Banana picking work has the characteristics of strong seasonality and labor-intensive. Banana picking in China is still in the stage of manual picking, and there is a lack of banana picking machinery, and the picking efficiency is low. At the same time, the device itself is mainly used by picking workers, and the labor cost still exists. The rural labor force in China is gradually shifting to other industries, and the population aging problem is serious, which leads to the gradual increase of the cost of manual picking. With the gradual and continuous increase of the demand for bananas in China and the rapid development of the banana industry, the shortcomings of traditional picking methods such as manual picking are increasingly prominent, and the problems of labor shortage, low picking efficiency, high cost and the like need to be solved. The automatic and intelligent banana picking robot has become an urgent need for the development of the banana industry. In order to promote the sustainable development of the banana industry, intelligent harvesting robots are used for banana harvesting during the ripening period of bananas, which can effectively improve the harvesting efficiency and reduce the labor intensity of farmers. As the core technology of banana picking robots, accurate recognition of banana cluster and banana peduncle is crucial.

[0003] At present, deep learning has achieved many achievements in the field of agricultural engineering. The use of deep learning models based on convolutional neural networks for automatic recognition of crop fruits and vegetables and then automatic picking has gradually become a future development trend. However, the current mainstream deep learning models are large in size, require a large amount of computing resources, and have a long recognition time, which are not suitable for deployment on mobile terminals and edge devices. In the banana picking field, the picking robot needs to accurately and quickly recognize the target crop. Therefore, there is an urgent need for a light-weight and small-sized deep learning target recognition model to realize light-weight computing of terminal devices.

[0004] In addition, the current banana recognition method mainly focuses on object recognition classification and positioning, and does not involve object segmentation, which will inevitably cause errors in automatic recognition and picking, thereby affecting the picking result. Today's detection results related to banana cluster recognition and fruit axis recognition basically adopt the form of label and rectangular bounding box, which correspond to image recognition and positioning, respectively. Although the rectangular box in this detection form can be fitted to the object with the smallest size after multiple rounds of training, for irregular objects, there are still a lot of background noise information in the rectangular box. Under the premise that image recognition cannot be 100% accurate, this rectangular box recognition form will cause the picking robot to misoperate and cut off the wrong objects. More importantly, for overlapping targets, the picking robot cannot specifically distinguish the type information, which is useless for picking work. Therefore, how to provide a more perfect model for the picking robot is still a problem to be solved. SUMMARY

[0005] The purpose of the present application is to provide a method for recognizing banana fruit axis and banana cluster based on a lightweight convolutional neural network, to recognize banana fruit axis and banana cluster based on an instance segmentation method Mask RCNN and a lightweight convolutional neural network Mobile Net, and to simultaneously improve and enhance the Mask RCNN network in three ways, to establish a model suitable for recognizing banana fruit axis and banana cluster, and suitable for the picking robot to accurately and quickly recognize banana fruit axis and banana cluster in the banana picking site.

[0006] To achieve the above purpose, the present application provides the following scheme:

[0007] A method for recognizing banana fruit axis and banana cluster based on a lightweight convolutional neural network, comprising the following steps:

[0008] Constructing a banana sample image dataset containing banana fruit axis and banana cluster;

[0009] Performing data enhancement on the banana sample image dataset;

[0010] Replacing the backbone network ResNet of the Mask RCNN network model with a Mobile-Net lightweight convolutional neural network, and improving and enhancing it to obtain a Mask-Mobile model, including: adding a bottom-up path in the FPN image feature pyramid network for enhancement, adding a fully connected prediction branch in the mask prediction branch, and using a dilated convolution in the FCN to increase the receptive field;

[0011] The Mask-Mobile model is pre-trained by using a COCO dataset, and the learning results of the Mask-Mobile model in the COCO dataset are migrated to the banana sample image dataset by a transfer learning method, and after retraining, a weight file generated by the Mask-Mobile model is obtained;

[0012] The Mask-Mobile model and the weight file generated by the Mask-Mobile model are deployed to a mobile terminal for identifying banana peduncles and banana clusters, so as to realize identification of banana peduncles and banana clusters in the field.

[0013] Further, the bottom-up path enhancement added in the FPN image feature pyramid network comprises:

[0014] In the FPN image feature pyramid network, an independent bottom-up path enhancement is added in the lateral direction, only the bottom layer is connected to the original FPN in the horizontal direction, and a structure layer with a layer number lower than 10 is formed.

[0015] Further, the full connection prediction branch added in the mask prediction branch comprises:

[0016] On the basis of the mask prediction branch of the full convolution network FCN, a full connection prediction branch is added in the PANet, the added full connection prediction branch is a full connection layer branched from the conv3 in the original Mask prediction branch, the added full connection prediction branch is first reduced in dimension through two 3x3 convolutions, then expanded into a one-dimensional vector through a full connection layer, then the vector is used to predict the class-agnostic foreground or background mask, and then a reshape operation is performed to restore the mask to a 28x28 Feature Map, matching the original mask size, finally, the output masks of the mask prediction branch and the full connection prediction branch are fused to obtain the prediction result.

[0017] Further, the size of the convolution kernel of the dilated convolution is 3x3, and the dilation rate is 2 in the object mask prediction branch.

[0018] Further, the banana sample image dataset containing banana peduncles and banana clusters comprises:

[0019] Images of banana clusters, banana peduncles and blue bags containing bananas in the orchard are collected;

[0020] A plurality of images with different light intensities, different target quantities, different occlusion degrees and different resolutions are collected respectively, and the images are divided into a training set, a validation set and a test set according to a ratio of 8:1:1;

[0021] The boundary regions of the collected image of the banana cluster, banana fruit axis and blue bag are labeled by using a labelme software, and the rest is background, a json file is generated, and then the json file is converted by using a dataset generation tool of the labelme to generate an object mask picture, and the banana sample image dataset is obtained.

[0022] According to the specific embodiments provided by the present application, the following technical effects are disclosed: the method for identifying banana fruit axis and banana cluster based on a lightweight convolutional neural network provided by the present application combines a MobileNet lightweight convolutional neural network with an instance segmentation method Mask RCNN, replaces the backbone network ResNet of Mask RCNN with a MobileNet lightweight network, and performs three improvements of image feature extraction FPN, mask prediction branch and hollow convolution on the Mask RCNN network, to obtain a Mask-Mobile model, through a transfer learning method, the learning results of the Mask-Mobile model on a COCO dataset are transferred to a banana image dataset, after retraining, a weight file is obtained, and finally the identification model is deployed to a mobile terminal, to realize rapid and accurate identification of banana fruit axis and banana cluster on a banana picking site. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced below, and obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0024] Figure 1 is a flowchart of the method for identifying banana fruit axis and banana cluster based on a lightweight convolutional neural network according to the embodiments of the present application;

[0025] Figure 2 is a schematic diagram of bottom-up path enhancement according to the embodiments of the present application;

[0026] Figure 3 is a schematic diagram of improved mask prediction branch of banana fruit axis and banana cluster according to the embodiments of the present application;

[0027] Figure 4 is a schematic diagram of hollow convolution according to the embodiments of the present application;

[0028] Figure 5 is a schematic diagram of Mask-Mobile model according to the embodiments of the present application. DETAILED DESCRIPTION

[0029] Clearly, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0030] The object of the present application is to provide a method for identifying banana peduncle and banana cluster based on a lightweight convolutional neural network, to identify banana peduncle and banana cluster based on an instance segmentation method Mask RCNN and a lightweight convolutional neural network Mobile Net, and to simultaneously perform three improvements and enhancements on the Mask RCNN network, to establish a model suitable for identifying banana peduncle and banana cluster, and suitable for a picking robot to accurately and quickly identify banana peduncle and banana cluster in a banana picking site.

[0031] In order to make the above-mentioned objects, features and advantages of the present application more apparent and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0032] As shown in Figure 1 The method for identifying banana peduncle and banana cluster based on a lightweight convolutional neural network provided by the present application comprises the following steps:

[0033] Constructing a banana sample image dataset containing banana peduncle and banana cluster;

[0034] Performing data enhancement on the banana sample image dataset;

[0035] Replacing the backbone network ResNet of the Mask RCNN network model with a Mobile-Net lightweight convolutional neural network, and performing improvement and enhancement to obtain a Mask-Mobile model, as shown in Figure 5 The Mask-Mobile model comprises: adding a bottom-up path enhancement in the FPN image feature pyramid network, adding a fully connected prediction branch in the mask prediction branch, and using a dilated convolution to increase the receptive field in the FCN; the Mobile-Net lightweight convolutional neural network is combined with the instance segmentation method Mask RCNN through the above three improvement methods;

[0036] The Mask-Mobile model is pre-trained using a COCO dataset, and the learning results of the Mask-Mobile model in the COCO dataset are transferred to the banana sample image dataset through a transfer learning method; after retraining, a weight file generated by the Mask-Mobile model is obtained;

[0037] The Mask-Mobile model and the generated weight file thereof are deployed to a mobile terminal for identifying banana peduncles and banana clusters, and based on the special growth environment of bananas and the positional relationship between the peduncles and the banana clusters, the bananas peduncles and the banana clusters in the field can be identified.

[0038] Among them, the Mask-Mobile model is obtained by combining the MobileNet lightweight convolutional neural network with the instance segmentation method Mask RCNN through three improved methods.

[0039] The first improvement is the improvement of the image feature extraction FPN.

[0040] The original Mask RCNN uses the FPN image feature pyramid to extract features from the feature image, and Mask RCNN combines ResNet with FPN to complete the first part of the bottom-up process using the deep residual structure. In ResNet, each feature extraction stage corresponds to a level of FPN, and the last layer of features of each stage is selected as the corresponding features in the FPN of the corresponding level. In the second part, the top-down process uses up-sampling to enlarge the small feature map located at the top layer, so that its size and the feature Figure 1 In order to combine high-level semantic features and accurate positioning ability of the bottom layer, FPN proposes a lateral connection structure similar to the residual network. However, there are problems in the FPN network: the path between the high-level features and the low-level features is long, and the shallow features need to pass through dozens or even hundreds of network layers in ResNet to reach the top layer, which increases the difficulty of accessing accurate positioning information, and the shallow feature information is lost seriously. Therefore, a bottom-up path enhancement is added in the FPN image feature pyramid network, as shown in Figure 2 , thereby reducing the number of structure layers, specifically including:

[0041] An independent bottom-up path enhancement is added in the lateral direction according to PANet, and only the bottom layer has a lateral connection with the original FPN. The number of structure layers is less than 10. In this way, the shallow features can be quickly transmitted to the top layer along the bottom-up path enhancement through the lateral connection of the original FPN at the bottom, the number of layers passed is greatly reduced, and the shallow banana feature information can be better preserved.

[0042] The second improvement is to improve the banana peduncle and banana cluster mask prediction branch, as shown in Figure 3 .

[0043] The full connection prediction branch is added in the mask prediction branch, which can improve the prediction quality, specifically including:

[0044] On the basis of the mask prediction branch of the full convolutional network FCN, a full connection prediction branch is added by imitating the PANet, the added full connection prediction branch is a full connection layer branched from the conv3 in the original mask prediction branch, the added full connection prediction branch is first reduced in dimension through two 3x3 convolutions, then is expanded into a one-dimensional vector through a full connection layer, then a mask of a class-agnostic foreground or background is predicted through the vector, and then the mask is restored to a 28x28 feature map through a reshape operation to match the original mask size, finally, the output masks of the mask prediction branch and the full connection prediction branch are fused to obtain a prediction result through digital addition.

[0045] The third improvement is to use a dilated convolution, as shown in Figure 4

[0046] Since the original Mask RCNN model uses a semantic segmentation branch FCN to predict the mask of an image, the original input image is compressed to a 28x28 size feature map for pixel prediction, which causes a certain amount of loss of image feature information, and the dilated convolution is a variant of the standard convolutional neural network, which can effectively control the receptive field and handle large-scale changes of the object without introducing additional calculations, and the larger the value of the receptive field obtained by the convolution indicates that the range of the original image that it can access is larger, which means that it can contain more global and higher semantic level features; on the contrary, the smaller the value is, the more the features it contains tend to be local and detailed. Therefore, the value of the receptive field can be used to roughly judge the abstraction level of each layer. If the convolution kernel of each layer is of a 3x3 size, the receptive field will be different when the dilation rate is set for each layer, that is, multi-scale information is obtained. The dilation rate of 2 is selected for the object mask prediction branch. Since the dilated convolution does not affect the size of the feature map, it can utilize the multi-scale information while avoiding the information loss caused by down-sampling, and can better extract the multi-dimensional features of the banana stem and banana cluster image, so the dilated convolution is used for improvement in the present application.

[0047] The banana sample image dataset containing banana stems and banana clusters is constructed, including:

[0048] Images of banana clusters, banana stems and blue bags containing bananas in the orchard are collected;

[0049] A plurality of images with different light intensities, different target quantities, different occlusion degrees and different resolutions are collected respectively, and the images are divided into a training set, a validation set and a test set according to a ratio of 8:1:1;

[0050] ​The boundary regions of the collected image of banana bunches, banana fruit axes, and blue bags are labeled by using the labelme software, and the rest is background, to generate a json file. Then the json file is converted by using the dataset generation tool of the labelme to generate an object mask picture, and the banana sample image dataset is obtained.

[0051] In specific embodiments, bananas from a banana plantation base in Zhanjiang City, Guangdong Province are used as samples to collect the banana sample image dataset. Specifically, the camera used is an Intel RealSense Depth Camera D435, which has a maximum resolution of 1280x720. The objects collected are banana bunches, banana fruit axes, and blue bags containing bananas in the orchard. The collection time is between 10 am and 4 pm on sunny and cloudy days. To ensure the diversity of data samples, 600 images with resolutions of 1280x720 and 640x480 are collected under different light intensities (strong light, normal light, and weak light), different target quantities, and different degrees of occlusion, and are divided into a training set of 480 images, a validation set of 60 images, and a test set of 60 images according to a ratio of 8:1:1. The boundary regions of bananas, fruit axes, and bags in the collected images are manually labeled by using the labelme software, and the rest is background, to generate a json file. Then the json file is converted by using the dataset generation tool of the labelme to generate an object mask picture.

[0052] To improve the effect of the network training model and the generalization ability of the model, and to prevent model overfitting, a data augmentation method is used to increase the number of dataset samples. For example, the banana sample image dataset is subjected to data augmentation, including: selecting 540 images from the training set and the validation set for data augmentation, such as flipping, Gaussian noise, contrast change, sharpening, changing color, and other random combination methods, generating five enhanced sub-pictures for each picture, a total of 2700 new pictures, and a total of 3240 training and test pictures.

[0053] For example, the Mask-Mobile model is pre-trained using the COCO dataset, and the learning results of the Mask-Mobile model on the COCO dataset are transferred to the banana sample image dataset by using a transfer learning method. After retraining, the weight file generated by the Mask-Mobile model is obtained, including:

[0054] The training adopts Intel(R) Xeon(R) Platinum 8255C CPU, 2.5GHz frequency, 39GB running memory, and kernel 11. The graphics card selected is GeForce RTX 2080Ti, and the video memory is 11GB. In order to reduce the network model training time, the picture is scaled to 1024x1024 in proportion. The number of photos in each training batch is 1, the initial learning rate is set to 0.001, the momentum factor is 0.9, the weight decay coefficient is 0.0001, and the training round is 200 rounds. The training process is carried out in the way of transfer learning, the pre-trained weight of COCO data set is loaded, the model trained is properly improved, and is applied to banana sample identification. The process of training is to input an image and output an image of the same size, to obtain the category, position and mask of the target object in the image.

[0055] The Mask-Mobile model adds a less-layered bottom-up path enhancement, a fully connected prediction branch and a hollow convolution on the basis of the original Mask RCNN, and replaces the original backbone network. The Mask-Mobile model and the trained weight file can be deployed on, for example, an Nvidia Jetson Nano embedded device, and after the picture is taken on the spot of banana picking and is transmitted into the device, the on-the-spot rapid identification can be carried out.

[0056] Through experiments, the accuracy, recall rate, mAP and identification time of different models are shown in Table 1, and the model size is shown in Table 2.

[0057] Table 1 Evaluation results of different models

[0058]

[0059] Note: M is the Mask RCNN model, IM is the Mask-Mobile model, R101 and Mobile are the ResNet101 and MobileNetv1 networks respectively.

[0060] From Table 1, it can be seen that data enhancement can obviously improve the accuracy of the network model, and can well increase the generalization ability and robustness of the model. The single data enhancement expands the number of data samples, and focuses on data enhancement related to light, so the average detection accuracy is obviously improved compared with the network self-data enhancement.

[0061] Table 2 Comparison of sizes of various models

[0062]

[0063]

[0064] Note: The parameter size is calculated by Tensorflow.

[0065] From the experiment, it can be known that the evaluation indexes of the single enhanced M-Mobile (namely Mask-Mobile) are all optimal, the recognition time is relatively short, the recognition accuracy is the highest, the model is the smallest, and the task requirements can be met, so it is more reasonable to use the single enhanced M-Mobile as the banana peduncle recognition lightweight model.

[0066] The method for identifying banana peduncle and banana cluster based on the lightweight convolutional neural network provided in the application combines MobileNet lightweight convolutional neural network and instance segmentation method Mask RCNN, replaces the backbone network ResNet of Mask RCNN with MobileNet lightweight network, and performs three improvements of image feature extraction FPN, mask prediction branch and hollow convolution on the Mask RCNN network, obtains a Mask-Mobile model, migrates the learning achievements of the Mask-Mobile model in a COCO data set to a banana image data set through a transfer learning method, obtains a weight file after retraining, and finally deploys the identification model to a mobile terminal to realize fast and accurate identification of banana peduncle and banana cluster on a banana picking site.

[0067] The principles and implementation manners of the application are described by applying specific examples in the present application, and the above description of the examples is only used to help understand the method of the application and the core idea thereof; meanwhile, according to the idea of the application, the specific implementation manners and application ranges will be changed by those skilled in the art. In conclusion, the content of the present application should not be understood as a limitation of the application.

Claims

1.A method for identifying banana peduncle and banana cluster based on a lightweight convolutional neural network, characterized in that, The method comprises the following steps: constructing a banana sample image dataset comprising banana peduncles and banana bunches; performing data enhancement on the banana sample image dataset; replacing the backbone network ResNet of the Mask RCNN network model with a Mobile-Net lightweight convolutional neural network and performing improvement and enhancement to obtain a Mask-Mobile model, comprising: adding a bottom-up path enhancement in the FPN image feature pyramid network, adding a fully connected prediction branch in the mask prediction branch, and using a hollow convolution to increase the receptive field in the FCN; pre-training the Mask-Mobile model using a COCO dataset, migrating the learning results of the Mask-Mobile model on the COCO dataset to the banana sample image dataset through a transfer learning method, and obtaining a weight file generated by the Mask-Mobile model after retraining; deploying the Mask-Mobile model and the weight file generated thereby to a mobile terminal for identifying banana peduncles and banana bunches, so as to realize identification of banana peduncles and banana bunches in the field; wherein the bottom-up path enhancement in the FPN image feature pyramid network comprises: in the FPN image feature pyramid network, an independent bottom-up path enhancement is added in the lateral direction in imitation of PANet, only the bottom layer is connected to the original FPN in the horizontal direction, and a structure layer with a layer number lower than 10 is formed; the fully connected prediction branch added in the mask prediction branch comprises: a fully connected prediction branch is added in the mask prediction branch of the fully convolutional network FCN in imitation of PANet, the added fully connected prediction branch is a fully connected layer branched from the conv3 in the original Mask prediction branch, the added fully connected prediction branch is first reduced in dimension through two 3x3 convolutions, then expanded into a one-dimensional vector through a fully connected layer, then the vector is used to predict a class-agnostic foreground or background mask, and then a reshape operation is performed to restore the mask to a 28x28 Feature Map, matching the original mask size, finally, the output masks of the mask prediction branch and the fully connected prediction branch are fused by numerical addition to obtain a prediction result. 2.The method of identifying banana peduncle and banana cluster based on lightweight convolutional neural network according to claim 1, characterized in that, The size of the convolution kernel of the hollow convolution is 3x3, and the dilation rate is 2 in the object mask prediction branch. 3.The method of claim 1, wherein the method further comprises: determining a number of the banana stems and the number of the banana bunches in the image based on the number of the banana stems and the number of the banana bunches in the image. The banana sample image dataset comprising banana peduncles and banana bunches is constructed by: collecting images of banana bunches, banana peduncles and blue bags containing bananas in an orchard; collecting multiple images with different light intensities, different target quantities, different occlusion degrees and different resolutions, and dividing them into a training set, a validation set and a test set according to a ratio of 8:1:1; labeling the boundary regions of banana bunches, banana peduncles and blue bags in the collected images using the labelme software, and the rest as background, generating a json file, and then converting the json file into an object mask picture through a dataset generation tool provided by the labelme software to obtain the banana sample image dataset.

Citation Information

Patent Citations

  • Orchard banana detection method, system and device based on deep neural network and medium

    CN112597897A

  • Lightweight open-set landmark identification method for mobile terminal

    CN112818893A