A tobacco leaf grading method based on the Faster R-CNN network

Through the improved Faster R-CNN network model, the ROI Align and Inception network structure are used to solve the problem of feature similarity and insufficient data in tobacco leaf grading, and high-precision and efficient tobacco leaf grading are achieved.

CN113159083BActive Publication Date: 2025-07-29GUIZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011426985.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-09
Publication Date
2025-07-29
Estimated Expiration
2040-12-09

AI Technical Summary

Technical Problem

The existing tobacco leaf grading methods have poor processing effects on image features without obvious results, and there is a problem of low grading accuracy, especially because the characteristics between different levels of tobacco leaf parts and insufficient data sets.

Method used

The tobacco leaf grading method based on Faster R-CNN network was adopted, and the parameters of the VGG16 network model were adjusted, the ROI Pooling was improved to ROI Align, and the Inception network structure was introduced, the Faster R-CNN network model was established, and the deep learning framework caffe was trained.

Benefits of technology

It improves the accuracy and recognition speed of tobacco leaf grading, reduces the manual design process, improves the universality and grading accuracy of the model, and solves problems such as regional mismatch and gradient dispersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113159083B_ABST
    Figure CN113159083B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of computer image processing, and specifically to a tobacco leaf grading method based on the Faster R-CNN network. The method includes the following steps: (1) Collect tobacco leaf images and establish a tobacco leaf image dataset for tobacco leaf grade classification; (2) Based on the VGG16 network model, adjust the parameters of the model, improve the region of interest pooling of the model to ROI Align, remove three convolutional layers of the 8th, 12th, and 15th layers, and introduce the Inception network structure to establish a Faster R-CNN network model; (3) Use the deep learning framework caffe as an experimental platform to train the tobacco leaf image dataset with the Faster R-CNN network. The improved tobacco leaf grading algorithm of the present invention not only has a fast network convergence speed in classifier training, but also has advantages such as high recognition rate and fast recognition speed in recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image processing, and particularly to a tobacco leaf grading method based on the Faster R-CNN network. Background Art

[0002] As a major agricultural economic crop in China, the quality evaluation and grading of tobacco leaves play a crucial role. As one of the main raw materials of tobacco products, the quality of tobacco leaves is the key to affecting the quality stability of later tobacco products. Using computer vision to grade tobacco leaves can not only solve the disadvantages of traditional manual grading methods, such as high labor intensity, strong subjectivity, and low work efficiency, but also stabilize the grading accuracy and qualification rate. However, the existing tobacco leaf grading methods have strong dependence on image acquisition, preprocessing, and feature extraction, especially for the poor processing effect of images with unclear features.

[0003] With the development of deep learning, the convolutional neural network (CNN) has good feature extraction ability and generalization ability. The detection target not only has a fast detection speed, but also has a high accuracy of the detection model. However, with the increase in the number of layers of the convolutional neural network, while bringing high precision, there are also problems such as gradient dispersion, gradient disappearance, difficulty in optimizing the network model, and suppression of the convergence of shallow network parameters, resulting in poor training effects. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a tobacco leaf grading method based on the Faster R-CNN network, which solves the technical problems such as the over-similar features between different grades of tobacco leaf parts and the low grading accuracy caused by insufficient data sets, and the recognition accuracy of tobacco leaf classification.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A tobacco leaf grading method based on the Faster R-CNN network includes the following steps:

[0006] S1. Collect tobacco leaf images and establish a tobacco leaf image data set for tobacco leaf grade classification;

[0007] S2. Set the CNN network structure, based on the VGG16 network model, adjust the parameters of the model, improve the region of interest pooling of the model to ROI Align, remove 3 convolutional layers of the 8th layer, 12th layer, and 15th layer, and introduce the Inception network structure to establish a Faster R-CNN network model;

[0008] S3. Use the deep learning framework caffe as an experimental platform to train the tobacco leaf image data set with the Faster R-CNN network.

[0009] Preferably, the specific process of step S1 is as follows:

[0010] S11. Construction of the tobacco leaf image acquisition device: Design a tobacco leaf acquisition box. Inside the box, there is a loading platform with adjustable height. A camera is fixed on the top of the box, and position light sources are placed on both sides. The entire inside of the box is pasted with black anti-reflection stickers;

[0011] S12. Selection and image acquisition of tobacco leaf image samples: Take the upper, middle, and lower parts of the tobacco leaves as the original training samples and test sets;

[0012] S13. Establish a tobacco leaf grading image data set: Crop the original samples and name them uniformly, and perform 90-degree, 180-degree, 270-degree horizontal mirroring and vertical mirroring amplification on the processed samples to form an augmented training data set;

[0013] S14. Establish a PASCAL VOC data set: Establish a tobacco leaf image data set according to the PASCAL VOC2007 standard data set format. The entire tobacco leaf image data set consists of training images, test images, and validation images.

[0014] Preferably, the specific process of adjusting the parameters of the model in step S2 is as follows: Adjust the image size, learning rate, mini-batch, and RPN network parameters of the VGG16 network model to obtain the model parameters.

[0015] Preferably, the specific process of step S3 is as follows:

[0016] S31. Extract 80% of the data from the established database as the training set samples, and the remaining 20% of the data as the validation set samples;

[0017] S32. Adopt a four-step alternating training method to train the two networks of RPN and Faster R-CNN;

[0018] S33. Set the model parameters: The total number of training iterations is 4.4×10 6 , the mini-batch size is 128, the momentum is 0.9, the weight_decay is 5×10 -4 , the maximum number of iterations is 1.2×10 5 . The number of training times in the first and second stages of RPN is both 1.2×10 5 , and the number of training times in the first and second stages of Fast R-CNN is both 10 6 , where the learning rate in the first stage of RPN and Fast R-CNN is set to 10 -4 , and the learning rate in the second stage is set to 10 -3 .

[0019] Preferably, the camera model in step S11 is MV-VD078SM / SC, the light source model is YX-BL64238K strip LED lamp, and the light source intensity is controlled by a controller with the model number YX-APC24300-2.

[0020] Preferably, the specific network parameters of the Inception network structure in step S2 are 128#1×1, 128#3×3reduce, 128#3×3, 64#5×5, 24#5×5reduce, 64#pool proj, where 128#3×3reduce and 24#5×5reduce represent the 1×1 dimensionality reduction layer filters added before the 3×3 and 5×5 convolutional layers.

[0021] The present invention provides a tobacco leaf grading method based on the Faster R-CNN network. Compared with the prior art, it has the following beneficial effects:

[0022] (1) The tobacco leaf grading method based on the Faster R-CNN network collects tobacco leaf images through an image acquisition box and a camera, establishing a tobacco leaf image dataset that can effectively train a convolutional neural network, providing a data source for subsequent algorithm design and model training based on depth video images.

[0023] (2) The tobacco leaf grading method based on the Faster R-CNN network proposes a research on the detection and grading algorithm of tobacco leaf images based on the VGG16 model, directly driving the self-learning of features and their expression relationships by the data itself, which is beneficial to extracting the inherent information of the data itself, avoiding complex manual design processes, improving the universality of the model, and reducing the tobacco leaf grading cost.

[0024] (3) The tobacco leaf grading method based on the Faster R-CNN network trains the dataset by adjusting the parameters of the fully connected layer of the VGG16 network model. On this basis, ROI Pooling is improved to ROI Align, and then the 8th, 12th, and 15th convolutional layers of the network model are removed. At the same time, the Inception network structure is introduced as the final grading model of this article, solving problems such as regional mismatch, excessive parameters, gradient dispersion, and increased computational complexity, and improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is the technical route flowchart of the present invention.

[0026] Figure 2 is the framework diagram of the VGG16 model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0028] The present invention will be further described below in conjunction with the accompanying drawings and examples.

[0029] Embodiment 1

[0030] Figure 1 This is the technical flow chart of the present invention, and the technical solution of the present invention includes 4 parts;

[0031] The first part is the production of the tobacco leaf image acquisition device and the establishment of the tobacco leaf image data set. In order to capture complete tobacco leaves, a tobacco leaf image acquisition device is independently designed and produced according to the leaf size of the tobacco leaves; the collected tobacco leaf samples are preprocessed, the data set is labeled to obtain the original training set and test set, and the original training set is amplified to prepare the set. Finally, the labeled training set and test set constitute the cigarette grading database to provide data support for subsequent model training and testing;

[0032] The second part is to construct a tobacco leaf grading network algorithm model. According to the established tobacco leaf image data set, by training the image data set, the training effects of three training models, namely VGG16, VGG_CNN_M1024, and ZF, under the Faster R-CNN model framework are compared; then the network model is trained and analyzed from several aspects such as the number and size of the convolutional kernels, the number of pooling layers, the pooling method and size, the number of convolutional layers and fully connected layers, and the activation function;

[0033] The third part is to design a Faster R-CNN tobacco leaf grading model. The model with the best training effect in the second part is used as the basic network model of the improved algorithm. Then, by adjusting the network parameters of the fully connected layer, ROI Pooling is replaced with ROIAlign and the new tobacco leaf grading algorithm of the Inception network structure is introduced;

[0034] The fourth part is to train the data set with the improved algorithm and compare it with other classic models.

[0035] The specific implementation is as follows:

[0036] Step 1: Tobacco leaf image acquisition, data preprocessing, and database establishment;

[0037] Step 2: Compare the accuracies of the three training models of VGG16, VGG_CNN_M1024, and ZF, and select the optimal model;

[0038] The loss function of the convolutional neural network is as follows:

[0039]

[0040] Where n is the nth data sample; z is the number of nodes in the output layer; t is the correct training sample value; y is the output value of network training;

[0041] During the training process, in order to make the error smaller, during this process, along with the adjustment of the weight parameters, the adjustment of the weight parameters can be expressed as:

[0042]

[0043]

[0044]

[0045] Where: ΔW l represents the weight parameter of the lth layer; η represents the learning rate; δ represents the residual; b represents the bias;

[0046] Based on the above three training models, set the same training parameters to train and analyze the tobacco leaf image dataset;

[0047] According to the analysis of the training results, select to use the basic neural network VGG16, add a 3*3 convolutional kernel, reduce the weight parameters, and increase the network depth;

[0048] Step 3: Based on the VGG16 network model, propose a Faster R-CNN network model, and improve the convergence speed and recognition accuracy of the network model by optimizing the training parameters of the model, the number of convolutional layers of the model, and the problem of regional feature mismatch of the model;

[0049] Step 4: Use the training set to train the Faster R-CNN network model.

[0050] The specific method for establishing the database in the said Step 1 includes:

[0051] 1), Select tobacco leaf samples. The tobacco leaf variety is Yunyan 87. The tobacco leaf samples are selected by on-site experts to be representative and information-rich tobacco leaves, ensuring that the tobacco leaf samples cover as many characteristic information as possible, achieving good and different. The collected tobacco leaf samples are: lower tobacco leaves X2F, middle tobacco leaves C2F, C3F, C4F, upper tobacco leaves B1F, B2F, B3F; among them, X represents lower tobacco leaves, C represents middle tobacco leaves, B represents upper tobacco leaves, and F represents orange;

[0052] 2) Acquisition of tobacco leaf images. During the image acquisition process, to ensure the stability of the image acquisition system and the consistency of the acquired samples, when collecting images each time, the distance from the lens to the stage, the illumination intensity, the camera focal length, etc. remain unchanged; before collection, the optimal illumination brightness needs to be determined; the pixel size of the image is 1024×768, the pixel size is 4.65μm, and the storage format is set to BMP;

[0053] 3) Making of the dataset. The original dataset is cropped to a size of 608×342, and the storage format of the image is.jpg. The images are uniformly named and consist of 6-digit numbers, for example: 000001.jpg, with the serial numbers connected; the processed original samples are respectively flipped vertically and rotated clockwise by 90°, 180°, and 270° for amplification to form the amplified training dataset. The amplified dataset totals 29,416 tobacco leaf images; according to the POSCAL VOC2007 dataset format, with the help of the open-source software Label-image, the tobacco leaf dataset is labeled to generate xml-class labels.

[0054] The selection of the convolutional neural network model in the second step specifically includes:

[0055] 1) Using the training samples in the first step as training data and the test set as model performance test data;

[0056] 2) Using a hardware platform with 32GB of memory, a GPU of the GeForce GTX 1060 model, and a CPU of the Intel Core i7-8700k model, and the Ubuntu16.04 operating system, training three network models, namely VGG16, VGG_CNN_M1024, and ZF, on the Caffe deep learning framework;

[0057] 3) During the training process, keep the structural parameters of the training model consistent, set the mini-batch size to 128, the momentum to 0.9, the dropout rate to 0.5, the weight decay coefficient to 0.0005, and the maximum number of iterations to 8×10 5 times, and the learning rate to 10 -4 .

[0058] The improved Faster R-CNN algorithm based on the VGG16 network model and its training in the third step specifically include:

[0059] 1) Optimization of model parameters. Using the VGG16 network model as a pre-training model, based on the idea of trial and error, adjust the size of the input image, the learning rate, the mini-batch, the RPN network parameters, etc. multiple times, set the total number of training times each time to 280,000 times, and conduct multiple pre-training and verification analyses on the model to obtain better model parameters;

[0060] 2) Calibration of the region of interest pooling. When quantifying the candidate box positions and each small grid position in the original ROI pooling in the original network, the two quantization operations may cause deviations in the candidate box positions, resulting in the problem of region mismatch; the proposed ROI Align well solves the problem of region mismatch (mis-alignment) caused by the two quantizations in the ROI Pooling operation; ROI Align cancels the quantization operation process and uses the bilinear interpolation method to obtain the image values at pixel points with floating-point coordinates, thus transforming the entire feature aggregation process into a continuous operation;

[0061] 3) Improve the feature extraction part of the network structure. Remove the 3 convolutional layers of the 8th, 12th, and 15th layers from the original VGG16 network model. At the same time, in order to match the original network structure, introduce the Inception network structure.

[0062] The specific steps of using the training set to train the Faster R-CNN network model in Step 4 include:

[0063] 1) Use the deep learning framework caffe as the experimental platform. Extract 80% of the data from the established database as the training set samples, and the remaining 20% of the data as the validation set samples. Adopt the four-step alternating training method to train the two networks of RPN and Faster R-CNN;

[0064] 2) For the RPN network, set three area scales of {64 2 , 2 128 2},

[0065] 6 and aspect ratios of {1:1, 1:2, 2:1} of the anchor points; -4

[0066] 5 3) Set the model parameters: The total number of training iterations is 4.4×10 5 , the mini-batch size is 128, the momentum is 0.9, the weight_decay is 5×10 6 , the maximum number of iterations is 1.2×10 -4 ; The number of training times in the first and second stages of RPN is 1.2×10 -3 , and the number of training times in the first and second stages of Fast R-CNN is 10 6

[0066] 4) Verify the classification ability of the VGG16 model based on the Faster R-CNN framework in the design, and compare it with existing classical neural networks.

Claims

1. A tobacco leaf grading method based on the Faster R-CNN network, characterized in that, It includes the following steps: S1. Collect tobacco leaf images and establish a tobacco leaf image dataset for tobacco leaf grade classification; S2. Set the CNN network structure, adjust the parameters of the VGG16 network model, improve the region of interest pooling of the model to ROIAlign, remove three convolutional layers of the 8th, 12th, and 15th layers, and introduce the Inception network structure to establish a Faster R-CNN network model; S3. Use the deep learning framework caffe as the experimental platform to train the tobacco leaf image dataset with the Faster R-CNN network; The specific process in step S1 is as follows: S11. Construction of the tobacco leaf image acquisition device: Design a tobacco leaf acquisition box. There is a loading platform inside the box, and the height of the loading platform is adjustable; a camera is fixed on the top of the box, and position light sources are placed on both sides. The inside of the box is entirely pasted with black anti-reflection stickers; S12. Selection of tobacco leaf image samples and image acquisition: Take the upper, middle, and lower parts of the tobacco leaves as the original training samples and test sets; S13. Establish a tobacco leaf grading image dataset: Crop and uniformly name the original samples, and perform 90-degree, 180-degree, 270-degree horizontal mirroring and vertical mirroring amplification on the processed samples to form an augmented training dataset; S14. Establish a PASCAL VOC dataset: Establish a tobacco leaf image dataset according to the PASCAL VOC2007 standard dataset format. The entire tobacco leaf image dataset consists of training images, test images, and validation images; The camera model is MV-VD078SM / SC, and the light source model is YX-BL64238K strip LED lights. The light source intensity is controlled by a controller with the model YX-APC24300-2; The specific process of adjusting the model parameters in step S2 is as follows: Adjust the image size, learning rate, mini-batch, and RPN network parameters of the VGG16 network model to obtain the model parameters; The specific process of step S3 is as follows: S31. Extract 80% of the data from the established database as training set samples, and the remaining 20% of the data as validation set samples; S32. Adopt a four-step alternating training method to train two networks, namely RPN and Faster R-CNN; S33. Set model parameters: The total number of training iterations is 4.4×10 6 , the mini-batch size is 128, the momentum is 0.9, and the weight_decay is 5×10 -4 , the maximum number of iterations is 1.2×10 5 ; The number of training times for the first and second stages of RPN is 1.2×10 5 , and the number of training times for the first and second stages of Fast R-CNN is 10 6 , where the learning rate for the first stage of RPN and Fast R-CNN is set to 10 -4 , and the learning rate for the second stage is set to 10 -; The specific network parameters of the Inception v1 network structure in step S2 are 128#1×1, 128#3×3reduce, 128#3×3, 64#5×5, 24#5×5reduce, 64#pool proj, where 128#3×3reduce and 24#5×5reduce represent the 1×1 dimensionality reduction layer filters added before the 3×3 and 5×5 convolutional layers.

Citation Information

Patent Citations

  • Tobacco leaf grading method based on hyperspectral image and deep learning algorithm

    CN106326899A

  • Vehicle target detection method

    CN108009509A