Peach orchard fruit maturity identification method based on improved Moilenetv3 model

By improving the Mobilenetv3 model, using technical means such as slice splicing, global attention module and DO-DConv convolution, the problem of insufficient accuracy and stability of fruit maturity judgment in traditional methods is solved, and high-precision and robust fruit maturity recognition is achieved.

CN120032252APending Publication Date: 2025-05-23NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510155939.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional manual judgment of fruit ripening is time-consuming and laborious, and the deep learning classification model lacks accuracy and stability in Taoyuan fruit ripening.

Method used

Improve the Mobilenetv3 model, replace the original first layer downsampling through slice stitching operation, replace the SE attention module with a global attention module, and replace the DO-DConv convolution with depth convolution to improve recognition accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and stability of fruit maturity recognition, can accurately identify the maturity of fruits, and is suitable for real-time detection systems for fruit maturity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032252A_ABST
    Figure CN120032252A_ABST
Patent Text Reader

Abstract

The invention discloses a peach orchard fruit maturity identification method based on an improved Moilenetv3 model. The method comprises the following steps: (1) constructing a data set for peach orchard fruit maturity identification and preprocessing the data set; (2) selecting a Mobilenetv3 network model and carrying out certain improvement on the Mobilenetv3 network model; and (3) selecting proper parameters, and inputting data for training into the improved target network for training. And (4) classifying and identifying an image for testing by using the trained target network. The peach orchard fruit maturity classification method has the advantages that the Mobilenetv3 network is improved, and the peach orchard fruit maturity classification precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to fields such as image processing and deep learning, and in particular to a peach orchard fruit maturity recognition method based on an improved Mobilenetv3 model. Background Art

[0002] With the continuous advancement of agricultural science and technology and the rapid development of computer vision technology, image recognition is increasingly used in the agricultural field. In peach orchard management, accurate judgment of fruit maturity is an important link to ensure fruit quality, optimize picking time and improve economic benefits. However, traditional manual classification is time-consuming and labor-intensive, and deep learning classification can improve efficiency. Summary of the invention

[0003] In order to accurately identify the maturity of fruits, the present invention provides a peach orchard fruit maturity recognition method based on an improved Mobilenetv3 model, which can ensure high fruit maturity recognition accuracy, and has good stability and strong robustness, and can be applied to a fruit maturity recognition system.

[0004] In order to achieve the purpose of the present invention, the present invention provides a peach orchard fruit maturity recognition method based on an improved Mobilenetv3 model, comprising the following steps: Step 1: Construct a fruit maturity recognition dataset and perform preprocessing Step 2: Build an improved Mobilenetv3 network model Step 3: Select appropriate parameters and input the training set images into the improved network model for training Step 4: Input the test set into the trained network model for detection and recognition Compared with the prior art, the present invention has at least the following beneficial effects: The present invention performs peach orchard fruit maturity recognition based on the improved Mobilenetv3, combines the recognition of fruit maturity with deep learning, can provide an effective method for real-time detection of fruit maturity recognition, and can well identify the fruit maturity status.

[0005] The present invention replaces the original first-layer downsampling of the Mobilenetv3 network model with a slice splicing operation; replaces the original SE attention module with a global attention module; and replaces the original deep convolution with a DO-DConv convolution. The improved network has greatly improved recognition accuracy and can accurately identify the maturity of the fruit. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1This is a flow chart of a peach orchard fruit maturity recognition method based on an improved Mobilenetv3 model provided by an embodiment of the present invention.

[0007] Figure 2 Schematic diagram of Mosaic data enhancement in an embodiment of the present invention.

[0008] Figure 3 2 is a schematic diagram of a slice splicing operation in an embodiment of the present invention.

[0009] Figure 4 Schematic diagram of DO-DConv convolution in an embodiment of the present invention.

[0010] Figure 5 Schematic diagram of a global attention module in an embodiment of the present invention.

[0011] Figure 6 Schematic diagram of a channel attention submodule in an embodiment of the present invention.

[0012] Figure 7 Schematic diagram of the spatial attention submodule in an embodiment of the present invention.

[0013] Figure 8 It is a partial recognition effect diagram in the embodiment of the present invention.

[0014] Fig. 9 It is a schematic diagram of an improved Block in an embodiment of the present invention.

[0015] Fig.10 It is a schematic diagram of an improved network in an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution will be clearly and completely described below in combination with the present application.

[0017] See also Figure 1 The present invention provides a peach orchard fruit maturity recognition method based on an improved Mobilenetv3 model, comprising the following steps: Step 1: Construct a fruit maturity recognition dataset and perform preprocessing.

[0018] Step 1 specifically includes the following sub-steps: Step 1.1: The peach orchard fruits mainly come from the experimental peach orchard of Beijing Forestry and Fruit Research Institute. The peach species collected are mainly the hairy peaches cultivated by the Institute of Forestry and Fruit Research. The dataset has been cleaned and the quality is relatively high; the total number is about 2,500 images, of which the training set is about 2,000 peach images, and the test set is about 500 peach images. There are 4 maturity categories: rp (ripe), hp (half-ripe), up (unripe).

[0019] Step 1.2: Perform data augmentation on the data.

[0020] In some embodiments of the present invention, a Mosaic method is used to enhance the data. Figure 2 , randomly select 4 pictures in the data set, scale, rotate, and arrange them to form a new picture. This not only greatly increases the number of pictures input to the network, but also speeds up the training speed, achieving the effect of data enhancement.

[0021] Step 2: Build an improved Mobilenetv3 network model.

[0022] The improvement of the network model includes the following steps: 1. The first layer of the MobileNetV3 network uses a normal convolution with a stride of 2 for convolution operations. During this process, the spatial size of the image changes from 224×224 to 112×112, which is similar to downsampling the image. However, for the peach maturity dataset, only peaches are shown in the picture, and the difference in peach color may cause distortion in maturity recognition. This operation will cause a large amount of effective information in the picture to be lost, reducing the accuracy of maturity recognition.

[0023] In response to this problem, the present invention proposes an image slicing and splicing operation to replace the first layer operation of the original network. In the first step, the image is sliced, and a value is selected for every other pixel on the feature map of each channel, so as to ensure that four complementary images can be obtained on each channel. In this way, it can be guaranteed that no information is lost during the slicing process, and the four images together constitute a comprehensive representation of the original image. The second step is the splicing operation, which puts the three images of the same pixel position obtained on different channels together, and then organizes the four groups of images into a feature map with 12 channels and a size of 112×112. After these two steps, the information in the W and H spaces of the image is concentrated on the channel. After that, the sliced ​​and spliced ​​image is input into the next layer module, and the number of input feature map channels of the next layer network is also adjusted to 12. The schematic diagram of the slicing and splicing operation process is shown in the figure. Figure 3 shown.

[0024] 2. The MobileNetV3 network makes extensive use of deep separable convolutions in the inverted bottleneck residual module, which greatly reduces the amount of calculation. However, the accuracy of the network is lower than that of the neural network using ordinary convolution. This article uses DO-DConv (Depthwise over-parameterized depthwise convolution) deep convolution as channel convolution to make the network converge to a better parameter combination to improve accuracy. DO-DConv deep convolution updates the parameters of two convolution kernels simultaneously during network training; in the inference phase, these parameters are merged into the parameters of a deep convolution kernel and saved. This operation enables DO-DConv deep convolution to make full use of the information of emotional features in the training phase and does not increase the inference time in the inference phase.

[0025] like Figure 4 The figure shows a convolution operation of the DO-DConv convolution kernel on a sliding window area. D in the figure is the additional depth convolution kernel, and W is the original depth convolution kernel. M and N in the figure are the sizes of the convolution kernel in two directions in space. The figure shows that M×N=4; C in Represents the number of channels of the input feature map; W is the weight tensor of the convolution kernel, P is the small area of ​​a sliding window, and O is the output of this convolution operation. The orange area has M×N values, which are multiplied and added (dot product operation) with the M×N values ​​of the corresponding orange area in P to obtain the value of the orange area of ​​O, and so on to obtain the yellow and blue values. After the deep convolution kernel D is dot-producted with P, it is dot-producted with the original deep convolution kernel W to obtain the output of the convolution operation. During the network training phase, the parameters of the two deep convolution kernels will be updated. However, in the inference phase, these parameters will be updated according to Figure 4 The parameters of a deep convolution kernel are merged into a form and saved. During inference, calculations are performed based on the merged parameters. Therefore, the amount of calculation during inference is the same as that of ordinary deep convolution, and the speed will not be slowed down by adding a convolution kernel. This effectively reduces the computational complexity during inference and improves the inference speed. Using DO-DConv convolution to replace the channel convolution in the depthwise separable convolution not only does not increase the amount of calculation during network inference, but also because there is one more deep convolution during training, it can converge to a better parameter combination and improve the accuracy of the network.

[0026] 3. The global attention module replaces the original SE attention module. The channel shuffle global attention module is as follows Figure 5 shown.

[0027] The SE attention module used in the inverted bottleneck residual module in the MobileNetV3 network focuses on the different weight contributions assigned on the channel scale, but lacks attention to the spatial scale information of the feature map. This paper uses the Global Attention Mechanism (GAM) to replace the SE module in the inverted bottleneck residual structure to achieve dual attention to the channel and spatial scale of the feature map. Compared with other attention modules such as CBAM, the GAM attention module enhances the retention of information in channels and space, and improves the performance of the network by reducing information loss and amplifying the global interaction between feature map information.

[0028] The GAM structure is as follows Figure 5 As shown in FIG, it includes two parts, namely the channel attention submodule and the spatial attention submodule. At the same time, in practical applications, the sigmoid activation function in the GAM attention submodule is replaced with the hard-sigmoid activation function.

[0029] The channel attention submodule structure is as follows Figure 6 As shown in the figure. The input dimensions of the feature map include width W, height H, and number of channels C. First, the information in the input feature map is arranged three-dimensionally in three dimensions for preservation. Next, a dimension conversion operation is performed. After that, two multi-layer perceptrons (MLPs) are used to enhance the association between cross-channel and spatial dimensions. Finally, the data is converted to the original dimension again and processed through the hard-sigmoid activation function to obtain the final output result.

[0030] The structure of the spatial attention submodule is as follows Figure 7 As shown in the figure. The input dimensions of the feature map include width W, height H, and number of channels C. r is the reduction ratio, which controls the capacity and overhead of the two attention branches. Information fusion in the feature map space is achieved through two convolutional layers in the spatial attention submodule. At the same time, since there is no pooling operation, the feature map is further preserved. For the input feature map, a 7×7 convolutional layer is first used to reduce the number of channels to reduce the amount of calculation, and then a 7×7 convolutional layer is used to increase the number of channels to ensure that the number of channels of the input and output are consistent. Finally, the output is processed by the hard-sigmoid activation function.

[0031] The formulas for input features, intermediate states, and output features are: F 2 =M c (F 1 )×F 1 F 3 =M s (F 2 )×F 2 In the formula, F 1 is the input feature, M c (F 1 ) is the output feature of the channel attention submodule, F 2 is the intermediate feature, M s (F 2 ) is the output feature of the spatial attention submodule, F 3 is the output feature.

[0032] Step 3: Select appropriate learning rate, batch size, epoch, etc., input the training set images into the improved network model for training, and obtain the trained target network model.

[0033] The input size of the image is set to 224*224, and the initial parameters are set to: the learning rate is set to 0.001, the batch size is set to 32 due to the memory limit of the graphics card, and the number of epoch iterations is set to 100. The learning rate is decayed to 0.9 times the original value after each epoch iteration. The learning rate is initially set to a relatively large value of 0.001 in order to make the decline faster at the beginning of training, and the subsequent learning rate gradually decays to make the model converge as soon as possible. After multiple iterations, the loss value will tend to stabilize, and the network parameter model at this time is saved.

[0034] Step 4: Input the image to be tested into the trained network model for classification and recognition.

[0035] Finally, the peach maturity recognition model was used for testing to verify the recognition effect of the improved peach maturity recognition model on peach maturity. Figure 8 It shows the partial recognition effect. The improved network is shown as follows Fig. 9 The Block diagram is shown as Fig.10 shown.

Claims

1. A peach orchard fruit maturity recognition method based on an improved Mobilenetv3 model. Its characteristics include the following steps: Construct a suitable fruit maturity recognition dataset and divide it into training set and test set; An improved Mobilenetv3 network model is constructed, in which the original first-layer downsampling of the Mobilenetv3 network model is replaced with a slice reorganization operation; the original SE attention module is replaced with a global attention module; and the original deep convolution is replaced with DO-DConv convolution; Use the training set to train the improved Mobilenetv3 network model; The image to be tested is input into the trained improved Mobilenetv3 network model to obtain the classification and recognition results of peach orchard fruit maturity.