Corn seed form classification method based on residual network model

Through the corn seed morphology classification method based on the residual network model, combined with active learning and conditional generation adversarial network generation data, the problems of gradient disappearance, insufficient classification accuracy, serious overfitting and poor generalization ability in the existing technology are solved, and higher classification accuracy and better generalization ability are achieved.

CN120107667APending Publication Date: 2025-06-06HUNAN AGRI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510170698.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, there are problems such as gradient disappearance, insufficient classification accuracy, serious overfitting and poor generalization ability, especially in the maize seed morphology classification task.

Method used

The corn seed morphology classification method based on the residual network model is adopted, and the corn seed image data set is constructed, and the residual network model is fine-tuned using ImageNet pre-trained residual network model, and combined with active learning and conditional generation adversarial network to generate image samples of scarce categories, data augmentation and model training are carried out to reduce the risk of overfitting.

Benefits of technology

It effectively solves the problems of data imbalance and insufficient diversity, improves the generalization ability of the model, reduces the risk of overfitting, improves classification accuracy, and avoids the problem of gradient vanishing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107667A_ABST
    Figure CN120107667A_ABST
Patent Text Reader

Abstract

According to the corn seed form classification method based on the residual network model, the residual network model is designed, input is allowed to be directly transmitted to a subsequent layer through jump connection, and the problem of gradient disappearance is avoided. A corn seed image sample is generated by combining active learning and a conditional generative adversarial network, and the problems of data imbalance and insufficient diversity are effectively solved. A residual network pre-trained through an ImageNet data set is introduced as a reference model, and the model adapts to a corn seed classification task through fine tuning. Dropout, an L2 regularization technology and an early stop method are combined, and the over-fitting risk of the classification model is reduced. And an RMSprop optimizer is used for dynamically adjusting the learning rate according to the change of the training process, so that an ideal optimization effect is achieved. A classification result shows that the corn seed classification model provided by the invention has remarkable advantages in the aspects of overall accuracy, balanced performance of all categories, low misclassification rate and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image classification, and in particular, is a corn seed morphology classification method based on a residual network model. Background Art

[0002] The classification of corn seed morphology is of great significance in agricultural research and production, especially in the fields of breeding, planting management and harvest prediction. The traditional method of corn seed classification mainly relies on manual experience. Classifiers identify seed types by observing the appearance characteristics of seeds, such as color, shape, size, etc. Although this method played an important role in early agricultural production, it is inefficient and the classification results are easily affected by human subjective factors, resulting in low classification accuracy. At the same time, when faced with large-scale seed processing, manual classification methods are difficult to meet the speed and accuracy requirements of modern agriculture.

[0003] With the development of computer vision and artificial intelligence technology, deep neural networks have gradually become the mainstream technology in the field of image classification because of their ability to automatically extract image features and perform efficient classification. As the number of network layers increases, the difficulty of training will also increase, mainly due to the problem of gradient vanishing or gradient exploding. Existing technologies also have problems such as limited generalization ability of the model, unstable optimization process and overfitting. The limited generalization ability is due to the insufficient amount and diversity of training data, which leads to poor performance of the model on unseen data. The instability of the optimization process is due to the fact that the setting of the learning rate is usually fixed and fails to be dynamically adjusted according to the changes in the training process, resulting in unsatisfactory optimization results, and often slow convergence or falling into the local optimal solution. Overfitting is often due to the lack of effective regularization methods. Therefore, it is important to find a corn seed morphology classification method that solves the gradient vanishing problem, has high classification accuracy, reduces the risk of overfitting and improves the generalization ability of the model. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] Based on this, the present invention provides a corn seed morphology classification method based on a residual network model to solve the problems of gradient disappearance, insufficient classification accuracy, serious overfitting and poor generalization ability in the prior art.

[0006] (II) Technical solution

[0007] In order to achieve the above object, the present invention provides a corn seed morphology classification method based on a residual network model, comprising:

[0008] S1: construct a corn seed image dataset, and divide the dataset into a training set, a validation set, and a test set;

[0009] S2: Using the residual network model pre-trained by the ImageNet dataset as a benchmark model, fine-tuning the parameters of the benchmark model to adapt to the corn seed image dataset; setting an optimizer and a loss function to train the images in the training set; and evaluating the performance of the fine-tuned model by the early stopping method on the validation set; obtaining the trained fine-tuned model as a corn seed classification model; specifically comprising:

[0010] S201: The residual network pre-trained with the ImageNet dataset is used as the baseline model;

[0011] The residual network model mainly goes through five stages from input image to output. The first stage includes a convolution layer, a batch normalization layer, a RELU activation function and a maximum pooling layer connected in sequence; the convolution layer has 64 7*7 convolution kernels, and the step size of the convolution kernel is 2; the kernel of the maximum pooling layer is 3*3, and the step size is 2; the second to fifth stages have a residual module 1 and several residual modules 2 in structure; the number of residual modules 2 included in the second to fifth stages is 2, 3, 5 and 2 respectively; in addition, after the fifth stage, average pooling and full connection are performed, and regression is realized by softmax;

[0012] The left side of the residual module 1 first reduces the dimension of the feature image of the input residual module 1 through a 1*1 convolution, then performs a 3*3 convolution operation, and finally restores the dimension through a 1*1 convolution, and each convolution is followed by batch normalization and RELU; in addition, after performing a 1*1 convolution and batch normalization on the input feature image on the right side of the residual module 1, the batch normalization layer output of the second 1*1 convolution on the left side of the residual module 1 is jumped and added;

[0013] The left side of the residual module 2 first reduces the dimension of the feature image of the input residual module 2 through a 1*1 convolution, then performs a 3*3 convolution operation, and finally restores the dimension through a 1*1 convolution. Each convolution is followed by batch normalization and RELU; in addition, on the right side of the residual module 2, the feature image of the input residual module 2 is jump-connected to the batch normalization layer output of the second 1*1 convolution on the left side of the residual module 2 for addition;

[0014] S202: Freeze some layers of the benchmark model to prevent the some layers from changing during the fine-tuning process; fine-tune the unfrozen layers of the benchmark model to adapt to the corn seed image dataset;

[0015] S203: Setting the optimizer and loss function;

[0016] S204: During the training process, the validation set is used to evaluate the performance of the fine-tuned model by an early stopping method, and is visualized using a learning curve;

[0017] S3: Input the test set into the trained fine-tuning model to obtain the classification result of the test set.

[0018] Optionally, S1 includes:

[0019] S101: Acquire corn seed images containing 5 different forms using a high-resolution camera to construct a data subset X;

[0020] S102: Combining active learning and conditional generative adversarial networks to generate corn seed images of scarce categories to construct data subset Y; specifically including:

[0021] Step 1: Use the generator G of cGAN to generate fake corn seed images similar to real corn seed images;

[0022] Step 2: Selecting some of the fake corn seed images according to the active learning strategy;

[0023] Active learning uses the intermediate model of the discriminator D of the cGAN in the training process to predict the fake corn seed image, and obtains the classification probability distribution of each fake corn seed image sample; calculates the classification entropy value of each image sample, and the entropy value is defined as follows:

[0024]

[0025] Among them, C is the number of categories, p i (x) is the probability that the image sample belongs to the i-th category;

[0026] Sort by entropy value from large to small, select the top 20% of image samples and input them into cGAN for the next step of training;

[0027] Step 3: inputting the partial fake corn seed images and the real corn seed images into the discriminator D of cGAN for training;

[0028] Step 4: After cGAN iterative training, synthetic corn seed images are obtained to construct the data subset Y;

[0029] S103: Randomly select corn seed images from the data subset X and the data subset Y for data enhancement to obtain a data subset Z;

[0030] S104: The data subsets X, Y and Z are integrated to construct a corn seed image dataset, and the dataset is divided into a training set, a validation set and a test set.

[0031] Optionally, both the generator G and the discriminator D of cGAN use the Adam optimizer, and the learning rate is set to 0.0002.

[0032] Optionally, the preset number of iterative training rounds of cGAN is 5000.

[0033] Optionally, the data enhancement method in S103 includes: rotation, flipping, cropping and scaling.

[0034] Optionally, the unfrozen layer of the benchmark model in S202 is the last fully connected layer; a Dropout layer is added after the last fully connected layer, and the probability of random zeroing is 0.5.

[0035] Optionally, the optimizer in S203 is an RMSprop optimizer, and the specific parameters in pytorch are: learning rate learning_rate=0.0005, smoothing constant alpha=0.9, epsilon=1e-6 added to the denominator to increase the stability of numerical calculation, centered=False, weight decay weight_decay=1e-4, momentum factor momentum=0.9;

[0036] The loss function is cross entropy loss.

[0037] Optionally, in S204, the classification accuracy is set as the verification set performance monitoring indicator; a tolerance round P=10 is defined, and when the verification set performance has not improved for P consecutive rounds, early stopping is triggered; if the verification set performance improves within the tolerance round, the counter is reset.

[0038] (III) Beneficial effects

[0039] It can be seen from the above technical solution that the corn seed morphology classification method based on the residual network model proposed in the present invention has the following beneficial effects:

[0040] 1. Combining active learning and conditional generative adversarial networks to generate corn seed image samples effectively solves the problems of data imbalance and insufficient diversity, provides a better training data foundation for the corn seed classification model, and improves the generalization ability.

[0041] 2. The residual network pre-trained with the ImageNet dataset is introduced as the baseline model. The residual module of this model can effectively improve the ability to extract complex features of corn seeds, and through fine-tuning, the model can better adapt to the specific corn seed classification task. At the same time, the residual module allows the input to be directly passed to the subsequent layers through skip connections, so that even if the network is very deep, the gradient can be directly propagated, thus avoiding the problem of gradient disappearance.

[0042] 3. By combining Dropout, L2 regularization technology and early stopping method, the risk of overfitting of classification models can be significantly reduced. Dropout randomly discards some neurons to prevent the model from over-relying on certain specific features; while L2 regularization limits model parameters to prevent excessive model complexity caused by too many parameters; the early stopping method is introduced during the training process to monitor the model in real time based on the performance of the validation set. When the accuracy of the validation set is no longer improved, the training is stopped in advance to avoid the model overfitting the training data.

[0043] 4. Use the RMSprop optimizer, an adaptive learning rate optimization algorithm, to dynamically adjust the learning rate according to changes in the training process to achieve the ideal optimization effect.

[0044] 5. The classification results of the test set in the corn seed classification model show that this application not only has a high overall accuracy rate, but also has significant advantages in terms of balanced performance among various categories and low misclassification rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0046] Figure 1 is a flow chart of an embodiment of the present invention;

[0047] Figure 2 These are images of corn seeds in five different forms according to an embodiment of the present invention;

[0048] Figure 3 This is a structural diagram of a residual network model according to an embodiment of the present invention;

[0049] Figure 4 1 is a structural diagram of residual module 1 and residual module 2 according to an embodiment of the present invention;

[0050] Figure 5 is the accuracy curve of the training set and the verification set of the embodiment of the present invention;

[0051] Figure 6 This is a confusion matrix heat map of the classification results of the test set of the embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] The present invention provides a corn seed classification method based on a residual network model, the process is as follows: Figure 1 As shown, the specific steps include:

[0054] S1: construct a corn seed image dataset, and divide the dataset into a training set, a validation set, and a test set;

[0055] S101: Acquire corn seed images containing 5 different forms using a high-resolution camera to construct a data subset X;

[0056] Figure 2 There are 5 different forms of corn seed images. Figure 2 (a)-(e) are semi-dental, thick round, round, oblate and dental, respectively.

[0057] S102: Combine active learning and conditional generative adversarial networks to generate corn seed images of scarce categories to construct data subset Y;

[0058] Uneven data distribution is a common problem in corn seed classification tasks. In reality, the number of seed samples of different varieties or forms varies greatly, and the small number of samples in scarce categories seriously affects the classification performance of the classification model on these categories. In this embodiment, the number of round grain samples is small and belongs to the scarce category. Therefore, this embodiment uses an active learning algorithm to locate image samples with low model confidence or misclassification, and uses these image samples to guide the conditional generative adversarial network (cGAN) to generate more image samples to expand the data set and enhance the learning ability of the classification model on scarce categories. The specific steps are:

[0059] (1) Use the generator G of cGAN to generate fake corn seed images that are similar to real corn seed images;

[0060] In this embodiment, the real corn seed image is acquired by a high-resolution camera.

[0061] (2) Select some fake corn seed images based on the active learning strategy;

[0062] Active learning uses the intermediate model (i.e. not the optimal weight model) of the discriminator D during the training process to predict the fake corn seed images and obtain the classification probability distribution of each image sample. Then the classification entropy value of each image sample is calculated (the entropy value reflects the uncertainty of the classification, and the larger the entropy value, the greater the classification uncertainty). The entropy value is defined as follows:

[0063]

[0064] Where C is the number of categories. In this embodiment, the value of C is 5. i (x) is the probability that the image sample belongs to the i-th class.

[0065] According to the entropy value sorting from large to small, the top 20% of image samples are selected and input into cGAN for the next step of training.

[0066] In addition to sorting by entropy value, we can also combine the confidence method to select image samples with the highest classification probability lower than the set threshold of 0.6 and input them into cGAN for the next step of training, that is, max(p i (x))<0.6.

[0067] (3) Some fake corn seed images and real corn seed images are input into the discriminator D of cGAN for training.

[0068] In the above process, cGAN generates diversified samples with category-specific features through conditional input (i.e., category label), including generator G and discriminator D. The input of generator G is a random noise vector and a category label, and the output is an image with the same dimension as the real corn seed image. In this embodiment, the dimension of the image is 224×224. The structure of generator G is a combination of a fully connected layer and a convolutional layer, and the Tanh activation function is used to ensure that the output image value range is [-1,1]. The loss function of generator G is:

[0069]

[0070] Among them, y represents the category label; z represents the noise vector, from the noise distribution P z It is obtained by sampling from the image, usually with Gaussian distribution or uniform distribution; G(z,y) indicates that the generator receives the noise vector and the category label as input and generates image samples related to the category label; P z represents the distribution of noise data; D(G(z,y) represents the authenticity and category consistency score of the generated image sample G(z,y) and its category y by the discriminator D; stands for mathematical expectation.

[0071] The input of the discriminator D is the fake corn seed image samples and real corn seed images selected by active learning, as well as the corresponding category labels, and the output is the authenticity probability of the image and the category consistency judgment. The structure of the discriminator D is that the convolution layer extracts features and the fully connected layer judges the authenticity and category. The loss function of the discriminator D is:

[0072]

[0073] Among them, x represents the real image sample, which is distributed from the real image data P x The sample is obtained from x represents the distribution of real image training data; D(x,y) represents the authenticity and category consistency score of the discriminator D for the real sample x and its corresponding category y.

[0074] Both the generator G and the discriminator D use the Adam optimizer, and the learning rate is set to 0.0002.

[0075] (4) After iterative training of cGAN, synthetic corn seed images are obtained to construct the data subset Y.

[0076] In this embodiment, the preset number of iterative training rounds of cGAN is 5000.

[0077] S103: Randomly select corn seed images from data subset X and data subset Y for data enhancement to obtain data subset Z;

[0078] Specifically, the same number of corn seed images are randomly selected from data subsets X and Y for data augmentation (such as rotation, flipping, cropping and scaling, etc.) to obtain data subset Z, so as to increase the diversity of corn seed image samples.

[0079] S104: fusing the three data subsets X, Y and Z to construct a corn seed image dataset, and dividing the dataset into a training set, a validation set and a test set;

[0080] In this embodiment, the number of corn seed images of five different shapes, namely, semi-dent, thick round, round grain, oblate and dent, in the training set, validation set and test set are 1260, 360 and 180 respectively.

[0081] S2: Use the residual network pre-trained by the ImageNet dataset as the benchmark model, fine-tune the parameters of the benchmark model to adapt to the corn seed image dataset; set the optimizer and loss function to train the images in the training set; and use the early stopping method to evaluate the performance of the fine-tuned model on the validation set; obtain the trained fine-tuned model as the corn seed classification model; the specific steps are as follows:

[0082] S201: The residual network pre-trained with the ImageNet dataset is used as the baseline model;

[0083] The overall structure of the residual network is as follows Figure 3As shown, there are mainly five stages from input image to output. In this embodiment, the number of channels of the input image is 3, and the height and width are both 244, so it is represented by (3,244,244). The first stage includes a convolutional layer, a batch normalization layer, a RELU activation function and a maximum pooling layer connected in sequence; the convolutional layer has 64 7*7 convolution kernels, and the step size of the convolution kernel is 2; the kernel of the maximum pooling layer is 3*3, and the step size is 2; the first stage obtains an output with a shape of (64,56,56). The second to fifth stages have a residual module 1 and several residual modules 2 in structure; the number of residual modules 2 contained in the second to fifth stages is 2, 3, 5 and 2 respectively. In addition, after the fifth stage, average pooling and full connection are performed, and regression is achieved using softmax.

[0084] The structures of residual module 1 and residual module 2 are as follows Figure 4 As shown. The left side of residual module 1 first reduces the dimension of the input feature image through 1*1 convolution, then performs a 3*3 convolution operation, and finally restores the dimension through 1*1 convolution, and each convolution is followed by batch normalization and RELU. In addition, after performing 1*1 convolution and batch normalization on the input feature image on the right, it is jump-connected to the batch normalization layer output of the second 1*1 convolution on the left for addition. The left side of residual module 2 first reduces the dimension of the input feature image through 1*1 convolution, then performs a 3*3 convolution operation, and finally restores the dimension through 1*1 convolution, and each convolution is followed by batch normalization and RELU. In addition, on the right, the input feature image is jump-connected to the batch normalization layer output of the second 1*1 convolution on the left for addition.

[0085] The number of channels of the input (feature image (C, W, W)) and the output (feature image (C1*4, W / S, W / S)) of residual module 1 is different; the number of channels of the input (feature image (C, W, W)) and the output (feature image (C, W, W)) of residual module 2 is the same. Among them, S is equivalent to the step size in the convolution layer, which is used to indicate whether the convolution layer is downsampled. When S is 1, no downsampling is performed, and the input size and output size are equal; when S is 2, downsampling is performed, and the input size is twice the output size. C represents the number of input channels, and C1 represents the number of feature maps output by the convolution layer, that is, the number of output channels. The equality of C and C1 indicates that the 1*1 convolution layer on the left does not reduce the number of channels, and C=2*C1 indicates that the 1*1 convolution layer on the left reduces the number of channels. W represents the input size, that is, the length and width. The main functions of the two 1*1 convolutional layers on the left side of residual module 1 and residual module 2 are to reduce the number of channels and restore the number of channels, respectively, so that the number of input and output channels of the 3*3 convolutional layer in the middle can be smaller, so the number of parameters is much less than directly using a 3*3 convolution kernel, and the efficiency is higher. Residual module 1 and residual module 2 add and output the forward propagation of the neural network and the jump connection of the input (or a convolution of the input), so that the output result will not be worse than before the propagation, avoiding the problem of gradient disappearance.

[0086] The ImageNet dataset is one of the most commonly used datasets for image classification in the field of deep learning. The residual network has learned many useful features after being pre-trained on the ImageNet dataset. Therefore, this embodiment uses the residual network pre-trained on the ImageNet dataset as the benchmark model.

[0087] S202: Freeze some layers of the benchmark model to prevent the some layers from changing during the fine-tuning process; and fine-tune the unfrozen layers of the benchmark model to adapt to the corn seed image dataset;

[0088] In this example, the convolutional layers in front of the baseline model are frozen, and the last fully connected layer is fine-tuned to match the number of categories in the corn seed image dataset. The purpose of this is to retain the common features learned by the baseline model while adapting the model to the corn seed classification task. As training progresses, more layers can be gradually unfrozen and jointly trained to help further fine-tune the model to better adapt it to the corn seed image dataset.

[0089] A Dropout layer is added to the unfrozen layer (i.e., the unfrozen layer) (e.g., after the fully connected layer) to reduce the excessive dependence of the residual network on certain specific features by randomly discarding some neurons. In this embodiment, the probability of random zeroing is 0.5.

[0090] S203: Setting the optimizer and loss function;

[0091] In this embodiment, the RMSprop optimizer is used, and the specific parameters in pytorch are: learning rate learning_rate = 0.0005, smoothing constant alpha = 0.9, epsilon = 1e-6 added to the denominator to increase the stability of numerical calculations, centered = False, weight decay (ie, L2 regularization coefficient) weight_decay = 1e-4, and momentum factor momentum = 0.9.

[0092] The loss function is cross entropy loss.

[0093] S204: During the training process, the performance of the fine-tuned model is evaluated on the validation set using the early stopping method and visualized using a learning curve.

[0094] Early stopping is a common technique to prevent model overfitting. During the training process, when the performance of the validation set no longer improves in multiple consecutive training rounds, stop training in advance to avoid overfitting of the model to the training data. In this embodiment, the performance of the validation set (the classification accuracy of the validation set) is set as a monitoring indicator. Define a tolerance round P. When the performance of the validation set has not improved for P consecutive rounds, early stopping is triggered; if the performance of the validation set improves within the tolerance round, the counter is reset. In this embodiment, P=10.

[0095] Draw the accuracy curves (i.e., learning curves) of the training set and validation set to determine the convergence and overfitting risk of the model, such as Figure 5 shown.

[0096] S3: Input the test set into the trained fine-tuning model to obtain the classification result of the test set.

[0097] The confusion matrix is ​​a common tool for evaluating the performance of classification models. It displays the classification results of the model in the form of a matrix, especially which categories it performs well in and which categories it performs poorly in. In this embodiment, the test set is input into the trained fine-tuning model, and the heat map of the confusion matrix is ​​drawn after the test set is classified. Figure 6 As shown, the horizontal axis is the predicted value and the vertical axis is the actual value.

[0098] The classification effect evaluation indicators of the corn seed classification model calculated according to the confusion matrix are shown in Table 1-3:

[0099] Table 1 Precision, recall, F1 score and number of samples of classification results of five corn seeds in the test set

[0100]

[0101]

[0102] Table 2 Accuracy and number of samples of test set classification results

[0103]

[0104] Table 3 Precision, recall, F1 score and number of samples of macro average and weighted average of classification results of test set

[0105]

[0106] As shown in Table 1-3, the classification accuracy of the five corn seed test sets is 0.89, indicating that the classification model can correctly classify most corn seed samples as a whole, showing the effectiveness of the model in handling multi-category classification tasks. The macro-average and weighted average accuracy, precision, recall, and F1 scores are all no less than 0.94, indicating that the model performs relatively balanced across different categories, with no category significantly lower than other categories. The model performs particularly well in the round grain category, with precision, recall, and F1 scores all at 1.00, meaning that the model achieves perfect classification in this category, with no misclassification or missed classification.

[0107] The corn seed classification model proposed in the embodiment of the present invention shows significant advantages in terms of overall accuracy, balanced performance of various categories, and low misclassification rate. High precision and recall rate ensure the reliability and effectiveness of the model in practical applications, while balanced performance indicators ensure the stability and consistency of the model when processing multi-category classification tasks. These advantages make the classification method of the present invention highly competitive and practical in the field of corn seed classification.

[0108] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A corn seed morphology classification method based on a residual network model, comprising: S1: construct a corn seed image dataset, and divide the dataset into a training set, a validation set, and a test set; S2: using the residual network model pre-trained by the ImageNet dataset as a benchmark model, fine-tuning the parameters of the benchmark model to adapt to the corn seed image dataset; setting an optimizer and a loss function to train the images in the training set; At the same time, the performance of the fine-tuned model is evaluated by the early stopping method on the validation set; the trained fine-tuned model is obtained as a corn seed classification model; specifically comprising: S201: The residual network pre-trained with the ImageNet dataset is used as the baseline model; The residual network model mainly goes through five stages from input image to output. The first stage includes a convolution layer, a batch normalization layer, a RELU activation function and a maximum pooling layer connected in sequence; the convolution layer has 64 7*7 convolution kernels, and the step size of the convolution kernel is 2; the kernel of the maximum pooling layer is 3*3, and the step size is 2; the second to fifth stages have a residual module 1 and several residual modules 2 in structure; the number of residual modules 2 included in the second to fifth stages is 2, 3, 5 and 2 respectively; in addition, after the fifth stage, average pooling and full connection are performed, and regression is realized by softmax; The left side of the residual module 1 first reduces the dimension of the feature image of the input residual module 1 through a 1*1 convolution, then performs a 3*3 convolution operation, and finally restores the dimension through a 1*1 convolution, and each convolution is followed by batch normalization and RELU; in addition, after performing a 1*1 convolution and batch normalization on the input feature image on the right side of the residual module 1, the batch normalization layer output of the second 1*1 convolution on the left side of the residual module 1 is jumped and added; The left side of the residual module 2 first reduces the dimension of the feature image of the input residual module 2 through a 1*1 convolution, then performs a 3*3 convolution operation, and finally restores the dimension through a 1*1 convolution. Each convolution is followed by batch normalization and RELU; in addition, on the right side of the residual module 2, the feature image of the input residual module 2 is jump-connected to the batch normalization layer output of the second 1*1 convolution on the left side of the residual module 2 for addition; S202: Freeze some layers of the benchmark model to prevent the some layers from changing during the fine-tuning process; fine-tune the unfrozen layers of the benchmark model to adapt to the corn seed image dataset; S203: Setting the optimizer and loss function; S204: During the training process, the validation set is used to evaluate the performance of the fine-tuned model by an early stopping method, and is visualized using a learning curve; S3: Input the test set into the trained fine-tuning model to obtain the classification result of the test set.

2. The method according to claim 1, characterized in that S1 includes: S101: Acquire corn seed images containing 5 different forms using a high-resolution camera to construct a data subset X; S102: Combining active learning and conditional generative adversarial networks to generate corn seed images of scarce categories to construct data subset Y; specifically including: Step 1: Use the generator G of cGAN to generate fake corn seed images similar to real corn seed images; Step 2: Selecting some of the fake corn seed images according to the active learning strategy; Active learning uses the intermediate model of the discriminator D of the cGAN in the training process to predict the fake corn seed image, and obtains the classification probability distribution of each fake corn seed image sample; calculates the classification entropy value of each image sample, and the entropy value is defined as follows: Among them, C is the number of categories, p i (x) is the probability that the image sample belongs to the i-th category; Sort by entropy value from large to small, select the top 20% of image samples and input them into cGAN for the next step of training; Step 3: inputting the partial fake corn seed images and the real corn seed images into the discriminator D of cGAN for training; Step 4: After cGAN iterative training, synthetic corn seed images are obtained to construct the data subset Y; S103: Randomly select corn seed images from the data subset X and the data subset Y for data enhancement to obtain a data subset Z; S104: The data subsets X, Y and Z are integrated to construct a corn seed image dataset, and the dataset is divided into a training set, a validation set and a test set.

3. The method according to claim 2, characterized in that Both the generator G and the discriminator D of cGAN use the Adam optimizer, and the learning rate is set to 0.0002.

4. The method according to claim 2, characterized in that: The default number of iterative training rounds for cGAN is 5000.

5. The method according to claim 2, characterized in that: The data enhancement method in S103 includes: rotation, flipping, cropping and scaling.

6. The method according to claim 1, characterized in that The unfrozen layer of the benchmark model in S202 is the last fully connected layer; a Dropout layer is added after the last fully connected layer, and the probability of random zeroing is 0.

5.

7. The method according to claim 1, characterized in that The optimizer in S203 is the RMSprop optimizer, and the specific parameters in pytorch are: learning rate learning_rate=0.0005, smoothing constant alpha=0.9, epsilon=1e-6 added to the denominator to increase the stability of numerical calculation, centered=False, weight decay weight_decay=1e-4, momentum factor momentum=0.9; The loss function is cross entropy loss.

8. The method according to claim 1, characterized in that In S204, the classification accuracy is set as the verification set performance monitoring indicator; a tolerance round P=10 is defined, and when the verification set performance does not improve for P consecutive rounds, early stopping is triggered; If the validation set performance improves within the tolerance round, reset the counter.

Citation Information

Patent Citations

  • Garbage classification behavior recognition algorithm system

    CN117068598A

  • Spray image classification and quality detection method based on VGG-16 model

    CN118351366A