Lightweight jujube variety identification method based on knowledge distillation and transfer learning

Through knowledge distillation and transfer learning methods, a lightweight jujube variety recognition model is constructed, which solves the problems of low identification efficiency and large parameters in the existing technology, and realizes efficient jujube variety recognition on small mobile devices.

CN120260031AInactive Publication Date: 2025-07-04SHIJIAZHUANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127321.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing jujube variety recognition methods have problems such as low recognition efficiency, low accuracy and difficulty in applying on small mobile devices. In particular, the traditional manual recognition method consumes time and labor-intensive and has high requirements for molecular marker equipment. The recognition model based on the CNN model has a large amount of parameters and takes up a large memory space.

Method used

The lightweight jujube variety recognition method based on knowledge distillation and transfer learning is adopted. By constructing teacher models and student models, using transfer learning to train the teacher models and use knowledge distillation technology to guide student model training, lightweight jujube variety recognition is achieved.

Benefits of technology

While maintaining a high recognition rate, the amount of model parameters is reduced, making the lightweight model suitable for small mobile devices, improving recognition efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260031A_ABST
    Figure CN120260031A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight jujube variety identification method based on knowledge distillation and transfer learning, and the method comprises the steps: firstly collecting 34 jujube variety images in a natural environment, and generating a jujube image data set after cutting, preprocessing and data expansion; then, determining a teacher model and a student model, taking the teacher model and the student model trained on the ImageNet data set as a pre-training network, and training the teacher model by using the jujube data set to obtain an optimal teacher model; the teacher model guides training to obtain a high-performance lightweight student model; and cutting a to-be-detected jujube image, inputting the cut to-be-detected jujube image into the student model, and obtaining an identification result through forward calculation. According to the method for identifying the jujube variety by using the knowledge distillation technology, the problem that the traditional network model is large in parameter and long in reasoning calculation time is solved, and relatively high performance is considered while the model parameters are greatly reduced, so that the jujube variety identification method based on deep learning is more suitable for being embedded into small mobile equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of jujube variety identification and the field of computer vision, and particularly relates to a lightweight jujube variety identification method based on knowledge distillation and transfer learning. Background Art

[0002] Jujube (Ziziphus jujuba Mill.) is a plant of the family Rhamnaceae and is widely distributed in more than 50 countries in Asia, Europe and America. Jujube is one of the most important temperate crops in East Asia, especially in China. Jujube has a cultivation history of thousands of years in China and is currently cultivated throughout the country. Among them, Hebei, Shandong, Shanxi, Shaanxi, Xinjiang Uygur Autonomous Region and Henan are the main jujube production areas. During the long cultivation process of jujube trees, many variants have emerged. So far, more than 900 jujube varieties have been recorded, and more than 500 germplasm resources are collected and preserved in the National Fruit Tree Germplasm Jujube Variety Resource Nursery in Shanxi Province. Therefore, the research on jujube variety identification can promote the breeding and popularization of excellent jujube qualities, and also facilitate breeders to construct and manage the jujube germplasm resource database.

[0003] At present, the methods for jujube variety identification include artificial identification method, molecular marker method and image recognition method based on CNN model. The artificial identification method mainly identifies the variety through the characteristics of jujube fruits such as shape, color and maturity period. However, this process not only requires the identifier to have certain professional knowledge and experience of jujube varieties, but also takes a lot of time and energy, and is easily affected by subjective factors, resulting in low identification efficiency and accuracy. The molecular marker method is an identification based on the laboratory environment, with high reliability, but it has high requirements for the technology of operators and the equipment of experimental instruments, and takes a long time and has a large workload to analyze experimental data. The network model constructed by the image recognition method based on CNN model is large in volume, has a large number of parameters and occupies a large amount of memory space, and is difficult to be applied to small mobile devices. In the actual scenario, users hope to use portable small devices to obtain jujube varieties. Therefore, it is of great significance to study a jujube variety identification model with small volume and high performance for real-time jujube variety identification. Summary of the Invention

[0004] The purpose of the present invention is to address the challenges in the above-mentioned jujube variety identification problem, and propose a lightweight jujube variety identification method based on transfer learning and knowledge distillation technology in the natural environment, using the knowledge of complex models to guide the training of lightweight models, so as to obtain a lightweight jujube variety identification model with high performance.

[0005] For this reason, the technical solution adopted by the present invention is as follows:

[0006] A lightweight jujube variety identification method based on knowledge distillation and transfer learning, comprising the following steps:

[0007] S1: Collect jujube image data and build a jujube image dataset;

[0008] S1-1: The jujube image dataset contains m jujube varieties, a total of n images, and each jujube image is labeled with its real variety label; and its shape label is labeled according to its shape characteristics, a total of k shapes; finally, the jujube images with variety and shape labels are classified and sorted to generate a jujube image dataset containing m varieties;

[0009] S1-2: normalizing the jujube fruit image constructed in step S1-1, and transforming the original data into a corresponding standard form;

[0010] S1-3: The jujube image dataset is divided into two parts: a training set and a validation set. The training set is used to train the network model, and the validation set is used to verify the model recognition effect.

[0011] S2: Determine the teacher model and student model;

[0012] S3: Use transfer learning method to train the teacher model to obtain a suitable teacher model;

[0013] S4: Initialize the student model;

[0014] S5: Use the knowledge distillation method to use the trained teacher model to guide the student model and obtain the optimal student model;

[0015] S6: Input the jujube images in the validation set into the student model for testing, and obtain the jujube image shape and variety labels through forward propagation calculation. This completes the jujube variety recognition.

[0016] Furthermore, in step S2, the teacher model is a ResNet 50 network with high accuracy, good stability and large model, and the student model is a MobileNetV3 large network with a smaller model;

[0017] Furthermore, step S3 includes:

[0018] S3-1: The backbone network structure of the teacher model is ResNet 50, which is a shared layer composed of convolutional layers, batch normalization, ReLU, and 4 bottleneck modules. The last two layers are replaced by classification layers to implement a multi-task framework, namely the shape classification layer and the category classification layer. The shape classification layer contains two fully connected layers FC1-1 and FC1-2 to complete the shape classification task, with parameters of 300×1 and 10×1 respectively, and the category classification layer contains two fully connected layers FC2-1 and FC2-2 to complete the category classification task, with parameters of 300×1 and 34×1 respectively;

[0019] The teacher model uses transfer learning to load the pre-trained weights on ImageNet, which is a large visual database for visual object recognition software research. The original ResNet50 structure is replaced with an improved fully connected layer to complete the initialization of the teacher model. When a jujube fruit image is input into the teacher model, what the teacher model outputs are the probabilities of the image for each shape category and variety category;

[0020] S3-2: Input a jujube fruit image. The final output of the teacher model includes the shape label and variety label. The features obtained after the image passes through the shared layer are respectively transformed into a 300-dimensional shape feature vector and a variety feature vector after passing through the fully connected layers FC1_1 and FC1_2. On the one hand, the variety feature vector combined with the shape feature vector passes through FC2_1 to obtain a 10-dimensional output vector, and the shape label corresponding to the largest element of the vector is the shape of the current image; on the other hand, the variety feature vector passes through FC2_2 to output a 34-dimensional vector, and the variety label corresponding to the largest element of the vector is the variety of the current image;

[0021] The teacher model uses the cross-entropy loss function CrossEntropyLoss to update the weights. The shape loss function is expressed as:

[0022] L shape_t =CE(y shape , label shape )

[0023] where label shape is the true shape label of the current image, and y shape is the shape value output by the teacher model;

[0024] The variety loss function is expressed as:

[0025] L species_t =CE(y species , label species )

[0026] where label species is the true variety label of the current image, and y species is the variety value output by the teacher model;

[0027] The total loss function of the teacher model is expressed as:

[0028] L Loss_t =γ*L shape_t +(1 - γ)*L species_t

[0029] where γ is an adjustable regularization coefficient;

[0030] Optimize the loss function using the Stochastic Gradient Descent (SGD) algorithm, and update the parameters of the teacher model:

[0031] W l ,B l = SGD(L Loss_t ,W l ,b l ,lr)

[0032] where lr represents the learning rate, SGD represents the SGD function, w l ,b l represent the parameters before the update of the l-th layer, and W l ,B l are the parameters after the update of the l-th layer.

[0033] S3-3: Repeat the process of S3-2 for each jujube image in the training set, update the parameters of the teacher model, and obtain the trained teacher model;

[0034] S3-4: Use each jujube image in the test set as the input, the corresponding shape and category labels as the output, compare the actual labels with the model output labels, calculate the recognition rate, repeat this process, and finally select the model with the highest recognition rate as the optimal teacher model.

[0035] Furthermore, the steps in S4 include:

[0036] The backbone network structure of the student model is MobileNetV3 large, which consists of 1 convolutional layer, 15 bneck structures, 1 convolutional layer, 1 Pooling layer, and 1 convolutional layer. Among them, the 4th - 6th and 11th - 15th structures of the bneck structure introduce the lightweight attention mechanism SE module, and at the same time introduce the h-swish activation function, and replace the last two layers with classification layers to implement a multi-task framework, which is divided into a shape classification layer and a category classification layer. The shape classification layer contains two fully connected layers to complete the shape classification task, with parameters of 300×1 and 10×1 respectively, and the category classification layer part is two fully connected layers to complete the category classification task, with parameters of 300×1 and 4×1 respectively;

[0037] The student model uses the transfer learning method to load the pre-trained weights on ImageNet, and replaces the original MobileNetV3 large network with an improved fully connected layer to complete the initialization of the student model; when a jujube fruit image is input into the student model, what the student model outputs is the probability of the image belonging to each shape category and category.

[0038] Furthermore, the steps in S5 include:

[0039] S5-1: Set the teacher model obtained in S3 to the evaluation (eval) mode and do not participate in backpropagation.

[0040] S5-2: Input the jujube fruit images in the training set into the teacher model and the initialized student model obtained in S4 for forward calculation simultaneously.

[0041] The calculation methods of the shape loss and the category loss are the same. That is, first calculate the loss between the classification result and the true label using cross-entropy, and then calculate the gap loss between the outputs of the teacher model and the student model using KL divergence; then weight and sum the two losses to obtain the shape / category loss; finally, weight and sum the shape loss and the category loss to obtain the total loss value.

[0042] Calculate the cross-entropy loss function between the shape output of the student model and the true shape label value, that is:

[0043]

[0044] where x i represents the true shape label of the i-th image, represents the probability value of the student model outputting the j-th shape category;

[0045] Calculate the KL divergence between the shape output value of the student model and the shape output value of the teacher model, that is:

[0046]

[0047] where T represents the adjustable temperature parameter; represents the probability value of the teacher model outputting the j-th shape category;

[0048] Then the total shape loss function of the student model is:

[0049] L total_shape = α * L CE_shape +(1 - α) * L KD_shape

[0050] where α is the adjustable regularization coefficient;

[0051] Similarly, calculate the cross-entropy loss function between the category output of the student model and the true category label value, that is:

[0052]

[0053] where y i represents the true category label of the i-th image, represents the probability value of the student model outputting the j-th category;

[0054] Calculate the KL divergence between the output values of the student model's categories and the output values of the teacher model's categories, i.e.:

[0055]

[0056] where T represents the adjustable temperature parameter, represents the probability value of the teacher model outputting the j-th category;

[0057] Then the total loss function of the student model's categories is:

[0058] L total_species =β * L CE_species +(1 - β) * L KD_species

[0059] where β is the adjustable regularization coefficient;

[0060] Finally, the total loss of the student model includes the shape loss and the category loss. The total loss function for training the student model is:

[0061] L Loss_s =λ * L total_shape +(1 - λ) * L total_species

[0062] where λ is the adjustable regularization coefficient.

[0063] S5-3: Use the Stochastic Gradient Descent (SGD) algorithm for backpropagation to optimize the total loss function L Loss_s , and update the parameters of the student model:

[0064] W l , B l =SGD(L Loss_s , w l , b l , lr)

[0065] where lr represents the learning rate, SGD represents the SGD function, w l , b l represents the parameters before the update of the l-th layer, and W l , B l are the parameters after the update of the l-th layer.

[0066] S5-4: Repeat the process of S5-2 and S5-3 for each jujube image in the training set, update the parameters of the student model, and obtain the trained student model;

[0067] S5-5: Use each jujube image in the test set as input, with the corresponding shape and variety labels as output. Compare the actual labels with the model output labels, calculate the recognition rate, and repeat this process. Finally, select the model with the highest recognition rate as the optimal student model.

[0068] Compared with the prior art, the advantages of the present invention are as follows:

[0069] The jujube variety recognition method of the present invention uses knowledge distillation technology to transfer the knowledge of complex models to lightweight student models, making up for the disadvantages of low performance of lightweight networks, so that lightweight models have the advantages of both high recognition rate and small number of parameters, and are thus more suitable for deployment on small mobile devices. Brief Description of the Drawings

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0071] Figure 1 The model structure diagram constructed for the jujube variety recognition method of the present invention;

[0072] Figure 2 The flow schematic diagram of the jujube variety recognition method of the present invention;

[0073] Figure 3 The flow schematic diagram of the jujube variety recognition method based on knowledge distillation of the present invention. Detailed Embodiments

[0074] The present invention provides a jujube variety recognition method based on knowledge distillation and transfer learning technologies in a natural environment. By constructing a convolutional neural network model, as Figure 1 shown, after training it, the classification of 34 jujube varieties is realized.

[0075] As Figure 2 shown, a jujube variety recognition method based on knowledge distillation and transfer learning technologies in a natural environment includes the following steps:

[0076] S1: Data preparation: Collect jujube fruit image data and construct a jujube image data set;

[0077] S2: Determine the teacher model and the student model;

[0078] S3: Use the transfer learning method to train the teacher model to obtain a suitable teacher model;

[0079] S4: Initialize the student model;

[0080] S5: Use the knowledge distillation method to guide the student model with the trained teacher model to obtain the optimal student model;

[0081] S6: Input the date jujube image to be recognized into the student model for testing. The student model is set to the eval evaluation mode, and the shape and variety labels of the jujube image are obtained through forward propagation calculation. Thus, the jujube variety recognition is completed;

[0082] Among them, the step S1 includes:

[0083] S1-1: The jujube image data used in this embodiment is from the National Jujube Cultivar Breeding Base in Cang County, Hebei Province. The image dataset contains 34 jujube varieties, namely Yuanling Xiaozao, Yuanlingzao, Xuancheng Yuanzao, Xiaozizao, Xiangfen Yuanzao, Suyuanling, Dongzao, Guangyang Dazao, Xiaozao, Zaoshu Tangzao, Yueyazao, Chaoyang Yuanzao, Dali Longzao, Duanli Changhong, Goutou Xiaozao, Huizao, Chuanling, Lianxian Muzao, Nanjingzao, Ruanhezao, Shengxian Baipu, Damuzao, Guantanzao, Jinzao, Jinsi Xiaozao, Linze Dazao, Xiangzao, Lajiaozao, Chahuzao, Hulu Changhong, Mopanzao, Fengmiguan, Jidanzao, Longxuzao, with a total of 3,099 images. They were taken under natural light in the orchard using an Android mobile phone and a Nikon D7500 single-lens reflex digital camera, with resolutions of 2976×3968 and 2784×1856 respectively. The camera was in the automatic exposure mode, and the photos were stored in the computer in JPEG format. The collection time was when the jujube fruits were in the mature stage. When taking pictures, an appropriate distance was selected between the lens and the jujube fruits to obtain clear data. Single or multiple separated jujube fruits were photographed, and they were taken under sunny, cloudy, front-light, and back-light environments respectively.

[0084] Crop the collected images to only include images with the shape of a complete jujube fruit. On the one hand, because background noise will cause certain interference to variety recognition; on the other hand, because the task of the present invention is to recognize the jujube images that have been cropped by people, the similarity of the data can be improved after cropping.

[0085] According to the jujube germplasm resource description specifications and the opinions of consulting jujube experts, the 34 jujube varieties collected are divided into 10 categories according to their shape characteristics, namely round, oblong, cylindrical, pepper-shaped, teapot-shaped, gourd-shaped, disk-shaped, oval, flat-round, and dragon-beard-shaped. The number of jujube varieties included in each category is shown in Table 1. Store the jujube images and their corresponding shape and variety labels;

[0086] Table 1 The number of jujube fruits with similar shapes

[0087]

[0088]

[0089] S1-2: Divide the dataset: Divide the jujube image dataset into two parts according to 8:2, namely the training set and the validation set;

[0090] Among them, in the step S2, the teacher model is selected as the ResNet 50 network with higher accuracy, better stability and larger model, and the student model is selected as the MobileNetV3 large network with a smaller model;

[0091] Among them, the step S3 includes:

[0092] S3-1: The backbone network structure of the teacher model is ResNet 50, including convolutional layers, batch normalization, ReLU, and 4 bottleneck modules, and a multi-task framework is implemented by replacing the last two layers with classification layers, which are divided into a shape classification layer and a variety classification layer. Among them, the shape classification layer contains two fully connected layers FC1_1 and FC2_1 to complete the shape classification task, and the parameters are 300×1 and 10×1 respectively. The variety classification layer part is two fully connected layers FC1_2 and FC2_2 to complete the variety classification task, and the parameters are 300×1 and 34×1 respectively;

[0093] The teacher model uses the transfer learning method to load the pre-trained weights on ImageNet, and uses the improved fully connected layer to replace the original RestNet50 structure to complete the initialization of the teacher model; when the jujube fruit image is input into the teacher model, what the teacher model outputs is: the probabilities of the image for each shape category and variety category;

[0094] S3-2: As Figure 3 shown, input a jujube fruit image, and the final output of the teacher model includes the shape label and the variety label. The features obtained after the image passes through the shared layer are respectively passed through the fully connected layers FC1_1 and FC1_2 to become a 300-dimensional shape feature vector and a variety feature vector. On the one hand, the variety feature vector combined with the shape feature vector passes through FC2_1 to obtain a 10-dimensional output vector, and the shape label corresponding to the largest element of the vector is the shape of the current image; on the other hand, the variety feature vector passes through FC2_2 to output a 34-dimensional vector, and the variety label corresponding to the largest element of the vector is the variety of the current image.

[0095] The teacher model uses the cross-entropy loss function CrossEntropyLoss to update the weights, and the shape loss function is expressed as

[0096] L shape_t =CE(y shape ,l abe l shape )

[0097] Among them, label shapeis the true shape label of the current image, y shape is the shape value output by the teacher model;

[0098] The category loss function is expressed as

[0099] L species_t = CE(y species , label specied )

[0100] where label species is the true category label of the current image, y species is the category value output by the teacher model;

[0101] The total loss function of the teacher model is expressed as:

[0102] L Loss_t = γ * L shape_t + (1 - γ) * L species_t

[0103] where γ is an adjustable regularization coefficient with a value of 0.5;

[0104] Use the Stochastic Gradient Descent (SGD) algorithm to optimize the loss function and update the parameters of the teacher model:

[0105] W l , B l = SGD(L loss_t , w l , b l , lr)

[0106] where lr represents the learning rate, SGD represents the SGD function, w l , b l represents the parameters before the update of the l-th layer, W l , B l are the parameters after the update of the l-th layer;

[0107] Furthermore, when training the teacher model, the number of iterations Epoch is 50, the learning rate is set to 0.01, and the cosine annealing algorithm is used to automatically decay the learning rate every Epoch. At the 50th Epoch, it drops to 1 / 1000 of the initial learning rate. According to the video memory size, the batch size is set to 64. When using the SGD algorithm to optimize the loss function, the historical gradient weight coefficient (momentum) and the decay factor (weight decay) are introduced, with values of 0.9 and 5e-4 respectively.

[0108] S3-3: Repeat the process of S3-2 for each jujube image in the training set, update the parameters of the teacher model, and obtain the trained teacher model;

[0109] S3-4: Use each jujube image in the test set as input, and the corresponding shape and category labels as output. Compare the actual labels with the model output labels, calculate the recognition rate, and repeat this process. Finally, select the model with the highest recognition rate as the optimal teacher model;

[0110] Among them, the steps of S4 include:

[0111] The backbone network structure of the student model is MobileNetV3 large, which consists of a convolutional layer, 15 bneck structures (the 4th - 6th and 11th - 15th structures introduce the lightweight attention mechanism SE module), as well as a convolutional layer, a Pooling layer, and a convolutional layer. At the same time, the h-swish activation function is introduced, and the last two layers are replaced with classification layers to implement the multi-task framework, which is divided into a shape classification layer and a category classification layer. The shape classification layer contains two fully connected layers to complete the shape classification task, with parameters of 300×1 and 10×1 respectively. The category classification layer consists of two fully connected layers to complete the category classification task, with parameters of 300×1 and 34×1 respectively;

[0112] The student model uses the transfer learning method to load the pre-trained weights on ImageNet, and replaces the original MobileNetV3 large network with an improved fully connected layer to complete the initialization of the student model; when a jujube fruit image is input into the student model, what the student model outputs are the probabilities of the image belonging to each shape category and category.

[0113] Among them, the steps of S5 include:

[0114] S5-1: Set the teacher model obtained in S3 to the evaluation eval mode and do not participate in backpropagation;

[0115] S5-2: Input the jujube fruit images in the training set into the teacher model and the initialized student model obtained in S4 for forward operation at the same time;

[0116] The calculation methods of the shape loss and the category loss are the same. First, use cross-entropy to calculate the loss between the classification result and the true label, and then use KL divergence to calculate the gap loss between the logits output by the teacher model and the student model; then, weight and sum the two losses to obtain the shape / category loss; finally, weight and sum the shape loss and the category loss to obtain the total loss value;

[0117] Calculate the cross-entropy loss function between the shape output of the student model and the true shape label value, that is:

[0118]

[0119] Among them, x i represents the true shape label of the i-th image, and represents the probability value that the student model outputs the j-th shape category;

[0120] Calculate the KL divergence between the shape output value of the student model and the shape output value of the teacher model, that is:

[0121]

[0122] where T represents the adjustable temperature parameter; represents the probability value that the teacher model outputs the j-th shape category;

[0123] Then the total shape loss function of the student model is:

[0124] L total_shape = α * L CE_shape + (1 - α) * L KD_shape

[0125] where α is the adjustable regularization coefficient;

[0126] Similarly, calculate the cross-entropy loss function between the category output of the student model and the true category label value, that is:

[0127]

[0128] Among them, y i represents the true category label of the i-th image, and represents the probability value that the student model outputs the j-th category;

[0129] Calculate the KL divergence between the category output value of the student model and the category output value of the teacher model, that is:

[0130]

[0131] where T represents the adjustable temperature parameter, and represents the probability value that the teacher model outputs the j-th category;

[0132] Then the total category loss function of the student model is:

[0133] L total_species = β * L CE_species + (1 - β) * L KD_species

[0134] where β is the adjustable regularization coefficient;

[0135] Finally, the total loss of the student model includes the shape loss and the category loss. The total loss function for training the student model is as follows:

[0136] L Loss_s = λ * L total_shape + (1 - λ) * L total_species

[0137] where λ is an adjustable regularization coefficient.

[0138] S5-3: Use the Stochastic Gradient Descent (SGD) algorithm for backpropagation to optimize the total loss function L Loss_s , and update the parameters of the student model:

[0139] W l , B l = SGD(L Loss_t , w l , b l , lr)

[0140] where lr represents the learning rate, SGD represents the SGD function, w l , b l represent the parameters before updating for the l-th layer, and W l , B l are the parameters after updating for the l-th layer;

[0141] Furthermore, in this embodiment, k = 10, m = 34, the temperature coefficient of knowledge distillation is set to T = 5, the values of α, β, and λ are all 0.5, the number of epochs for training the student model is 50, the learning rate is set to 0.01, and the cosine annealing algorithm is used to automatically decay the learning rate every epoch, which drops to 1 / 1000 of the initial learning rate at the 50th epoch. According to the video memory size, the batch size is set to 64. When using the SGD algorithm to optimize the loss function, the historical gradient weight coefficient (momentum) and the decay factor (weight decay) are introduced, with values of 0.9 and 5e-4 respectively.

[0142] S5-4: Repeat the process of S5-2 and S5-3 for each date image in the training set to update the parameters of the student model and obtain the trained student model;

[0143] S5-5: Use each date image in the test set as the input, the corresponding shape and category labels as the output, compare the actual labels with the model output labels, calculate the recognition rate, and repeat this process. Finally, select the model with the highest recognition rate as the optimal student model;

[0144] The software and hardware environment of this embodiment:

[0145] Hardware Configuration

[0146] ①CPU: 2.90GHz Intel Core i7-10700;

[0147] ②Memory: 16G;

[0148] ③Graphics Card: NVIDIA GeForce GTX 1080Ti GPU.

[0149] Software Configuration

[0150] ①Deep Learning Framework: Pytorch;

[0151] ②System: Windows 10;

[0152] ③Library Files: CUDA Library, Cudnn Library, etc.

[0153] The above shows a preferred embodiment of the present invention, and should not limit other embodiments. Any changes and variations made within the principles and concepts of the present invention should be within the protection scope of the present invention.

[0154] Table 2 Classification Effect of Jujube Variety Recognition Model

[0155]

[0156] As can be seen from Table 2, the Resnet50 model has a high accuracy rate of 83.41% on the jujube variety dataset, but it has a large number of model parameters and high requirements for computing resources. The lightweight network MobileNetV3-large has a lower recognition rate, but a small model size and short calculation time. By combining the knowledge distillation technology with the lightweight network MobileNetV3-large, the accuracy of the jujube variety recognition model has been significantly improved. The accuracy rate has increased by 5.84% compared to MobileNetV3-large, and the number of model parameters remains basically unchanged; compared with ResNet50, the accuracy rate has decreased by 0.95%, while the model size is only 1 / 8 of ResNet50.

[0157] The above results show the effectiveness of the lightweight jujube variety recognition model based on knowledge distillation used in the present invention, and this model is suitable for subsequent deployment to small mobile devices.

[0158] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. A lightweight jujube variety recognition method based on knowledge distillation and transfer learning, characterized in that The following steps are involved: S1: Collect jujube image data and build a jujube image dataset; S1-1: The jujube image dataset has a total of n images, and each jujube image is labeled with its true type label, including m jujube varieties, and its shape label is labeled according to its shape characteristics, with a total of k shapes; S1-2: normalizing the jujube image dataset constructed in step S1-1, and transforming the original data into a corresponding standard form; S1-3: The jujube image dataset is divided into two parts: a training set and a validation set. The training set is used to train the network model, and the validation set is used to verify the model recognition effect. S2: Determine the teacher model and student model; S3: Use transfer learning method to train the teacher model to obtain a suitable teacher model; S4: Initialize the student model; S5: Use the knowledge distillation method to use the trained teacher model to guide the student model and obtain the optimal student model; S6: Input the jujube images in the validation set into the student model for testing, and obtain the jujube image shape and variety labels through forward propagation calculation. This completes the jujube variety recognition.

2. According to claim 1, a lightweight jujube variety identification method based on knowledge distillation and transfer learning is characterized in that: In step S2, the teacher model uses ResNet 50 as the backbone network, and the student model uses MobileNetV3large as the backbone network.

3. According to a lightweight jujube variety identification method based on knowledge distillation and transfer learning according to claim 1, it is characterized in that: The corresponding step S3 includes: S3-1: The backbone network structure of the teacher model is ResNet 50, which consists of a convolutional layer, batch normalization, ReLU, and 4 bottleneck modules. The last two layers are replaced by a shape classification layer and a type classification layer. The shape classification layer contains two fully connected layers to complete the shape classification task, with parameters of 300×1 and 10×1 respectively. The type classification layer contains two fully connected layers to complete the type classification task, with parameters of 300×1 and 34×1 respectively. The teacher model uses the transfer learning method to load the pre-trained weights on ImageNet, and then uses the RestNet50 network with the shape classification layer and the type classification layer as the last two layers to complete the teacher model initialization; S3-2: Input a jujube fruit image. After the image data passes through the shape classification layer described in S3-1, a 10-dimensional output vector is obtained. The shape label corresponding to the largest element of the vector is the shape of the current image. After the image data passes through the type classification layer described in S3-1, a 34-dimensional output vector is obtained. The type label corresponding to the largest element of the vector is the type of the current image. The final output of the teacher model includes shape labels and category labels; S3-3: Repeat the process of S3-2 for each jujube image in the training set, update the teacher model parameters, and obtain the trained teacher model; S3-4: Take each jujube image in the test set as input, with the corresponding shape and variety labels as output. Compare the actual labels with the model output labels, calculate the recognition rate, and repeat this process. Finally, select the model with the highest recognition rate as the optimal teacher model.

4. A lightweight jujube variety recognition method based on knowledge distillation and transfer learning according to claim 3, characterized in that The teacher model uses the cross-entropy loss function CrossEntropyLoss to update the weights. The shape loss function is expressed as: L shape_t = CE(y shape , label shape ) Among them, label shape is the ground truth shape label of the current image, and y shape is the shape value output by the teacher model; The variety loss function is expressed as: L species_t = CE(y species , label species ) Among them, label species is the true class label of the current image, and y species is the class value output by the teacher model; The total loss function of the teacher model is expressed as: L Loss_t = γ * L shape_t + (1 - γ) * L species_t where γ is an adjustable regularization coefficient; Use the Stochastic Gradient Descent (SGD) algorithm to optimize the loss function and update the parameters of the teacher model: W l ,B l = SGD(L Loss_t ,w l ,b l ,lr) Among them, lr represents the learning rate, SGD represents the SGD function, w l , b l represents the parameters before the update of the l-th layer, W l , B l are the parameters after the update of the l-th layer.

5. A lightweight jujube variety recognition method based on knowledge distillation and transfer learning according to claim 1, characterized in that The steps in S4 include: The backbone network structure of the student model is MobileNetV3 large, which consists of 1 convolutional layer, 15 bneck structures, 1 convolutional layer, 1 Pooling layer, and 1 convolutional layer. Among them, the SE module of the lightweight attention mechanism is introduced into the 4th - 6th and 11th - 15th structures of the bneck structure; At the same time, the h-swish activation function is introduced, and the last two layers are replaced with a shape classification layer and a variety classification layer. The shape classification layer contains two fully connected layers to complete the shape classification task, with parameters 300×1 and 10×1 respectively. The variety classification layer consists of two fully connected layers to complete the variety classification task, with parameters 300×1 and 4×1 respectively; The student model uses the transfer learning method to load the pre-trained weights on ImageNet, and uses the RestNet50 network with the last two layers replaced by a shape classification layer and a variety classification layer to complete the initialization of the student model. When a jujube fruit image is input into the student model, what the student model outputs is the probability of the image belonging to each shape category and variety category.

6. A lightweight jujube variety recognition method based on knowledge distillation and transfer learning according to claim 1, characterized in that The steps in S5 include: S5-1: Set the teacher model obtained in S3 to the evaluation eval mode and do not participate in backpropagation; S5-2: Input the jujube fruit images in the training set into the teacher model and the initialized student model obtained in S4 for forward operations simultaneously; Among them, the calculation methods of the shape loss and the variety loss are the same; first, use cross-entropy to calculate the loss between the classification result and the true label, then calculate the gap loss between the logits output by the teacher model and the student model through the KL divergence, and then weight and sum the two losses to obtain the shape loss and the variety loss; finally, weight and sum the shape loss and the variety loss to obtain the total loss value; S5-3: Use the Stochastic Gradient Descent (SGD) algorithm for backpropagation to optimize the total loss function L Loss_s , and update the parameters of the student model: W l ,B l = SGD(L Loss_s ,w l ,b l ,lr) Among them, lr represents the learning rate, SGD represents the SGD function, w l , b l represents the parameters before the update of the l-th layer, W l , B l are the parameters after the update of the l-th layer. S5-4: Repeat the process of S5-2 and S5-3 for each jujube image in the training set, update the parameters of the student model, and obtain the trained student model; S5-5: Take each jujube image in the test set as input, and the corresponding shape and variety labels as output. Compare the actual labels with the model output labels, calculate the recognition rate, and repeat this process. Finally, select the model with the highest recognition rate as the optimal student model.

7. A lightweight jujube variety recognition method based on knowledge distillation and transfer learning according to claim 6, wherein the calculation method of the total loss value in the said S5-2 is as follows: Calculate the cross-entropy loss function between the shape output of the student model and the true shape label value, that is: Among them, x i represents the true shape label of the i-th image, and represents the probability value that the student model outputs the j-th shape category; Calculate the KL divergence between the shape output value of the student model and the shape output value of the teacher model, that is: where T represents an adjustable temperature parameter; represents the probability value that the teacher model outputs the j-th shape category; Then the total shape loss function of the student model is: L total_shapee = α * L CE_shape + (1 - α) * L KD_shape where α is an adjustable regularization coefficient; Similarly, calculate the cross-entropy loss function between the variety output of the student model and the true variety label value, that is: Among them, y i represents the true class label of the i-th image, and represents the probability value that the student model outputs the j-th class category; Calculate the KL divergence between the variety output value of the student model and the variety output value of the teacher model, that is: where T represents an adjustable temperature parameter, represents the probability value of the teacher model outputting the j-th category; Then the total variety loss function of the student model is: L total_species = β * L CE_species + (1 - β) * L KD_species where β is an adjustable regularization coefficient; Finally, the total loss of the student model includes shape loss and variety loss, and the total loss function for training the student model is: L Loss_s = λ * L total_shape + (1 - λ) * L total_species where λ is an adjustable regularization coefficient.