A Knowledge Distillation Kiwifruit Visual Counting Method with a 1+2+1 Distillation Mode
By using the knowledge distillation method of 1+2+1 distillation mode in kiwi fruit counting, the YOLOv11 network model and related attention modules are used to optimize the student model, and the problems of insufficient data set coverage, insufficient network processing capabilities and large amount of knowledge distillation calculation in the existing technology are solved, and efficient and accurate kiwi fruit counting is achieved.
Patent Information
- Application Number
- CN202510336159.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The prior art has problems such as insufficient data set coverage in kiwi counting, network limitations in handling occlusion, small targets and multi-scales, and large amounts of calculations and long time for knowledge distillation methods.
The knowledge distillation method of the 1+2+1 distillation mode is adopted, and knowledge distillation training is carried out by constructing student models, teacher models and teaching assistant models, and using the YOLOv11 network model, SEAM attention module, SBA adaptive attention module and SOEP module to perform knowledge distillation training to optimize the performance of the student model.
It reduces the work intensity of fruit farmers, improves the counting efficiency of kiwi fruit, saves labor costs, and improves the accuracy of counting.
Smart Images

Figure CN119887743B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and relates to a knowledge distillation kiwifruit visual counting method with a 1+2+1 distillation mode. Background Art
[0002] As a temperate and subtropical fruit with rich nutrition and unique flavor, kiwifruit is deeply favored by consumers and growers. However, its growth characteristics in the natural environment, such as fruit overlapping, branch occlusion, and light changes, pose great challenges to yield statistics. Traditional manual counting methods are not only time-consuming, laborious, and costly, but also difficult to ensure accuracy. Therefore, developing a convenient and efficient kiwifruit counting technology has become an urgent problem to be solved in digital agricultural production.
[0003] The rapid development of deep learning technology provides new possibilities for solving this problem. A series of algorithms designed specifically for fruit detection, recognition, and localization have been introduced into the agricultural field. These methods, with their reliable performance, strong generalization ability, and excellent environmental adaptability, provide strong support for precision agriculture. As the core means to achieve accurate crop counting, computer vision technology can significantly reduce labor input, improve the accuracy of yield statistics, and at the same time help growers optimize production decisions, increase yields, and reduce production costs, thus promoting the sustainable development of agricultural production.
[0004] Although existing research has made progress in the kiwifruit counting task, there are still some limitations that need to be improved. First, it is necessary to expand the dataset to cover more diverse kiwifruit varieties and types. Second, there is still room for improvement in the network's research on handling occlusion, small targets, and multi-scales. Finally, the knowledge distillation methods used have deficiencies such as long distillation time and large computational complexity. Summary of the Invention
[0005] The purpose of the present invention is to overcome the limitations of the prior art and propose a knowledge distillation kiwifruit visual counting method with a 1+2+1 distillation mode. The 1+2+1 distillation mode refers to the way of knowledge distillation training of 1 teacher model and 2 assistant models for 1 student model. By this mode, the student model is optimized. Through this method, the work intensity of fruit farmers can be reduced, the counting efficiency of kiwifruit can be improved, and the labor cost can be saved.
[0006] The present invention is realized through the following solutions. A knowledge distillation kiwifruit visual counting method with a 1+2+1 distillation mode includes:
[0007] Constructing a student model with the YOLOv11 network model;
[0008] Taking the SEAM attention module as the detection head of the YOLOv11 network model to obtain a teacher model;
[0009] Add the SBA adaptive attention module to the neck network of the YOLOv11 network model to splice features of different scales, obtaining the teaching assistant model I; the SBA adaptive attention module includes two calibration attention units. After the predicted feature maps output by the two calibration attention units are spliced, convolution is performed to obtain the output of the SBA adaptive attention module.
[0010] Add the SOEP module to the neck network of the YOLOv11 network model to obtain the teaching assistant model II. The SOEP module includes the CSP-OmniKernel module and the spatial depth conversion convolution.
[0011] Use the teacher model to distill the teaching assistant model I to obtain the distillation loss function I; use the teacher model to distill the teaching assistant model II to obtain the distillation loss function II.
[0012] Use the sum of the distillation loss function I and the distillation loss function II to distill and train the student model to generate the distillation loss function III. Then, perform self-distillation on the student model, and use the self-distilled student model to detect and count the input kiwifruit images.
[0013] Further preferably, in the neck network of the teaching assistant model I, first use the SBA adaptive attention module to gradually splice the features of each layer of the backbone network from deep to shallow. The spliced features are fused through the feature fusion module. Then, use the SBA adaptive attention module to gradually splice the fused features from shallow to deep. The spliced features are fused through the feature fusion module, and the fused features of different layers are used as the inputs of different detection heads.
[0014] Further preferably, the SBA adaptive attention module includes two calibration attention units. Both input features are processed by the two calibration attention units. After the predicted feature maps output by the two calibration attention units are spliced, 3×3 convolution is performed to obtain the output of the SBA adaptive attention module.
[0015] Further preferably, the SOEP module includes the CSP-OmniKernel module and the spatial depth conversion convolution (SPDConv). In the neck network of the teaching assistant model II, sample the shallow features and the deep features, then process them by the spatial depth conversion convolution, and then splice the features output by the spatial depth conversion convolution with the shallow features. The spliced features are further processed by the CSP-OmniKernel module.
[0016] Further preferably, the teacher model is used to train the collected kiwifruit images to generate the knowledge soft targets of the teacher model, and then the soft labels of the teacher model are generated from the knowledge soft targets of the teacher model. Finally, the distillation loss function I is generated through the soft labels of the teacher model and the soft labels of the teaching assistant model I; the distillation loss function II is generated through the soft labels of the teacher model and the soft labels of the teaching assistant model II.
[0017] Further preferably, the student model after distillation training generates knowledge soft targets for the input kiwifruit images, then generates soft labels from the knowledge soft targets, and calculates the loss of the student model. Finally, hard targets are obtained, and the kiwifruit images are predicted using the hard targets to obtain the recognition frames, and the recognition frames are counted to obtain the number of kiwifruits.
[0018] Further preferably, the student model is trained using the kiwifruit image dataset to generate knowledge soft targets, and then soft labels are generated from the knowledge soft targets.
[0019] The present invention uses a teacher model. The teacher model adds a SEAM module to the head. As the teacher model, it can improve the recognition of the student model for occluded objects in the target object. Then, in order to bridge the gap between the teacher model and the student model, a teaching assistant model is introduced. The teaching assistant model adds an SBA module to the neck network of the YOLOv11 network model. As the teaching assistant model I, it can improve the recognition of the student model for small targets; the CSP-OmniKernel module is added to the neck network of the YOLOv11 network model as the teaching assistant model II to improve the recognition of the module for small targets. Finally, self-distillation operation is performed on the student model to improve the performance of the student model. This method is helpful for counting kiwifruits and has important theoretical significance and practical value in the field of agricultural science. Description of the Drawings
[0020] Figure 1 is the flowchart of the method of the present invention.
[0021] Figure 2 is an example diagram of the kiwifruit fruit image dataset of the present invention. (a) is an image of a mature fruit under strong light, (b) is an image of an immature fruit under strong light, (c) is an image of a branch occlusion, (d) is an image of a leaf occlusion, (e) is an image of a small target, and (f) is an image of multiple scales.
[0022] Figure 3 is the architecture diagram of the YOLOv11 network model.
[0023] Figure 4 is the architecture diagram of the teacher model YOLOv11-SEAM.
[0024] Figure 5 is the architecture diagram of YOLOv11-SBA.
[0025] Figure 6 It is the architecture diagram of YOLOv11-SOEP.
[0026] Figure 7 It is the schematic diagram of the SBA adaptive attention module. Specific implementation manners
[0027] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0028] As Figure 1 shown, a knowledge distillation kiwifruit visual counting method in a 1+2+1 distillation mode includes the following steps one to eight.
[0029] Step 1: Collect kiwifruit images and establish a kiwifruit image dataset;
[0030] Manually label the kiwifruit images to ensure that the kiwifruits are all within the closed rectangular bounding boxes. Finally, 4336 kiwifruit images with a resolution of 640×640 pixels (as Figure 2 shown) are selected as the total kiwifruit image dataset and randomly divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. Table 1 shows the detailed information of the division of the kiwifruit image dataset;
[0031] Table 1
[0032]
[0033] Step 2: Construct a student model using the YOLOv11 network model, train the student model using the kiwifruit image dataset to generate knowledge soft targets, and then generate soft labels from the knowledge soft targets;
[0034] Use the YOLOv11 network model and modify the model training parameters and yaml configuration file therein to make it a student model module;
[0035] The structure of the YOLOv11 network model is as Figure 3As shown in the figure, it includes a backbone network, a neck network, and a head network. The backbone network includes a first convolutional module, a second convolutional module, a first feature fusion module, a third convolutional module, a second feature fusion module, a fourth convolutional module, a third feature fusion module, a fifth convolutional module, a fourth feature fusion module, a spatial pyramid pooling module, and a channel and spatial attention module connected in sequence; the neck network includes a first upsampling module, a first splicing module, a fifth feature fusion module, a second upsampling module, a second splicing module, a sixth feature fusion module, a sixth convolutional module, a third splicing module, a seventh feature fusion module, a seventh convolutional module, and an eighth feature fusion module. The head network includes a first detection head, a second detection head, and a third detection head. The features output by the channel and spatial attention module are processed by the first upsampling module and then spliced with the features output by the third feature fusion module in the first splicing module. The spliced features are subjected to feature fusion by the fifth feature fusion module. The fused features output by the fifth feature fusion module are processed by the second upsampling module and then spliced with the features output by the second feature fusion module in the second splicing module. The spliced features are subjected to feature fusion by the sixth feature fusion module. The fused features output by the sixth feature fusion module serve as the input to the first detection head. The fused features output by the sixth feature fusion module are also processed by the sixth convolutional module. The features output by the sixth convolutional module are spliced with the fused features output by the fifth feature fusion module in the third splicing module. The spliced features are input into the seventh feature fusion module for feature fusion. The fused features output by the seventh feature fusion module serve as the input to the second detection head. The fused features output by the seventh feature fusion module are also processed by the seventh convolutional module. The features output by the seventh convolutional module are spliced with the features output by the channel and spatial attention module in the fourth splicing module. The spliced features are input into the eighth feature fusion module for feature fusion. The fused features output by the eighth feature fusion module serve as the input to the third detection head.
[0036] The main processing flow of the YOLOv11 network model is as follows: First, the image is adjusted to a size of 640×640×3 through image preprocessing, and then the image is input into the backbone network for feature extraction. The result of feature extraction is to generate features with dimensions of 80×80×256, 40×40×512, and 20×20×1024. Subsequently, the three feature layers are subjected to feature fusion within the neck network (Neck) to form three new fused features. Finally, all the fused features are sent to the three detection heads to complete the object detection task.
[0037] It should be noted that for the numbers such as "first, second, third" of each module in this application, they are only for facilitating the description of the model architecture. The module structures with different numbers are the same or similar, only the parameters are different. The structure of the YOLOv11 network model is a publicly available prior art and will not be further elaborated here.
[0038] Step 3: Use the SEAM attention module as the detection head of the YOLOv11 network model to obtain the teacher model YOLOv11-SEAM, as Figure 4 shown;
[0039] Due to the occlusion of obstacles, problems such as incomplete features and misjudgment of the number of fruits are likely to occur. In order to achieve multi-scale kiwifruit detection, enhance the saliency of kiwifruit in the image, and suppress the interference of the background, the SEAM attention module is introduced. The robustness and generalization ability of the model are enhanced through multi-view feature fusion and consistency regularization. The SEAM attention module starts with a depthwise separable convolution operation with residual connection, and then is divided according to the number of channels. Although this method reduces the number of parameters and knows the importance of each channel, it ignores the information interaction between channels. To solve this problem, point (1×1) convolution is used to combine the outputs of different depthwise separable convolutions. Then, a two-layer fully connected network is used to enable all channels to communicate with each other. The loss is compensated by learning the relationship between occluded kiwifruit and unoccluded kiwifruit. The exponential function is borrowed to process the original importance scores learned by the fully connected layer, mapping the values from [0, 1] to [1, e], where e is the natural constant. This mapping relationship improves the tolerance of position errors. Finally, the output of the SEAM attention module is multiplied by the attention of the original features, making the teacher model more effective for the occlusion of kiwifruit.
[0040] Step 4: Add the SBA adaptive attention module to the neck network of the YOLOv11 network model to splice features of different scales, and obtain the YOLOv11-SBA network model as the teaching assistant model Ⅰ;
[0041] The structure of YOLOv11-SBA is as Figure 5As shown, the backbone network is the same as that of the YOLOv11 network model. The head network has 4 detection heads. In the neck network, the features output by the third feature fusion module, the spatial pyramid pooling module, and the channel and spatial attention module are concatenated through the first SBA adaptive attention module. The features concatenated by the first SBA adaptive attention module are fused in the fifth feature fusion module. The fused features output by the fifth feature fusion module and the features output by the second feature fusion module are concatenated in the second SBA adaptive attention module. The features concatenated by the second SBA adaptive attention module are fused in the sixth feature fusion module. The fused features output by the sixth feature fusion module and the features output by the first feature fusion module are concatenated in the third SBA adaptive attention module. The features concatenated by the third SBA adaptive attention module are fused in the ninth feature fusion module. The fused features output by the ninth feature fusion module and the features concatenated by the second SBA adaptive attention module are concatenated in the fourth SBA adaptive attention module. The features concatenated by the fourth SBA adaptive attention module are fused in the seventh feature fusion module. The fused features output by the seventh feature fusion module and the features concatenated by the first SBA adaptive attention module are concatenated in the fifth SBA adaptive attention module. The features concatenated by the fifth SBA adaptive attention module are fused in the tenth feature fusion module. The fused features output by the tenth feature fusion module and the features output by the spatial pyramid pooling module and the channel and spatial attention module are concatenated in the sixth SBA adaptive attention module. The features concatenated by the sixth SBA adaptive attention module are fused in the eighth feature fusion module; The fused features output by the ninth feature fusion module, the fused features output by the seventh feature fusion module, the features concatenated by the fifth SBA adaptive attention module, and the fused features output by the eighth feature fusion module are selected as the inputs for the 4 detection heads of the head network.
[0042] Different from previous fusion methods, the SBA adaptive attention module includes two calibration attention units (RAU). After the predicted feature maps output by the two calibration attention units are concatenated, 3×3 convolution is performed to obtain the output of the SBA adaptive attention module.
[0043] The calibration attention unit (RAU) performs mutual adaptation expression on the input features before fusion. As Figure 7 shown, in the two calibration attention units (RAU), deep and shallow information is input in different ways to make up for the lack of shallow semantic information and the lack of spatial boundary information in the deep layer. The calibration attention unit combines different features and perfects the incomplete features. and are two input features. Through two linear mappings and activation functions, the input features are processed, and the number of channels is reduced to 32 to obtain the feature map and , and then through dot multiplication and reverse operations, the uncertain and approximate measurements are transformed into a complete prediction map. The processing process of the calibration attention unit is expressed as:
[0044] ;
[0045] ;
[0046] where is the predicted feature map output by one of the calibration attention units; are two input features, is the deep semantic information of 40×40×512 dimensions, is the shallow semantic information of 80×80×256 dimensions, is the feature map obtained after processing, is the feature map obtained after processing. and are two sigmoid activation functions, and ⊙ represents element-wise multiplication.
[0047] The fusion process of the SBA adaptive attention module is expressed as:
[0048] ;
[0049] where represents a 3×3 convolution, followed by a batch normalization and a RELU activation layer, is the predicted feature map output by another calibration attention unit, and Concat represents the concatenation operation. Z represents the feature map output by the SBA adaptive attention module, with a dimension of 80×80×256.
[0050] Step 5: Add the SOEP module to the neck network of the YOLOv11 network model to obtain the YOLOv11-SOEP network model, which is used as the teaching assistant model II;
[0051] As Figure 6As shown, the SOEP module includes the CSP-OmniKernel module and the Spatial Depth Conversion Convolution (SPDConv). In the neck network of the YOLOv11 network model, the features output by the second feature fusion module and the fifth feature fusion module are sampled by the second upsampling module, and then processed by the Spatial Depth Conversion Convolution (SPDConv). The features output by the Spatial Depth Conversion Convolution are then concatenated with the features output by the second feature fusion module in the second concatenation module. The concatenated features are then processed by the CSP-OmniKernel module, and the features output by the CSP-OmniKernel module are fused through the sixth feature fusion module.
[0052] Furthermore, in the object detection task using YOLOv11, generally, it is a bit difficult to process small targets on the normal detection layer. The traditional improvement method is to add a second convolution module to improve the detection ability of small targets. However, this also brings some problems, such as an increase in computational complexity and more time-consuming post-processing after adding the second convolution module. To solve the problems and improve efficiency, the neck network is improved. Based on the CSP module and the OmniKernel module, the CSP-OmniKernel module is created, and the SOEP module is proposed. Instead of directly adding a second convolution module, we use the second feature fusion module to obtain features rich in small target information through the Spatial Depth Conversion Convolution (SPDConv), and then fuse them with the sixth feature fusion module.
[0053] Step 6: Use the teacher model to distill the teaching assistant model I. First, use the teacher model to train the collected kiwifruit images to generate the knowledge soft targets of the teacher model. Then, generate the soft labels of the teacher model from the knowledge soft targets of the teacher model. Finally, generate the distillation loss function I through the soft labels of the teacher model and the soft labels of the teaching assistant model I;
[0054] Step 7: Use the teacher model to distill the teaching assistant model II. First, use the teacher model to train the collected kiwifruit images to generate the knowledge soft targets of the teacher model. Then, generate the soft labels of the teacher model from the knowledge soft targets of the teacher model. Finally, generate the distillation loss function II through the soft labels of the teacher model and the soft labels of the teaching assistant model II;
[0055] Step 8: Perform self-distillation on the student model. Use the student model after distillation training to detect and count the input kiwifruit images. The student model is trained by distillation using the sum of distillation loss function Ⅰ and distillation loss function Ⅱ to generate distillation loss function Ⅲ, and then self-distillation operation is performed on the student model to generate distillation loss function Ⅳ. The student model after self-distillation training generates knowledge soft targets for the input kiwifruit images, then generates soft labels from the knowledge soft targets, calculates the loss of the student model, and finally obtains hard targets. Use the hard targets to predict the kiwifruit images to obtain recognition frames, and count the recognition frames, which is the number of kiwifruits.
[0056] To improve the accuracy of the model, more networks and parameters are usually introduced. The result is that the model becomes more complex, but complex models are not suitable for deployment on devices with limited computing power. Knowledge distillation can train complex models by softening the output, thereby improving the accuracy of the model. However, in most cases, the learning ability of the student model is low and it cannot observe many changes. An assistant model is introduced to solve such problems. The learning ability and scale of the assistant model are between the teacher model and the student model, which can help the student model learn better.
[0057] In this application, YOLOv11-SEAM with a large number of parameters and high accuracy is selected as the teacher model for pre-training. Secondly, YOLOv11-SBA and YOLOv11-SOEP network models are used as assistant model Ⅰ and assistant model Ⅱ respectively. The teacher model is used to distill assistant model Ⅰ to obtain distillation loss function Ⅰ. Then, the teacher model is used to distill assistant model Ⅱ to obtain distillation loss function Ⅱ. The student model is trained by distillation using the sum of distillation loss function Ⅰ and distillation loss function Ⅱ to generate distillation loss function Ⅲ, and then self-distillation operation is performed on the student model to generate distillation loss function Ⅳ, so that the student model can obtain more knowledge and improve the generalization ability of the student model. The training results are shown in Table 2. Among them, T-S-S means that the teacher model directly performs knowledge distillation on the student model, and then the student model performs self-distillation operation. T-TA1-S-S means that before the teacher model performs knowledge distillation on the student model, it is necessary to first perform knowledge distillation on assistant model Ⅰ, and then assistant model Ⅰ performs knowledge distillation on the student model, and finally self-distillation operation is performed. T-TA2-S-S means that before the teacher model performs knowledge distillation on the student model, it is necessary to first perform knowledge distillation on assistant model Ⅱ, and then assistant model Ⅱ performs knowledge distillation on the student model, and finally self-distillation operation is performed. The evaluation parameters P represent precision, R represent recall rate, mAP50 represents average precision, and mAP50:95 represents a series of IoU thresholds from 0.5 to 0.95.
[0058] Table 2
[0059]
[0060] To further explore the performance advantages and disadvantages of the present invention and current excellent object detection algorithms, in this embodiment, YOLOv11, YOLOv10, YOLOv9, and YOLOv8 were trained in the same experimental environment. As shown in Table 3, the present invention exhibits better performance. Compared with other single-stage object detection algorithms, its mean average precision (mAP50) is 1.3%, 4.1%, 2.5%, and 1.9% higher than that of YOLOv11, YOLOv10, YOLOv9, and YOLOv8 respectively; its precision is 1.1%, 2.9%, 2.2%, and 0.1% higher respectively. These results highlight the advantages of the present invention.
[0061] Table 3
[0062]
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A knowledge distillation kiwi fruit visual counting method based on a 1+2+1 distillation model, characterized in that: include: Build a student model using the YOLOv11 network model; The SEAM attention module is used as the detection head of the YOLOv11 network model to obtain the teacher model; An SBA adaptive attention module is added to the neck network of the YOLOv11 network model to splice features of different scales, thereby obtaining the teaching assistant model I; the SBA adaptive attention module includes two calibration attention units, and the predicted feature maps output by the two calibration attention units are spliced and convolved to obtain the output of the SBA adaptive attention module; A SOEP module is added to the neck network of the YOLOv11 network model to obtain the teaching assistant model II, wherein the SOEP module includes a CSP-OmniKernel module and a spatial depth conversion convolution; The teacher model is used to distill the teaching assistant model I to obtain the distillation loss function I; the teacher model is used to distill the teaching assistant model II to obtain the distillation loss function II; The student model is distilled and trained using the sum of distillation loss function I and distillation loss function II to generate distillation loss function III. The student model is then self-distilled and the input kiwi images are detected and counted using the student model trained by self-distillation.
2. The knowledge distillation kiwifruit visual counting method according to claim 1, characterized in that: In the neck network of the teaching assistant model I, the features of each layer of the backbone network are first spliced from deep to shallow through the SBA adaptive attention module, and the spliced features are fused through the feature fusion module; then the fused features are spliced from shallow to deep through the SBA adaptive attention module, and the spliced features are fused through the feature fusion module. The fused features of different levels are used as inputs of different detection heads.
3. The knowledge distillation kiwifruit visual counting method according to claim 1, characterized in that: In the neck network of the assistant model II, shallow features and deep features are sampled and then processed by spatial depth conversion convolution. The features output by the spatial depth conversion convolution are then concatenated with the shallow features, and the concatenated features are processed by the CSP-OmniKernel module.
4. The knowledge distillation kiwifruit visual counting method according to claim 1, characterized in that: The collected kiwifruit images are trained using the teacher model to generate the knowledge soft targets of the teacher model, and then the soft labels of the teacher model are generated from the knowledge soft targets. Finally, the distillation loss function I is generated by the soft labels of the teacher model and the soft labels of the teaching assistant model I; the distillation loss function II is generated by the soft labels of the teacher model and the soft labels of the teaching assistant model II.
5. The knowledge distillation kiwifruit visual counting method according to claim 1, characterized in that: The student model after distillation training generates knowledge soft targets for the input kiwifruit images, and then generates soft labels from the knowledge soft targets, and calculates the student model loss to finally obtain hard targets. The hard targets are used to predict the kiwifruit images to obtain recognition boxes. The recognition boxes are counted, which is the number of kiwifruits.
6. The knowledge distillation kiwifruit visual counting method according to claim 1, characterized in that: The student model is trained using the kiwi image dataset to generate knowledge soft targets, which are then used to generate soft labels.
7. The knowledge distillation kiwifruit visual counting method according to claim 1, characterized in that: The processing process of the calibration attention unit is expressed as: ; in, The predicted feature map output by one of the calibrated attention units; are input features, and is the sigmoid activation function, and ⊙ represents point-by-point multiplication.
8. The knowledge distillation kiwifruit visual counting method according to claim 7, characterized in that: The fusion process of the SBA adaptive attention module is expressed as: ; in represents a 3×3 convolution, is the predicted feature map output by another calibrated attention unit, Concat represents the concatenation operation; Z represents the feature map output by the SBA adaptive attention module.
Citation Information
Patent Citations
Method for knowledge distillation and model generation
WO2023210914A1
Knowledge distillation based neural network training method, device, and storage medium
WO2023212997A1