A lightweight method for identifying the growth status of economic fruit trees based on PKD-YOLO

By combining the YOLOv8s model, channel pruning algorithm and knowledge distillation algorithm, a lightweight economic forestry and fruit growth state recognition method based on PKD-YOLO was constructed, which solved the problem of single detection varieties and large model parameters, and achieved the lightweight and real-time improvement of the model.

CN118941957BActive Publication Date: 2025-06-06TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410998858.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-06-06
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

In the prior art, there are problems such as single detection varieties and large number of model parameters, which are difficult to meet the real-time requirements, and edge computing hardware cannot run large-scale deep neural network models.

Method used

A lightweight economical forest and fruit growth state recognition method based on PKD-YOLO is adopted, and a lightweight growth state recognition model is constructed by combining the YOLOv8s model, channel pruning algorithm and knowledge distillation algorithm.

Benefits of technology

The problem of variety singleness and large number of model parameters is solved, the model is lightweight is realized, the parameter quantity and model size are reduced, and the accuracy of the model after pruning is restored, improving real-time and feasibility of edge computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118941957B_ABST
    Figure CN118941957B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of deep learning, and discloses a lightweight economic fruit growth state recognition method based on PKD-YOLO, comprising the following steps: constructing an economic fruit growth state picture data set; using the LabelImg annotation tool to annotate the data set; designing a channel pruning strategy to obtain a P-YOLO model, and the specific operations include sparse training, preliminary pruning, iterative pruning, secondary iterative pruning, and model fine-tuning of the YOLOv8s model; performing knowledge distillation on the pruned model, using the YOLOv8s model as a teacher model and the P-YOLO model as a student model, and obtaining a PKD-YOLO model after knowledge distillation training, thereby solving the problem of reduced accuracy caused by pruning. The PKD-YOLO model of the present invention reduces the number of parameters by 81.33%, the model size by 80.75%, the reasoning time by 52.17%, and the FPS by 1.38 times without affecting the accuracy of the original model, thereby achieving the lightweight of the model, which is of great significance for guiding the lightweight of complex models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and in particular relates to a lightweight economic fruit growth status recognition method based on PKD-YOLO. Background Art

[0002] The versatility of computer vision makes it a technical tool suitable for many fields, including precision agriculture. In the field of precision agriculture and smart farms, deep learning technology can more effectively solve the problems of insufficient robustness and generalization than other machine vision technologies, and has become a research focus. Using deep learning technology to obtain the growth status of fruits in real time and accurately, to provide more precise guidance for agricultural activities such as flower thinning, vegetable thinning, and fruit thinning, to more effectively manage orchards, and thus to increase the yield of cash crops, product quality, and economic benefits.

[0003] At present, scholars at home and abroad have carried out extensive research on the detection of economic forest fruits, but they mainly focus on the research of a certain kind of economic forest fruit, and most of them only focus on the maturity and quality of the fruit, etc. There are few studies on the phenological period or growth status identification of multiple types of economic forest fruits. With the increase in the variety of fruits planted in orchards, there is an urgent need for a method that can simultaneously identify the growth status of multiple fruits. Therefore, the present invention studies the growth status detection method of three economic forest fruits: apples, pears, and cherries, and introduces computer vision technology and deep learning technology to achieve it. A large number of scholars use deep learning models to identify the growth status of economic forest fruits, achieving high accuracy, but the deep learning model has a complex structure and a large number of parameters, which makes it difficult to meet real-time requirements, and edge computing hardware cannot run large deep neural network models. Summary of the invention

[0004] In view of the problems of single variety detection and large number of model parameters in the above-mentioned prior art, the present invention provides a lightweight economic forest and fruit growth status recognition method based on PKD-YOLO. By combining the YOLOv8s model, channel pruning algorithm and knowledge distillation algorithm, a lightweight growth status recognition model is constructed to solve the problems of single variety and large number of model parameters.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] A lightweight economic fruit growth status recognition method based on PKD-YOLO includes the following steps:

[0007] S1. Develop observation standards for the growth status of three economic fruit trees: apples, pears and cherries;

[0008] S2, collecting economic forest and fruit growth status images, and preprocessing the collected images to construct an economic forest and fruit growth status data set;

[0009] S3. Design a channel pruning strategy to improve the YOLOv8s model and obtain a preliminary lightweight model P-YOLO;

[0010] S4. Improve the P-YOLO model to obtain the final lightweight model PKD-YOLO;

[0011] S5. Compare the performance of the PKD-YOLO model with other object detection models.

[0012] The method for formulating the observation standard of the growth status of economic fruit trees in S1 is as follows: apple trees, pear trees and cherry trees are woody plants, apples and pears are pome fruits, cherries are stone fruits, and the flowering periods of pome fruits and stone fruits are slightly different. Therefore, the growth status of apples is divided into eight stages: budding stage, leaf expansion stage, initial flowering stage, full flowering stage, flower falling stage, young fruit stage, fruit expansion stage, and fruit maturity stage; the growth status of pears is divided into eight stages: bud expansion stage, bud opening stage, initial flowering stage, full flowering stage, flower falling stage, young fruit stage, fruit expansion stage, and fruit maturity stage; the growth status of cherries is divided into eight stages: budding stage, calyx exposure stage, petal exposure stage, initial flowering stage, full flowering stage, flower falling stage, fruit setting stage, and fruit maturity stage.

[0013] The data set in S2 includes an automatically collected data set and a manually collected data set;

[0014] The automatic acquisition data set includes 8136 apple images. An automatic acquisition system is set up in the orchard, mainly including power supply equipment, an Olympus E450 camera and a probe. The camera is set at a height of 5 meters from the ground, and the probe is set 0.5 meters below the camera. The solar panel is to prevent sudden power outages of the power supply equipment. At the same time, a lightning rod is installed on the top of the observation equipment to prevent thunderstorms. The pictures taken by the camera will be automatically transmitted to the computer via a wireless network. The observation equipment is set to full-time automatic mode and automatically takes a picture every five minutes.

[0015] The manually collected data set includes 22,562 images, including 12,444 images of apples, 5,990 images of pears, and 4,128 images of cherries. The collection scheme is to manually use a Huawei nova 9 mobile phone with a resolution of 4000×3000 for shooting. In order to collect data under different light intensities, the shooting time period is set to 10:00-11:00, 14:00-15:00, and 18:00-19:00 every day. Three fruit trees are photographed each time, and ten pictures are collected for each fruit tree, including four full-view pictures of the tree, three pictures of branches, and three pictures of flowers or fruits.

[0016] The preprocessing methods in S2 are: image labeling, data enhancement and division. The image labeling is performed using labelImg software. The data enhancement is to randomly select one of the following methods for each image: rotation, adding Gaussian noise, salt and pepper noise, and brightness change. The data set is divided into a training set, a validation set and a test set in a ratio of 8:1:1 for subsequent model training.

[0017] The specific construction method of the preliminary lightweight model P-YOLO in S3 is as follows: it includes three steps: sparse training, channel pruning, and model fine-tuning. First, the YOLOv8s model that has been pre-trained on the COCO dataset is sparsely trained using the economic forestry dataset. The sparse training is achieved by adding an L1 regularization penalty term to the original loss function. The formula is as follows:

[0018]

[0019] In the above formula, the first term is the training loss of the network, (x, y) is the training input and target, W is the trainable parameter in the network, the second term is the L1 regular constraint term of the γ coefficient of the BN layer, that is, g(s) = |s|, and λ is the penalty factor used to balance the first and second terms;

[0020] The channel pruning strategy is as follows: the penalty factors are selected as 0.01, 0.005 and 0.003 for sparse training, and the performance of the model after sparse training is compared; while ensuring that most of the β parameters of the BN layer are small enough, the optimal sparsity rate is selected according to the mAP value, and the model after sparse training is initially pruned. The pruning ratio is selected as 70%, 80%, and 90%. The importance score of each channel is represented by the γ coefficient in the BN layer. After sorting the importance scores, the global threshold of the importance score is obtained according to the pruning ratio, and unimportant channels will be pruned. However, if the importance scores of all channels in some layers are less than the global threshold, all channels of this layer will be pruned, which will cause the connection between layers to be disconnected. Therefore, a local threshold is introduced to ensure that the number of channels retained in each layer is not less than 8. The specific implementation is that the initial value of the local threshold is equal to the global threshold, but when the number of channels retained in this layer is less than 8, the local threshold is updated to 0.9 times the current local threshold. Until the number of channels retained by the layer is greater than or equal to 8 under the condition of the local threshold, fine-tuning is performed for 100 epochs after each pruning; after the initial channel pruning, the best model is selected for iterative pruning based on the comprehensive mAP value, the reduced parameter amount and the model size. The proportion retained in each round of pruning is 0.1 times that of the previous round. After each round of pruning, fine-tuning is performed for 200 epochs. When the parameter amount no longer decreases, the iteration is stopped; after the iterative pruning is completed, the best model is selected for secondary iterative pruning to retain at least two channels for each layer. Specifically, when the number of channels retained by the layer is less than 2, the local threshold is updated to 0.9 times the current local threshold, until the number of channels retained by the layer is greater than or equal to 2 under the condition of the local threshold. Similarly, the proportion retained in each round of pruning is 0.1 times that of the previous round. Fine-tuning is performed for 300 epochs after each pruning. When the parameter amount decreases less and the mAP value decreases too much, the iteration is stopped to obtain the P-YOLO model.

[0021] The specific construction method of the final lightweight model PKD-YOLO in S4 is: use the pre-pruning YOLOv8s model to perform knowledge distillation on the pruned P-YOLO model.

[0022] The indicators for comparing different models in S5 include parameters, model size, mean average precision (mAP), inference time (inference), and frames per second (FPS). The formula is as follows:

[0023]

[0024]

[0025] In the above formula, Precision is the precision, Recall is the recall, AP is the area enclosed by the PR curve of a single category and the coordinate axis, and mAP is the average AP of each category. TP (True positive) refers to the number of positive samples predicted as positive samples; FP (Flase positive) refers to the number of negative samples predicted as positive samples; FN (Flase negatives) refers to the number of positive samples predicted as negative samples; C is the number of predicted categories; AP(c) is the AP of the cth category sample.

[0026] Other target detection models in S5 include YOLOv7 model, YOLOv8n model, and YOLOv8s model.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] The present invention obtains the P-YOLO model by combining the YOLOv8s model, the channel pruning algorithm and the economic forest fruit data set, thereby solving the problem of product singularity and the problem of large model parameters; the PKD-YOLO model is obtained by combining the P-YOLO model and the knowledge distillation algorithm, thereby realizing the recovery of the model accuracy after pruning. In addition, the pruning strategy of the present invention designs three processes: preliminary pruning, iterative pruning and secondary iterative pruning. The preliminary pruning realizes the high efficiency of rapid pruning of a large number of channels by proportional one-time pruning, and the iterative pruning and secondary iterative pruning realize the flexibility and adaptability of iterative pruning; and the pruning strategy of the present invention designs a local threshold, thereby solving the problem of disconnection between layers caused by pruning according to the global threshold, and at the same time, more accurate pruning is performed by setting the minimum number of channels retained in each layer. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0030] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0031] Figure 1The present invention is a flowchart of a lightweight economic fruit growth status recognition method based on PKD-YOLO.

[0032] Figure 2 This is a comparison chart of the number of channels of the PKD-YOLO model of the present invention and the original model YOLOv8s.

[0033] Figure 3 This is a comparison chart of the normal training and knowledge distillation training of the P-YOLO model. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than to limit the claims of the present invention. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0035] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0036] In order to solve the problem of variety homogeneity and the problem of large number of parameters of deep learning models, this embodiment proposes a lightweight economic fruit growth status recognition method based on PKD-YOLO, and collects data on economic fruit (apples, pears, cherries) in the orchard of the Agricultural Meteorological Experimental Station in Linyi County, Yuncheng City to solve the problem of variety homogeneity; the YOLOv8s model is improved by combining the channel pruning algorithm to obtain the P-YOLO model to solve the problem of large number of parameters; the P-YOLO model is improved by combining the knowledge distillation algorithm to obtain the PKD-YOLO model to restore the accuracy loss caused by the channel pruning algorithm.

[0037] The overall process of this embodiment is as follows Figure 1 As shown, the specific implementation steps are as follows:

[0038] Step 1: Data preparation

[0039] With reference to the phenological cycles and observation standards of woody plants in the "Specifications for Agricultural Meteorological Observation", observation standards for the growth status of three economic fruit trees, apple, pear and cherry, suitable for this method were formulated.

[0040] The observation standards for the phenological cycles of apples, pears, and cherries are as follows: apple trees, pear trees, and cherry trees are woody plants, apples and pears are pome fruits, and cherries are stone fruits. The flowering periods of pome fruits and stone fruits are slightly different, so the growth status of apples is divided into eight stages: bud expansion period, leaf expansion period, initial flowering period, full flowering period, flower drop period, young fruit period, fruit expansion period, and fruit maturity period, as shown in Table 1; the growth status of pears is divided into eight stages: bud expansion period, bud opening period, initial flowering period, full flowering period, flower drop period, young fruit period, fruit expansion period, and fruit maturity period, as shown in Table 2; the growth status of cherries is divided into eight stages: bud expansion period, calyx exposure period, petal exposure period, initial flowering period, full flowering period, flower drop period, fruit setting period, and fruit maturity period, as shown in Table 3.

[0041] Table 1 Apple growth status and observation standards

[0042]

[0043] Table 2 Pear growth status and observation standards

[0044]

[0045]

[0046] Table 3 Cherry growth status and observation standards

[0047]

[0048] The data set of this embodiment includes an automatically collected data set and a manually collected data set. The data comes from the orchard of the Agricultural Meteorological Experiment Station in Linyi County, Yuncheng City, Shanxi Province, and specifically includes:

[0049] The automatically collected data set includes 8,136 apple images. An automatic collection system is set up in the orchard, which mainly includes power supply equipment, an Olympus E450 camera and a probe. The camera is set at a height of 5 meters from the ground, and the probe is set 0.5 meters below the camera. The solar panels are used to prevent sudden power outages in the power supply equipment. At the same time, a lightning rod is installed on the top of the observation equipment to prevent thunderstorms. The pictures taken by the camera will be automatically transmitted to the computer via a wireless network. The observation time is from March 2023 to November 2023. The observation equipment is set to full-time automatic mode, and a picture is automatically taken every five minutes.

[0050] The manually collected dataset includes 22,562 images, including 12,444 images of apples, 5,990 images of pears, and 4,128 images of cherries. The collection scheme is to manually shoot with a Huawei nova 9 mobile phone with a resolution of 4000×3000. In order to collect data under different light intensities, the shooting time period is set to 10:00-11:00, 14:00-15:00, and 18:00-19:00 every day. Three fruit trees were photographed each time, and ten pictures were collected for each fruit tree, including four full-view pictures of the tree, three pictures of branches, and three pictures of flowers or fruits.

[0051] The preprocessing methods for the above collected images are: image annotation, data enhancement and division. LabelImg software is used for annotation. Data enhancement is to randomly select one of the following methods for each image: rotation, adding Gaussian noise, salt and pepper noise, and brightness change. The data set is divided into training set, validation set and test set in a ratio of 8:1:1 for subsequent model training.

[0052] Step 2: Channel Pruning

[0053] A channel pruning strategy is designed to prune unimportant channels in the model, thereby reducing the number of parameters, compressing the model size, and achieving the lightweight YOLOv8s model to obtain the P-YOLO model.

[0054] The common pruning strategy is to prune in proportion at one time. The pruning strategy of the present invention designs three processes: preliminary pruning, iterative pruning, and secondary iterative pruning. The preliminary pruning realizes the high efficiency of quickly pruning a large number of channels by pruning in proportion at one time. The iterative pruning and secondary iterative pruning realize the flexibility and adaptability of iterative pruning. The pruning strategy of the present invention designs a local threshold, which solves the problem of disconnection between layers caused by pruning according to the global threshold, and at the same time, more accurate pruning is performed by setting the minimum number of channels retained in each layer.

[0055] Channel pruning includes three stages: sparse training, channel pruning, and model fine-tuning. First, the YOLOv8s model that has been pre-trained on the COCO dataset is sparsely trained using the economic forestry dataset. In the sparse training stage, an L1 regularization penalty term is added to the original loss function. The formula is as follows:

[0056]

[0057] In the above formula, the first term is the training loss of the network, (x, y) is the training input and target, W is the trainable parameter in the network, and the second term is the L1 regular constraint term of the γ coefficient of the BN layer, that is, g(s) = |s|, λ is the penalty factor used to balance the first and second terms.

[0058] The penalty factors of 0.01, 0.005 and 0.003 were selected for sparse training. After the three parameter sparse training, most of the β coefficients of the BN layer were small enough. When the penalty factor was 0.003, the mAP value of the model decreased the least compared with the original model, as shown in Table 4. Therefore, the penalty factor of 0.003 was selected.

[0059] Table 4 Sparse training results

[0060]

[0061] The model after sparse training is initially pruned, and the pruning ratio is selected as 70%, 80%, and 90%. The importance score of each channel is represented by the γ coefficient in the BN layer. After sorting the importance scores, the global threshold of the importance score is obtained according to the pruning ratio. Unimportant channels will be pruned, but if the importance scores of all channels in some layers are less than the global threshold, all channels of the layer will be pruned, which will cause the connection between layers to be disconnected. Therefore, a local threshold is introduced to ensure that the number of channels retained in each layer is a multiple of 8. The specific implementation is that the initial value of the local threshold is equal to the global threshold, but when the number of channels retained by the layer is less than 8, the local threshold is updated to 0.9 times the current local threshold until the number of channels retained by the layer is greater than or equal to 8 under the condition of the local threshold. Fine-tuning is performed for 100 epochs after each pruning.

[0062] The preliminary pruning results are shown in Table 5. When the pruning rate is 90%, the number of parameters is reduced by 77.74% and the model size is compressed by 77.24%. Because the model accuracy will be restored through knowledge distillation later, the model pruning stage focuses on the reduction of the number of parameters and model size, and pays less attention to the loss of mAP value. Considering the reduction of the number of parameters and the loss of mAP value, the model with 90% pruning is selected as the final model of the preliminary pruning stage, recorded as Prune9 model.

[0063] Table 5 Preliminary pruning results

[0064]

[0065]

[0066] The final model Prune9 of the preliminary pruning phase is used as the input model of the iterative pruning phase. The retention ratio of each round of pruning is 0.1 times that of the previous round. After each round of pruning, fine-tuning is performed for 200 epochs. When the parameter quantity no longer decreases, the iteration is stopped.

[0067] The results of iterative pruning are shown in Table 6. In the fourth round of iteration, the number of model parameters and the reduction ratio of model size no longer decreased. Therefore, no iteration was performed after the fourth round. The model of the fourth round was used as the final model of the iterative pruning link, recorded as the Prune94 model.

[0068] Table 6 Iterative pruning results

[0069]

[0070] The final model of the iterative pruning phase is used as the input model of the secondary iterative pruning, and at least two channels are retained for each layer. Specifically, when the number of channels retained by the layer is less than 2, the local threshold is updated to 0.9 times the current local threshold, until the number of channels retained by the layer is greater than or equal to 2 under the condition of the local threshold. Similarly, the retention ratio of each round of pruning is 0.1 times that of the previous round, and fine-tuning is performed for 300 epochs after each pruning.

[0071] The results of the second iteration pruning are shown in Table 7. The second iteration only reduces the number of parameters by 0.73% compared with the first iteration, but the mAP value drops by 38.18%. It is not worthwhile to reduce the mAP value by 38.18% in order to reduce the number of parameters by 0.73%. We comprehensively weigh the indicators of the number of parameters and the mAP value. Therefore, the model of the first iteration is selected as the final model P-YOLO of channel pruning, that is, the number of parameters is reduced by 81.33%, the model size is reduced by 80.75%, and the mAP value is lost by 13.26%. The lost mAP value will be recovered in the subsequent knowledge distillation stage.

[0072] Table 7 Secondary iterative pruning results

[0073]

[0074] Step 3: Knowledge Distillation

[0075] The P-YOLO model obtained in step 2 has reduced 81.33% of the parameters compared to the YOLOv8s model, but the mAP value has also been lost to a certain extent. In order to restore the mAP value of the P-YOLO model, the P-YOLO model is improved based on the knowledge distillation algorithm to obtain the PKD-YOLO model.

[0076] The knowledge distillation algorithm uses the knowledge of the teacher model to guide the student model to improve the accuracy of the student model. In this embodiment, the YOLOv8s model in step 2 is used as the teacher model and the P-YOLO model is used as the student model. The original score of the YOLOv8s model before making the final decision represents the model's prediction confidence for each category. The original score is used as a supervisory signal to train the P-YOLO model. By minimizing the difference between the YOLOv8s model and the P-YOLO model at the original score level, the performance of the P-YOLO model is improved to obtain the PKD-YOLO model.

[0077] Table 8 Knowledge distillation results

[0078]

[0079] The parameter T is set to 5, α is set to 0.5, and epoch is set to 500. The training results are shown in Table 8. The PKD-YOLO model is the final model of this step. Compared with the P-YOLO model, the mAP of the PKD-YOLO model has recovered 13.13%; compared with the original YOLOv8s model, the number of parameters of the PKD-YOLO model has dropped from 11135987 to 2078835, and the model size has dropped from 21.66MB to 4.17MB, and the mAP value has hardly dropped. Figure 2 This is a comparison chart of the number of channels of the PKD-YOLO model of the present invention and the original model YOLOv8s. Compared with the original model, the PKD-YOLO model has cut off a large number of channels. Figure 3 This is a comparison chart of normal training and knowledge distillation training of the P-YOLO model. Knowledge distillation training achieves higher accuracy and converges faster than normal training.

[0080] Step 4: Model performance evaluation

[0081] In order to verify the effectiveness of the lightweight economic fruit growth status identification method proposed in this embodiment, the YOLOv7, YOLOv8n, YOLOv8s, and PKD-YOLO models were trained and performance evaluated using the same data set and training parameters. The hardware facilities were a desktop computer equipped with an Intel i9-10900X CPU and an NVIDIA GeForce RTX 3090 24GB GPU. The software tools included CUDA 11.3, Python 3.9, and Pytorch 1.10.0. The results are shown in Table 9. Compared with the original model YOLOv8s, the mAP value of the PKD-YOLO model obtained in this embodiment has almost no decrease, that is, without affecting the model accuracy, the number of parameters is reduced by 81.33%, the model size is reduced by 80.75%, the reasoning time is reduced by 52.17%, and the FPS is increased by 1.38 times, thereby achieving the lightweight of the model; compared with other models, the PKD-YOLO model has the highest mAP, the smallest number of parameters and model size, the fastest reasoning time, and the highest FPS, thus proving that the method of this embodiment is effective.

[0082] Table 9 Model performance comparison results

[0083]

[0084] Only the preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention, and various changes should be included in the protection scope of the present invention.

Claims

1. A lightweight economic fruit growth status recognition method based on PKD-YOLO, characterized in that: The following steps are involved: S1. Develop observation standards for the growth status of three economic fruit trees: apples, pears and cherries; S2, collecting economic forest and fruit growth status images, and preprocessing the collected images to construct an economic forest and fruit growth status data set; S3. Design a channel pruning strategy to improve the YOLOv8s model and obtain a preliminary lightweight model P-YOLO; The specific construction method of the preliminary lightweight model P-YOLO in S3 is as follows: it includes three steps: sparse training, channel pruning, and model fine-tuning. First, the YOLOv8s model that has been pre-trained on the COCO dataset is sparsely trained using the economic forestry dataset. The sparse training is achieved by adding an L1 regularization penalty term to the original loss function. The formula is as follows: In the above formula, the first term is the training loss of the network, (x, y) is the training input and target, W is the trainable parameter in the network, the second term is the L1 regular constraint term of the γ coefficient of the BN layer, that is, g(s) = |s|, and λ is the penalty factor used to balance the first and second terms; The channel pruning strategy is as follows: the penalty factors are selected as 0.01, 0.005 and 0.003 for sparse training, and the performance of the model after sparse training is compared; while ensuring that most of the β parameters of the BN layer are small enough, the optimal sparsity rate is selected according to the mAP value, and the model after sparse training is initially pruned. The pruning ratio is selected as 70%, 80%, and 90%. The importance score of each channel is represented by the γ coefficient in the BN layer. After sorting the importance scores, the global threshold of the importance score is obtained according to the pruning ratio, and unimportant channels will be pruned. However, if the importance scores of all channels in some layers are less than the global threshold, all channels of this layer will be pruned, which will cause the connection between layers to be disconnected. Therefore, a local threshold is introduced to ensure that the number of channels retained in each layer is not less than 8. The specific implementation is that the initial value of the local threshold is equal to the global threshold, but when the number of channels retained in this layer is less than 8, the local threshold is updated to 0.9 times the current local threshold. Until the number of channels retained by the layer is greater than or equal to 8 under the condition of the local threshold, fine-tuning is performed for 100 epochs after each pruning; after the initial channel pruning, the best model is selected for iterative pruning based on the comprehensive mAP value, the reduced parameter amount and the model size. The proportion retained in each round of pruning is 0.1 times that of the previous round. After each round of pruning, fine-tuning is performed for 200 epochs. When the parameter amount no longer decreases, the iteration is stopped; after the iterative pruning is completed, the best model is selected for secondary iterative pruning to retain at least two channels for each layer. Specifically, when the number of channels retained by the layer is less than 2, the local threshold is updated to 0.9 times the current local threshold, until the number of channels retained by the layer is greater than or equal to 2 under the condition of the local threshold. Similarly, the proportion retained in each round of pruning is 0.1 times that of the previous round. Fine-tuning is performed for 300 epochs after each pruning. When the parameter amount decreases less and the mAP value decreases too much, the iteration is stopped to obtain the P-YOLO model; S4. Improve the P-YOLO model to obtain the final lightweight model PKD-YOLO; S5. Compare the performance of the PKD-YOLO model with other object detection models.

2. According to claim 1, a lightweight economic forest and fruit growth state recognition method based on PKD-YOLO is characterized in that: The method for formulating the observation standard of the growth status of economic fruit trees in S1 is as follows: apple trees, pear trees and cherry trees are woody plants, apples and pears are pome fruits, cherries are stone fruits, and the flowering periods of pome fruits and stone fruits are slightly different. Therefore, the growth status of apples is divided into eight stages: budding stage, leaf expansion stage, initial flowering stage, full flowering stage, flower falling stage, young fruit stage, fruit expansion stage, and fruit maturity stage; the growth status of pears is divided into eight stages: bud expansion stage, bud opening stage, initial flowering stage, full flowering stage, flower falling stage, young fruit stage, fruit expansion stage, and fruit maturity stage; the growth status of cherries is divided into eight stages: budding stage, calyx exposure stage, petal exposure stage, initial flowering stage, full flowering stage, flower falling stage, fruit setting stage, and fruit maturity stage.

3. According to claim 1, a lightweight economic forest fruit growth state recognition method based on PKD-YOLO is characterized in that: The data set in S2 includes an automatically collected data set and a manually collected data set; The automatic acquisition data set includes 8136 apple images. An automatic acquisition system is set up in the orchard, mainly including power supply equipment, an Olympus E450 camera and a probe. The camera is set at a height of 5 meters from the ground, and the probe is set 0.5 meters below the camera. The solar panel is to prevent sudden power outages of the power supply equipment. At the same time, a lightning rod is installed on the top of the observation equipment to prevent thunderstorms. The pictures taken by the camera will be automatically transmitted to the computer via a wireless network. The observation equipment is set to full-time automatic mode and automatically takes a picture every five minutes. The manually collected data set includes 22,562 images, including 12,444 images of apples, 5,990 images of pears, and 4,128 images of cherries. The collection scheme is to manually use a Huawei nova 9 mobile phone with a resolution of 4000×3000 for shooting. In order to collect data under different light intensities, the shooting time period is set to 10:00-11:00, 14:00-15:00, and 18:00-19:00 every day. Three fruit trees are photographed each time, and ten pictures are collected for each fruit tree, including four full-view pictures of the tree, three pictures of branches, and three pictures of flowers or fruits.

4. According to claim 1, a lightweight economic forest fruit growth state recognition method based on PKD-YOLO is characterized in that: The preprocessing methods in S2 are: image labeling, data enhancement and division. The image labeling is performed using labelImg software. The data enhancement is to randomly select one of the following methods for each image: rotation, adding Gaussian noise, salt and pepper noise, and brightness change. The data set is divided into a training set, a validation set and a test set in a ratio of 8:1:1 for subsequent model training.

5. According to claim 1, a lightweight economic forest fruit growth state recognition method based on PKD-YOLO is characterized in that: The specific construction method of the final lightweight model PKD-YOLO in S4 is: use the pre-pruning YOLOv8s model to perform knowledge distillation on the pruned P-YOLO model.

6. A lightweight economic forest fruit growth state recognition method based on PKD-YOLO according to claim 1, characterized in that: The indicators for comparing different models in S5 include parameters, model size, mean average precision (mAP), inference time (inference), and frames per second (FPS). The formula is as follows: In the above formula, Precision is the precision, Recall is the recall, AP is the area enclosed by the PR curve of a single category and the coordinate axis, and mAP is the average AP of each category. TP refers to the number of positive samples predicted as positive samples; FP refers to the number of negative samples predicted as positive samples; FN refers to the number of positive samples predicted as negative samples; C is the number of predicted categories; AP(c) is the AP of the cth category sample.

7. The method for identifying the growth status of light-weight economic fruit trees based on PKD-YOLO according to claim 1, characterized in that: Other target detection models in S5 include YOLOv7 model, YOLOv8n model, and YOLOv8s model.

Citation Information

Patent Citations

  • Cicada tea image detection method in natural scene based on improved YOLO model

    CN117690065A

  • Crop growth period prediction method and device

    CN117789037A