Light-weight tea disease target detection method based on TeaDisease LiteNet
Through the lightweight tea disease object detection method based on TeaDiseaseLiteNet, feature extraction is used using MobileNetV3 and SPPF modules, and feature fusion is combined with Slim-neck_AKConv and iRMB_EMA modules, the problems of high computing complexity and low detection accuracy in the existing technology are solved, and efficient and accurate tea disease detection is achieved, which is suitable for equipment with resource limitations.
Patent Information
- Application Number
- CN202411976426.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
The existing tea disease detection technology has the problems of high computational complexity, large model parameters, and difficulty in efficient operation on equipment with limited resource-constrained, and the detection accuracy and stability of traditional image processing methods are limited.
The lightweight tea disease object detection method based on TeaDiseaseLiteNet was adopted, and feature extraction was performed through MobileNetV3 and fast spatial pyramid pooling SPPF module, cross-scale feature fusion was performed by combining Slim-neck_AKConv module and iRMB_EMA attention mechanism module, and the model was optimized using Shape-IoU regression loss function and SlideLoss_EMA classification loss function.
It realizes that while ensuring detection accuracy, the calculation complexity and parameter volume of the model are reduced, the applicability in mobile devices and edge computing environments is improved, and the accuracy and recall of tea disease detection is significantly improved. It is suitable for resource-constrained environments and equipment.
Smart Images

Figure CN120047818A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for target detection of tea diseases. More specifically, it relates to a lightweight method for target detection of tea diseases based on TeaDiseaseLiteNet, belonging to the technical field of tea disease detection. Background Art
[0002] Existing tea disease detection technologies mainly rely on traditional manual experience and image processing methods. Although these methods have played a certain role in tea production and quality control, with the expansion of tea planting areas and the diversification of tea varieties, the efficiency and accuracy of manual detection methods have gradually shown deficiencies. Although traditional image processing technologies can achieve automatic identification of tea diseases to a certain extent, their recognition effects are easily affected by external factors such as light and angle, and the detection speed is relatively slow, making it difficult to meet the needs of large-scale tea planting and production.
[0003] In recent years, object detection algorithms based on deep learning have made certain progress in the field of tea disease detection. The YOLO (You Only Look Once) series of algorithms have become a popular choice in tea disease detection research due to their significant advantages in speed and accuracy. YOLOv8 further improves the detection performance and model flexibility by introducing new architecture designs and optimization techniques, making these models show good application prospects in the rapid identification of tea diseases. However, although these models have high detection accuracy, they usually come with a large number of parameters and high computational requirements, which limits their application in mobile devices and edge computing environments, especially in agricultural scenarios that require real-time response, making it difficult to achieve efficient operation.
[0004] The main deficiencies of current tea disease detection technologies include the following aspects: First, the detection accuracy and stability of traditional image processing methods are limited, making it difficult to accurately handle complex and variable disease characteristics; Second, although the YOLO series and other deep learning models perform well in terms of accuracy and speed, their high computational complexity and large number of parameters make it difficult for them to be effectively applied in resource-constrained environments; Finally, there are still deficiencies in dedicated models for tea disease detection, and the existing models still need to be further improved in terms of applicability and generality.
[0005] Therefore, there is an urgent need to develop a lightweight and efficient tea disease detection method to reduce the computational complexity and the number of parameters of the model while ensuring the detection accuracy, thereby improving its applicability in mobile devices and edge computing environments. This can not only promote the wide application of intelligent agricultural technologies in tea production but also significantly improve the economic benefits of the tea industry. Summary of the Invention
[0006] To solve the above-mentioned problems in the prior art, the present invention provides a lightweight tea disease object detection method based on TeaDiseaseLiteNet, which has the technical characteristics of high computational complexity, large model parameter quantity, and difficulty in efficiently running on resource-constrained devices existing in the existing tea disease detection model, etc.
[0007] To achieve the above object, the present invention is realized through the following technical solutions:
[0008] A lightweight tea disease object detection method based on TeaDiseaseLiteNet of the present invention is characterized in that the method includes the following steps:
[0009] S1. Obtain a dataset of tea disease pictures;
[0010] S2. Perform feature extraction through the backbone network Backbone Network, and use the lightweight MobileNetV3 and the fast spatial pyramid pooling SPPF module to achieve efficient multi-scale feature expression; the neck network Neck Network receives the feature maps output from the Backbone Network for cross-scale feature fusion. Take the Slim-neck_AKConv module as the neck network, and use the iRMB_EMA attention mechanism module, and integrate high and low-level feature maps through upsampling, splicing and convolution operations to provide high-quality input for the detection head; the head network Head Network completes the prediction of the target bounding box and category based on the feature maps output by the Neck Network; calculate the loss in combination with the loss function Loss Function and optimize the network parameters, use the Shape-IoU regression loss function, and use the SlideLoss_EMA classification loss function to obtain the initial target detection model;
[0011] S3. Train the initial target detection model based on the target dataset.
[0012] S4. Use the trained TeaDiseaseLiteNet network model to detect tea diseases and obtain the tea disease detection results.
[0013] Preferably, S1 specifically includes: collecting tea disease pictures in a real environment, cropping the images, and then uniformly scaling the images; dividing the tea disease dataset into a training set, a validation set, and a test set, and then performing rectangular box annotation on the tea disease dataset; performing data augmentation on the tea disease dataset.
[0014] The collected tea disease pictures include images under different scenarios, different angles and different lighting conditions;
[0015] Among them, the different scenarios refer to tea plants in different geographical locations, climate conditions, and different growth stages; the different angles refer to taking pictures of the front, back, side, and tilt angles of the leaves; the different lightings refer to taking images under various light conditions such as strong light, shadow, natural light, and scattered light.
[0016] Preferably, S2 specifically includes: First, the input image enters the backbone network Backbone Network, and the initial convolutional layer performs preliminary feature extraction on the original image, reducing the image size to 1 / 2 of the input. On this basis, the backbone network performs step-by-step feature extraction through the MobileNetV3 (series) module, which includes a total of 11 lightweight depthwise separable convolutional modules. The feature maps are gradually downsampled through processing in each layer, and feature maps of different scales from P2 / 4 to P5 / 3 are gradually generated; finally, through the SPPF module, the backbone network performs feature aggregation on the high-order feature maps, and can capture global context information through pooling operations of different scales, and outputs the final feature maps to the Neck Network for multi-scale feature fusion.
[0017] Preferably, the Neck Network in S2 receives multi-scale feature maps output from the Backbone Network, including the output of the SPPF module and the P4 / 1 and P5 / 3 feature maps at earlier levels; the first processing unit of the data stream is the upsampling UpSample module. The upsampling UpSample module performs upsampling on the feature maps output by the SPPF to match its resolution with the shallower P5 / 3 feature maps. The upsampled feature maps and the P5 / 3 feature maps are concatenated through the Concat splicing module to achieve the preliminary fusion of low-level detail information and high-level semantic information;
[0018] The concatenated feature maps enter the VoV-GSCSP module for further processing and reconstruction; VoV-GSCSP fully fuses the concatenated multi-scale features through its internal feature recombination and information transmission mechanism, and enhances the feature expression ability; after being processed by the VoV-GSCSP module, the data stream continues to perform upsampling upward, upsampling the feature maps to a resolution matching the P4 / 1 scale, and the upsampled feature maps and the P4 / 1 feature maps are concatenated again in the Concat module to complete the second cross-scale feature fusion;
[0019] After the second splicing is completed, the feature map is further processed by the VoV-GSCSP module. On this basis, the data stream enters the iRMB_EMA attention mechanism module, which highlights key information through feature reweighting and suppresses irrelevant or redundant information, thereby improving the global perception ability and expression effect of the feature map. The feature map processed by the attention mechanism enters the AKConv module. AKConv further optimizes the spatial distribution and channel information of the features through efficient convolution units to ensure the information consistency of the feature map after multi-scale fusion;
[0020] The processed feature map is concatenated with the feature maps of other upsampling paths in the Concat module to form a multi-level fused intermediate feature. The concatenated feature map is processed again by the VoV-GSCSP module to refine and reconstruct the feature map, further fuse the feature information, and improve the expressiveness of the final feature map. The feature map then flows into the second AKConv module, where the convolution operation is used to further reorganize and optimize the features to ensure high-quality output features.
[0021] Finally, the data stream is concatenated for the third time through the Concat module, combined with the feature information of the previous level, to form a more efficient multi-scale fusion feature again. After the final processing of the VoV-GSCSP module, all cross-scale feature fusion and reconstruction operations are completed, generating the final output multi-scale feature map.
[0022] Preferably, the head network Head Network in S2 completes the specific prediction of target detection based on the feature map output by the Neck Network. The network uses a multi-scale detection head Detect to process feature maps of different resolutions to ensure that targets of different sizes can be detected; each detection head is responsible for outputting two key information: 4*reg_max is used for bounding box regression of the target to predict the precise position and size of the target; num_class is used to predict the category confidence to determine whether there is a target at each pixel position and its category.
[0023] Preferably, the loss function module in S2 is used to evaluate the error between the network prediction result and the true label, and guide the network to perform back-propagation optimization; for the bounding box regression task, Shape-IoU is used as the loss function; in the classification task, the network uses SlideLoss_EMA to calculate the classification error to ensure that the network can accurately distinguish targets of different categories.
[0024] Preferably, the SlideLoss_EMA loss function is expressed as:
[0025]
[0026] Among them, Li is the base loss term for each sample, given by the loss function loss_fcn; w mod (IoU i ) is the modulation weight calculated according to IoUi, used to adjust the weight of each sample in the loss, and its formula is expressed as:
[0027]
[0028] where μ is dynamically updated through the EMA mechanism, and the update rule is:
[0029] μt = d × μt-1 + (1 - d) × IoUt
[0030] where μt is the mean IoU after the t-th update; IoU t is the IoU value of the current sample; d is the decay coefficient, and its calculation formula is:
[0031]
[0032] where γ is a constant, τ is a hyperparameter controlling the EMA decay rate. Through this decay coefficient d, the model can gradually reduce the weight of historical IoU values, so as to focus more attention on the latest IoU changes.
[0033] Beneficial effects: By introducing the lightweight backbone network MobileNetV3, the feature extraction process is optimized, enabling the model to maintain efficient feature extraction capabilities while reducing computational complexity; the Slim-neck_AKConv module not only effectively enhances the multi-scale feature fusion ability but also further reduces the number of model parameters, thereby improving computational efficiency and resource utilization; combined with the iRMB_EMA module, the detection accuracy of small target tea diseases is significantly improved, and the detection ability of the model in complex backgrounds is enhanced. By introducing the Shape-IoU regression loss function and the SlideLoss_EMA classification loss function; the present invention not only improves the accuracy of bounding box prediction but also improves the classification ability for difficult-to-detect samples. Experimental results show that the present invention performs excellently in the tea disease detection task, has higher detection accuracy and recall rate, and significantly reduces computational resource consumption. It is especially suitable for resource-constrained environments and devices, and can provide strong technical support for the early detection and precise prevention and control of tea diseases, promoting the development of smart agriculture. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flowchart of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention.
[0035] Figure 2It is the structural diagram of the neural network model of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention.
[0036] Figure 3 It is the Slim-neck_AKConv structural diagram of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention.
[0037] Figure 4 It is the iRMB_EMA structural diagram of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention.
[0038] Figure 5 It is the model training result diagram of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention.
[0039] Figure 6 It is the schematic diagram of the tea disease recognition result of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] The English terms involved in the accompanying drawings of this application are general technical terms in the field (such as neural network, convolutional network). For further explanation, some of the English terms involved are interpreted as follows: Backbone Network main network, Neck Network neck network, Head Network head network, Loss Function loss function, Conv convolution, SPPF fast spatial pyramid pooling, UpSample upsampling, Concat splicing, Detect multi-scale detection head.
[0042] As Figure 1 shown is a specific embodiment of a lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention, and the steps are as follows:
[0043] S1. Obtain a dataset of tea disease images: Collect tea disease images in a real environment, crop the images, and then uniformly scale the images; divide the tea disease dataset into a training set, a validation set, and a test set, and then perform rectangular box annotation on the tea disease dataset; perform data augmentation on the tea disease dataset;
[0044] S2. Extract features through the Backbone Network, and use the lightweight MobileNetV3 and SPPF modules to achieve efficient multi-scale feature expression; perform cross-scale feature fusion through the Neck Network, use the Slim-neck_AKConv module as the Neck network, and use the iRMB_EMA attention mechanism module. Combine upsampling, splicing, and convolution operations to integrate high- and low-level feature maps and provide high-quality input for the detection head; the Head Network completes the prediction of the target bounding box and category through a multi-scale detection head; calculate the loss and optimize the network parameters in combination with the Loss Function. Use the Shape-IoU regression loss function and the SlideLoss_EMA classification loss function to obtain an initial object detection model;
[0045] S3. Train the initial object detection model based on the target dataset;
[0046] S4. Use the trained TeaDiseaseLiteNet network model to detect tea diseases and obtain tea disease detection results.
[0047] The neural network model structure diagram of the present invention is as Figure 2 shown. The model improves the backbone network based on MobileNetV3, designs the Slim-neck_AKConv to improve the Neck network, designs the iRMB_EMA module and adds it to the small target layer, introduces the Shape-IoU regression loss function, and designs the SlideLoss_EMA classification loss function for use:
[0048] 1. The collected tea disease images include images under different scenarios, different angles, and different lighting conditions;
[0049] Among them, the different scenarios refer to tea trees in different geographical locations, climate conditions, and different growth stages; the different angles refer to shooting from the front, back, side, and tilt angles of the leaves; the different lighting refers to shooting images under various lighting conditions such as strong light, shadow, natural light, and scattered light.
[0050] 2. First, the input image enters the Backbone Network. After passing through the initial convolutional layer, the original image is initially feature-extracted, and the image size is reduced to 1 / 2 of the input. On this basis, the network performs step-by-step feature extraction through a series of MobileNetV3 modules, including a total of 11 lightweight depthwise separable convolutional modules. The feature maps are gradually downsampled through processing in each layer, and feature maps of different scales from P2 / 4 to P5 / 3 are gradually generated. Finally, through the SPPF module, the backbone network aggregates features of the high-order feature maps, captures global context information through pooling operations of different scales, and outputs the final feature maps to the Neck Network for multi-scale feature fusion.
[0051] 3. The Neck Network receives the multi-scale feature maps output from the Backbone Network, including the output of the SPPF module and the P4 / 1 and P5 / 3 feature maps from earlier levels. As Figure 3 shown, the first processing unit of the data stream is the UpSample module, which upsamples the feature maps output by the SPPF to match the resolution of the shallower P5 / 3 feature maps. The upsampled feature maps and the P5 / 3 feature maps are concatenated through the Concat module to achieve the initial fusion of low-level detail information and high-level semantic information.
[0052] Subsequently, the concatenated feature maps enter the VoV-GSCSP module for further processing and reconstruction. VoV-GSCSP fully fuses the concatenated multi-scale features through its internal feature recombination and information transmission mechanisms, and enhances the feature expression ability. After being processed by this module, the data stream continues to perform UpSample upward, upsampling the feature maps to a resolution matching the P4 / 1 scale. The upsampled feature maps and the P4 / 1 feature maps are concatenated again in the Concat module to complete the second cross-scale feature fusion.
[0053] After the second concatenation is completed, the feature maps are further processed through the VoV-GSCSP module. On this basis, the data stream enters the iRMB_EMA attention mechanism module. As Figure 4 shown, key information is highlighted by reweighting the features, and irrelevant or redundant information is suppressed, thereby enhancing the global perception ability and expression effect of the feature maps. The feature maps processed by the attention mechanism enter the AKConv module, and AKConv further optimizes the spatial distribution and channel information of the features through efficient convolutional units to ensure the information consistency of the feature maps after multi-scale fusion.
[0054] Next, the processed feature map is concatenated with the feature maps of other upsampling paths in the Concat module to form a multi-level fused intermediate feature. The concatenated feature map is processed again by the VoV-GSCSP module to refine and reconstruct the feature map, further fuse the feature information, and improve the expressiveness of the final feature map. Subsequently, the feature map flows into the second AKConv module, and the features are further reorganized and optimized through convolution operations to ensure high-quality output features.
[0055] Finally, the data stream is concatenated for the third time through the Concat module, combined with the feature information of the previous level, to form a more efficient multi-scale fusion feature again. After the final VoV-GSCSP module processing, all cross-scale feature fusion and reconstruction operations are completed, generating the final output multi-scale feature map.
[0056] 4. Head Network completes the specific prediction of target detection based on the feature map output by the Neck network. The network uses a multi-scale detection head Detect to process feature maps of different resolutions to ensure that targets of different sizes can be detected. Each detection head is responsible for outputting two key information: 4*reg_max is used for bounding box regression of the target to predict the precise position and size of the target; num_class is used for predicting the category confidence to determine whether there is a target at each pixel position and its category.
[0057] 5. The Loss Function module is used to evaluate the error between the network prediction result and the true label, and guide the network to perform back-propagation optimization. For the bounding box regression task, Shape-IoU is used as the loss function. In the classification task, the network uses SlideLoss_EMA to calculate the classification error to ensure that the network can accurately distinguish between different categories of objects.
[0058] Among them, the SlideLoss_EMA loss function is expressed as:
[0059]
[0060] Among them, L i is the basic loss term for each sample, given by the loss function loss_fcn; w mod (IoU i ) is the modulation weight calculated according to IoUi, which is used to adjust the weight of each sample in the loss. Its formula is expressed as:
[0061]
[0062] Among them, μ is dynamically updated through the EMA mechanism, and the update rule is:
[0063] μt = d × μt-1 + (1 - d) × IoUt
[0064] where μt is the mean IoU after the t-th update; IoU t is the IoU value of the current sample; d is the decay coefficient, and its calculation formula is:
[0065]
[0066] where γ is a constant, τ is a hyperparameter that controls the EMA decay rate. Through this decay coefficient d, the model can gradually reduce the weight of historical IoU values, so as to pay more attention to the latest IoU changes.
[0067] The tea disease image dataset used in the lightweight tea disease target detection method based on TeaDiseaseLiteNet of the present invention aims to support the detection and classification of tea diseases by the model under different environmental conditions. The image shooting environments in the dataset are diverse, including different lighting conditions and weather changes, reflecting the real field conditions, thereby enhancing the practical application ability of the model. A specific implementation manner: The images are collected using the high-definition camera of an Apple iPhone 14 Pro Max smartphone, and the initial resolution is 4032×3024 pixels. In order to ensure data consistency and improve the accuracy of the model, multi-step preprocessing is performed on the collected original images.
[0068] First, the images are cropped to remove irrelevant background information, retain the main part of the tea disease, and ensure that the main part is located at the center of the image. Then, the images are scaled to standardize the size of all images to 640×640 pixels. The bilinear interpolation method is used to retain the details and quality of the images, ensuring that the tea disease features are clearly visible. In addition, a quality check is performed on the scaled images to avoid distortion or blurring. Subsequently, the dataset is divided into a training set and a validation set in a ratio of 8:2.
[0069] To further enhance the diversity and scale of the dataset, data augmentation techniques are adopted. By performing random transformations such as rotation, scaling, flipping, and brightness adjustment on the original images, new training samples are generated. These transformations not only increase the number of samples in the dataset but also improve the robustness of the model in detecting tea diseases under different environments and conditions. For example, randomly rotating the images can simulate the disease manifestations of tea at different angles, and adjusting the brightness can simulate different lighting conditions, thereby ensuring that the model has better generalization ability in practical applications.
[0070] After the above processing and data augmentation, a tea disease dataset containing 6,564 images was finally constructed, among which the training set contained 5,210 images and the validation set contained 1,354 images. The dataset includes 11 tea classifications: Tea Mosquito bug infested leaf, Red Spider infested tea leaf, Black rot of tea, Leafrust of tea, White spot of tea, Algal leaf spot, Grey blight, Brown blight, Helopeltis, other unclassifiable Diseases, and healthy Tea leaf.
[0071] After 300 epochs of model training iteration for the lightweight tea disease object detection method based on TeaDiseaseLiteNet of the present invention, the model proposed by the present invention has achieved remarkable results on both the training set and the validation set, as Figure 5 shown.
[0072] Among them, box_loss measures the error of the model in predicting the position and size of the target bounding box, cls_loss reflects the error of the model in predicting the target category, and dfl_loss evaluates the distributed aggregation loss of the bounding box localization. As the training progresses, these three losses all decrease rapidly and tend to be stable, indicating that the accuracy of the model in target localization and classification is continuously improving and the error is significantly reduced.
[0073] In terms of evaluation metrics, the steady increase and high-level maintenance of precision and recall show the excellent performance of the model in predicting and capturing positive samples. In addition, mAP50 and mAP50-95 measure the average precision under the IoU thresholds of 50% and multiple thresholds respectively. Both reached a high level in the early stage of training and showed good generalization ability and robustness under different detection difficulties.
[0074] The model proposed by the present invention shows good convergence and detection performance in all indicators, proving its practical application value and technical advantages in the tea disease detection task. Specifically, the schematic diagram of the tea disease recognition result is as Figure 6 shown. The model proposed by the present invention shows significant advantages in disease detection, has a high confidence level, and at the same time significantly improves the detection accuracy while ensuring lightweight, achieving the best balance between performance and resource consumption. Its superior performance and reliability make it have higher practical value in practical applications, can more accurately identify diseases, support disease monitoring and prevention in agricultural production, and reflect the perfect combination of the model in terms of efficiency and accuracy.
[0075] The above are only embodiments of the present invention, and do not thereby limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall equally be included within the patent protection scope of the present invention.
Claims
1. A lightweight tea disease target detection method based on TeaDiseaseLiteNet, characterized in that The method comprises the following steps: S1, obtain the dataset of tea disease images; S2, feature extraction is performed through the backbone network Backbone Network, and efficient multi-scale feature expression is achieved by using lightweight MobileNetV3 and fast spatial pyramid pooling SPPF modules; the neck network Neck Network receives the feature map output from the Backbone Network for cross-scale feature fusion, with the Slim-neck_AKConv module as the neck network, and uses the iRMB_EMA attention mechanism module, combined with upsampling, splicing and convolution operations to integrate high- and low-level feature maps to provide high-quality input for the detection head; the head network HeadNetwork completes the prediction of the target bounding box and category through the multi-scale detection head based on the feature map output by the NeckNetwork; the loss function Loss Function is combined to calculate the loss and optimize the network parameters, using the Shape-IoU regression loss function and the SlideLoss_EMA classification loss function to obtain the initial target detection model; S3: training the initial target detection model based on the target data set. S4, use the trained TeaDiseaseLiteNet network model to detect tea diseases and obtain tea disease detection results.
2. The lightweight tea disease target detection method based on TeaDiseaseLiteNet according to claim 1, characterized in that: S1 specifically includes: collecting tea disease pictures in a real environment, cropping the images, and then uniformly scaling the images; dividing the tea disease dataset into a training set, a validation set, and a test set, and then marking the tea disease dataset with rectangular frames; and performing data enhancement on the tea disease dataset.
3. The lightweight tea disease target detection method based on TeaDiseaseLiteNet according to claim 1, characterized in that: S2 specifically includes: first, the input image enters the backbone network Backbone Network, and the initial convolution layer performs preliminary feature extraction on the original image, reducing the image size to 1 / 2 of the input. On this basis, the backbone network performs step-by-step feature extraction through the MobileNetV3 module, which includes a total of 11 layers of lightweight depth-separable convolution modules. The feature map is processed and gradually downsampled in each layer, and feature maps of different scales from P2 / 4 to P5 / 3 are gradually generated; finally, after the SPPF module, the backbone network performs feature aggregation on the high-order feature map, and can capture global context information through pooling operations of different scales, and output the final feature map to NeckNetwork for multi-scale feature fusion.
4. The lightweight tea disease target detection method based on TeaDiseaseLiteNet according to claim 3, characterized in that: The Neck Network in S2 receives multi-scale feature maps output by the Backbone Network, including the output of the SPPF module and the P4 / 1 and P5 / 3 feature maps of the earlier layers. The first processing unit of the data stream is the UpSample module, which upsamples the feature map output by the SPPF to match its resolution with the shallower P5 / 3 feature map. The upsampled feature map is concatenated with the P5 / 3 feature map through the Concat module to achieve a preliminary fusion of low-level detail information and high-level semantic information. The spliced feature map enters the VoV-GSCSP module for further processing and reconstruction; VoV-GSCSP fully integrates the spliced multi-scale features and enhances the feature expression capability through its internal feature reorganization and information transmission mechanism; after being processed by the VoV-GSCSP module, the data flow continues to UpSample upward, and the feature map is further upsampled to a resolution matching the P4 / 1 scale. The upsampled feature map is spliced again with the P4 / 1 feature map in the Concat module to complete the second cross-scale feature fusion; After the second splicing is completed, the feature map is further processed by the VoV-GSCSP module. On this basis, the data stream enters the iRMB_EMA attention mechanism module, which highlights key information through feature reweighting and suppresses irrelevant or redundant information, thereby improving the global perception ability and expression effect of the feature map. The feature map processed by the attention mechanism enters the AKConv module. AKConv further optimizes the spatial distribution and channel information of the features through efficient convolution units to ensure the information consistency of the feature map after multi-scale fusion; The processed feature map is concatenated with the feature maps of other upsampling paths in the Concat module to form a multi-level fused intermediate feature. The concatenated feature map is processed again by the VoV-GSCSP module to refine and reconstruct the feature map, further fuse the feature information, and improve the expressiveness of the final feature map. The feature map then flows into the second AKConv module, where the convolution operation is used to further reorganize and optimize the features to ensure high-quality output features. Finally, the data stream is concatenated for the third time through the Concat module, combined with the feature information of the previous level, to form a more efficient multi-scale fusion feature again. After the final processing of the VoV-GSCSP module, all cross-scale feature fusion and reconstruction operations are completed, generating the final output multi-scale feature map.
5. The lightweight tea disease target detection method based on TeaDiseaseLiteNet according to claim 1, characterized in that: The head network HeadNetwork in S2 completes the specific prediction of target detection based on the feature map output by Neck Network. The network uses a multi-scale detection head Detect to process feature maps of different resolutions to ensure that targets of different sizes can be detected; each detection head is responsible for outputting two key information: 4*reg_max is used for bounding box regression of the target to predict the precise position and size of the target; num_class is used to predict the category confidence to determine whether there is a target at each pixel position and its category.
6. The lightweight tea disease target detection method based on TeaDiseaseLiteNet according to claim 1, characterized in that: The loss function module in S2 is used to evaluate the error between the network prediction result and the true label, and guide the network to perform back-propagation optimization; for the bounding box regression task, Shape-IoU is used as the loss function; in the classification task, the network uses SlideLoss_EMA to calculate the classification error to ensure that the network can accurately distinguish different categories of targets.
7. The lightweight tea disease target detection method based on TeaDiseaseLiteNet according to claim 6, characterized in that: The SlideLoss_EMA loss function is expressed as: Among them, L i is the basic loss term for each sample, given by the loss function loss_fcn; w mod (IoU i ) is the modulation weight calculated according to IoUi, which is used to adjust the weight of each sample in the loss. Its formula is expressed as: Among them, μ is dynamically updated through the EMA mechanism, and the update rule is: μt=d×μt-1+(1-d)×IoUt Among them, μt is the mean IoU after the tth update; IoUt is the IoU value of the current sample; d is the attenuation coefficient, which is calculated as follows: Among them, γ is a constant, τ is a hyperparameter that controls the EMA decay rate. Through the decay coefficient d, the model can gradually reduce the weight of the historical IoU value, so as to focus more attention on the latest IoU changes.
Citation Information
Cited By
Tea disease detection method based on lightweight YOLOv8 model
CN120472325A
Tea disease detection method based on lightweight YOLOv8 model
CN120472325B
Road disease detection method and system based on unmanned aerial vehicle optical remote sensing image
CN120579032A
Dual-scale gesture detection method and device and storage medium
CN121236798A
Camellia oleifera tree disease detection method and system based on unmanned aerial vehicle and YOLO algorithm
CN121708485A