Low-parameter real-time tea tender shoot detection method based on network model
By using a low-parameter network model and a multi-layer feature extraction module in the tea bud detection model, combining the attention mechanism and Focal-IoU_Inner loss function, the problems of high computational complexity and insufficient detection accuracy of the existing models are solved, and efficient and real-time tea bud detection is achieved.
Patent Information
- Application Number
- CN202411976416.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
The existing tea bud detection model has high computational complexity and large model parameters, making it difficult to operate efficiently on equipment with resource-constrained resources, and the detection accuracy and real-time performance of small-target tea buds need to be further optimized.
A low-parameter real-time tea bud detection method based on network model is adopted, feature extraction and downsampling are performed through multi-layer standard convolutional layer and C2f_PConv module, combined with fast spatial pyramid pooling and SimAM_Slice attention mechanism, multi-scale features are integrated, and network parameters are optimized using Focal-IoU_Inner loss function.
It realizes tea bud detection that operates efficiently on resource-constrained equipment, improves detection accuracy and real-time performance, and has good generalization ability and adaptability.
Smart Images

Figure CN120047817A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting tea buds, and more specifically, to a method for real-time detection of tea buds with low parameters based on a network model, belonging to the field of tea bud detection. Background Art
[0002] In recent years, as an important cash crop worldwide, the quality and yield of tea have great significance for the global agricultural development. However, due to the complex tea garden environment and variable working conditions, the traditional manual picking method is inefficient and difficult to meet the needs of large-scale and high-efficiency production. Therefore, intelligent picking technology has gradually become a research hotspot, especially the application of target detection technology based on artificial intelligence in tea agriculture, which has promoted the automation of tea bud recognition and picking tasks. Although deep learning technology has made remarkable progress in the field of tea target detection, most of the existing models have problems such as high computational complexity and excessive resource requirements, and are not suitable for use in resource-constrained devices and environments. At the same time, the detection accuracy and real-time performance of the model for small target tea buds still need to be further optimized. Therefore, researching a lightweight target detection model to achieve real-time performance and high efficiency while ensuring detection accuracy has become the key to solving the above problems. Summary of the Invention
[0003] The purpose of the present invention is to solve the above technical problems and provide a method for real-time detection of tea buds with low parameters based on a network model, which can solve the problems existing in the existing tea bud detection models, such as high computational complexity, large model parameters, and difficulty in efficiently running on resource-constrained devices.
[0004] A method for real-time detection of tea buds with low parameters based on a network model according to the present invention includes the following steps:
[0005] S1. First, collect tea bud images in a real environment and crop the images; then, divide the image dataset into a training set, a validation set, and a test set; next, perform rectangular box annotation on the dataset to clarify the positions of the tea buds; finally, expand the dataset through data augmentation methods;
[0006] S2. Input the image and first perform feature extraction and downsampling on the input image through multiple standard convolutional layers Conv and the C2f_PConv module. Then, fuse multi-level context information and highlight key feature channels through the Fast Spatial Pyramid Pooling SPPF module and the SimAM_Slice module. Next, use upsampling and concatenation operations to integrate high-level semantic features and low-level fine-grained features to form a multi-scale feature representation. Finally, output regression and classification prediction results through the Detect layer of the detection head on feature maps of different scales, and optimize the network parameters by combining the corresponding loss functions CLSLoss and Focal-IoU_Inner to obtain the initial object detection model;
[0007] S3. Based on the constructed object dataset above, train the initial object detection model. Through the annotation information in the dataset, train the model to automatically learn the features and distributions of tea buds;
[0008] S4. Use the trained TeaBudLiteNet network model to detect tea buds, and realize identifying the positions of tea buds in the input image according to the input image and generating corresponding detection results.
[0009] Preferably, the tea bud images collected in S1 include images with different shapes, colors, and sizes obtained randomly;
[0010] Among them, the different shapes refer to the morphological differences of the buds at different growth stages, varieties, and environments; the different colors refer to the color changes of the buds at different growth stages or under different lighting conditions; the different sizes refer to the size differences of the buds at different growth stages and environments.
[0011] Preferably, S2 specifically includes: the input image data is preprocessed and then passes through several standard convolutional layers Conv and the C2f_PConv module in the form of a standard tensor to achieve preliminary feature extraction and downsampling; the standard convolutional layer extracts low-level features such as edges and textures through a learnable convolutional kernel, and at the same time improves the distinguishability of the features under the action of a non-linear activation function; the C2f_PConv module introduces a partial convolution strategy based on the CSP structure, divides the input feature channels into different branches, and performs local convolution and feature concatenation before and after fusion, thereby reducing parameter redundancy and enhancing the diversity of feature expression.
[0012] Preferably, in S2, when the feature reaches a predetermined depth, the network introduces an SPPF module to perform multi-scale aggregation of context information on the feature; the SPPF uses pooling windows of multiple scales to extract global context information at different levels on the same feature map and splices these features in the channel dimension; the feature output from the SPPF then enters the SimAM_Slice module for attention enhancement; the SimAM_Slice weights the feature map in the channel and spatial dimensions, making the network pay more attention to specific target regions and key feature channels while suppressing redundant backgrounds and useless features.
[0013] Preferably, in S2, the network enters the multi-scale feature fusion and upsampling stage: the high-level feature is upsampled in the spatial dimension through the upsampling UpSample operation to match its resolution with the corresponding low-level feature; then the concatenation operation Concat is used to splice the upsampled high-level feature and the early low-level feature in the channel dimension; the fused feature map is further integrated and screened through the C2f_PConv module, so as to still maintain a strong feature expression ability after fusion; by repeating the upsampling-concatenation-C2f_PConv operation, the network establishes rich feature hierarchies at multiple scales.
[0014] Preferably, the feature map that has undergone multiple fusions and refinements in S2 is input to the Detect layer of the detection head, and prediction results are output at different scales; each detection head will parse the feature to obtain the spatial localization parameters and class prediction results of the target; a corresponding mechanism is used to locate the targets that are difficult to locate during the regression process, thereby improving the accuracy of bounding box prediction.
[0015] Preferably, the Focal-IoU_Inner mechanism is used to locate the targets that are difficult to locate during the regression process.
[0016] Preferably, the Focaler-IoU_Inner loss function is:
[0017] L Focaler-IoU_Inner = 1 - IoU Focaler-IoU_Inner
[0018] where IoU Focaler-IoU_Inner The calculation formula is:
[0019]
[0020] where IoU inner is the intersection over union defined by Inner-IoU, and the intersection and union are calculated using the auxiliary box, and its formula is:
[0021]
[0022] Among them, inter inner is the overlapping area of the GT box and the predicted box calculated through the auxiliary box, and its calculation formula is:
[0023] inter inner = max(0, min(bgt r , br) - max(bgt l , bl)) · max(0, min(bgt b , bb) - max(bgt t , bt))
[0024] Among them, bgt r , bgt l , bgt t , bgt b are the boundary coordinates of the auxiliary GT box; br, bl, bt, bb are the boundary coordinates of the auxiliary predicted box;
[0025] Among them, μnion inner is the total area of the GT box and the predicted box minus the overlapping area, and its calculation formula is:
[0026] union inner = (wgt · hgt) · ratio 2 + (w · h) · ratio 2 - inter inner
[0027] Among them, wgt, hgt are the width and height of the GT box; w, h are the width and height of the predicted box; ratio is the scaling ratio of the auxiliary box for adjusting the size of the box.
[0028] Beneficial effects: By combining the efficiency of PConv and the non - linear representation ability of the C2f module, the present invention achieves a good balance between computational efficiency and model performance; using the SimAM_Slice attention mechanism, by balancing the weighted enhancement of large and small target features, it makes full use of local feature distribution information and improves the adaptability of the module to target detection and recognition tasks; introducing the Focaler - IoU_Inner regression loss function, dynamically adjusting the attention to samples, and accelerating convergence through auxiliary bounding boxes, it has strong generalization ability and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is the flowchart of a low - parameter real - time tea bud detection method based on a network model of the present invention.
[0030] Figure 2 is the structural diagram of the neural network model of a low - parameter real - time tea bud detection method based on a network model of the present invention.
[0031] Figure 3 It is the C2f_PConv structure diagram of a low-parameter real-time tea tender shoot detection method based on a network model of the present invention.
[0032] Figure 4 It is the SimAM_Slice structure diagram of a low-parameter real-time tea tender shoot detection method based on a network model of the present invention.
[0033] Figure 5 It is the model training result diagram of a low-parameter real-time tea tender shoot detection method based on a network model of the present invention.
[0034] Figure 6 It is the schematic diagram of the tea tender shoot recognition result of a low-parameter real-time tea tender shoot detection method based on a network model of the present invention. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] The English terms involved in the attached drawings of this application are common technical terms in the field (such as neural network, convolutional network). For further explanation, some of the English terms involved are interpreted as follows: Backbone Network main network, Neck Network neck network, Head Network head network, Loss Function loss function, Conv convolution, SPPF fast spatial pyramid pooling, UpSample upsampling, Concat splicing, Detect multi-scale detection head.
[0037] The main technical problem to be solved by the present invention is to provide a low-parameter real-time tea tender shoot detection method based on a network model, which can solve the problems existing in the existing tea tender shoot detection models, such as high computational complexity, large model parameter quantity, and difficulty in efficiently running on resource-constrained devices.
[0038] As Figure 1 shown, a low-parameter real-time tea tender shoot detection method based on a network model of the present invention comprises the following steps:
[0039] S1. First, collect tea bud images in the real environment and crop the images. Subsequently, divide the image dataset into a training set, a validation set, and a test set. Then, perform rectangular box annotation on the dataset to clarify the positions of the tea buds. Finally, augment the dataset through data augmentation methods;
[0040] S2. Input the image and first perform feature extraction and downsampling on the input image through multiple Conv and C2f_PConv layers. Then, fuse multi-level context information and highlight key feature channels through SPPF and SimAM_Slice. Next, use upsampling and concatenation operations to integrate high-level semantic features and low-level fine-grained features to form a multi-scale feature representation. Finally, output regression and classification prediction results through the Detect layer of the detection head on feature maps of different scales, and optimize the network parameters by combining the corresponding loss functions CLSLoss and Focal-IoU_Inner to obtain an initial object detection model;
[0041] S3. Based on the above constructed target dataset, train the initial object detection model, and let the model automatically learn the features and distributions of the tea buds through the annotation information in the dataset;
[0042] S4. Use the trained TeaBudLiteNet network model to detect the tea buds. This model can identify the positions of the tea buds in the input image and generate corresponding detection results.
[0043] This invention uses the C2f_PConv module, adds the SimAM_Slice attention mechanism, and introduces the Focaler-IoU_Inner regression loss function:
[0044] 1. The C2f_PConv module (as Figure 3 shown)
[0045] This invention proposes the C2f_PConv module. Based on the traditional C2f module, replace the Bottleneck unit in it with a more efficient new operator PConv (Partial Convolution) to construct a faster and more efficient neural network structure.
[0046] The input image data, after preprocessing, passes through several standard convolutional layers Conv and C2f_PConv modules in the form of a standard tensor to achieve preliminary feature extraction and downsampling. The standard convolutional layer extracts low-level features such as edges and textures through a learnable convolutional kernel, and at the same time enhances the discriminability of features under the action of a non-linear activation function. The C2f_PConv module introduces a partial convolution strategy based on the CSP structure, divides the input feature channels into different branches, and performs local convolution and feature splicing before and after fusion, thereby reducing parameter redundancy and enhancing the diversity of feature representation.
[0047] 2. SimAM_Slice Attention Mechanism (as Figure 4 shown)
[0048] The present invention proposes a SimAM_Slice attention mechanism, which improves SimAM by introducing a slicing operation, further enhancing the feature extraction ability for small targets while ensuring the performance stability of large target features.
[0049] When the features reach a certain (specifically, it can be preset according to the actual situation) depth, the network introduces an SPPF module to perform multi-scale aggregation of context information on the features. The SPPF uses pooling windows of multiple scales to extract different levels of global context information on the same feature map and splices these features in the channel dimension. The features output from the SPPF then enter the SimAM_Slice module for attention enhancement. SimAM_Slice weights the feature map in the channel and spatial dimensions, making the network pay more attention to specific target regions and key feature channels while suppressing redundant backgrounds and useless features.
[0050] 3. Focaler-IoU_Inner Loss Function
[0051] The present invention proposes a loss function Focaler-IoU_Inner, which combines the advantages of Focaler-IoU and Inner-IoU, providing an efficient optimization strategy with both sample focusing and dynamic regression capabilities for object detection tasks.
[0052] The feature maps that have undergone multiple fusions and refinements are input to the Detect layer of the detection head, and prediction results are output at different scales. Each detection head parses the features to obtain the spatial localization parameters and class prediction results of the target. Difficult-to-localize targets during the regression process are localized through mechanisms such as Focal-IoU_Inner, thereby improving the accuracy of bounding box prediction.
[0053] The Focaler-IoU_Inner loss function can be expressed as:
[0054] L Focaler-IoU_Inner= 1 - IoU Focaler-IoU_Inner
[0055] where IoU Focaler-IoU_Inner is calculated as follows:
[0056]
[0057] where IoU inner is the intersection over union defined for Inner - IoU, and the intersection and union are calculated using the auxiliary box, and its formula is:
[0058]
[0059] where inter inner is the overlapping area of the GT box and the predicted box calculated through the auxiliary box, and its calculation formula is:
[0060] inter inner = max(0, min(bgt r, br) - max(bgt l, bl))·max(0, min(bgt b, bb) - max(bgt t, bt))
[0061] where bgt r , bgt l , bgt t , bgt b are the boundary coordinates of the auxiliary GT box; br, bl, bt, bb are the boundary coordinates of the auxiliary predicted box.
[0062] where μnion inner is the total area of the GT box and the predicted box minus the overlapping area, and its calculation formula is:
[0063] union inner = (wgt·hgt)·ratio 2 + (w·h)·ratio 2 - inter inner
[0064] where wgt, hgt are the width and height of the GT box; w, h are the width and height of the predicted box; ratio is the scaling ratio of the auxiliary box for adjusting the size of the box.
[0065] The tea shoot image dataset constructed in this invention covers diverse environments, ensuring the detection and classification capabilities of the model in different scenarios. The image acquisition time covers morning, noon, afternoon, and evening, and different weather conditions such as sunny, cloudy, overcast, and rainy days are included, fully reflecting the complex environment in the field. All images were taken with an Apple iPhone 14 Pro Max, using the 48-megapixel main camera, f / 1.78 aperture, 24 mm focal length, and automatically adjusting the exposure to adapt to light changes. The initial resolution is 4032×3024 pixels, ensuring stable image quality under different lighting conditions.
[0066] To ensure data consistency and improve model accuracy, this study performed multi-step preprocessing on the original images. First, the images were cropped to remove irrelevant background information, ensuring that the main body of the tea shoots was located in the center of the image. After strict image data cleaning, the impact of distorted or blurred image data on the training of the network model was avoided. Finally, the dataset was divided into a training set, a validation set, and a test set in a 7:2:1 ratio to ensure the scientificity and reliability of model training, validation, and testing.
[0067] To further enhance the diversity and scale of the dataset, data augmentation was performed on the collected dataset in this embodiment. By performing random transformations such as rotating, scaling, flipping, and adjusting the brightness of the original images, new training samples were generated. These transformations not only increased the number of samples in the dataset but also improved the robustness of the model in detecting tea shoots under different environments and conditions. For example, randomly rotating the images can simulate the performance of tea leaves at different angles, and adjusting the brightness can simulate different lighting conditions, thus ensuring better generalization ability of the model in practical applications.
[0068] After the above processing and data augmentation, a tea shoot dataset containing 4655 images was finally constructed, where the training set contains 3258 images, the validation set contains 931 images, and the test set contains 466 images. It contains one label of tea shoots.
[0069] After 300 epochs of model training iteration for a low-parameter real-time tea shoot detection method based on a network model in this invention, the model proposed in this invention achieved remarkable results on both the training set and the validation set, as Figure 5 shown.
[0070] Among them, box_loss is used to measure the deviation between the predicted bounding box and the ground truth bounding box, cls_loss evaluates the accuracy of target class prediction, and dfl_loss reflects the performance of the model in the refined prediction of the bounding box. Judging from the curves, the three types of losses decrease rapidly during the training stage and tend to be stable after about 200 epochs. The validation loss shows a consistent downward trend but is slightly slower than the training loss, and the gap between the two is small, indicating that the model has good convergence, stable training, and no obvious overfitting. The model has good generalization ability in the tasks of bounding box localization, classification, and distribution optimization, and the training strategy is effective, laying a foundation for the practical application of the model.
[0071] In terms of evaluation metrics, precision and recall reflect the discrimination and recall ability of the model for positive samples, and mAP50 and mAP50-95 indicate the overall detection performance of the model under different IOU thresholds. The curves show that the model performs well under the current hyperparameter and dataset configurations and is suitable for further practical applications.
[0072] The model proposed in the present invention shows good convergence and detection performance in all indicators, demonstrating its practical application value and technical advantages in the task of tea bud detection.
[0073] The schematic diagram of the tea bud recognition result of the low-parameter real-time tea bud detection method based on a network model according to the present invention is as Figure 6 shown. The model proposed in the present invention shows significant advantages in tea bud detection, has a high confidence level, and at the same time significantly improves the detection accuracy while ensuring lightweight, achieving the best balance between performance and resource consumption. It performs well in terms of the detection accuracy and robustness of tea buds in a complex tea garden environment, providing more reliable technical support for subsequent large-scale and efficient tea bud picking and quality monitoring.
[0074] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A low-parameter real-time tea bud detection method based on a network model, characterized in that The method comprises the following steps: S1, first, collect tea bud images in a real environment and crop the images; then, divide the image dataset into training set, validation set and test set; then, annotate the dataset with rectangular frames to clarify the location of tea buds; finally, expand the dataset through data enhancement methods; S2, input image and first pass the input image through multi-layer standard convolution layer Conv and C2f_PConv module for feature extraction and downsampling, then through fast spatial pyramid pooling SPPF module and SimAM_Slice module to fuse multi-level context information and highlight key feature channels, then use upsampling and splicing operations to integrate high-level semantic features with low-level fine-grained features to form a multi-scale feature representation, finally output regression and classification prediction results through the detection head Detect layer on feature maps of different scales, and optimize the network parameters in combination with the loss function CLSLoss and Focal-IoU_Inner to obtain the initial target detection model; S3, based on the target data set constructed above, the initial target detection model is trained, and the training model automatically learns the characteristics and distribution of tea buds through the annotation information in the data set; S4, uses the trained TeaBudLiteNet network model to detect tea buds, recognizes the position of tea buds in the image according to the input image, and generates corresponding detection results.
2. A low-parameter real-time tea bud detection method based on a network model according to claim 1, characterized in that: The tea bud images collected in S1 include randomly obtained images of different shapes, colors, and sizes; Among them, the different forms refer to the morphological differences of the sprouts in different growth stages, varieties and environments; the different colors refer to the color changes of the sprouts in different growth stages or lighting conditions; the different sizes refer to the size differences of the sprouts in different growth stages and environments.
3. A low-parameter real-time tea bud detection method based on a network model according to claim 1 or 2, characterized in that: S2 specifically includes: the input image data is preprocessed and passed through several standard convolutional layers Conv and C2f_PConv modules in the form of standard tensors to achieve preliminary feature extraction and downsampling; the standard convolutional layer extracts low-level features through learnable convolution kernels, and improves the distinguishability of features under the action of nonlinear activation functions; the C2f_PConv module introduces a partial convolution strategy based on the CSP structure, divides the input feature channel into different branches, and performs local convolution and feature splicing before and after fusion, thereby reducing parameter redundancy and improving the diversity of feature expression.
4. The low-parameter real-time tea bud detection method based on a network model according to claim 3, characterized in that: In S2, when the feature reaches the predetermined depth, the network introduces the SPPF module to perform multi-scale aggregation of contextual information on the features; SPPF uses pooling windows of multiple scales to extract global contextual information at different levels on the same feature map, and splices these features in the channel dimension; the features output from SPPF then enter the SimAM_Slice module for attention enhancement; SimAM_Slice weights the feature map in the channel and spatial dimensions, allowing the network to pay more attention to specific target areas and key feature channels, while suppressing redundant background and useless features.
5. A low-parameter real-time tea bud detection method based on a network model according to claim 1 or 4, characterized in that: In S2, the network enters the multi-scale feature fusion and upsampling stage: the high-level features are upsampled in the spatial dimension through the UpSample operation, so that their resolution matches the corresponding shallow features; The Concat operation is then used to concatenate the upsampled high-level features with the early shallow features in the channel dimension; the fused feature map is further integrated and filtered through the C2f_PConv module, so that it can still maintain strong feature expression capabilities after fusion; by repeating upsampling-concatenation-C2f_PConv operations, the network establishes a rich feature hierarchy at multiple scales.
6. A low-parameter real-time tea bud detection method based on a network model according to claim 5, characterized in that: The feature maps in S2 that have been fused and refined multiple times are input into the detection head Detect layer, and the prediction results are output at different scales. The features are analyzed by each detection head to obtain the spatial positioning parameters and category prediction results of the target. The targets that are difficult to locate in the regression process are located through the corresponding mechanism, thereby improving the accuracy of bounding box prediction.
7. A low-parameter real-time tea bud detection method based on a network model according to claim 6, characterized in that: The Focal-IoU_Inner mechanism is used to locate targets that are difficult to locate during the regression process.
8. According to a low-parameter real-time tea bud detection method based on a network model as claimed in claim 7: the Focaler-IoU_Inner loss function is: LFocaler-IoU_Inner=1-IoUFocaler-IoU_Inner in, IoU Focaler-IoU_Inner The calculation formula is: Among them, IoU inner The intersection-union ratio defined for Inner-IoU uses the auxiliary box to calculate the intersection and union. The formula is: Among them, inter inner is the overlapping area of the GT box and the prediction box calculated by the auxiliary box, and its calculation formula is: inter inner =max(0,min(bgt r, br)-max(bgt l, bl))·max(0,min(bgt b, bb)-max(bgt t, bt)) Among them, bgt r ,bgt l ,bgt t ,bgt b are the boundary coordinates of the auxiliary GT box; br, bl, bt, bb are the boundary coordinates of the auxiliary prediction box; Among them, μnion inner The total area of the GT box and the predicted box minus the overlapping area is calculated as follows: union inner =(wgt·hgt)·ratio 2 +(w·h)·ratio 2 -inter inner Among them, wgt,hgt are the width and height of the GT box; w,h are the width and height of the prediction box; ratio is the scaling factor of the auxiliary box used to adjust the size of the box.
Citation Information
Cited By
Digitized in-situ yield estimation method and model for Yinghong No.9 tea garden
CN121147759A
A digital in-situ yield estimation method and device for yinghong No. 9 tea garden
CN121147759B