BTM-YOLOv8-based peanut pod detection method and system

By improving the BTM-YOLOv8 network and combining the BiFormer module and MPDIoU loss function, the accuracy and robustness issues of peanut pod detection in complex environments have been solved, achieving efficient and accurate pod detection and grading, which is suitable for smart agriculture.

CN121884124APending Publication Date: 2026-04-17福建省农业科学院数字农业研究所
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
福建省农业科学院数字农业研究所
Filing Date
2025-12-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing peanut pod detection technologies are poorly adapted to interference factors such as changes in light intensity, pod adhesion, size differences, and complex backgrounds in actual production, resulting in decreased recognition accuracy and robustness, making it difficult to meet the requirements for high-precision grading.

Method used

A peanut pod detection method based on BTM-YOLOv8 is adopted. By collecting and expanding historical image datasets, BiFormer module and triple attention module are introduced, and MPDIoU loss function is combined to optimize bounding box localization accuracy, construct a lightweight network architecture, and perform real-time detection in industrial cameras.

Benefits of technology

It significantly improves the accuracy, efficiency, and generalization ability of peanut pod detection, reduces false positives and false negatives, supports agricultural grading and yield assessment, and is suitable for automated production lines and smart agriculture scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884124A_ABST
    Figure CN121884124A_ABST
Patent Text Reader

Abstract

The invention provides a BTM-YOLOv8-based peanut pod detection method and system in the technical field of peanut pod intelligent detection and image recognition. The method comprises the steps of S1, collecting a large number of historical peanut pod images to construct a data set, and dividing the data set into a training set, a verification set and a test set; s2, creating a peanut pod detection model, and setting a loss function of the peanut pod detection model as an MPDIOU function; s3, training the peanut pod detection model through the training set and the loss function, verifying and testing the trained peanut pod detection model through the verification set and the test set in sequence, and deploying the peanut pod detection model passing the test to peanut pod detection equipment; and S4, collecting a real-time peanut pod image by the peanut pod detection equipment, inputting the real-time peanut pod image into the peanut pod detection model to obtain a peanut pod detection result, and displaying the peanut pod detection result. The method has the advantages that the accuracy, efficiency and generalization ability of peanut pod detection are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent detection and image recognition technology for peanut pods, and specifically to a peanut pod detection method and system based on BTM-YOLOv8. Background Technology

[0002] Peanuts (Arachis hypogaea L.) are a globally important oilseed and food crop, widely cultivated in over 100 countries, particularly in Asia and Africa. Post-harvest grading and sorting of peanuts is a crucial step in commercial processing, significantly impacting the quality and price of peanuts. It is also an essential step in oil extraction, food processing, breeding research, and pre-consumption treatment. Efficient sorting helps promote the diversified utilization of peanut products, reduce post-harvest losses, and improve resource utilization efficiency. Currently, manual sorting suffers from low efficiency and high labor intensity, while traditional mechanical sorting easily damages peanut pods and has limited grading accuracy. Therefore, developing efficient and precise intelligent peanut quality detection and classification technologies plays a vital role in improving the quality and efficiency of the peanut industry. Intelligent detection technology can not only significantly improve sorting efficiency and reduce labor costs but also effectively control missed and false detections through high-precision identification, ensuring the consistency of peanut quality on the market, thereby enhancing product added value and market competitiveness.

[0003] In recent years, with the rapid development of computer vision technology, deep learning has demonstrated outstanding performance in image recognition tasks, significantly improving the accuracy and efficiency of automatic recognition and continuously expanding its application boundaries. From security monitoring and medical image analysis to autonomous driving and industrial visual inspection, deep learning, with its powerful feature learning and high-precision recognition capabilities, has shown broad transformative potential in multiple technological fields. In the field of smart agriculture, deep learning has become one of the key technological supports, and related research has made significant progress. For example, Chen et al. (2024) proposed the CES-YOLOv8 network, an improvement based on YOLOv8, for strawberry maturity detection, achieving an accuracy of 88.2%; Megalingam et al. (2024) used an integrated fuzzy and deep learning model (IFDM) to classify coconut maturity, achieving an accuracy of 86.3%; Chen et al. (2024) addressed the weak multi-scale target recognition ability in passion fruit disease detection by constructing the YOLOv8-MDN-Tiny model, which improved precision and recall by 1.5% and 6.0% respectively compared to YOLOv8s; Zhao et al. (2022) based on Faster... R-CNN achieves disease identification of strawberry leaves, flowers and fruits in complex backgrounds, with an average detection precision (mAP) of 92.18% for 7 types of diseases and a single image detection time of 229ms; Liang et al. (2024) proposed an improved YOLOv5 tomato pruning point segmentation model (IPS-YOLO), with an accuracy, recall and average precision of 93.6%, 86.1% and 91.2%, respectively.

[0004] In fruit quality inspection, deep learning offers a new technological approach to address the shortcomings of traditional methods, such as low efficiency, limited applicability, and poor robustness. Through automatic feature extraction, model structure optimization, innovative training strategies, and data augmentation, deep learning significantly improves the accuracy and efficiency of quality inspection. For example, Moallem et al. (2017) used computer vision and support vector machines (SVM) to classify Golden Delicious apples, achieving an accuracy of 92.5% in classifying intact and defective apples and 89.2% in three-level classification tasks; Yang et al. (2021) used an improved deep convolutional network to classify 12 peanut varieties, with an average recognition accuracy of 96.7%; Li et al. (2018) analyzed features such as peanut length-to-width ratio, histogram of oriented gradients (HOG), and Hu invariant moments, combined with SVM to classify single, double, and triple peanuts, with the length-to-width ratio feature achieving the highest classification accuracy of 96.72%; Momin et al. (2017) used image processing algorithms to identify cracked pods, contaminated beans, and defective beans, with accuracies of 96%, 75%, and 98%, respectively; Olgun et al. (2016) based on Dense... SIFT features and SVM classifiers were used to classify wheat grains with an accuracy of 88.33%. Koklu et al. (2021) compared the classification effects of various deep learning models on five rice varieties, with the CNN model achieving 100% classification accuracy. Guo et al. (2024) combined deep module combinatorial optimization (DMCO) with hyperspectral imaging technology to improve the detection accuracy of aflatoxin B1 (AFB1) in peanut kernels, with the optimal combination showing an average absolute error of 0.945 on the validation set. Huang et al. (2024) proposed the lightweight peanut quality detection model LE-YOLO to address the problem of low efficiency in traditional screening, which performed well in both accuracy and speed.

[0005] While the aforementioned methods have achieved some success in specific tasks, existing peanut pod detection technologies still have significant shortcomings: most models are built based on laboratory environments and are poorly adaptable to interference factors in actual production, such as changes in light intensity, pod adhesion, size differences, and complex backgrounds, leading to decreased recognition accuracy and robustness. Furthermore, existing research largely focuses on variety identification or pod count, with weak capabilities for fine-grained detection of appearance quality, making it difficult to meet the demands of high-precision grading. Therefore, a peanut pod detection scheme that can simultaneously achieve high accuracy, high efficiency, and strong generalization ability has yet to emerge, hindering the practical application of intelligent grading.

[0006] Therefore, how to provide a peanut pod detection method and system based on BTM-YOLOv8 to improve the accuracy, efficiency and generalization ability of peanut pod detection has become an urgent technical problem to be solved. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a peanut pod detection method and system based on BTM-YOLOv8, so as to improve the accuracy, efficiency and generalization ability of peanut pod detection.

[0008] In a first aspect, the present invention provides a method for detecting peanut pods based on BTM-YOLOv8, comprising the following steps: Step S1: Collect a large number of historical peanut pod images, perform sample augmentation and annotation on each of the historical peanut pod images to construct a dataset, and divide the dataset into a training set, a validation set and a test set; Step S2: Create a peanut pod detection model based on the improved YOLOv8 network to output peanut pod detection results, and set the loss function of the peanut pod detection model as the MPDIoU function; Step S3: Train the peanut pod detection model using the training set and loss function, and verify and test the trained peanut pod detection model using the validation set and test set in sequence. Deploy the peanut pod detection model that passes the test to the peanut pod detection device. Step S4: The peanut pod detection device acquires real-time peanut pod images, inputs the real-time peanut pod images into the deployed peanut pod detection model to obtain peanut pod detection results, and displays the peanut pod detection results through a graphical user interface.

[0009] Furthermore, step S1 specifically includes: A large number of historical peanut pod images were collected, and each of these historical peanut pod images underwent sample augmentation operations including rotation, brightness adjustment, Gaussian blur, contrast adjustment, and random translation. The peanut pod bounding boxes, seed counts, and peanut-level annotations were then performed on each of the historical peanut pod images after the sample augmentation operations using Labelimg. A dataset was constructed based on the annotated historical peanut pod images, and the dataset was randomly divided into a training set, a validation set, and a test set in a ratio of 7:1:2.

[0010] Furthermore, in step S2, the peanut pod detection model is constructed based on a backbone network, a neck network, and a detection head; the backbone network, the neck network, and the detection head are connected sequentially. The backbone network includes a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a fifth convolutional module, a first BiFormer module, a second BiFormer module, a first cross-stage feature fusion module, a second cross-stage feature fusion module, a triple attention module, and a fast spatial pyramid pooling module. The first convolution module, the second convolution module, the first BiFormer module, the third convolution module, the second BiFormer module, the fourth convolution module, the first cross-stage feature fusion module, the fifth convolution module, the second cross-stage feature fusion module, the triple attention module, and the fast spatial pyramid pooling module are connected in sequence. The neck network includes a first upsampling module, a second upsampling module, a first connection module, a second connection module, a third connection module, a fourth connection module, a third cross-stage feature fusion module, a fourth cross-stage feature fusion module, a fifth cross-stage feature fusion module, a sixth cross-stage feature fusion module, a sixth convolution module, and a seventh convolution module. The fast spatial pyramid pooling module, the first upsampling module, the first connection module, the third cross-stage feature fusion module, the second upsampling module, the second connection module, the fourth cross-stage feature fusion module, the sixth convolution module, the third connection module, the fifth cross-stage feature fusion module, the seventh convolution module, the fourth connection module, and the sixth cross-stage feature fusion module are connected in sequence. The input of the first connection module is connected to the output of the first cross-stage feature fusion module; the input of the second connection module is connected to the output of the second BiFormer module; the input of the third connection module is connected to the output of the third cross-stage feature fusion module; and the input of the fourth connection module is connected to the output of the fast spatial pyramid pooling module. The detection head includes a first detection sub-head, a second detection sub-head, and a third detection sub-head. The input of the first detection sub-head is connected to the output of the fourth cross-stage feature fusion module and the output of the sixth convolution module. The input of the second detection sub-head is connected to the output of the fifth cross-stage feature fusion module. The input of the third detection sub-head is connected to the output of the sixth cross-stage feature fusion module. The first, second, and third detection sub-heads are used for predicting peanut pod bounding boxes, seed quantity, and peanut level, respectively.

[0011] Furthermore, step S3 specifically includes: The peanut pod detection model is trained using the training set. After each training cycle, the validation loss value of the loss function is calculated using the validation set to prevent overfitting. Training is terminated when the validation loss value converges. After training, the peanut pod detection model is evaluated using the test set that was not used in the training. Performance metrics including precision, recall, F1 score, and mean precision are calculated. Once the performance metrics meet the preset deployment requirements, the peanut pod detection model is deployed to the peanut pod detection device.

[0012] Furthermore, step S4 specifically includes: The peanut pod detection equipment uses an industrial camera to capture real-time images of peanut pods on the conveyor belt. After preprocessing the real-time peanut pod images, they are input into a deployed peanut pod detection model for inference, resulting in peanut pod detection results that include the peanut pod bounding box, the number of kernels, and the peanut grade. The peanut pod detection results are displayed in real-time through a graphical user interface and are also stored.

[0013] Secondly, the present invention provides a peanut pod detection system based on BTM-YOLOv8, comprising the following modules: The dataset construction module is used to collect a large number of historical peanut pod images, perform sample augmentation and annotation on each of the historical peanut pod images to construct a dataset, and divide the dataset into a training set, a validation set and a test set. The peanut pod detection model creation module is used to create a peanut pod detection model based on an improved YOLOv8 network to output peanut pod detection results. The loss function of the peanut pod detection model is set to the MPDIoU function. The peanut pod detection model training module is used to train the peanut pod detection model using the training set and loss function, and to verify and test the trained peanut pod detection model using the validation set and test set in sequence, and to deploy the peanut pod detection model that passes the test to the peanut pod detection device. The peanut pod detection module is used by the peanut pod detection equipment to acquire real-time peanut pod images, input the real-time peanut pod images into the deployed peanut pod detection model to obtain peanut pod detection results, and display the peanut pod detection results through a graphical user interface.

[0014] Furthermore, the dataset construction module is specifically used for: A large number of historical peanut pod images were collected, and each of these historical peanut pod images underwent sample augmentation operations including rotation, brightness adjustment, Gaussian blur, contrast adjustment, and random translation. The peanut pod bounding boxes, seed counts, and peanut-level annotations were then performed on each of the historical peanut pod images after the sample augmentation operations using Labelimg. A dataset was constructed based on the annotated historical peanut pod images, and the dataset was randomly divided into a training set, a validation set, and a test set in a ratio of 7:1:2.

[0015] Furthermore, in the peanut pod detection model creation module, the peanut pod detection model is constructed based on a backbone network, a neck network, and a detection head; the backbone network, the neck network, and the detection head are connected sequentially. The backbone network includes a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a fifth convolutional module, a first BiFormer module, a second BiFormer module, a first cross-stage feature fusion module, a second cross-stage feature fusion module, a triple attention module, and a fast spatial pyramid pooling module. The first convolution module, the second convolution module, the first BiFormer module, the third convolution module, the second BiFormer module, the fourth convolution module, the first cross-stage feature fusion module, the fifth convolution module, the second cross-stage feature fusion module, the triple attention module, and the fast spatial pyramid pooling module are connected in sequence. The neck network includes a first upsampling module, a second upsampling module, a first connection module, a second connection module, a third connection module, a fourth connection module, a third cross-stage feature fusion module, a fourth cross-stage feature fusion module, a fifth cross-stage feature fusion module, a sixth cross-stage feature fusion module, a sixth convolution module, and a seventh convolution module. The fast spatial pyramid pooling module, the first upsampling module, the first connection module, the third cross-stage feature fusion module, the second upsampling module, the second connection module, the fourth cross-stage feature fusion module, the sixth convolution module, the third connection module, the fifth cross-stage feature fusion module, the seventh convolution module, the fourth connection module, and the sixth cross-stage feature fusion module are connected in sequence. The input of the first connection module is connected to the output of the first cross-stage feature fusion module; the input of the second connection module is connected to the output of the second BiFormer module; the input of the third connection module is connected to the output of the third cross-stage feature fusion module; and the input of the fourth connection module is connected to the output of the fast spatial pyramid pooling module. The detection head includes a first detection sub-head, a second detection sub-head, and a third detection sub-head. The input of the first detection sub-head is connected to the output of the fourth cross-stage feature fusion module and the output of the sixth convolution module. The input of the second detection sub-head is connected to the output of the fifth cross-stage feature fusion module. The input of the third detection sub-head is connected to the output of the sixth cross-stage feature fusion module. The first, second, and third detection sub-heads are used for predicting peanut pod bounding boxes, seed quantity, and peanut level, respectively.

[0016] Furthermore, the peanut pod detection model training module is specifically used for: The peanut pod detection model is trained using the training set. After each training cycle, the validation loss value of the loss function is calculated using the validation set to prevent overfitting. Training is terminated when the validation loss value converges. After training, the peanut pod detection model is evaluated using the test set that was not used in the training. Performance metrics including precision, recall, F1 score, and mean precision are calculated. Once the performance metrics meet the preset deployment requirements, the peanut pod detection model is deployed to the peanut pod detection device.

[0017] Furthermore, the peanut pod detection module is specifically used for: The peanut pod detection equipment uses an industrial camera to capture real-time images of peanut pods on the conveyor belt. After preprocessing the real-time peanut pod images, they are input into a deployed peanut pod detection model for inference, resulting in peanut pod detection results that include the peanut pod bounding box, the number of kernels, and the peanut grade. The peanut pod detection results are displayed in real-time through a graphical user interface and are also stored.

[0018] The advantages of this invention are: 1. By collecting a large number of historical peanut pod images, sample augmentation and annotation are performed on each historical peanut pod image to construct a dataset, which is then divided into training, validation, and test sets. Next, a peanut pod detection model is created based on an improved YOLOv8 network to output peanut pod detection results, with the MPDIoU function set as the loss function. The peanut pod detection model is then trained using the training set and the loss function, and validated and tested sequentially using the validation and test sets. The successfully tested peanut pod detection model is deployed to a peanut pod detection device. The peanut pod detection device collects real-time peanut pod images and inputs them into the peanut pod detection model to obtain the peanut pod detection results, which are displayed through a graphical user interface. The system displays peanut pod detection results. At the data level, a sample augmentation strategy incorporating operations such as rotation and brightness adjustment effectively simulates complex changes in real-world working environments, providing a rich variety of training samples for the peanut pod detection model and fundamentally enhancing its generalization ability. At the model level, an improved YOLOv8 network is used to introduce the BiFormer module and a triple attention module, enabling it to more accurately capture multi-scale targets and fine-grained features in complex backgrounds. The MPDIoU loss function is combined to optimize bounding box localization accuracy, significantly improving detection accuracy. Furthermore, leveraging YOLOv8's lightweight architecture and multi-scale feature fusion design, and after rigorous training validation and industrial deployment, the system ultimately greatly improves the accuracy, efficiency, and generalization ability of peanut pod detection.

[0019] 2. By improving the YOLOv8 network architecture, a BiFormer module, a cross-stage feature fusion module, and a triple attention module are introduced. These components enhance the model's ability to extract features from peanut pods, especially in complex backgrounds (such as changes in lighting or occlusion), enabling more accurate bounding box localization. Simultaneously, the MPDIoU loss function is used to optimize the bounding box regression process, reducing localization errors and thus improving detection precision and recall. This design makes the model more reliable in detecting peanut pods in agricultural environments, reducing false positives and false negatives, and improving overall performance.

[0020] 3. During the dataset construction phase, various sample augmentation operations (such as rotation, brightness adjustment, Gaussian blur, etc.) were employed, and detailed annotations were used to mark the bounding boxes, seed quantity, and level of peanut pods. This effectively increased data diversity and prevented model overfitting. In addition, the dataset was divided into training, validation, and test sets proportionally, and the loss value was monitored using the validation set during training to ensure that the model was deployed after convergence. This systematic training process improved the model's generalization ability, enabling it to adapt to peanut detection needs in different scenarios and reducing the risk of overfitting.

[0021] 4. The trained model is deployed into a peanut pod inspection device. Real-time images are acquired via an industrial camera and preprocessed, enabling efficient real-time inference. The inspection head is designed with a multi-head structure, capable of simultaneously predicting bounding boxes, kernel count, and peanut grade. This provides comprehensive inspection information, supporting applications such as agricultural grading and yield assessment. The integrated graphical user interface visualizes the results, facilitating real-time monitoring and decision-making, enhancing the system's usability and user experience. It is particularly suitable for automated production lines or smart agriculture scenarios.

[0022] 5. The meticulous design of the backbone network, neck network, and detection head in the network architecture, such as the cross-stage feature fusion module and the fast spatial pyramid pooling module, optimizes feature transfer and scale adaptation, and reduces computational redundancy. The introduction of the BiFormer module enhances the attention mechanism and improves the detection sensitivity of small targets (such as peanut pods). The overall model is based on the lightweight characteristics of YOLOv8, ensuring inference speed and meeting real-time requirements. This structural innovation reduces computational resource consumption while maintaining high accuracy, which is beneficial for deployment on resource-constrained embedded devices.

[0023] 6. By improving the network architecture (such as introducing the BiFormer module, cross-stage feature fusion, and triple attention module) and the MPDIoU loss function, the detection accuracy and robustness are significantly improved, enabling it to effectively cope with light variations and occlusion problems in complex agricultural environments. Simultaneously, through diverse sample augmentation and a systematic dataset partitioning and training process, the model's generalization ability and training efficiency are optimized, preventing overfitting. Real-time detection and multi-task integration are also achieved, simultaneously outputting bounding boxes, seed count, and peanut-level information, and visualizing the results through a graphical user interface, enhancing practicality and user experience. Furthermore, its lightweight design and structured innovation ensure high computational efficiency, facilitating deployment on embedded devices, while a complete verification, testing, and deployment process guarantees reliability and scalability, providing an efficient and reliable solution for intelligent agriculture. Attached Figure Description

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0025] Figure 1 This is a flowchart of a peanut pod detection method based on BTM-YOLOv8 according to the present invention.

[0026] Figure 2 This is a schematic diagram of the structure of a peanut pod detection system based on BTM-YOLOv8 according to the present invention.

[0027] Figure 3 This is a schematic diagram of the peanut pod detection model of the present invention.

[0028] Figure 4 This is a schematic diagram of the structure of the BiFormer module of the present invention.

[0029] Figure 5 This is a schematic diagram of the structure of the triple attention module of the present invention.

[0030] Figure 6 This is a flowchart illustrating the triple attention module of the present invention.

[0031] Figure 7 This is a schematic diagram of the MPDIoU function of this invention.

[0032] Figure 8 This is a schematic diagram of the graphical user interface of the present invention. Detailed Implementation

[0033] The overall approach of the technical solution in this application is as follows: At the data level, a sample augmentation strategy including rotation and brightness adjustment is adopted to effectively simulate the complex changes in the real working environment, providing a rich variety of training samples for the peanut pod detection model and fundamentally enhancing its generalization ability; at the model level, a BiFormer module and a triple attention module are introduced based on the improved YOLOv8 network, enabling it to more accurately capture multi-scale targets and fine-grained features in complex backgrounds, and the bounding box localization accuracy is optimized by combining the MPDIoU loss function, thereby significantly improving the detection accuracy; at the same time, relying on the lightweight architecture and multi-scale feature fusion design of YOLOv8 itself, and after rigorous training verification and industrial deployment, the accuracy, efficiency, and generalization ability of peanut pod detection are improved.

[0034] Please refer to Figures 1 to 8 As shown, a preferred embodiment of the peanut pod detection method based on BTM-YOLOv8 of the present invention includes the following steps: Step S1: Collect a large number of historical peanut pod images, perform sample augmentation and annotation on each of the historical peanut pod images to construct a dataset, and divide the dataset into a training set, a validation set and a test set; Step S2: Create a peanut pod detection model based on the improved YOLOv8 network to output peanut pod detection results, and set the loss function of the peanut pod detection model as the MPDIoU function; The improvements to the YOLOv8 network include the introduction of a BiFormer module into the backbone network, which flexibly adjusts attention weights based on the features of the input image through dynamic sparse attention. A triple attention module is added above the fast spatial pyramid pooling module, and the loss function is replaced with the MPDIoU function.

[0035] Step S3: Train the peanut pod detection model using the training set and loss function, and verify and test the trained peanut pod detection model using the validation set and test set in sequence. Deploy the peanut pod detection model that passes the test to the peanut pod detection device. Step S4: The peanut pod detection device acquires real-time peanut pod images, inputs the real-time peanut pod images into the deployed peanut pod detection model to obtain peanut pod detection results, and displays the peanut pod detection results through a graphical user interface.

[0036] This invention first introduces a Bi-Level Routing Attention (BRA) module into the backbone network of the peanut pod detection model. This module assigns different levels of attention to different locations or features, capturing long-distance dependencies while maintaining efficient computation. This improves the model's ability to extract features from smaller, densely packed peanuts and reduces false detections. Second, a Triplet Attention Module is introduced to enable interaction between different dimensions, enhancing the model's ability to model dependencies between channel and spatial dimensions. Finally, the MPDIoU function simplifies distance metric calculation, allowing bounding box regression to focus more on scale matching optimization. The improved BTM-YOLOv8 (BTM stands for BiFormer + Triplet Attention Module + MPDIoU) achieves a precision of 98.40%, a recall of 96.20%, an F1 score of 97.29%, and an mAP50 of 99.00%, representing improvements of 3.9%, 2.4%, 1.2%, and 3.14% respectively compared to the original YOLOv8 network. It also demonstrated high accuracy in tests of conveyor-type peanut pod detection equipment.

[0037] Step S1 specifically involves: A large number of historical peanut pod images were collected, and each of these historical peanut pod images underwent sample augmentation operations including rotation, brightness adjustment, Gaussian blur, contrast adjustment, and random translation. The peanut pod bounding boxes, seed counts, and peanut-level annotations were then performed on each of the historical peanut pod images after the sample augmentation operations using Labelimg. A dataset was constructed based on the annotated historical peanut pod images, and the dataset was randomly divided into a training set, a validation set, and a test set in a ratio of 7:1:2.

[0038] The number of seeds in a peanut pod is an important characteristic for measuring its fruit structure. Based on the number of seeds contained in the pod, peanut pods can be divided into single-seed pods, double-seed pods, triple-seed pods, and multi-seed pods. Single-seed pods are usually smaller in size, but the seeds inside are more plump. Double-seed pods and triple-seed pods contain two and three peanut seeds, respectively. Multi-seed pods contain four or more peanut seeds.

[0039] To enhance the diversity and scale of the dataset, improve the convergence speed and detection accuracy of the model during training, and increase the model's generalization ability, sample expansion is performed.

[0040] In step S2, the peanut pod detection model is constructed based on a backbone network, a neck network, and a head; the backbone network, neck network, and head are connected sequentially. The neck network is based on a bidirectional feature pyramid (BiFPN-PAN), which dynamically aggregates multi-resolution features of the backbone network through weighted bidirectional cross-scale connections. The top-down path fuses deep semantic information to improve small target detection performance, while the bottom-up path enhances shallow detail features to optimize localization accuracy. The detection head employs a decoupled design, separating classification and regression tasks into independent branches. The classification branch uses binary cross-entropy loss (BCE Loss) for multi-label prediction, while the regression branch combines distributed focus loss (DFL) and complete intersection-union loss (CIoU Loss) to model bounding box coordinates using discrete probability distributions. Anchor-free mechanisms are used to directly predict target center point offset and scale factor, and a dynamic task-aligned label assignment strategy is employed for end-to-end optimization.

[0041] The backbone network includes a first convolutional module (Conv), a second convolutional module, a third convolutional module, a fourth convolutional module, a fifth convolutional module, a first BiFormer module, a second BiFormer module, a first cross-stage feature fusion module (C2f), a second cross-stage feature fusion module, a triple attention module (Triplet AttentionModule), and a fast spatial pyramid pooling module (SPPF). The first convolution module, the second convolution module, the first BiFormer module, the third convolution module, the second BiFormer module, the fourth convolution module, the first cross-stage feature fusion module, the fifth convolution module, the second cross-stage feature fusion module, the triple attention module, and the fast spatial pyramid pooling module are connected in sequence. The BiFormer module is a novel visual Transformer architecture. Its core lies in the Bi-Level Routing Attention (BRA) mechanism. This mechanism uses dynamic sparse attention to flexibly adjust attention weights according to the features of the input image, assigning different levels of attention to different positions or features. This allows it to capture long-distance dependencies while maintaining efficient computation, solving the problem of excessive computational complexity of global self-attention in traditional Transformers, reducing computational load, and improving the efficiency and performance of prediction tasks.

[0042] Specifically, the BiFormer module's two-layer routing attention mechanism first processes the input feature map... Divided into S×S non-overlapping regions, each region contains The eigenvectors are represented by the following formula: ; Where R represents a real number; H represents the height of the feature map; W represents the width of the feature map; C represents the number of channels in the feature map; S represents the number of regions to be divided on each edge; patchify() represents the patching operation; Indicates the size of each region block; Indicates integer division; Then, a region-level average pooling operation is performed on the query and key to obtain the region-level query. and regional key : ; ; Where mean() represents the average pooling operation; dim represents the dimension; Q represents the original query tensor obtained by linear projection; and K represents the original key tensor obtained by linear projection. Next, construct an inter-regional affinity map. : ; Where T represents transpose; Pruning the affinity graph between regions to retain the top-k connections of each region yields the routing index matrix. Then, based on the key tensors and value tensors collected from the routing index matrix: ; ; Where N represents the range of index values; k represents the number of top-k connections retained in each region; This represents the key tensor collected based on the routing index matrix; V represents the value tensor collected based on the routing index matrix; V represents the original value tensor obtained through linear projection; gather() represents collecting data at the corresponding positions from the original tensor based on the index matrix; Finally, attention operations are applied to the collected key-value tensors: ; ; Where dwconv() represents depthwise separable convolution, used to introduce local context information; A represents the attention weight matrix; softmax() represents the normalized exponential function; and output represents the output of the attention module. The BiFormer module's two-layer routing mechanism can effectively distinguish the boundary features of adhering peanut kernels and suppress the interference of background noise (such as soil particles or broken shells) through adaptive weights, thereby improving the model's ability to extract features from smaller, denser peanuts and reducing false detections. Furthermore, its sparse attention mechanism is also more adaptable to high-resolution inputs in real-time detection.

[0043] To extract cross-dimensional interactions of image information and enrich feature representations, a triple attention module is introduced. This module is lightweight, parameter-free, and allows for cross-dimensional interaction. Its core functionality lies in calculating attention weights by modeling the cross-dimensional interaction between channels and space through rotation, pooling, and convolution. The triple attention module consists of three parallel branches. The first two branches handle the cross-dimensional interaction between the channel dimension C and either the spatial dimension H or W. The third branch establishes spatial attention. The outputs of all three branches are summed using a simple averaging method, overcoming the shortcomings of traditional attention mechanisms (such as SE and CBAM) that neglect inter-dimensional coupling.

[0044] The triple attention module captures cross-dimensional information through three parallel branches to calculate attention weights, given an input feature map. The three branches of the triple attention module process it independently.

[0045] In the first branch, an interaction is established between the height (H) and channel (C) dimensions by rotating the input feature map X counterclockwise by 90° along the height dimension, resulting in... (Dimensions are H×W×C); after z-pooling, it is simplified to (2×H×C in dimension), a 1×H×C intermediate tensor is generated through k×k convolutional layers and batch normalization. Attention weights are then generated using a sigmoid activation function and applied to... And restore the original dimensions by rotating 90° clockwise.

[0046] In the second branch, the interaction between the width and channel dimensions is captured, and the input feature map is... By rotating 90° counterclockwise along the width dimension, we obtain... (Dimensions are H×C×W), after z-pooling, it is simplified to (2×H×W in dimension), through k×k convolutional layers and batch normalization, a 1×W×H intermediate tensor is generated. Attention weights are then generated using a sigmoid activation function and applied to... And restore the original dimensions by rotating 90° clockwise.

[0047] The third branch is used to capture the input feature map. To address the spatial dependency, a z-pooling operation is performed along the channel dimension to obtain... (2×H×W in dimension), after k×k convolutional layers and batch normalization, an intermediate tensor of 1×H×W is generated. Attention weights are generated by the sigmoid activation function and directly applied to the original input.

[0048] For the input feature map The process of obtaining fine attention from triple attention and applying tensor y is as follows: ; ; Where σ represents the sigmoid activation function; This represents a standard two-dimensional convolutional layer defined by the kernel size k in the three branches of triple attention; This represents the three cross-dimensional attention weights calculated in triple attention; This indicates a 90° clockwise rotation to maintain the input shape (C×H×W); This represents the tensor obtained after rotating the input feature map in different directions; Indicates to Reshape to perform cross-dimensional attention computation; the underline indicates the inverse operation.

[0049] The neck network includes a first upsampling module, a second upsampling module, a first connection module, a second connection module, a third connection module, a fourth connection module, a third cross-stage feature fusion module, a fourth cross-stage feature fusion module, a fifth cross-stage feature fusion module, a sixth cross-stage feature fusion module, a sixth convolution module, and a seventh convolution module. The fast spatial pyramid pooling module, the first upsampling module, the first connection module, the third cross-stage feature fusion module, the second upsampling module, the second connection module, the fourth cross-stage feature fusion module, the sixth convolution module, the third connection module, the fifth cross-stage feature fusion module, the seventh convolution module, the fourth connection module, and the sixth cross-stage feature fusion module are connected in sequence. The input of the first connection module is connected to the output of the first cross-stage feature fusion module; the input of the second connection module is connected to the output of the second BiFormer module; the input of the third connection module is connected to the output of the third cross-stage feature fusion module; and the input of the fourth connection module is connected to the output of the fast spatial pyramid pooling module. The detection head includes a first detection sub-head, a second detection sub-head, and a third detection sub-head. The input of the first detection sub-head is connected to the output of the fourth cross-stage feature fusion module and the output of the sixth convolution module. The input of the second detection sub-head is connected to the output of the fifth cross-stage feature fusion module. The input of the third detection sub-head is connected to the output of the sixth cross-stage feature fusion module. The first, second, and third detection sub-heads are used for predicting peanut pod bounding boxes, seed quantity, and peanut level, respectively.

[0050] The original YOLOv8 network uses the CIoU function to calculate the localization loss. Although CIoU considers the intersection area of ​​bounding boxes, the distance between center points, and the aspect ratio of the bounding boxes, it uses different measures of aspect ratio instead of the actual difference between width and confidence, which slows down the model's convergence speed. To address the regression problem of overlapping and non-overlapping bounding boxes, and to consider both center point distance and width and height deviations, the MPDIoU function is introduced to replace the original CIoU function.

[0051] The MPDIoU function minimizes the Euclidean distance between the top-left and bottom-right corners of the predicted bounding box and the ground truth bounding box, combining this with IoU to form a new similarity metric. This eliminates the need for complex bounding box calculations, simplifying the computation process and thus improving model convergence speed and regression accuracy. Its expression is: ; ; ; ; Where A represents the predicted bounding box; B represents the ground truth bounding box; This indicates the coordinates of the top-left corner of the prediction box; This indicates the coordinates of the bottom right corner of the prediction box; Represents the coordinates of the top-left corner of the true bounding box; The coordinates of the bottom right corner of the ground truth bounding box are given; w represents the width of the smallest bounding rectangle that contains both the predicted and ground truth bounding boxes; h represents the height of the smallest bounding rectangle that contains both the predicted and ground truth bounding boxes. This represents the square of the Euclidean distance between the top-left corner of the predicted bounding box and the ground truth bounding box; This represents the square of the Euclidean distance between the bottom right corner of the predicted bounding box and the ground truth bounding box.

[0052] Step S3 specifically involves: The peanut pod detection model is trained using the training set. After each training cycle, the validation loss value of the loss function is calculated using the validation set to prevent overfitting. Training is terminated when the validation loss value converges. After training, the peanut pod detection model is evaluated using the test set that was not used in the training. Performance metrics including precision, recall, F1 score, and mean precision are calculated. Once the performance metrics meet the preset deployment requirements, the peanut pod detection model is deployed to the peanut pod detection device.

[0053] Step S4 specifically involves: The peanut pod detection equipment uses an industrial camera (model optional HT-GE1000C-T-CL, 10-megapixel resolution, equipped with two stepless dimming LED light sources) to capture real-time peanut pod images of the conveyor belt. After preprocessing the real-time peanut pod images, they are input into the deployed peanut pod detection model for inference, resulting in peanut pod detection results that include the peanut pod bounding box, number of kernels, and peanut grade. The peanut pod detection results are displayed in real-time through a graphical user interface and are also stored.

[0054] A preferred embodiment of the peanut pod detection system based on BTM-YOLOv8 of the present invention includes the following modules: The dataset construction module is used to collect a large number of historical peanut pod images, perform sample augmentation and annotation on each of the historical peanut pod images to construct a dataset, and divide the dataset into a training set, a validation set and a test set. The peanut pod detection model creation module is used to create a peanut pod detection model based on an improved YOLOv8 network to output peanut pod detection results. The loss function of the peanut pod detection model is set to the MPDIoU function. The improvements to the YOLOv8 network include the introduction of a BiFormer module into the backbone network, which flexibly adjusts attention weights based on the features of the input image through dynamic sparse attention. A triple attention module is added above the fast spatial pyramid pooling module, and the loss function is replaced with the MPDIoU function.

[0055] The peanut pod detection model training module is used to train the peanut pod detection model using the training set and loss function, and to verify and test the trained peanut pod detection model using the validation set and test set in sequence, and to deploy the peanut pod detection model that passes the test to the peanut pod detection device. The peanut pod detection module is used by the peanut pod detection equipment to acquire real-time peanut pod images, input the real-time peanut pod images into the deployed peanut pod detection model to obtain peanut pod detection results, and display the peanut pod detection results through a graphical user interface.

[0056] This invention first introduces a Bi-Level Routing Attention (BRA) module into the backbone network of the peanut pod detection model. This module assigns different levels of attention to different locations or features, capturing long-distance dependencies while maintaining efficient computation. This improves the model's ability to extract features from smaller, densely packed peanuts and reduces false detections. Second, a Triplet Attention Module is introduced to enable interaction between different dimensions, enhancing the model's ability to model dependencies between channel and spatial dimensions. Finally, the MPDIoU function simplifies distance metric calculation, allowing bounding box regression to focus more on scale matching optimization. The improved BTM-YOLOv8 (BTM stands for BiFormer + Triplet Attention Module + MPDIoU) achieves a precision of 98.40%, a recall of 96.20%, an F1 score of 97.29%, and an mAP50 of 99.00%, representing improvements of 3.9%, 2.4%, 1.2%, and 3.14% respectively compared to the original YOLOv8 network. It also demonstrated high accuracy in tests of conveyor-type peanut pod detection equipment.

[0057] The dataset construction module is specifically used for: A large number of historical peanut pod images were collected, and each of these historical peanut pod images underwent sample augmentation operations including rotation, brightness adjustment, Gaussian blur, contrast adjustment, and random translation. The peanut pod bounding boxes, seed counts, and peanut-level annotations were then performed on each of the historical peanut pod images after the sample augmentation operations using Labelimg. A dataset was constructed based on the annotated historical peanut pod images, and the dataset was randomly divided into a training set, a validation set, and a test set in a ratio of 7:1:2.

[0058] The number of seeds in a peanut pod is an important characteristic for measuring its fruit structure. Based on the number of seeds contained in the pod, peanut pods can be divided into single-seed pods, double-seed pods, triple-seed pods, and multi-seed pods. Single-seed pods are usually smaller in size, but the seeds inside are more plump. Double-seed pods and triple-seed pods contain two and three peanut seeds, respectively. Multi-seed pods contain four or more peanut seeds.

[0059] To enhance the diversity and scale of the dataset, improve the convergence speed and detection accuracy of the model during training, and increase the model's generalization ability, sample expansion is performed.

[0060] In the peanut pod detection model creation module, the peanut pod detection model is constructed based on a backbone network, a neck network, and a head; the backbone network, neck network, and head are connected sequentially. The neck network is based on a bidirectional feature pyramid (BiFPN-PAN), which dynamically aggregates multi-resolution features of the backbone network through weighted bidirectional cross-scale connections. The top-down path fuses deep semantic information to improve small target detection performance, while the bottom-up path enhances shallow detail features to optimize localization accuracy. The detection head employs a decoupled design, separating classification and regression tasks into independent branches. The classification branch uses binary cross-entropy loss (BCE Loss) for multi-label prediction, while the regression branch combines distributed focus loss (DFL) and complete intersection-union loss (CIoU Loss) to model bounding box coordinates using discrete probability distributions. Anchor-free mechanisms are used to directly predict target center point offset and scale factor, and a dynamic task-aligned label assignment strategy is employed for end-to-end optimization.

[0061] The backbone network includes a first convolutional module (Conv), a second convolutional module, a third convolutional module, a fourth convolutional module, a fifth convolutional module, a first BiFormer module, a second BiFormer module, a first cross-stage feature fusion module (C2f), a second cross-stage feature fusion module, a triple attention module (Triplet AttentionModule), and a fast spatial pyramid pooling module (SPPF). The first convolution module, the second convolution module, the first BiFormer module, the third convolution module, the second BiFormer module, the fourth convolution module, the first cross-stage feature fusion module, the fifth convolution module, the second cross-stage feature fusion module, the triple attention module, and the fast spatial pyramid pooling module are connected in sequence. The BiFormer module is a novel visual Transformer architecture. Its core lies in the Bi-Level Routing Attention (BRA) mechanism. This mechanism uses dynamic sparse attention to flexibly adjust attention weights according to the features of the input image, assigning different levels of attention to different positions or features. This allows it to capture long-distance dependencies while maintaining efficient computation, solving the problem of excessive computational complexity of global self-attention in traditional Transformers, reducing computational load, and improving the efficiency and performance of prediction tasks.

[0062] Specifically, the BiFormer module's two-layer routing attention mechanism first processes the input feature map... Divided into S×S non-overlapping regions, each region contains The eigenvectors are represented by the following formula: ; Where R represents a real number; H represents the height of the feature map; W represents the width of the feature map; C represents the number of channels in the feature map; S represents the number of regions to be divided on each edge; patchify() represents the patching operation; Indicates the size of each region block; Indicates integer division; Then, a region-level average pooling operation is performed on the query and key to obtain the region-level query. and regional key : ; ; Where mean() represents the average pooling operation; dim represents the dimension; Q represents the original query tensor obtained by linear projection; and K represents the original key tensor obtained by linear projection. Next, construct an inter-regional affinity map. : ; Where T represents transpose; Pruning the affinity graph between regions to retain the top-k connections of each region yields the routing index matrix. Then, based on the key tensors and value tensors collected from the routing index matrix: ; ; Where N represents the range of index values; k represents the number of top-k connections retained in each region; This represents the key tensor collected based on the routing index matrix; V represents the value tensor collected based on the routing index matrix; V represents the original value tensor obtained through linear projection; gather() represents collecting data at the corresponding positions from the original tensor based on the index matrix; Finally, attention operations are applied to the collected key-value tensors: ; ; Where dwconv() represents depthwise separable convolution, used to introduce local context information; A represents the attention weight matrix; softmax() represents the normalized exponential function; and output represents the output of the attention module. The BiFormer module's two-layer routing mechanism can effectively distinguish the boundary features of adhering peanut kernels and suppress the interference of background noise (such as soil particles or broken shells) through adaptive weights, thereby improving the model's ability to extract features from smaller, denser peanuts and reducing false detections. Furthermore, its sparse attention mechanism is also more adaptable to high-resolution inputs in real-time detection.

[0063] To extract cross-dimensional interactions of image information and enrich feature representations, a triple attention module is introduced. This module is lightweight, parameter-free, and allows for cross-dimensional interaction. Its core functionality lies in calculating attention weights by modeling the cross-dimensional interaction between channels and space through rotation, pooling, and convolution. The triple attention module consists of three parallel branches. The first two branches handle the cross-dimensional interaction between the channel dimension C and either the spatial dimension H or W. The third branch establishes spatial attention. The outputs of all three branches are summed using a simple averaging method, overcoming the shortcomings of traditional attention mechanisms (such as SE and CBAM) that neglect inter-dimensional coupling.

[0064] The triple attention module captures cross-dimensional information through three parallel branches to calculate attention weights, given an input feature map. The three branches of the triple attention module process it independently.

[0065] In the first branch, an interaction is established between the height (H) and channel (C) dimensions by rotating the input feature map X counterclockwise by 90° along the height dimension, resulting in... (Dimensions are H×W×C); after z-pooling, it is simplified to (2×H×C in dimension), a 1×H×C intermediate tensor is generated through k×k convolutional layers and batch normalization. Attention weights are then generated using a sigmoid activation function and applied to... And restore the original dimensions by rotating 90° clockwise.

[0066] In the second branch, the interaction between the width and channel dimensions is captured, and the input feature map is... By rotating 90° counterclockwise along the width dimension, we obtain... (Dimensions are H×C×W), after z-pooling, it is simplified to (2×H×W in dimension), through k×k convolutional layers and batch normalization, a 1×W×H intermediate tensor is generated. Attention weights are then generated using a sigmoid activation function and applied to... And restore the original dimensions by rotating 90° clockwise.

[0067] The third branch is used to capture the input feature map. To address the spatial dependency, a z-pooling operation is performed along the channel dimension to obtain... (2×H×W in dimension), after k×k convolutional layers and batch normalization, an intermediate tensor of 1×H×W is generated. Attention weights are generated by the sigmoid activation function and directly applied to the original input.

[0068] For the input feature map The process of obtaining fine attention from triple attention and applying tensor y is as follows: ; ; Where σ represents the sigmoid activation function; This represents a standard two-dimensional convolutional layer defined by the kernel size k in the three branches of triple attention; This represents the three cross-dimensional attention weights calculated in triple attention; This indicates a 90° clockwise rotation to maintain the input shape (C×H×W). This represents the tensor obtained after rotating the input feature map in different directions; Indicates to Reshape to perform cross-dimensional attention computation; the underline indicates the inverse operation.

[0069] The neck network includes a first upsampling module, a second upsampling module, a first connection module, a second connection module, a third connection module, a fourth connection module, a third cross-stage feature fusion module, a fourth cross-stage feature fusion module, a fifth cross-stage feature fusion module, a sixth cross-stage feature fusion module, a sixth convolution module, and a seventh convolution module. The fast spatial pyramid pooling module, the first upsampling module, the first connection module, the third cross-stage feature fusion module, the second upsampling module, the second connection module, the fourth cross-stage feature fusion module, the sixth convolution module, the third connection module, the fifth cross-stage feature fusion module, the seventh convolution module, the fourth connection module, and the sixth cross-stage feature fusion module are connected in sequence. The input of the first connection module is connected to the output of the first cross-stage feature fusion module; the input of the second connection module is connected to the output of the second BiFormer module; the input of the third connection module is connected to the output of the third cross-stage feature fusion module; and the input of the fourth connection module is connected to the output of the fast spatial pyramid pooling module. The detection head includes a first detection sub-head, a second detection sub-head, and a third detection sub-head. The input of the first detection sub-head is connected to the output of the fourth cross-stage feature fusion module and the output of the sixth convolution module. The input of the second detection sub-head is connected to the output of the fifth cross-stage feature fusion module. The input of the third detection sub-head is connected to the output of the sixth cross-stage feature fusion module. The first, second, and third detection sub-heads are used for predicting peanut pod bounding boxes, seed quantity, and peanut level, respectively.

[0070] The original YOLOv8 network uses the CIoU function to calculate the localization loss. Although CIoU considers the intersection area of ​​bounding boxes, the distance between center points, and the aspect ratio of the bounding boxes, it uses different measures of aspect ratio instead of the actual difference between width and confidence, which slows down the model's convergence speed. To address the regression problem of overlapping and non-overlapping bounding boxes, and to consider both center point distance and width and height deviations, the MPDIoU function is introduced to replace the original CIoU function.

[0071] The MPDIoU function minimizes the Euclidean distance between the top-left and bottom-right corners of the predicted bounding box and the ground truth bounding box, combining this with IoU to form a new similarity metric. This eliminates the need for complex bounding box calculations, simplifying the computation process and thus improving model convergence speed and regression accuracy. Its expression is: ; ; ; ; Where A represents the predicted bounding box; B represents the ground truth bounding box; This indicates the coordinates of the top-left corner of the prediction box; This indicates the coordinates of the bottom right corner of the prediction box; Represents the coordinates of the top-left corner of the true bounding box; The coordinates of the bottom right corner of the ground truth bounding box are given; w represents the width of the smallest bounding rectangle that contains both the predicted and ground truth bounding boxes; h represents the height of the smallest bounding rectangle that contains both the predicted and ground truth bounding boxes. This represents the square of the Euclidean distance between the top-left corner of the predicted bounding box and the ground truth bounding box; This represents the square of the Euclidean distance between the bottom right corner of the predicted bounding box and the ground truth bounding box.

[0072] The peanut pod detection model training module is specifically used for: The peanut pod detection model is trained using the training set. After each training cycle, the validation loss value of the loss function is calculated using the validation set to prevent overfitting. Training is terminated when the validation loss value converges. After training, the peanut pod detection model is evaluated using the test set that was not used in the training. Performance metrics including precision, recall, F1 score, and mean precision are calculated. Once the performance metrics meet the preset deployment requirements, the peanut pod detection model is deployed to the peanut pod detection device.

[0073] The peanut pod detection module is specifically used for: The peanut pod detection equipment uses an industrial camera (model optional HT-GE1000C-T-CL, 10-megapixel resolution, equipped with two stepless dimming LED light sources) to capture real-time peanut pod images of the conveyor belt. After preprocessing the real-time peanut pod images, they are input into the deployed peanut pod detection model for inference, resulting in peanut pod detection results that include the peanut pod bounding box, number of kernels, and peanut grade. The peanut pod detection results are displayed in real-time through a graphical user interface and are also stored.

[0074] In summary, the advantages of this invention are: 1. By collecting a large number of historical peanut pod images, sample augmentation and annotation are performed on each historical peanut pod image to construct a dataset, which is then divided into training, validation, and test sets. Next, a peanut pod detection model is created based on an improved YOLOv8 network to output peanut pod detection results, with the MPDIoU function set as the loss function. The peanut pod detection model is then trained using the training set and the loss function, and validated and tested sequentially using the validation and test sets. The successfully tested peanut pod detection model is deployed to a peanut pod detection device. The peanut pod detection device collects real-time peanut pod images and inputs them into the peanut pod detection model to obtain the peanut pod detection results, which are displayed through a graphical user interface. The system displays peanut pod detection results. At the data level, a sample augmentation strategy incorporating operations such as rotation and brightness adjustment effectively simulates complex changes in real-world working environments, providing a rich variety of training samples for the peanut pod detection model and fundamentally enhancing its generalization ability. At the model level, an improved YOLOv8 network is used to introduce the BiFormer module and a triple attention module, enabling it to more accurately capture multi-scale targets and fine-grained features in complex backgrounds. The MPDIoU loss function is combined to optimize bounding box localization accuracy, significantly improving detection accuracy. Furthermore, leveraging YOLOv8's lightweight architecture and multi-scale feature fusion design, and after rigorous training validation and industrial deployment, the system ultimately greatly improves the accuracy, efficiency, and generalization ability of peanut pod detection.

[0075] 2. By improving the YOLOv8 network architecture, a BiFormer module, a cross-stage feature fusion module, and a triple attention module are introduced. These components enhance the model's ability to extract features from peanut pods, especially in complex backgrounds (such as changes in lighting or occlusion), enabling more accurate bounding box localization. Simultaneously, the MPDIoU loss function is used to optimize the bounding box regression process, reducing localization errors and thus improving detection precision and recall. This design makes the model more reliable in detecting peanut pods in agricultural environments, reducing false positives and false negatives, and improving overall performance.

[0076] 3. During the dataset construction phase, various sample augmentation operations (such as rotation, brightness adjustment, Gaussian blur, etc.) were employed, and detailed annotations were used to mark the bounding boxes, seed quantity, and level of peanut pods. This effectively increased data diversity and prevented model overfitting. In addition, the dataset was divided into training, validation, and test sets proportionally, and the loss value was monitored using the validation set during training to ensure that the model was deployed after convergence. This systematic training process improved the model's generalization ability, enabling it to adapt to peanut detection needs in different scenarios and reducing the risk of overfitting.

[0077] 4. The trained model is deployed into a peanut pod inspection device. Real-time images are acquired via an industrial camera and preprocessed, enabling efficient real-time inference. The inspection head is designed with a multi-head structure, capable of simultaneously predicting bounding boxes, kernel count, and peanut grade. This provides comprehensive inspection information, supporting applications such as agricultural grading and yield assessment. The integrated graphical user interface visualizes the results, facilitating real-time monitoring and decision-making, enhancing the system's usability and user experience. It is particularly suitable for automated production lines or smart agriculture scenarios.

[0078] 5. The meticulous design of the backbone network, neck network, and detection head in the network architecture, such as the cross-stage feature fusion module and the fast spatial pyramid pooling module, optimizes feature transfer and scale adaptation, and reduces computational redundancy. The introduction of the BiFormer module enhances the attention mechanism and improves the detection sensitivity of small targets (such as peanut pods). The overall model is based on the lightweight characteristics of YOLOv8, ensuring inference speed and meeting real-time requirements. This structural innovation reduces computational resource consumption while maintaining high accuracy, which is beneficial for deployment on resource-constrained embedded devices.

[0079] 6. By improving the network architecture (such as introducing the BiFormer module, cross-stage feature fusion, and triple attention module) and the MPDIoU loss function, the detection accuracy and robustness are significantly improved, enabling it to effectively cope with light variations and occlusion problems in complex agricultural environments. Simultaneously, through diverse sample augmentation and a systematic dataset partitioning and training process, the model's generalization ability and training efficiency are optimized, preventing overfitting. Real-time detection and multi-task integration are also achieved, simultaneously outputting bounding boxes, seed count, and peanut-level information, and visualizing the results through a graphical user interface, enhancing practicality and user experience. Furthermore, its lightweight design and structured innovation ensure high computational efficiency, facilitating deployment on embedded devices, while a complete verification, testing, and deployment process guarantees reliability and scalability, providing an efficient and reliable solution for intelligent agriculture.

[0080] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A BTM-YOLOv8-based peanut pod detection method, characterized in that: Includes the following steps: Step S1: Collect a large number of historical peanut pod images, perform sample augmentation and annotation on each of the historical peanut pod images to construct a dataset, and divide the dataset into a training set, a validation set and a test set; Step S2: Create a peanut pod detection model based on the improved YOLOv8 network to output peanut pod detection results, and set the loss function of the peanut pod detection model as the MPDIoU function; Step S3: Train the peanut pod detection model using the training set and loss function, and verify and test the trained peanut pod detection model using the validation set and test set in sequence. Deploy the peanut pod detection model that passes the test to the peanut pod detection device. Step S4: The peanut pod detection device acquires real-time peanut pod images, inputs the real-time peanut pod images into the deployed peanut pod detection model to obtain peanut pod detection results, and displays the peanut pod detection results through a graphical user interface.

2. The BTM-YOLOv8-based peanut pod detection method according to claim 1, wherein: Step S1 specifically involves: A large number of historical peanut pod images were collected, and each of these historical peanut pod images underwent sample augmentation operations including rotation, brightness adjustment, Gaussian blur, contrast adjustment, and random translation. The peanut pod bounding boxes, seed counts, and peanut-level annotations were then performed on each of the historical peanut pod images after the sample augmentation operations using Labelimg. A dataset was constructed based on the annotated historical peanut pod images, and the dataset was randomly divided into a training set, a validation set, and a test set in a ratio of 7:1:

2.

3. The BTM-YOLOv8-based peanut pod detection method according to claim 1, wherein: In step S2, the peanut pod detection model is constructed based on a backbone network, a neck network, and a detection head; the backbone network, the neck network, and the detection head are connected in sequence. The backbone network includes a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a fifth convolutional module, a first BiFormer module, a second BiFormer module, a first cross-stage feature fusion module, a second cross-stage feature fusion module, a triple attention module, and a fast spatial pyramid pooling module. The first convolution module, the second convolution module, the first BiFormer module, the third convolution module, the second BiFormer module, the fourth convolution module, the first cross-stage feature fusion module, the fifth convolution module, the second cross-stage feature fusion module, the triple attention module, and the fast spatial pyramid pooling module are connected in sequence. The neck network includes a first upsampling module, a second upsampling module, a first connection module, a second connection module, a third connection module, a fourth connection module, a third cross-stage feature fusion module, a fourth cross-stage feature fusion module, a fifth cross-stage feature fusion module, a sixth cross-stage feature fusion module, a sixth convolution module, and a seventh convolution module. The fast spatial pyramid pooling module, the first upsampling module, the first connection module, the third cross-stage feature fusion module, the second upsampling module, the second connection module, the fourth cross-stage feature fusion module, the sixth convolution module, the third connection module, the fifth cross-stage feature fusion module, the seventh convolution module, the fourth connection module, and the sixth cross-stage feature fusion module are connected in sequence. The input of the first connection module is connected to the output of the first cross-stage feature fusion module; the input of the second connection module is connected to the output of the second BiFormer module; the input of the third connection module is connected to the output of the third cross-stage feature fusion module; and the input of the fourth connection module is connected to the output of the fast spatial pyramid pooling module. The detection head is provided with a first detection sub-head, a second detection sub-head, and a third detection sub-head; The input of the first detection sub-head is connected to the output of the fourth cross-stage feature fusion module and the output of the sixth convolution module; the input of the second detection sub-head is connected to the output of the fifth cross-stage feature fusion module; and the input of the third detection sub-head is connected to the output of the sixth cross-stage feature fusion module. The first detection subheader, the second detection subheader, and the third detection subheader are used for predicting the peanut pod bounding box, the number of kernels, and the peanut grade, respectively.

4. The BTM-YOLOv8-based peanut pod detection method according to claim 1, wherein: Step S3 specifically involves: The peanut pod detection model is trained using the training set, and after each training cycle, the validation loss value of the loss function is calculated using the validation set to prevent overfitting. Training is terminated when the validation loss value converges. After training, the peanut pod detection model is evaluated using the test set that was not used in the training. Performance metrics including precision, recall, F1 score, and mean precision are calculated. Once the performance metrics meet the preset deployment requirements, the peanut pod detection model is deployed to the peanut pod detection device.

5. The BTM-YOLOv8-based peanut pod detection method according to claim 1, wherein: Step S4 specifically involves: The peanut pod detection equipment uses an industrial camera to capture real-time images of peanut pods on the conveyor belt. After preprocessing the real-time peanut pod images, they are input into a deployed peanut pod detection model for inference, resulting in peanut pod detection results that include the peanut pod bounding box, the number of kernels, and the peanut grade. The peanut pod detection results are displayed in real-time through a graphical user interface and are also stored.

6. A peanut pod detection system based on BTM-YOLOv8, characterized in that: Includes the following modules: The dataset construction module is used to collect a large number of historical peanut pod images, perform sample expansion and annotation on each of the historical peanut pod images to construct a dataset, and divide the dataset into a training set, a validation set and a test set. The peanut pod detection model creation module is used to create a peanut pod detection model based on an improved YOLOv8 network to output peanut pod detection results. The loss function of the peanut pod detection model is set to the MPDIoU function. The peanut pod detection model training module is used to train the peanut pod detection model using the training set and loss function, and to verify and test the trained peanut pod detection model using the validation set and test set in sequence, and to deploy the peanut pod detection model that passes the test to the peanut pod detection device. The peanut pod detection module is used by the peanut pod detection equipment to acquire real-time peanut pod images, input the real-time peanut pod images into the deployed peanut pod detection model to obtain peanut pod detection results, and display the peanut pod detection results through a graphical user interface.

7. The peanut pod detection system based on BTM-YOLOv8 as described in claim 6, characterized in that: The dataset construction module is specifically used for: A large number of historical peanut pod images were collected, and each of these historical peanut pod images underwent sample augmentation operations including rotation, brightness adjustment, Gaussian blur, contrast adjustment, and random translation. The peanut pod bounding boxes, seed counts, and peanut-level annotations were then performed on each of the historical peanut pod images after the sample augmentation operations using Labelimg. A dataset was constructed based on the annotated historical peanut pod images, and the dataset was randomly divided into a training set, a validation set, and a test set in a ratio of 7:1:

2.

8. The peanut pod detection system based on BTM-YOLOv8 as described in claim 6, characterized in that: In the peanut pod detection model creation module, the peanut pod detection model is constructed based on a backbone network, a neck network, and a detection head; the backbone network, neck network, and detection head are connected in sequence. The backbone network includes a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a fifth convolutional module, a first BiFormer module, a second BiFormer module, a first cross-stage feature fusion module, a second cross-stage feature fusion module, a triple attention module, and a fast spatial pyramid pooling module. The first convolution module, the second convolution module, the first BiFormer module, the third convolution module, the second BiFormer module, the fourth convolution module, the first cross-stage feature fusion module, the fifth convolution module, the second cross-stage feature fusion module, the triple attention module, and the fast spatial pyramid pooling module are connected in sequence. The neck network includes a first upsampling module, a second upsampling module, a first connection module, a second connection module, a third connection module, a fourth connection module, a third cross-stage feature fusion module, a fourth cross-stage feature fusion module, a fifth cross-stage feature fusion module, a sixth cross-stage feature fusion module, a sixth convolution module, and a seventh convolution module. The fast spatial pyramid pooling module, the first upsampling module, the first connection module, the third cross-stage feature fusion module, the second upsampling module, the second connection module, the fourth cross-stage feature fusion module, the sixth convolution module, the third connection module, the fifth cross-stage feature fusion module, the seventh convolution module, the fourth connection module, and the sixth cross-stage feature fusion module are connected in sequence. The input of the first connection module is connected to the output of the first cross-stage feature fusion module; the input of the second connection module is connected to the output of the second BiFormer module; the input of the third connection module is connected to the output of the third cross-stage feature fusion module; and the input of the fourth connection module is connected to the output of the fast spatial pyramid pooling module. The detection head is provided with a first detection sub-head, a second detection sub-head, and a third detection sub-head; The input of the first detection sub-head is connected to the output of the fourth cross-stage feature fusion module and the output of the sixth convolution module; the input of the second detection sub-head is connected to the output of the fifth cross-stage feature fusion module; and the input of the third detection sub-head is connected to the output of the sixth cross-stage feature fusion module. The first detection subheader, the second detection subheader, and the third detection subheader are used for predicting the peanut pod bounding box, the number of kernels, and the peanut grade, respectively.

9. The peanut pod detection system based on BTM-YOLOv8 as described in claim 6, characterized in that: The peanut pod detection model training module is specifically used for: The peanut pod detection model is trained using the training set, and after each training cycle, the validation loss value of the loss function is calculated using the validation set to prevent overfitting. Training is terminated when the validation loss value converges. After training, the peanut pod detection model is evaluated using the test set that was not used in the training. Performance metrics including precision, recall, F1 score, and mean precision are calculated. Once the performance metrics meet the preset deployment requirements, the peanut pod detection model is deployed to the peanut pod detection device.

10. The peanut pod detection system based on BTM-YOLOv8 as described in claim 6, characterized in that: The peanut pod detection module is specifically used for: The peanut pod detection equipment uses an industrial camera to capture real-time images of peanut pods on the conveyor belt. After preprocessing the real-time peanut pod images, they are input into a deployed peanut pod detection model for inference, resulting in peanut pod detection results that include the peanut pod bounding box, the number of kernels, and the peanut grade. The peanut pod detection results are displayed in real-time through a graphical user interface and are also stored.