Oat grain counting method, system and equipment based on deep learning
Through the deep learning-based oat grain counting method, a model including a backbone network, a feature pyramid module, a rotating frame detection head and a counting module is built, which solves the problem of time-consuming and labor-intensive and low accuracy of traditional counting methods, and achieves efficient and accurate oat grain counting.
Patent Information
- Application Number
- CN202510265561.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The traditional oat grain counting method is time-consuming and labor-intensive, with low accuracy, and is difficult to meet the requirements of modern agriculture for efficiency and accuracy.
Using the deep learning-based oat grain counting method, a model including a backbone network, a feature pyramid module, a rotating frame detection head and a counting module is constructed by collecting and labeling the oat image data set, and training is carried out in combination with mass focus loss and rotating IoU loss.
It improves the accuracy and efficiency of oat grain counting, and can accurately detect and count oat ears in complex field environments, meeting the needs of modern agriculture for efficient and accurate.
Smart Images

Figure CN120198799A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural automation, and specifically relates to a method for counting oat grains based on deep learning, aiming to improve the efficiency and accuracy of oat yield assessment. Background Art
[0002] Oats (Avena sativa), as a nutritious food crop, have seen a rapid increase in demand in the global market in recent years due to their rich content of dietary fiber, protein, vitamins, and minerals. However, in the face of the growing food demand, there are many challenges in increasing the yield of oats. In the traditional agricultural management mode, the yield statistics and counting of oats mainly rely on manual labor. This method is not only time-consuming and laborious, vulnerable to human errors, but also requires professional knowledge, with high labor costs, and it is difficult to achieve large-scale and automated management. These factors limit the wide application of advanced technologies such as deep learning in agricultural scenarios and are difficult to meet the requirements of modern agriculture for efficiency and precision.
[0003] The accurate counting of oat grains is crucial in agricultural production because it can not only help evaluate the growth status of crops but also effectively predict the yield. However, traditional counting methods rely on detailed location information annotation, and existing oat grain counting methods have problems of being time-consuming and laborious and having low accuracy.
[0004] Therefore, how to improve the accuracy and efficiency of oat grain counting is an urgent problem to be solved in this field. Summary of the Invention
[0005] The purpose of the present invention is to address the deficiencies of the prior art by providing a method, system, and device for counting oat grains based on deep learning. The method for counting oat grains based on deep learning in the present invention collects an oat image dataset, annotates and performs data augmentation on the dataset; constructs a deep learning-based oat grain counting model; the model includes a backbone network, a feature pyramid module, a rotated bounding box detection head, and a counting module connected in sequence; trains the deep learning-based oat grain counting model according to the annotated and data-augmented oat image dataset to obtain a trained deep learning-based oat grain counting model; obtains an oat image to be counted, and based on the trained deep learning-based oat grain counting model, obtains the number of oat grains. The method for counting oat grains based on deep learning in the present invention performs multi-scale feature extraction and rotated bounding box detection, and combines quality focal loss and rotated IoU loss during the model training process, thereby solving the problems of low accuracy and efficiency in existing oat grain counting.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions:
[0007] The present invention provides a method for counting oat grains based on deep learning, which is characterized by including the following steps:
[0008] S1. Collect an oat image dataset, and perform annotation and data augmentation on the dataset;
[0009] S2. Construct a deep learning-based oat grain counting model; the model includes a backbone network, a feature pyramid module, a rotated box detection head, and a counting module connected in sequence;
[0010] Among them, the backbone network adopts the CSPNeXt network; the feature pyramid module adopts the CSPNeXtPAFPN network; the rotated box detection head RTMDetRotatedHead includes an angle encoder PseudoAngleCoder, a BatchNormalization strategy module RTMDetRotatedSepBNHeadModule, a dynamic label assigner BatchDynamicSoftLabelAssigner, a multi-level point generator MlvlPointGenerator, and a bounding box encoder DistanceAnglePointCoder; the counting module counts the number of bounding boxes according to the unique identifier of each image.
[0011] S3. Train the deep learning-based oat grain counting model according to the annotated and data-augmented oat image dataset to obtain a trained deep learning-based oat grain counting model; during the model training process, the Quality Focal Loss and the Rotated IoU Loss are combined.
[0012] S4. Obtain an oat image to be counted, and based on the trained deep learning-based oat grain counting model, obtain the number of oat grains.
[0013] Further, in step S1, the X-AnyLabeling-CPU tool is used for data annotation; the data augmentation operations include brightness and contrast adjustment, random rotation, and scaling transformation.
[0014] Further, in step S2, the CSPNeXt network gradually extracts features from the low layer to the high layer through a series of convolutional layers, residual modules, and pooling layers.
[0015] Further, in step S2, in the CSPNeXtPAFPN network, the feature maps of each scale are fused with the context features, and the features are aggregated through the top-down and bottom-up paths, and the feature pyramid module also performs multiple convolutional operations on the feature maps of each scale.
[0016] Further, in the rotated box detection head RTMDetRotatedHead, the angle encoder encodes the angle information of the rotated box. The BatchNormalization policy module assigns independent BN layers to rotated boxes of different angles or scales, enabling each BN layer to normalize specific rotated features. The output normalized feature maps are respectively input into the dynamic label allocator, the multi-level point generator, and the bounding box encoder. The dynamic label allocator dynamically assigns class labels and bounding box regression targets based on the rotated box and the ground truth annotation data. The multi-level point generator generates a set of predefined anchor points on the multi-scale feature maps, and these anchor points are converted by the bounding box encoder into the initial representation form of the rotated box, including the center point, width, height, and rotation angle.
[0017] Further, in step S3, according to the oat image dataset after annotation and data augmentation, the deep learning-based oat grain counting model is trained, specifically including:
[0018] S31. Input the oat image dataset after annotation and data augmentation into the backbone network to obtain multi-scale feature maps;
[0019] S32. Input the multi-scale feature maps into the feature pyramid module to aggregate features of different scales and obtain the fused multi-scale feature maps;
[0020] S33. Input the fused multi-scale feature maps into the rotated box detection head to obtain the output image;
[0021] S34. Input the output image into the counting module, and the counting module counts the number of bounding boxes according to the unique identifier of each image to obtain the number of oat grains.
[0022] Further, in step S3, the quality focus loss directly integrates the localization quality of the target into the classification loss. The rotated IoU loss is for rotated object detection. By calculating the overlapping area between the predicted box and the ground truth box, it optimizes the model's detection ability for rotated objects.
[0023] The present invention also proposes a deep learning-based oat grain counting system, which is characterized in that the oat grain counting system executes the deep learning-based oat grain counting method described above, including: an oat image acquisition module, an oat grain counting model construction module, an oat grain counting model training module, and an oat grain counting module;
[0024] The oat image acquisition module acquires the oat image dataset and performs annotation and data augmentation on the dataset;
[0025] An oat grain count model construction module constructs an oat grain count model based on deep learning; the model includes a backbone network, a feature pyramid module, a rotated bounding box detection head, and a counting module connected in sequence;
[0026] An oat grain count model training module trains the oat grain count model based on deep learning according to the labeled and data-augmented oat image dataset to obtain a trained oat grain count model based on deep learning;
[0027] An oat grain counting module obtains an oat image to be counted and gets the number of oat grains based on the trained oat grain count model based on deep learning.
[0028] The present invention also proposes a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.
[0029] Compared with the prior art, it has the following beneficial effects:
[0030] 1. The oat grain count method based on deep learning of the present invention collects an oat image dataset, labels and data-augments the dataset; constructs an oat grain count model based on deep learning, and the oat grain count model includes a backbone network, a feature pyramid module, a rotated bounding box detection head, and a counting module connected in sequence; trains the oat grain count model based on deep learning according to the labeled and data-augmented oat image dataset to obtain a trained oat grain count model based on deep learning; obtains an oat image to be counted and gets the number of oat grains based on the trained oat grain count model based on deep learning. The oat grain count method based on deep learning of the present invention performs multi-scale feature extraction and rotated bounding box detection, improving the accuracy and efficiency of oat grain counting.
[0031] 2. The oat grain count method based on deep learning of the present invention enhances the adaptability of the model in low-quality image scenarios through data augmentation operations, and also simulates the shooting environment under different lighting conditions, as well as the inclination state and scale change of the wheat ears in the natural distribution in the field, further improving the accuracy of subsequent oat grain counting.
[0032] 3. In the oat grain count method based on deep learning of the present invention, the backbone network in the oat grain count model adopts CSPNeXt, and through a series of convolutional layers, residual modules and pooling layers, features are gradually extracted from the low layer to the high layer, providing stable and rich image information for the subsequent object detection module and ensuring the high-performance performance of the model in the oat ear detection task.
[0033] 4. The oat grain counting method based on deep learning of the present invention. In the feature pyramid module of the oat grain counting model, CSPNeXtPAFPN is adopted. In CSPNeXtPAFPN, the feature maps of each scale are fused with the context features, thereby enhancing the robustness of the network in multi-scale object detection. This module aggregates features through top-down and bottom-up paths, enabling the model to maintain a high detection accuracy when processing oat spikes with significantly different sizes. The feature pyramid module also performs multiple convolutional operations on the feature maps of each scale to ensure the full fusion of features at different scales, thereby improving the performance of the model in complex field environments.
[0034] 5. The oat grain counting method based on deep learning of the present invention. The oat grain counting model includes a rotated bounding box detection head, which can more accurately capture tilted and overlapping targets by surrounding the oat spikes with rotated bounding boxes. In the rotated bounding box detection head, Pseudo Angle Coder and DistanceAnglePointCoder are used as rotated bounding box encoders, which can encode the angles and distances of the targets to help the model more precisely locate oat spikes at different angles. In addition, the rotated bounding box detection head adopts a separated BatchNormalization strategy, making the network more robust when processing rotated bounding boxes. By calculating the rotated IoU, the rotated bounding box detection head can better optimize the model output, ensuring that the detected bounding boxes highly coincide with the actual targets and improving the accuracy of oat grain counting.
[0035] 6. The oat grain counting method based on deep learning of the present invention. The oat grain counting model includes a counting module, which counts the number of bounding boxes according to the unique identifier of each image, converts the extracted features into the count of oat spikes, and realizes the efficient and accurate calculation of the number of oat spikes in the image without generating additional bounding boxes or density maps.
[0036] 7. The oat grain counting method based on deep learning of the present invention. During the training process of the oat grain counting model, QualityFocal Loss and RotatedIoU Loss are combined. QualityFocal Loss solves the inconsistency problem between the classification and localization tasks in traditional object detection by directly integrating the localization quality of the targets into the classification loss. RotatedIoU Loss is specifically designed for rotated object detection. By calculating the overlapping area between the predicted box and the ground truth box, it optimizes the detection ability of the model for rotated objects and further improves the accuracy of oat grain counting. Description of the Drawings
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0038] Figure 1 Schematic diagram of the oat grain counting method based on deep learning provided by the embodiment of the present invention.
[0039] Figure 2 Framework diagram of the oat grain counting method based on deep learning provided by the embodiment of the present invention
[0040] Figure 3 Schematic diagram of oat images under different light conditions provided by the embodiment of the present invention.
[0041] Figure 4 Schematic diagram of the oat image after data augmentation provided by the embodiment of the present invention.
[0042] Figure 5 Schematic diagram of the oat grain counting system based on deep learning provided by the embodiment of the present invention. Detailed implementation manners
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0044] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but it is not limited to the present invention.
[0045] The present invention proposes an oat grain counting method based on deep learning. As Figure 1 shown, the oat grain counting method based on deep learning includes the following steps S1 to S4. The framework diagram of this method is as Figure 2 shown.
[0046] S1. Collect an oat image dataset, and label and augment the dataset.
[0047] In terms of data collection, the present invention is based on an independently collected oat dataset and is committed to achieving high-precision detection and counting of oat ear numbers. The design of the dataset fully considers the complexity and diversity of the actual field scenarios, thus ensuring the high adaptability of the model in the real environment.
[0048] In the data collection stage, the present invention conducts data collection at the breeding base of the Qinghai-Tibet Plateau Research Institute to ensure the authenticity and representativeness of the data. The Nikon Z30 camera used provides high-quality image data with its excellent high resolution and imaging quality. To comprehensively capture the characteristics of oat ears in the natural growth environment, the present invention uses a single-lens reflex camera to shoot videos and extract images for data collection. Each video is 10 seconds long, and images are extracted at a frequency of 30 frames per second to construct a high-resolution image sequence. This method not only improves the efficiency of data collection but also ensures the diversity and representativeness of the data samples.
[0049] To make the oat ears more prominent and reduce background interference, a black background board is used during data collection. This background control strategy helps to more accurately identify and locate oat ears in the subsequent image processing stage. All image data are collected under similar environmental conditions, including lighting and climate conditions, to ensure data consistency. In addition, as Figure 3 shown, the present invention specifically collects data under different lighting conditions to simulate various complex scenarios that may occur in the field environment.
[0050] In terms of data annotation, in step S1, the X-AnyLabeling-CPU tool is used for data annotation.
[0051] Specifically, the present invention selects the X-AnyLabeling-CPU tool for data annotation work. This tool has a rotated bounding box annotation function, which is crucial for dealing with the situation of tilted oat ears. From the collected images, a total of 7,347 images are extracted for annotation, and finally 52,114 ear targets are annotated, with an average of about 7.09 ears per image. This high-quality annotation work not only improves the utilization value of the data but also provides accurate target position information for the model.
[0052] For data annotation, in step S1, the data augmentation operations include brightness and contrast adjustment, random rotation, and scaling transformation.
[0053] Specifically, to further improve the robustness and generalization ability of the model, this study performs various data augmentation operations on the dataset, such as Figure 4 shown, including brightness and contrast adjustment, random rotation, and scaling transformation. These operations not only enhance the model's adaptability in low-quality image scenarios but also simulate the shooting environment under different lighting conditions, as well as the tilted state and scale changes of ears in the natural distribution in the field.
[0054] S2. Construct an oat grain counting model based on deep learning; the model includes a backbone network, a feature pyramid module, a rotated bounding box detection head, and a counting module connected in sequence.
[0055] The present invention proposes a method for counting the number of oat ears based on the RTMDet framework (OatCountNet), aiming to achieve efficient and accurate automatic counting.
[0056] The model includes a backbone network, a feature pyramid module, a rotated box detection head, and a counting module connected in sequence.
[0057] Among them, the backbone network adopts the CSPNeXt network; the feature pyramid module adopts the CSPNeXtPAFPN network; the rotated box detection head RTMDetRotatedHead includes an angle encoder PseudoAngleCoder, a BatchNormalization strategy module RTMDetRotatedSepBNHeadModule, a dynamic label assigner BatchDynamicSoftLabelAssigner, a multi-level point generator MlvlPointGenerator, and a bounding box encoder DistanceAnglePointCoder; the counting module counts the number of bounding boxes according to the unique identifier of each image;
[0058] Specifically, the backbone network adopts the CSPNeXt network; the feature pyramid module adopts the CSPNeXtPAFPN network; the rotated box detection head RTMDetRotatedHead is used to detect oat ears at different angles. The rotated box detection head surrounds the oat ears through a rotated bounding box. In the rotated box detection head, Pseudo Angle Coder and Distance Angle PointCoder are used as the rotated box encoders; the rotated box detection head also includes a dynamic label assigner, a multi-level point generator, and a bounding box encoder, and the rotated box detection head adopts a separate BatchNormalization strategy. By calculating the rotated IoU, the rotated box detection head can better optimize the model output; the counting module counts the number of bounding boxes according to the unique identifier of each image.
[0059] CSPNeXt backbone network
[0060] In step S2, the CSPNeXt network gradually extracts features from the low layer to the high layer through a series of convolutional layers, residual modules, and pooling layers.
[0061] Specifically, the present invention selects CSPNeXt as the core network architecture to efficiently extract the underlying features of images. CSPNeXt is an improved version based on Cross Stage Partial Network (CSPNet). It optimizes the calculation process by adopting a partially connected method, reduces redundant calculations, and significantly improves the efficiency of the model in feature extraction. The core advantage of CSPNeXt lies in its design of Cross Stage Partial Connection, which uses a part of the feature map for residual learning while directly passing the other part to the next stage. This not only improves the efficiency of feature extraction but also reduces the number of model parameters, thereby effectively reducing the computational cost while maintaining high accuracy.
[0062] In the present invention, the hyperparameters of the CSPNeXt backbone network are adjusted to better adapt to the characteristics of oat spike images. In a specific embodiment, CSPNeXt is pre-trained on the ImageNet dataset, endowing the model with powerful initial feature extraction capabilities. In Oat CountNet, CSPNeXt gradually extracts features from low-level to high-level through a series of convolutional layers, residual modules, and pooling layers, providing stable and rich image information for the subsequent object detection module and ensuring high-performance in the oat spike detection task.
[0063] Specifically, the oat image is input into the CSPNeXt backbone network, and a multi-scale feature map is output, and the multi-scale feature map is input into the feature pyramid module.
[0064] Feature Pyramid Module
[0065] In step S2, in the CSPNeXt PAFPN network, the feature map of each scale is fused with the context features, and features are aggregated through top-down and bottom-up paths, and the feature pyramid module also performs multiple convolutional operations on the feature map of each scale.
[0066] To better process oat spike features of different scales, the present invention adopts CSPNeXt PAFPN (Path Aggregation Feature Pyramid Network) as the feature pyramid module. The role of the feature pyramid module is to fuse feature maps from different scales, enabling the network to capture key information in the image at scales from coarse-grained to fine-grained. Specifically, after the multi-scale feature map output by the CSPNeXt backbone network is input into the feature pyramid module CSPNeXt PAFPN, a fused multi-scale feature map is output, enhancing the detection ability for targets of different scales and providing a richer feature representation for the subsequent detection head.
[0067] Specifically, in CSPNeXtPAFPN, the feature maps of each scale are fused with the context features, thereby enhancing the robustness of the network in multi-scale object detection. This module aggregates features through top-down and bottom-up paths, enabling the model to maintain a high detection accuracy when dealing with oat spikes with significantly different sizes. The feature pyramid module also performs multiple convolutional operations on the feature maps of each scale to ensure the full fusion of features at different scales, thus improving the performance of the model in complex field environments.
[0068] Rotated bounding box detection head
[0069] In the actual field environment, oat spikes often have situations such as tilting and rotation. To solve the problem of insufficient accuracy of traditional detection methods when detecting tilted objects, the present invention introduces a rotated bounding box detection head (RTMDetRotatedHead), which is specifically used to detect oat spikes at different angles.
[0070] The rotated bounding box detection head surrounds the oat spike by a rotated bounding box (RBBox), which can more accurately capture tilted and overlapping objects. In the rotated bounding box detection head, we use PseudoAngle Coder and DistanceAnglePointCoder as rotated bounding box encoders, which can encode the angle and distance of the object to help the model more accurately locate oat spikes at different angles. In addition, the rotated bounding box detection head adopts a separate BatchNormalization (BN) strategy, making the network more robust when dealing with rotated bounding boxes. By calculating the rotated IoU (Intersection over Union), the rotated bounding box detection head can better optimize the model output to ensure that the detected bounding box highly coincides with the actual object.
[0071] The rotated bounding box detection head RTMDetRotatedHead includes an angle encoder PseudoAngleCoder, a BatchNormalization strategy module RTMDetRotatedSepBNHeadModule, a dynamic label allocator BatchDynamicSoftLabelAssigner, a multi-level point generator MlvlPointGenerator, and a bounding box encoder DistanceAnglePointCoder.
[0072] In the rotation box detection head of Oat CountNet, the angle encoder encodes the angle information of the rotation box to make it more suitable for network learning. The BatchNormalization policy module assigns independent BN layers to rotation boxes of different angles or scales, enabling each BN layer to normalize specific rotation features. The output normalized feature maps are respectively input into the dynamic label assigner, multi-level point generator, and bounding box encoder. The dynamic label assigner dynamically assigns class labels and bounding box regression targets based on the rotation box and ground truth data, reducing the dependence on precise annotations and improving the model's robustness to complex scenarios. The multi-level point generator generates a set of predefined anchor points on multi-scale feature maps, and these anchor points are converted by the bounding box encoder into the initial representation form of the rotation box, including the center point, width, height, and rotation angle.
[0073] Finally, the detection head network outputs the predicted rotation box, class confidence, and angle information based on the fused multi-scale feature maps, completing the detection of the tilt characteristics of the oat spike. Among them, the predicted rotation box includes the center point coordinates, width, height, and rotation angle; the class confidence is the confidence that each predicted box belongs to the target class, and the angle information represents the predicted rotation angle, which is used to describe the tilt degree of the target.
[0074] Angle Encoder (PseudoAngleCoder): For rotated object detection, angle encoding is crucial. By effectively encoding the angle information through PseudoAngleCoder, the detection accuracy of rotated oat spikes is improved. In agricultural scenarios, oat spikes may present different angles due to factors such as wind direction and planting methods. PseudoAngleCoder enables RTMDet to accurately identify and locate rotated oat spikes, maintaining high accuracy even in complex and variable field environments.
[0075] The role of the angle encoder is to encode the rotation angle information so that the model can more effectively learn and predict the rotation angle of the target.
[0076] Encoded rotation angle: By mapping the rotation angle θ to a continuous numerical range, the angle encoder outputs the encoded angle information θ^. The specific formula is:
[0077] θ^ = sin(θ) + cos(θ)
[0078] where θ is the actual rotation angle of the target.
[0079] Batch Normalization Strategy Module RTMDetRotatedSepBNHeadModule: In the rotated bounding box detection head of OatCountNet, the separate Batch Normalization (BN) strategy is adopted to better adapt to the tilt angles and scale variations of objects in the rotated object detection task. Specifically, this strategy assigns independent BN layers to rotated bounding boxes with different angles or scales, enabling each BN layer to normalize specific rotated features. For example, rotated bounding boxes can be grouped according to angle ranges (such as 0° - 45°, 45° - 90°, etc.) or scale sizes (such as small objects, medium objects, large objects), and each group uses an independent BN layer for normalization. This design can more precisely process the features of rotated bounding boxes with different angles and scales and can also effectively handle the large variations in object angles and scales in complex scenes.
[0080] Output of the Batch Normalization Strategy Module: The normalized feature map. The feature map after Batch Normalization processing has more stable statistical characteristics. The specific output is:
[0081] Fnorm = γ(σ2 + ∈F - μ) + β;
[0082] Where, Fnorm: represents the normalized feature map after Batch Normalization processing. γ: represents the scale factor, which is a learnable parameter used to scale the normalized feature map. σ2: represents the variance of the feature map. ∈: represents a very small constant used to avoid division by zero and ensure numerical stability. F: represents the original feature map. μ: represents the mean of the feature map. β: represents the shift, which is a learnable parameter used to translate the normalized feature map.
[0083] The output of the BatchNormalization strategy module is the normalized feature map Fnorm. These normalized feature maps have more stable statistical characteristics, which helps to improve the robustness and detection accuracy of the model. These feature maps are respectively input into the dynamic label assigner, multi-level point generator, and bounding box encoder.
[0084] BatchDynamicSoftLabelAssigner: This module dynamically assigns labels to the model, reducing the dependence on precise annotation while improving the robustness of the model. In object detection tasks, precise annotation is often time-consuming and costly. With the BatchDynamicSoftLabelAssigner, RTMDet can handle training data more flexibly, reducing the dependence on precise annotation, thus reducing the training cost and improving the generalization ability of the model.
[0085] Specifically, input the normalized feature map Fnorm into the BatchDynamicSoftLabelAssigner, assign class labels and bounding box labels to each predicted bounding box, and output the assigned labels.
[0086] The multi-level point generator (MlvlPointGenerator) of RTMDet can adapt to oat spikes of different sizes, which is crucial for improving the detection ability of the model in complex scenarios. By extracting features at different scales, RTMDet can better capture the detailed information of oat spikes and maintain high accuracy even when oat spikes are dense or partially occluded.
[0087] Specifically, input the normalized feature map Fnorm into the multi-level point generator, and output multi-scale anchor points and anchor positions; among them, anchor points of different scales are used to detect oat spikes of different sizes, and the anchor positions represent the position information of each anchor point on the feature map.
[0088] The bounding box encoder is a key component in the Oat CountNet model architecture, mainly used for processing rotated bounding box detection tasks. Its core function is to encode the location information of the target (including the center point coordinates, width, height, and rotation angle) so that the model can learn and predict the location and shape of the target more efficiently. Specifically, the encoder converts the rotation angle of the target into a format suitable for model learning through the PseudoAngleCoder. By mapping the angle information to a continuous numerical range, the model can learn and predict the rotation angle more effectively. At the same time, the DistanceAnglePointCoder is used to encode the center point location of the target and the size information of the bounding box. It converts the information such as the center point coordinates, width, and height of the target into a format that the model can process, and combines the rotation angle information to ensure that the model can accurately locate and describe the rotated target. During the encoding process, the input is the real bounding box information of the target, including the center point coordinates (x, y), width w, height h, and rotation angle θ. The encoder encodes the rotation angle θ into a continuous numerical value through the PseudoAngleCoder, and at the same time encodes the center point coordinates, width, and height into a format suitable for model learning through the DistanceAnglePointCoder. The output encoded bounding box information is used for model training. The decoding process is to convert the encoded bounding box information predicted by the model back to the actual bounding box position and shape for subsequent detection and counting. The decoder decodes the predicted rotation angle back to the actual angle value through the PseudoAngleCoder, and at the same time decodes the predicted center point coordinates, width, and height back to the actual bounding box position and size. The output decoded bounding box information is used for subsequent detection and counting tasks.
[0089] The labels assigned by the dynamic label allocator are used to guide the training process of the rotated bounding box detection head, helping the model better learn how to distinguish between targets and backgrounds. The multi-scale anchors and anchor positions output by the multi-level point generator provide detection bases at different scales for the rotated bounding box detection head, enabling the model to detect targets of different sizes. The encoded bounding box information output by the bounding box encoder provides the location and shape information of the target for the rotated bounding box detection head, enabling the model to more accurately predict the bounding box of the target. Through the collaborative work of these modules, the rotated bounding box detection head can output high-precision prediction results, including the predicted rotated bounding box, class confidence, and angle information, thus achieving efficient detection and counting of oat spikes.
[0090] Counting Module (CM)
[0091] The Counting Module (CountModule, CM) is used to efficiently and accurately count the number of oat spikes in an image. This module directly converts the extracted features into the count of oat spikes without generating additional bounding boxes or density maps, simplifying the post-processing workflow.
[0092] In a specific embodiment, the counting module counts the number of bounding boxes (bboxes) based on the unique identifier (image id) of each image. This process involves reducing the dimensionality of the features through a fully connected layer and using an average pooling layer to summarize the prediction results to reduce the counting errors caused by natural variations between samples. This method ensures accurate counting even when there are overlaps or occlusions between oat spikes.
[0093] Oat yield, as one of the key traits for measuring crop productivity, is mainly determined by three factors: the number of spikes per unit area, the number of grains per spike, and the grain weight. Therefore, this invention takes into account the density, diversity of oat spikes, and variations in the natural environment. Feature extraction and object detection are carried out through a convolutional neural network (CNN), which not only solves the problems of high annotation cost and noise caused by the dense position information of oat spikes, but also enhances the robustness of the model through techniques such as dynamic label assignment. To further improve the accuracy and efficiency of oat grain counting, the Counting Module (Count Module, CM), as the core component of the model. CM counts the number of bounding boxes (bboxes) based on the unique identifier (imageid) of each image, directly converting the extracted features into the count of oat spikes without generating additional bounding boxes or density maps. CM reduces the dimensionality of the features through a fully connected layer and uses an average pooling layer to summarize the prediction results to reduce the counting errors caused by natural variations between samples. During the training process, CM combines the Quality Focal Loss and the Rotated IoU Loss, which not only improves the classification accuracy but also enhances the ability of rotated object detection.
[0094] S3. Use the labeled and data-augmented oat image dataset to train the deep learning-based oat grain counting model to obtain a trained deep learning-based oat grain counting model; during the model training process, combine the Quality Focal Loss and the Rotated IoU Loss.
[0095] Step S3. Train the deep learning-based oat grain counting model according to the labeled and data-augmented oat image dataset, specifically including:
[0096] S31. Input the labeled and data-augmented oat image dataset into the backbone network to obtain multi-scale feature maps;
[0097] S32. Input the multi-scale feature map into a feature pyramid module to aggregate features of different scales and obtain a fused multi-scale feature map.
[0098] S33. Input the fused multi-scale feature map into a rotated bounding box detection head to obtain an output image.
[0099] S34. Input the output image into a counting module. The counting module counts the number of bounding boxes based on the unique identifier of each image to obtain the number of oat grains.
[0100] In step S3, during the model training process, the Quality Focal Loss and the Rotated IoU Loss are combined. The Quality Focal Loss directly integrates the localization quality of the target into the classification loss. The Rotated IoU Loss is for rotated object detection. By calculating the overlapping area between the predicted box and the ground truth box, it optimizes the model's detection ability for rotated objects.
[0101] Specifically, to further improve the performance of the model, the counting module combines the Quality Focal Loss (QFL) and the Rotated IoU Loss during the training process. The Quality Focal Loss resolves the inconsistency problem between the classification and localization tasks in traditional object detection by directly integrating the localization quality of the target into the classification loss. The Rotated IoU Loss is specifically designed for rotated object detection. By calculating the overlapping area between the predicted box and the ground truth box, it optimizes the model's detection ability for rotated objects.
[0102] The Quality Focal Loss (QFL) is an improved classification loss function. By combining the classification confidence and the localization quality (IoU), it addresses the deficiencies of traditional classification losses in handling imbalanced classification problems. QFL not only focuses on the correctness of classification but also considers the overlapping degree between the predicted box and the ground truth box, integrating the localization quality information into the classification loss. Specifically, QFL measures the localization quality by calculating the IoU value between the predicted box and the ground truth box and adjusts the weight of the classification loss according to the IoU value. If the overlapping degree between the predicted box and the ground truth box is high (high IoU value), the classification loss will be smaller; conversely, if the IoU value is low, the classification loss will be larger, thus prompting the model to pay more attention to targets with inaccurate localization. This design enables the model to optimize both the classification task and the localization task, reduces the weight of easily classifiable samples, and improves the model's adaptability to complex scenarios.
[0103] The Rotational IoU Loss (RIoU Loss) is a loss function used for rotational object detection, which measures the overlap degree between the predicted rotated bounding box and the ground truth rotated bounding box. Different from the traditional IoU calculation, the RIoU Loss takes into account the center point, width, height, and rotation angle of the rotated bounding box, and measures their overlap degree by calculating the ratio of the intersection area to the union area of the two rotated bounding boxes. By directly optimizing the positioning accuracy of the rotated bounding box, the RIoU Loss enables the predicted rotated bounding box to more accurately match the ground truth rotated bounding box, thereby improving the model's detection ability for inclined objects. This loss function is particularly suitable for handling the inclination and scale changes of objects such as oat spikes, and can significantly enhance the robustness of the model in complex field environments.
[0104] Specifically, the training process adopts the AdamW optimization algorithm and combines a linear learning rate scheduler, with the initial learning rate set to 0.00025 to achieve fast convergence and avoid overfitting. During the training process, the model optimizes the model parameters through an adaptive learning rate and weight decay mechanism to improve the stability and accuracy of the model.
[0105] In this way, the counting module not only improves the classification accuracy but also enhances the ability of rotational object detection, thus efficiently realizing the counting of oat spikes. This design reduces the need for complex post-processing, improves the overall efficiency of the model, and provides a new solution for the oat grain counting task.
[0106] S4. Obtain the oat image to be counted, and based on the trained deep learning-based oat grain counting model, obtain the number of oat grains.
[0107] The oat grain counting method RTMDet framework of the present invention is a CNN-based rotational object detection method designed specifically for handling the inclined characteristics of oat spikes. By introducing rotated bounding boxes and Rotational IoU Loss, RTMDet effectively addresses the challenges of oat spike recognition and counting in complex field environments. Our theoretical basis is built on multi-scale feature extraction and rotated bounding box detection, which not only improves the generalization ability of the model but also enhances its robustness in complex scenarios.
[0108] In a specific embodiment, the dataset is randomly divided into a training set (80%), a validation set (10%), and a test set (10%) to comprehensively evaluate the performance of the model and improve its generalization ability. At the same time, in order to comprehensively evaluate the performance of the oat grain counting method based on the RTMDet framework, we adopt three main evaluation metrics: Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Coefficient of Determination (R 2 ). These metrics can effectively measure the accuracy and stability of the model in the counting task.
[0109] In terms of the optimization strategy, the RTMDet model selects the AdamW optimization algorithm. Through AdamW, we have successfully achieved an effective balance between the model training convergence speed and the final performance. At the same time, to ensure the accuracy and robustness of the model in the oat ear counting task, the hyperparameters are optimized; the initial learning rate is set to 0.00025, and a linear learning rate scheduler is adopted. In addition, to improve the model's ability in dense object detection (such as overlapping or closely arranged oat ears), we choose the Quality Focal Loss as the main loss function. This loss function improves the model's accuracy in complex scenarios by reducing the attention to easily classified samples and increasing the weight of difficult-to-classify samples. In addition, we also introduce the Rotated IoU Loss to improve the accuracy of rotated bounding box detection and help the model better handle tilted oat ears.
[0110] Training process: All experiments are implemented in the MMDetection and MMRotate frameworks, which provide flexible configurations and powerful toolkits to support complex object detection tasks. We conducted the training on an NVIDIA A40 GPU and monitored the loss and accuracy during the training process to ensure the stable convergence of the model.
[0111] To verify the effectiveness of the proposed model in the oat ear counting task, we compared the proposed method with common object detection models, including YOLOv5, YOLOv8, RCNN, and the RTMDet of the present invention, and compared them in terms of multiple metrics such as accuracy, computational complexity, and the number of parameters. Through these comparisons, RTMDet performs excellently in the oat ear counting task, not only having high detection accuracy but also showing good computational efficiency. The advantages of RTMDet in terms of overall performance, computational efficiency, and resource consumption make it more suitable for this specific application scenario of oat ear counting, providing new ideas for future intelligent management in the agricultural field.
[0112] To further evaluate the performance of the model, we used multiple metrics to analyze the linear regression fitting degree of the RTMDet framework on the custom dataset. The results prove that the model can effectively predict the number of targets without relying on external position information. In addition, the mean absolute error (MAE) and root mean square error (RMSE) are calculated, and the results show that the model has high accuracy and low error levels during the prediction process.
[0113] Figure 5 This is a system for counting the number of oat grains based on deep learning provided by an embodiment of the present invention. As Figure 5As shown in the figure, the deep learning-based oat grain counting system includes: an oat image acquisition module, an oat grain counting model construction module, an oat grain counting model training module, and an oat grain counting module;
[0114] The oat image acquisition module acquires an oat image dataset, and annotates and performs data augmentation on the dataset;
[0115] The oat grain counting model construction module constructs a deep learning-based oat grain counting model; the model includes a backbone network, a feature pyramid module, a rotated bounding box detection head, and a counting module connected in sequence;
[0116] The oat grain counting model training module trains the deep learning-based oat grain counting model according to the annotated and data-augmented oat image dataset to obtain a trained deep learning-based oat grain counting model;
[0117] The oat grain counting module obtains an oat image to be counted and obtains the number of oat grains based on the trained deep learning-based oat grain counting model.
[0118] The above deep learning-based oat grain counting system can be implemented in the form of a computer program, and this computer program can run on a computer device.
[0119] The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.
[0120] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when these program instructions are executed, the processor can be made to execute a deep learning-based oat grain counting method.
[0121] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0122] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be made to execute a deep learning-based oat grain counting method.
[0123] Among them, the processor is used to run the computer program stored in the memory, and this program implements the deep learning-based oat grain counting method described in Embodiment 1.
[0124] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0126] The present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where when the computer program is executed by a processor, the processor executes a method for counting oat grains based on deep learning described in Embodiment 1.
[0127] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., all of which are computer-readable storage media that can store program codes.
[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
Claims
1. A method for counting oat grains based on deep learning, characterized in that: Includes steps: S1. Collect oat image dataset, annotate and enhance the dataset; S2. Constructing an oat grain counting model based on deep learning, wherein the model includes a backbone network, a feature pyramid module, a rotating frame detection head, and a counting module connected in sequence; The backbone network uses the CSPNeXt network; the feature pyramid module uses the CSPNeXtPAFPN network; the rotation box detection head RTMDetRotatedHead includes the angle encoder PseudoAngleCoder, the BatchNormalization strategy module RTMDetRotatedSepBNHeadModule, the dynamic label assigner BatchDynamicSoftLabelAssigner, the multi-level point generator MlvlPointGenerator and the bounding box encoder DistanceAnglePointCoder; the counting module counts the number of bounding boxes according to the unique identifier of each image; S3. Using the annotated and data-enhanced oat image dataset, the oat grain counting model based on deep learning is trained to obtain a trained oat grain counting model based on deep learning; in the model training process, the quality focal loss Quality Focal Loss and the rotation IoU loss Rotated IoU Loss are combined; S4. Obtain an image of oats to be counted, and obtain the number of oat grains based on a trained deep learning-based oat grain counting model.
2. The method according to claim 1, characterized in that In step S1, the X-AnyLabeling-CPU tool is used to label data; data enhancement operations include brightness and contrast adjustment, random rotation, and scaling transformation.
3. The method according to claim 1, characterized in that In step S2, the CSPNeXt network gradually extracts features from low layers to high layers through a series of convolutional layers, residual modules and pooling layers.
4. The method according to claim 1, characterized in that: Step S2, in the CSPNeXtPAFPN network, the feature maps of each scale are fused with the context features, the features are aggregated through top-down and bottom-up paths, and the feature pyramid module also performs multiple convolution operations on the feature maps of each scale.
5. The method according to claim 1, characterized in that In the rotation box detection head RTMDetRotatedHead, the angle encoder encodes the angle information of the rotation box, and the BatchNormalization strategy module allocates independent BN layers to rotation boxes of different angles or scales, so that each BN layer can normalize specific rotation features. The output normalized feature map is input into the dynamic label assigner, multi-level point generator, and bounding box encoder respectively; The dynamic label assigner dynamically assigns category labels and bounding box regression targets according to the rotated boxes and true annotated data, and the multi-level point generator generates a set of predefined anchor points on the multi-scale feature map, which are converted by the bounding box encoder into the initial representation of the rotated box, including the center point, width, height and rotation angle.
6. The method according to claim 1, characterized in that Step S3: training the oat grain counting model based on deep learning according to the annotated and data-enhanced oat image dataset, including: S31, inputting the annotated and data-enhanced oat image dataset into the backbone network to obtain a multi-scale feature map; S32, inputting the multi-scale feature map into a feature pyramid module, aggregating features of different scales, and obtaining a fused multi-scale feature map; S33, inputting the fused multi-scale feature map into a rotating frame detection head to obtain an output image; S34, inputting the output image into a counting module, and the counting module counts the number of bounding boxes according to the unique identifier of each image to obtain the number of oat grains.
7. The method according to claim 1, characterized in that In step S3, the quality focus loss directly integrates the positioning quality of the target into the classification loss, and the rotation IoU loss optimizes the model's detection ability for rotated targets by calculating the overlapping area between the predicted box and the true box for rotated target detection.
8. An oat grain counting system based on deep learning, characterized in that: The oat grain counting system executes the oat grain counting method based on deep learning as claimed in claim 1, comprising: an oat image acquisition module, an oat grain counting model construction module, an oat grain counting model training module and an oat grain counting module; Oat image acquisition module, which collects oat image datasets and annotates and enhances the datasets; An oat grain counting model building module is used to build an oat grain counting model based on deep learning; the model includes a backbone network, a feature pyramid module, a rotating box detection head, and a counting module connected in sequence; An oat grain counting model training module is used to train the oat grain counting model based on deep learning according to the annotated and data-enhanced oat image dataset to obtain a trained oat grain counting model based on deep learning; The oat grain counting module obtains the oat image to be counted and obtains the number of oat grains based on the trained deep learning-based oat grain counting model.
9. A computer device, characterized in that: The device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.