A wheat ear counting method based on image stitching
Patent Information
- Application Number
- CN202511253586.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-09-03
AI Technical Summary
计数模型使用CSPDarknet53提取图像的局部特征,并在特征提取过程中添加具有自注意力机制的Pyramid Pooling Transformer提取全局上下文信息,通过融合局部特征和全局上下文信息,提高麦穗计数的准确率,能够有效的适用于实际的田间麦穗计数”,但是上述文件中的技术方法存在数据集覆盖面窄、缺乏针对性的技术问题
[0038] 1. This invention constructs a Wheat-Stitch dataset specifically for wheat ears, including real images, synthetic images, and video frames, and covers complex agricultural scenes with multiple lighting conditions, large parallax, and texture repetition. This enables model training to closely resemble the actual farmland environment and solves the problems of narrow coverage and lack of specificity in existing datasets.
Smart Images

Figure CN121095206B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of precision agriculture technology, specifically to a method for counting wheat ears based on image stitching. Background Technology
[0002] The rapid development of precision agriculture requires accurate crop counting methods. Among them, the number of wheat ears is directly related to the assessment of wheat growth status and yield prediction. Currently, traditional wheat ear counting techniques mostly use local images from a single perspective for detection. Due to problems such as occlusion, changes in lighting, and texture repetition, the accuracy of the counting results is limited. Image stitching technology can stitch images taken from multiple angles and positions into a complete view, alleviating the above problems. However, existing traditional image stitching methods do not perform well in complex agricultural scenarios and are prone to serious stitching distortion, misidentification of target objects, and loss of details.
[0003] With the advancement of deep learning technology, unsupervised deep image stitching technology has made significant progress in handling large parallax and complex scene stitching, but it still has shortcomings in dealing with complex textures and large lighting variations in agricultural images.
[0004] The existing methods for counting wheat ears based on image stitching have the following drawbacks:
[0005] 1. Patent document CN118115846B discloses a field wheat ear counting method based on local-global feature fusion. The document states, "This invention discloses a field wheat ear counting method based on local-global feature fusion. This method first collects RGB image data of wheat ears, standardizes the size, and uses point annotations to preprocess the wheat ear dataset to establish labels, constructing a training dataset; then, the training dataset is input into the model for training, learning the wheat ear features of the wheat images and counting them. The counting model uses CSPDarknet53 to extract local features of the image, and adds a Pyramid Pooling Transformer with a self-attention mechanism to extract global contextual information during the feature extraction process. By fusing local features and global contextual information, the accuracy of wheat ear counting is improved, and it can be effectively applied to actual field wheat ear counting." However, the technical method in the aforementioned document suffers from narrow dataset coverage and a lack of specificity. Summary of the Invention
[0006] The purpose of this invention is to provide a wheat ear counting method based on image stitching to solve the technical problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for counting wheat ears based on image stitching, comprising the following steps:
[0008] S1. Construct the Wheat-Stitch dataset: Collect real wheat images using drones, combine them with synthetic images and video frames, covering scenes with multiple lighting conditions, multiple viewpoints, and different overlap rates.
[0009] S2. Design the UDIS2-Wheat stitching model: Based on the UDIS2 framework, a ResNet50 Siamese network with SE attention mechanism is used as the feature extraction module. The feature extraction module includes a shallow feature extraction stage and a deep feature extraction stage.
[0010] S3. Model Training: Supervised training is conducted using a comprehensive loss function, which includes overlapping region loss, inter-mesh loss, and intra-mesh loss.
[0011] S4. Image stitching: Input multi-angle wheat images into the trained UDIS2-Wheat model to generate a seamless panoramic image;
[0012] S5. Wheat Ear Detection and Counting: Based on the panoramic image, the YOLO target detection model is used to identify and count the number of wheat ears.
[0013] Preferably, the UAV in S1 uses DJI PHANTOM 4 PRO V2.0, with a flight altitude of 1.8m, an image overlap rate of 50% to 80%, and a resolution of 4000×3000 pixels.
[0014] Preferably, the SE attention mechanism described in S2 is embedded in the Bottleneck module of ResNet50.
[0015] Preferably, in S3, the overlapping region loss calculation involves the pixel alignment error of the stitched overlapping region;
[0016] Inter-mesh loss constraint on global geometric smoothness in non-overlapping regions;
[0017] The loss within the mesh maintains the consistency of the network's internal structure.
[0018] Preferably, the YOLOv8n model is used in S5, and detection is achieved through transfer learning and hyperparameter optimization.
[0019] Preferably, the model in S2 is first pre-trained on the UDIS-D dataset and then fine-tuned on the Wheat-Stitch dataset.
[0020] Preferably, when constructing the Wheat-Stitch dataset in S1, the original image is divided into grids, and image patches of effective overlapping regions are extracted to generate training sample image pairs. The grid size is 512×512 pixels, and the effective overlapping region is defined as the region with an overlap rate of ≥50%. The image patches are cropped arbitrarily to enhance data diversity.
[0021] Preferably, in the shallow feature extraction stage described in S2, the first three convolutional layers of ResNet50 are used to capture edge and local texture features, and the last two residual blocks are used to extract global semantic, spatial distribution and morphological features. The shallow and deep features are fused through skip connections and weighted by the SE attention mechanism with a weight coefficient of 0.7-0.9.
[0022] Preferably, the loss of the overlapping region in S3 is calculated using the L1 norm:
[0023]
[0024] in: Let N be the pixel alignment error in the overlapping region, and N be the total number of pixels in the overlapping region. This is the set of pixel coordinates for the overlapping region of the image. Let be the two-dimensional coordinates of the pixel in the image, where It is a horizontal coordinate. It is a vertical coordinate. The predicted image after deformation mesh transformation in coordinates Pixel value at that location, For the reference image in coordinates Pixel value at;
[0025] The inter-grid loss is defined as the Frobenius norm difference between the transformation matrices of adjacent grids:
[0026]
[0027] in Let E be the geometric smoothing constraint loss between networks, and E be the set of edges between adjacent networks. For the index of adjacent grid pairs, , Let be the 3×2 affine transformation matrices for the i-th and j-th grids, respectively. Let Frobenius norm be the matrix.
[0028] The mesh internal loss uses Huber loss to constrain local deformation:
[0029]
[0030] in This is due to the loss of consistency in the internal structure of the network. Here, k represents the number of control points uniformly sampled within the grid, and k is the index of the control point. Here are the predicted coordinates of the k-th control point. Let K be the true coordinates of the k-th control point. Huber's loss function:
[0031]
[0032] Where a and b are both input values. For threshold parameters, Let be the Euclidean distance between two coordinate vectors.
[0033] Preferably, in S5, the detection parameters of the YOLOv8n model are dynamically adjusted based on the PSNR and SSIM values of the stitched panoramic image:
[0034] When PSNR > 25 and SSIM > 0.75, a confidence threshold of 0.6 is used;
[0035] When the PSNR is 20-25 and the SSIM is 0.62-0.75, a confidence threshold of 0.5 is used.
[0036] When PSNR < 20 or SSIM < 0.62, a confidence threshold of 0.4 is used, combined with a non-maximum suppression loU threshold of 0.45.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. This invention constructs a Wheat-Stitch dataset specifically for wheat ears, including real images, synthetic images, and video frames, and covers complex agricultural scenes with multiple lighting conditions, large parallax, and texture repetition. This enables model training to closely resemble the actual farmland environment and solves the problems of narrow coverage and lack of specificity in existing datasets.
[0039] 2. The improved UDIS2-Wheat model significantly enhances the accuracy and stability of image stitching, especially in scenes with repetitive textures and large parallax. By extracting shallow and deep features in stages and fusing them with skip connections, the model's ability to distinguish wheat textures is strengthened, thereby reducing the mismatch rate in scenes with repetitive textures. Furthermore, the mathematically formulated comprehensive loss function precisely constrains geometric consistency, thereby improving the PSNR and SSIM of the stitched image and solving the problem of deformation accumulation caused by the overly abstract nature of existing loss functions.
[0040] 3. This invention significantly improves the detection accuracy of wheat ears by optimizing the image stitching method, and significantly reduces the phenomenon of missed or false detection of wheat ears caused by occlusion and other problems. At the same time, by adjusting the dynamic parameters based on PSNR and SSIM, the detection robustness is significantly improved. It can reduce the precision fluctuation range and technical error rate in low-quality stitched images, thus ensuring the accuracy and reliability of wheat ear counting.
[0041] 4. This invention is applicable to the rapid and accurate counting of wheat ears in large-area farmland, and can provide important technical support for the efficient management of precision agriculture;
[0042] 5. This invention enhances the effectiveness of the dataset and the generalization ability of the model by dividing the image into grids and extracting image patches from specific overlapping areas, thereby reducing training redundancy and improving stitching efficiency. Attached Figure Description
[0043] Figure 1 This is a flowchart of the overall process of wheat ear counting in this invention, which involves collecting field images by drone, generating a panoramic image through a depth image stitching model, and then combining it with target detection to achieve automatic identification and counting of wheat ears.
[0044] Figure 2 This is a schematic diagram of the segmentation of real image pairs in the Wheat-Stitch dataset of this invention, showing the process of extracting valid image blocks according to grid numbers and generating image pairs C and D from two original images of wheat ears in the field with overlapping areas, A and B;
[0045] Figure 3 This is a schematic diagram of the overall structure of the improved depth image stitching network UDIS2-Wheat of the present invention;
[0046] Figure 4 This is a diagram of the two-stage ResNet50-Siamese feature extraction network of this invention;
[0047] Figure 5 This is a structural diagram of the Bottleneck1 module of the present invention;
[0048] Figure 6 This is a structural diagram of the Bottleneck2 module of the present invention;
[0049] Figure 7 This is a bar chart comparing the PSNR index models of the four attention mechanisms of this invention.
[0050] Figure 8 This is a bar chart comparing the SSIM index models of the four attention mechanisms of this invention;
[0051] Figure 9 This is a line graph comparing the PSNR and SSIM of UDIS2 and UDIS2-Wheat in this invention;
[0052] Figure 10 Precision-PSNR curves for the four detection models of this invention;
[0053] Figure 11 Precision-SSIM curves for the four detection models of this invention;
[0054] Figure 12 This is a radar chart of the four evaluation indicators of this invention. Detailed Implementation
[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] See Figure 1 - Figure 12 This invention provides an embodiment: taking a wheat field at the Agronomic Experiment Station of Shandong Agricultural University Panhe Campus in Tai'an City, Shandong Province as the application scenario, the implementation process includes five stages: data collection, dataset construction, model training, image stitching, and wheat ear detection.
[0057] 1. Data Acquisition and Dataset Construction
[0058] Data Acquisition Equipment: A DJI PHANTOM 4 PRO V2.0 drone was used to acquire images during the wheat grain-filling to maturity stage;
[0059] Acquisition parameters: The acquired image resolution is 4000×3000 pixels, the flight altitude is 1.8m, the image overlap rate ranges from 50% to 80%, and the shooting conditions cover multiple angles, different lighting conditions, and multiple time periods;
[0060] Dataset Construction: The collected wheat image dataset is called the Wheat-Stitch dataset.
[0061] The data consists of real images: directly collected images of wheat ears in the field;
[0062] Synthetic image: An image of wheat ears generated through image enhancement;
[0063] Video frames: Extracting keyframes from drone aerial video;
[0064] Data processing:
[0065] Please see Figure 2 The original images A and B are divided into grids, and image patches with effective overlapping regions are extracted to generate paired training sample image pairs C and D in the dataset for model input. The image patches are enhanced by random cropping with a crop size of 512×512 pixels to improve data diversity.
[0066] 2. Construction and Training of UDIS2-Wheat Stitching Model
[0067] Model building:
[0068] Please see Figure 3 , Figure 4 , Figure 5 and Figure 6 Feature extraction module: Improved design, adopting a ResNet50 Siamese network that integrates SE attention mechanism;
[0069] Two-stage feature extraction:
[0070] Shallow feature extraction: The first three convolutional layers of ResNet50 are used to capture edge and local texture features;
[0071] Deep feature extraction: The last two residual blocks are used to extract global semantic, spatial distribution and morphological features;
[0072] Feature fusion: Shallow and deep features are fused through skip connections and weighted by the SE attention mechanism with a weight coefficient set to 0.8;
[0073] Evaluation module:
[0074] Warp network: Generates initial transformation parameters based on feature matching results;
[0075] Composition network: Optimizes stitching boundaries to generate seamless panoramic images;
[0076] Training strategy:
[0077] Loss function design:
[0078] Overlapping region loss: Calculate the pixel alignment error of the overlapping region during stitching, using the following formula:
[0079]
[0080] Inter-mesh loss: Ensures global geometric smoothness in non-overlapping regions, the formula is:
[0081]
[0082] Mesh Intra-Mesh Loss: To constrain the local geometric consistency within a single mesh, Huber loss is used, with the following formula:
[0083]
[0084] Huber parameters =1.0.
[0085] 3. Image stitching performance verification:
[0086] Please see Figure 7 , Figure 8 and Figure 9 Test the model using the Wheat-Stitch dataset:
[0087] Comparison of attention mechanisms: The SE attention mechanism outperforms CBAM, ECA and CA in both PSNR and SSIM metrics;
[0088] Model performance comparison: UDIS2-Wheat improves PSNR by an average of 2.1 and SSIM by 0.08, showing a significant advantage, especially in low-quality images.
[0089] 4. The impact of splicing quality on testing:
[0090] Please see Figure 10 and Figure 11 The stitched images are graded according to quality:
[0091] High quality: PSNR > 25, SSIM > 0.75;
[0092] Medium quality: PSNR 20-25, SSIM 0.62-0.75;
[0093] Low quality: PSNR < 20, SSIM < 0.62;
[0094] Detection process: YOLOv5n, YOLOv7 tiny, YOLOv8n and YOLOv11n target detection models were used to detect wheat ears, and the change trend of precision was recorded and analyzed.
[0095] Detection results: When PSNR>20 and SSIM>0.62, the detection performance of each model tends to stabilize, and Precision stabilizes above 0.89.
[0096] 5. Comparison of testing performance before and after splicing:
[0097] Please see Figure 12 The YOLOv8n model was used for wheat ear target detection, and the detection results were quantitatively compared using the Precision, Recall, F1-score, and mAP indices.
[0098] Wheat detection from the original image:
[0099] Due to obstruction, the detection was missed. The performance indicators are as follows:
[0100] Precision was 97.29%, Recall was 91.86%, F1-score was 94.49%, and mAP was 91.23%.
[0101] Image stitching detection:
[0102] Recover wheat ears that were missed due to occlusion in the original image to reduce missed detections;
[0103] The performance metrics for stitched images are:
[0104] Precision was 97.84%, Recall was 93.37%, F1-score was 95.55%, and mAP was 92.08%.
[0105] Radar chart comparison: The stitched image is superior to the original image in all four indicators;
[0106] Furthermore, by constructing a Wheat-Stitch dataset specifically for wheat ears, the problems of narrow coverage and lack of specificity in existing datasets are effectively solved.
[0107] The improved UDIS2-Wheat model significantly enhances the accuracy and stability of image stitching, especially in scenes with repetitive textures and large parallax.
[0108] By optimizing the image stitching method, the detection accuracy of wheat ears has been greatly improved, and the phenomenon of missed or false detection of wheat ears caused by occlusion and other problems has been significantly reduced, ensuring the accuracy and reliability of wheat ear counting.
[0109] It is suitable for rapid and accurate counting of wheat ears in large areas of farmland, and can provide important technical support for the efficient management of precision agriculture.
[0110] By dividing the data into grids and extracting image patches from specific overlapping regions, the effectiveness of the dataset and the generalization ability of the model are enhanced, training redundancy is reduced, and thus the stitching efficiency is improved.
[0111] By extracting shallow and deep features in stages and fusing them with skip connections, the model's ability to distinguish wheat textures is enhanced, thereby reducing the mismatch rate in scenes with repeated textures.
[0112] By using a mathematically formulated comprehensive loss function, geometric consistency is precisely constrained, thereby improving the PSNR and SSIM of stitched images and solving the problem of deformation accumulation caused by the overly abstract nature of existing loss functions.
[0113] By dynamically adjusting parameters based on PSNR and SSIM, detection robustness is significantly improved, and the Precision fluctuation range and technical error rate are reduced in low-quality mosaic images.
[0114] The working principle involves constructing a dedicated Wheat-Stitch dataset for wheat ears, including real images, synthetic images, and video frames. This dataset covers complex agricultural scenes with multiple lighting conditions, large parallax, and repetitive textures, enabling model training to closely resemble real farmland environments. This addresses the limitations of existing datasets, which often lack broad coverage and specificity. The improved UDIS2-Wheat model significantly enhances the accuracy and stability of image stitching, particularly in scenes with repetitive textures and large parallax. The optimized image stitching method drastically improves the accuracy of wheat ear detection, significantly reducing false negatives or missed detections due to occlusion and other issues. This ensures the accuracy and reliability of wheat ear counting, making it suitable for rapid and accurate wheat ear counting in large-area farmlands. Precision agriculture provides crucial technical support for efficient management. By using grid partitioning and extracting image patches from specific overlapping areas, it enhances the effectiveness of the dataset and the generalization ability of the model, reduces training redundancy, and thus improves stitching efficiency. Through staged extraction of shallow and deep features and skip connection fusion, it strengthens the model's ability to distinguish wheat textures, thereby reducing the mismatch rate in texture repetition scenes. Through a mathematically formulated comprehensive loss function, it precisely constrains geometric consistency, thereby improving the PSNR and SSIM of the stitched image and solving the deformation accumulation problem caused by the overly abstract nature of existing loss functions. By dynamically adjusting parameters based on PSNR and SSIM, it significantly improves detection robustness and reduces the precision fluctuation range and technical error rate in low-quality stitched images.
[0115] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
Claims
1. A method for counting wheat ears based on image stitching, characterized in that: Includes the following steps: S1. Construct the Wheat-Stitch dataset: Collect real wheat images using drones, combine them with synthetic images and video frames to cover scenes with multiple lighting conditions, multiple viewpoints and different overlap rates, and divide the original images into grids to extract image blocks of effective overlapping areas to generate training samples. S2. Design of UDIS2-Wheat stitching model: Based on UDIS2, a ResNet50 Siamese network with SE attention mechanism is used as the feature extraction module. The feature extraction module includes a shallow feature extraction stage and a deep feature extraction stage. Shallow and deep features are fused through skip connections and weighted by the SE attention mechanism. S3. Model Training: Supervised training is conducted using a comprehensive loss function, which includes overlapping region loss, inter-mesh loss, and intra-mesh loss. S4. Image stitching: Input multi-angle wheat images into the trained UDIS2-Wheat model to generate a seamless panoramic image; S5. Wheat Ear Detection and Counting: Based on the panoramic image, the YOLO target detection model is used to identify and count the number of wheat ears.
2. The wheat ear counting method based on image stitching according to claim 1, characterized in that: The drone in the S1 uses DJI PHANTOM 4 PRO V2.0, flies at an altitude of 1.8m, has an image overlap rate of 50% to 80%, and a resolution of 4000×3000 pixels.
3. The wheat ear counting method based on image stitching according to claim 1, characterized in that: The SE attention mechanism described in S2 is embedded in the Bottleneck module of ResNet50.
4. The wheat ear counting method based on image stitching according to claim 1, characterized in that: In S3, the overlapping region loss calculation shows the pixel alignment error of the stitched overlapping region. Inter-mesh loss constraint on global geometric smoothness in non-overlapping regions; The loss within the mesh maintains the consistency of the network's internal structure.
5. The wheat ear counting method based on image stitching according to claim 1, characterized in that: S5 uses the YOLOv8n model, and detection is achieved through transfer learning and hyperparameter optimization.
6. The wheat ear counting method based on image stitching according to claim 1, characterized in that: In S2, the model is first pre-trained on the UDIS-D dataset and then fine-tuned on the Wheat-Stitch dataset.
7. The wheat ear counting method based on image stitching according to claim 1, characterized in that: When constructing the Wheat-Stitch dataset in S1, the original images are divided into grids, and image patches of effective overlapping regions are extracted to generate training sample image pairs. The grid size is 512×512 pixels, and the effective overlapping region is defined as the region with an overlap rate of ≥50%. The image patches are cropped arbitrarily to enhance data diversity.
8. The wheat ear counting method based on image stitching according to claim 1, characterized in that: In S2, the shallow feature extraction stage uses the first three convolutional layers of ResNet50 to capture edge and local texture features, while the deep feature extraction stage uses the last two residual blocks to extract global semantic, spatial distribution and morphological features. The shallow and deep features are fused through skip connections and weighted by the SE attention mechanism with a weight coefficient of 0.7-0.
9.
9. The wheat ear counting method based on image stitching according to claim 1, characterized in that: The loss in the overlapping region described in S3 is calculated using the L1 norm: ; in: Let N be the pixel alignment error in the overlapping region, and N be the total number of pixels in the overlapping region. This is the set of pixel coordinates for the overlapping region of the image. Let be the two-dimensional coordinates of the pixel in the image, where It is a horizontal coordinate. It is a vertical coordinate. The predicted image after deformation mesh transformation in coordinates Pixel value at that location, For the reference image in coordinates Pixel value at; The inter-grid loss is defined as the Frobenius norm difference between the transformation matrices of adjacent grids: ; in Let E be the geometric smoothing constraint loss between networks, and E be the set of edges between adjacent networks. For the index of adjacent grid pairs, , Let be the 3×2 affine transformation matrices for the i-th and j-th grids, respectively. Let Frobenius norm be the matrix. The mesh internal loss uses Huber loss to constrain local deformation: ; in This is due to the loss of consistency in the internal structure of the network. Here, k represents the number of control points uniformly sampled within the grid, and k is the index of the control point. Here are the predicted coordinates of the k-th control point. Let K be the true coordinates of the k-th control point. Huber's loss function: ; Where a and b are both input values. For threshold parameters, Let be the Euclidean distance between two coordinate vectors.
10. The wheat ear counting method based on image stitching according to claim 5, characterized in that: In S5, the detection parameters of the YOLOv8n model are dynamically adjusted based on the PSNR and SSIM values of the panoramic image: When PSNR > 25 and SSIM > 0.75, a confidence threshold of 0.6 is used; When the PSNR is 20-25 and the SSIM is 0.62-0.75, a confidence threshold of 0.5 is used. When PSNR < 20 or SSIM < 0.62, a confidence threshold of 0.4 is used, combined with a non-maximum suppression loU threshold of 0.45.
Citation Information
Patent Citations
A method for counting wheat ear density in the field based on local-global feature fusion
CN118115846B
Wheat ear counting detection method based on small target detection and improved YOLOv5
CN116563205A
Field wheat ear density counting method based on local-global feature fusion
CN118115846A