Method for detecting grain number per ear, row number per ear and grain number per row per ear of corn

By constructing a corn ear grain counting model based on Transformer encoder and quadtree segmenter, combining feature extraction and principal component analysis methods, the detection problems of corn ear grain number, ear row number and ear row number in an open environment are solved, and efficient and accurate corn ear grain number identification and ear row number detection are achieved.

CN120298713AActive Publication Date: 2025-07-11HUAZHONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510280766.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect corn ear grains, ear rows and ear rows in open environments, especially in field environments, and its robustness is insufficient, manual observation is time-consuming and labor-intensive and error-prone. The detection accuracy and robustness of existing methods in natural field environments need to be improved.

Method used

The corn ear grain counting model based on the Transformer encoder and quadtree segmenter was used, combining feature extraction, query point initialization, quadtree segmentation and prediction heads, and the corn ear grain count was counted through the open environment image taken by the mobile phone, and the number of ear rows and ear grains were detected using principal component analysis and vector-based point search method.

Benefits of technology

It realizes the accurate identification and positioning of corn ear grain numbers in an open environment, improves the detection accuracy of corn ear rows and ear grain numbers, solves the robustness problem in the field environment, and has efficient and accurate detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298713A_ABST
    Figure CN120298713A_ABST
Patent Text Reader

Abstract

The invention discloses a corn ear grain number, ear row number and ear row grain number detection method, and belongs to the field of image detection.According to the corn ear grain number counting method, multi-scale features of a corn ear image are encoded through a Transform encoder, the long-distance dependency relationship and global context information between the features are enhanced, and the detection accuracy is improved. A decoder outputs feature representation of each query point and provides rich semantic information for a point query and Transform decoding module, the decoder outputs feature representation of each query point, rich semantic information is provided for classification and positioning tasks of prediction heads, background interference can be effectively eliminated, and counting precision is improved; according to the method for detecting the ear row number and the ear row grain number, common problems in half-row corn ear row detection, nonlinear corn ear row detection and the like are solved by utilizing clustering and a vector-based point searching updating iteration method, and row detection and ear row number detection can be efficiently and accurately realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image detection, and more specifically, relates to a method for detecting the number of kernels per ear, the number of rows per ear, and the number of kernels per row of corn ears. Background Art

[0002] Corn is one of the world's three major food crops and plays an important role in animal husbandry and industry. Analyzing and studying various traits of corn can assist in establishing the relationship between its traits and yield to obtain greater benefits. Corn yield is mainly determined by the number of ears, the number of kernels per ear, and the grain weight. Within a certain range, the increase in the number of ears and the number of kernels per ear often directly enhances the yield potential. The corn phenotype is usually determined manually through random sampling and visual observation. This method is not only cumbersome and time-consuming but also error-prone. At the same time, manual observation may damage the growth environment of corn. Manual observation cannot meet the needs of modern high-throughput phenotype analysis, and the technology of computer vision image analysis can be used to extract the corn phenotype.

[0003] With the popularization of plant phenotype platforms, the research on automated plant phenotype extraction based on digital images has increased year by year. The extraction of many plant phenotypes is inseparable from deep convolutional neural networks (CNNs). Some research has proposed a method for estimating the number of rows per ear of corn by using the cross-section in corn images. However, this research uses a flatbed scanner with a resolution of 600 dpi and is only applicable to small-scale detection in a simple background, and its robustness for detection in the field environment is not good. Another research proposed a neural network that uses mean absolute error and mean square error to evaluate the counting performance. Its local counting regression method is superior to other baseline methods, which is beneficial to improving the detection efficiency. However, its robustness in the natural field environment still needs to be improved. In recent years, some research has proposed a method for measuring the similarity of data by using a method based on distance metric learning, and at the same time, learning background network parameters, embedding space, and the feature distribution of each training data in this space. This method uses the PyTorch deep learning framework for model training and adopts data augmentation techniques such as random scaling, cropping, rotation, color jitter, and horizontal flipping to enhance the robustness of corn row detection in an open environment. However, there are still some angular deviations, resulting in certain errors in the detection and counting of the number of ears and rows of corn ears. Summary of the Invention

[0004] Aiming at the above defects or improvement requirements of the prior art, the present invention provides a method for detecting the number of kernels per ear, the number of rows per ear, and the number of kernels per row of corn ears. This method can accurately identify the number of kernels per ear from corn ear images in an open environment taken by mobile terminals such as mobile phones, and can also locate and detect the number of rows per ear and the number of kernels per row according to the results of kernel counting.

[0005] To achieve the above object, according to the first aspect of the present invention, a method for counting the number of kernels on a corn ear is provided, which is characterized by including:

[0006] Training stage:

[0007] Construct a corn ear kernel number counting model and train it with a data set to obtain a trained corn ear kernel number counting model; wherein, the data set includes corn ear images collected in living scenes, field scenes, and solid color scenes respectively and their annotation data, and the annotation data includes the number of kernels on the corn ear and the position coordinates of each kernel; the corn ear kernel number counting model includes:

[0008] A feature extraction module for extracting multi-scale feature maps of corn ear images;

[0009] A Transformer encoder for encoding the multi-scale feature maps;

[0010] A query point initialization module for processing the corn ear image to evenly distribute the original query points;

[0011] A quadtree splitter for splitting the original query points in the corn kernel dense area of the corn ear image into four query points according to the encoded feature map, and retaining the original query points in the corn kernel sparse area of the corn ear image;

[0012] A Transformer decoder for calculating the attention of each target query point according to the encoded multi-scale feature map to obtain the feature representation of each target query point; wherein, the target query points include the split query points and the retained original query points;

[0013] A prediction head, including a classification branch and a regression branch; the classification branch is used to judge whether the target query point corresponds to a corn kernel according to the feature representation of the target query point, and the regression branch is used to predict its position offset when the target query point corresponds to a corn kernel to obtain the position coordinates of the corn kernel;

[0014] Application stage:

[0015] Input the corn ear image to be detected into the trained corn ear kernel number counting model to obtain the counting result and position coordinates of the number of kernels on the corn ear.

[0016] According to the second aspect of the present invention, a method for detecting the number of rows and kernels per row of a corn ear is provided, including:

[0017] S1. Based on the point coordinates of all corn kernels in the corn ear image to be detected, perform principal component analysis on all points to calculate the first principal axis direction, and rotate all points so that the first principal axis is parallel to the x-axis, and sort them in ascending order of the x coordinate to form a point set;

[0018] S2. Crop the point set in the x direction according to a preset cropping ratio, and divide the point set into a first point set and a second point set; wherein, the points in the first point set are the points in the middle part of the point set, and the points in the second point set are the points in other positions of the point set;

[0019] S3. Perform K-Means clustering on the first point set, and take the clustering number with the highest silhouette coefficient as the optimal clustering number, and perform clustering analysis on the first point set according to the optimal clustering number to generate clustering labels, and assign the label of the row where each point in the first point set is located;

[0020] S4. Use the vector-based point searching method to find the points in the second point set that are in the same row as each row of points in the first point set and assign the label of the row where they are located until all points in the second point set are traversed;

[0021] Wherein, the vector-based point searching method is: take the points with the smallest and largest x coordinates in each row of points in the first point set as the left and right boundary points of this row respectively;

[0022] When searching for points to the left, set the initial vectors horizontally to the left starting from the left boundary point of each row respectively. Take the points in the first point set that are on the left side of the second point set in descending order of their x coordinates as the points to be processed in turn and perform the following operations: Connect the left boundary point of each row with the point to be processed to obtain a new vector, and judge whether the included angle between the initial vector of each row and the new vector is not within the preset angle range. If so, regard the point to be processed as a noise point, otherwise assign the label corresponding to the row where the point with the smallest included angle is located to the point to be processed, and use the corresponding new vector as the initial vector of the points in this row for the next operation; When searching for points to the right, do the same.

[0023] S5. Output the detection results of the number of corn ear rows and the number of kernels per row.

[0024] According to the third aspect of the present invention, there is provided an electronic device, including: a computer-readable storage medium and a processor;

[0025] The computer-readable storage medium is used to store executable instructions;

[0026] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method described in the first aspect or the second invention.

[0027] According to the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to execute the method described in the first aspect or the second aspect.

[0028] According to the fifth aspect of the present invention, a computer program product is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the method described in the first aspect or the second aspect is implemented.

[0029] Generally speaking, compared with the prior art, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0030] 1. The method for counting the number of grains per ear of corn provided by the present invention uses the method of point query quad-tree to count the grains of corn ears in an open environment. In an open environment, taking pictures of corn ears with a handheld device can accurately obtain the data and visualization results of the number statistics of corn kernels, excluding the interference of the background, and having accuracy and robustness.

[0031] 2. The method for detecting the number of rows and grains per row of corn ears provided by the present invention uses clustering and the point-finding update iteration method based on vectors to solve common problems in the detection of corn ear rows such as half rows and non-straight lines, and can efficiently and accurately achieve row detection and the detection of the number of grains per row of ears. Brief Description of the Drawings

[0032] Figure 1 is the overall flowchart of the method for detecting the number of grains per ear, the number of rows of ears, and the number of grains per row of ears of corn provided by the embodiment of the present invention;

[0033] Figure 2 is the schematic flowchart of the method for counting the number of grains per ear of corn provided by the embodiment of the present invention;

[0034] Figure 3 is the original picture of a corn ear in the field provided by the embodiment of the present invention;

[0035] Figure 4 is the original picture of a corn ear in a life scenario provided by the embodiment of the present invention;

[0036] Figure 5 is the original picture of a corn ear obtained by a trait scanner provided by the embodiment of the present invention;

[0037] Figure 6 is the labeled picture of a corn ear in the field provided by the embodiment of the present invention;

[0038] Figure 7 is the labeled picture of a corn ear in a life scenario provided by the embodiment of the present invention;

[0039] Figure 8 is the labeled picture of a corn ear taken in a machine provided by the embodiment of the present invention;

[0040] Figure 9 It is a graph of the counted results of the number of corn ears in the field provided by an embodiment of the present invention;

[0041] Figure 10 It is a graph of the counted results of the number of corn ears in a living scenario provided by an embodiment of the present invention;

[0042] Figure 11 It is a graph of the counted results of the number of corn ears obtained by a trait scanner provided by an embodiment of the present invention;

[0043] Figure 12 It is a specific flowchart of ear row number detection and ear row grain number detection provided by an embodiment of the present invention;

[0044] Figure 13 It is a graph of the detected results of the number of ear rows of corn in the field provided by an embodiment of the present invention;

[0045] Figure 14 It is a graph of the detected results of the number of ear rows of corn in a living scenario provided by an embodiment of the present invention;

[0046] Figure 15 It is a graph of the results of the number of ear rows of corn obtained by a trait scanner provided by an embodiment of the present invention.

[0047] Figure 16 It is a schematic diagram of a trait scoring instrument device and a collection process provided by an embodiment of the present invention; Figure 16 In (a), it is a schematic diagram of a corn ear trait scoring instrument for collecting pure color background images, Figure 16 In (b) and (c), they are schematic diagrams of containers for putting corn ears into the trait scoring instrument from different perspectives, Figure 16 In (d), it is a schematic diagram of inspecting corn ears. Detailed implementation manners

[0048] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0049] Although there are many current detection technologies related to the number of corn ears, due to the limitations of various methods or strategies, it is difficult to apply them in the actual field environment. There are relatively large gaps in the counting problems of phenotypic traits such as the total number of grains per ear and the number of ear rows of corn, and the robustness in the open environment has not reached the ideal effect.

[0050] Based on this, an embodiment of the present invention provides a method for counting the number of grains per corn ear, including:

[0051] Training stage:

[0052] Construct a corn ear grain number counting model and train it using a dataset to obtain a trained corn ear grain number counting model; wherein, the dataset includes corn ear images and their annotation data collected in living scenarios, field scenarios, and solid color scenarios respectively, and the annotation data includes the number of corn ear grains and the position coordinates of each corn ear grain; as Figure 2 shown, the corn ear grain number counting model includes:

[0053] A feature extraction module for extracting multi-scale feature maps of corn ear images;

[0054] A Transformer encoder for encoding the multi-scale feature maps;

[0055] A query point initialization module for processing the corn ear image to evenly distribute the original query points;

[0056] A quadtree splitter for splitting the original query points in the corn kernel dense area of the corn ear image into four query points according to the encoded feature map, and retaining the original query points in the corn kernel sparse area of the corn ear image;

[0057] A Transformer decoder for calculating the attention of each target query point according to the encoded multi-scale feature map to obtain the feature representation of each target query point; wherein, the target query points include the split query points and the retained original query points;

[0058] A prediction head, including a classification branch and a regression branch; the classification branch is used to judge whether the target query point corresponds to a corn ear grain according to the feature representation of the target query point, and the regression branch is used to predict the position offset of the target query point in the case that the target query point corresponds to a corn ear grain, and add the position offset to the position coordinates of the target query point to obtain the position coordinates of the corresponding corn ear grain;

[0059] Application stage:

[0060] Input the corn ear image to be detected into the trained corn ear grain number counting model to obtain the counting result and position coordinates of the corn ear grains.

[0061] The structure and training process of the corn ear grain number counting model are introduced below respectively.

[0062] First, collect a dataset containing corn ear images and their annotation data taken by mobile phones in living scenarios, field scenarios, and taken by a seed grader in solid color scenarios. The annotation data includes the center point position and quantity of corn ear grains, as Figures 4 - 9As shown. This dataset is used to train a deep learning network for the corn ear grain number counting model, and the network parameters are optimized to minimize the difference between the prediction result and the true annotation. The loss functions used in the training process include classification loss and localization loss. The network weights are updated through backpropagation until the model converges to learn the feature representation and spatial distribution law of corn ear grains. The classification loss uses the cross-entropy loss function to determine whether the query point is a corn ear grain; the localization loss uses the Smooth L1 loss function to optimize the difference between the predicted position and the true position. Through training, the model can learn the feature representation and spatial distribution law of corn ear grains, providing support for subsequent counting and localization tasks.

[0063] The corn ear grain number counting model includes a feature extraction module, a Transformer encoder, a query point initialization module, a quadtree splitter, a Transformer decoder, and a prediction head.

[0064] 1. Feature extraction module

[0065] Before using the feature extraction module to extract features from the corn ear image, it also includes: initial preprocessing of the corn ear image, including operations such as cropping, rotation, and normalization, to increase data diversity and improve generalization ability.

[0066] The feature extraction module is used to extract multi-scale features of the image. Any existing feature extraction module can be used, and preferably a feature extraction module based on a CNN backbone network (such as the VGG16 convolutional neural network) is used:

[0067] The preprocessed image is input into the feature extraction module based on the CNN backbone network to generate feature maps at multiple resolutions. The feature extraction module uses the VGG16 convolutional neural network as the backbone network to extract feature maps with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 layer by layer, providing comprehensive information from low-level local details to high-level global context for the corn ear grain number counting model. Subsequently, these feature maps are input into the encoder.

[0068] As Figure 1 shown, in order to obtain a large enough receptive field, the encoder encodes the preprocessed image through five encoding stages (1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32) of the VGG16 network to generate multi-scale feature maps. These feature maps provide rich local details and global context information for subsequent point queries and decoders.

[0069] 2. Encoder

[0070] The encoder is based on the Transformer architecture and is used to further extract and integrate context information. The output of the encoder is a set of encoded feature representations, which not only retain the visual information extracted by the CNN backbone network, but also enhance the long-range dependencies and global context information between features, providing richer semantic information for subsequent point queries and Transformer decoding modules.

[0071] As a preferred solution, the encoder adopts a progressive rectangular window attention mechanism to extract and integrate context information. The encoder first calculates self-attention within a larger rectangular window to capture the global context information in the image. As the number of encoder layers increases, the size of the rectangular window gradually decreases, enabling the model to focus on smaller local regions and capture more detailed corn kernel features. The self-attention calculation formula of the encoder is:

[0072]

[0073] where x l-1 is the output feature of the (l - 1)-th layer of the encoder, is the feature after rectangular window self-attention processing, and x l is the final output feature of the l-th layer of the encoder. RectWin-SA represents the rectangular window self-attention mechanism, FFN is the feed-forward network, and LN is the layer normalization. Through this progressive attention calculation method, the encoder can gradually understand the image content from global to local, and output a set of encoded feature representation maps F, which not only retain the visual information extracted by the CNN backbone network, but also enhance the long-range dependencies and global context information between features, providing richer semantic information for subsequent point queries and Transformer decoding modules.

[0074] 3. Query Point Initialization Module and Quadtree Splitter

[0075] The query point initialization module evenly distributes sparse query points on the corn ear image, and the initial distribution interval of the query points is K pixels.

[0076] The quadtree splitter analyzes the context features output by the encoder to determine which regions in the corn ear image are dense corn kernel regions and which are sparse regions. The quadtree splitter outputs a split map to decide which query points need to be further split. In the dense region, the quadtree nodes are recursively split, splitting a query point into four new query points to improve the counting accuracy in the dense region. In the sparse region, the original query point set is retained.

[0077] That is, as Figure 2As shown in the figure, the query point initialization module initializes a set of sparse query points on the corn ear image, and then inputs the initialized corn ear image into the quadtree splitter, which consists of an average pooling layer, a 1×1 convolutional layer, and a sigmoid activation function. The average pooling layer is used to reduce the spatial dimension of the feature map, the 1×1 convolutional layer is responsible for extracting the regional density features, and the sigmoid function maps the output value to the interval [0,1] to obtain the split map M s .

[0078] M S = σ(Conv 1×1 (AvgPool(F)))

[0079] where σ is the sigmoid function, Conv 1×1 is the 1×1 convolutional operation, and AvgPool is the average pooling operation. The closer the value of each element in M s is to 1, the higher the density of the corn kernels in the corresponding area, and the query points need to be split; on the contrary, the closer the value is to 0, the lower the area density and no splitting is required. According to M s , the corn kernel number counting model performs quadtree splitting on the query points in the dense area, that is, a query point is subdivided into four new query points, thereby increasing the query point density in the local area to more precisely capture the distribution of the corn kernels. Finally, a set of query points adapted to different density areas is obtained.

[0080] 4. Transformer Decoder

[0081] The generated set of query points and the context features output by the encoder are jointly input into the decoder. Based on the Transformer architecture, the decoder preferably uses the attention mechanism within a rectangular window (Progressive Rectangle Window Attention) to process the local features of the query points. Sparse query points calculate attention within a larger rectangular window, while dense query points calculate attention within a smaller rectangular window, thereby improving the counting accuracy without increasing the computational complexity. The decoder outputs the feature representations of each query point, and these feature representations fuse the context information and local features, providing rich semantic information for subsequent classification and localization tasks.

[0082] During the decoding phase of the Transformer, the decoder receives an input composed of a dot query quadtree and encoded image features and outputs the decoded dot queries. That is, by calculating the attention within a local rectangular window to infer the relationships between dot queries, it can determine whether each query point represents a corn kernel under the guidance of the image context and locate its position. Sparse query points calculate self-attention within a larger rectangular window, while dense query points are calculated within a smaller rectangular window. The attention calculation of the decoder involves not only self-attention but also cross-attention with encoder features. The attention calculation formula of the decoder is as follows:

[0083]

[0084] where \(z\) l-1 is the output feature of the \((l - 1)\)-th layer of the decoder, is the feature after self-attention and cross-attention processing, and \(z\) l is the final output feature of the \(l\)-th layer of the decoder. RectWin-SA represents the rectangular window self-attention mechanism, RectWin-CA represents the rectangular window cross-attention mechanism, and \(x\) N is the final output feature of the encoder. Through this local rectangular window attention calculation, the decoder can output the feature representations of each query point. These feature representations fuse local and global context information and provide rich semantic information for subsequent classification and localization tasks.

[0085] 5. Prediction Head

[0086] The feature representations output by the decoder are input into the Prediction Head. The Prediction Head classifies and locates each query point through a multi-layer perceptron (MLP) structure. The Prediction Head contains two main functional branches: the classification branch is used to determine whether each query point corresponds to a corn kernel and outputs the classification probability; the regression branch is used to predict the position offset of the query point to accurately locate the actual position of the corn kernel. Specifically, the classification branch maps the decoded features to the classification probability through an MLP network and uses the Sigmoid activation function to limit the output value within the range of [0, 1], representing the probability that the query point is a corn kernel. The regression branch also predicts the position offsets \((\Delta x, \Delta y)\) of the query points through an MLP network and adds these offsets to the original positions of the query points to calculate the final position coordinates \((x, y)\) of the corn kernels. Finally, accurate corn kernel counting and localization information are generated as the output results of the model.

[0087] Specifically, the classification branch maps the decoded features to classification probabilities through a multi-layer perceptron (MLP), and uses the Sigmoid activation function to limit the output value within the range of [0, 1], representing the probability that the query point is a corn kernel. The calculation formula for the classification probability is:

[0088] c i =σ(W c ·MLP(z decode )+b c )

[0089] where c i is the classification probability of the i-th query point, σ is the Sigmoid activation function, W c and b c are the weight and bias parameters of the classification branch, and z decode is the feature representation output by the decoder.

[0090] Similarly, the regression branch predicts the position offsets of the query points through a multi-layer perceptron (MLP), and adds these offsets to the original positions of the query points to calculate the final position coordinates of the corn kernels. The prediction formula for the position offsets is:

[0091] Δp i =W p ·MLP(z decode )+b p

[0092] where Δp i =(Δx i ,Δy i ) is the position offset, W p and b p are the weight and bias parameters of the regression branch, and (x i ,y i ) is the initial position of the query point.

[0093] In the output stage of the corn kernel number counting model, each query point is assigned a classification probability c i and a normalized pixel position p i =(x i +Δx i ,y i +Δy i ). The model sets a probability threshold (such as 0.5), determines the query points with classification probabilities greater than the threshold as corn kernels, and marks the positions of the corn kernels in the image according to their position information. Then, the corn kernel number counting model aggregates all the query points determined to be corn kernels to obtain the total number of corn kernels in the image. Finally, a corn kernel coordinate position file is generated, and the counting result of the corn kernel number is visually output, such as Figures 10 - 11as shown

[0094] An embodiment of the present invention provides a method for detecting the number of rows and grains per row of a corn ear, as Figure 12 shown, including:

[0095] S1. According to the point coordinates (x, y) of all corn kernels in the corn ear image to be detected, perform principal component analysis on all points to calculate the first principal axis direction, and rotate all points so that the first principal axis is parallel to the x-axis, and sort them in ascending order according to the x coordinate to form a point set.

[0096] Specifically, according to the point coordinates of all corn kernels in the corn ear image to be detected, perform principal component analysis (PCA) on all points, calculate the first principal axis direction, rotate all points until the first principal axis is parallel to the x-axis direction, and sort all points in ascending order according to the x coordinate to form a point set; wherein, the rotation process uses the calculated principal axis angle and the point set mean value;

[0097] The above PCA and rotation processing include:

[0098] Calculate the mean value and principal axis direction of the point set;

[0099] Construct a rotation matrix, rotate the point set to the horizontal position, and record the rotation angle and mean coordinate for subsequent recovery operations.

[0100] S2. Crop the point set in the x direction according to a preset cropping ratio, and divide the point set into a first point set and a second point set; wherein, the points in the first point set are the points located in the middle part of the point set, and the points in the second point set are the points located in other positions of the point set.

[0101] Specifically, crop the rotated point set in the x direction according to the set cropping ratio, and only retain the points in the middle part; wherein, the preset cropping ratio is between 0 and 1 and can be set according to requirements.

[0102] For example, the preset cropping ratio is 0.12. Crop the rotated point set according to the cropping ratio, find the center point of the x-axis coordinates of the point set, and based on this center point, expand the width of the cropping area to both sides to determine the starting and ending x coordinates (x1 and x2) of the cropping area, screen out the points located within the cropping area, and use the point set formed by them as the first point set, and the point set formed by the remaining points as the second point set.

[0103] S3. Cluster the first point set, and use the clustering number with the highest silhouette coefficient as the optimal clustering number, and perform clustering analysis on the first point set according to the optimal clustering number to generate clustering labels, and assign the label of the row where each point in the first point set is located.

[0104] Specifically, for example, set the maximum number of clusters to 10 and the minimum number of clusters to 2. Any clustering algorithm can be used here, such as the K-Means clustering algorithm. Use the K-Means clustering algorithm to cluster the cropped point set into 2 to 10 clusters respectively, and select the number of clusters with the highest silhouette coefficient among all clustering results as the optimal number of clusters for the cropped point set; use this optimal number of clusters to perform clustering analysis on the standardized point set, generate clustering labels, and assign the label of the row where each point is located to each point.

[0105] S4. Use the vector-based point search method to find the points in the second point set that are in the same row as each row of points in the first point set and assign the label of the row where they are located until all points in the second point set are traversed;

[0106] Among them, the vector-based point search method is: take the points with the minimum and maximum x coordinates in each row of points in the first point set as the left and right boundary points of that row respectively;

[0107] When searching for points to the left, respectively set the initial vectors starting from the left boundary point of each row and horizontally to the left. Arrange the points in the first point set on the left side of the second point set in descending order of their x coordinates and use them as the points to be processed in turn and perform the following operations: Connect the left boundary point of each row with the point to be processed to obtain a new vector, and judge whether the angles between the initial vector of each row and the new vector are not within the preset angle range. If so, regard the point to be processed as a noise point; otherwise, assign the label corresponding to the row of the point with the smallest angle to the point to be processed, and use the corresponding new vector as the initial vector of the points in that row for the next operation. When searching for points to the right, do the same, that is, respectively set the initial vectors starting from the right boundary point of each row and horizontally to the right. Arrange the points in the first point set on the right side of the second point set in ascending order of their x coordinates and use them as the points to be processed in turn and perform the following operations: Connect the right boundary point of each row with the point to be processed to obtain a new vector, and judge whether the angles between the initial vector of each row and the new vector are not within the preset angle range. If so, regard the point to be processed as a noise point; otherwise, assign the label corresponding to the row of the point with the smallest angle to the point to be processed, and use the corresponding new vector as the initial vector of the points in that row for the next operation.

[0108] Specifically, the vector-based point searching method is as follows: For one of the clusters of the points completed in S3, set the two points with the largest and smallest x coordinates in this cluster as the right boundary point and the left boundary point. When searching for points to the left, set an initial vector horizontally to the left (hereinafter referred to as the original vector) for each left boundary point. For each point in the second point set and to the left of the first point set, perform the following operations in the order of decreasing x coordinates: Connect the left boundary point of each cluster to this point to obtain a new vector. Whether the angle between the new vector and the original vector is within the set angle error (for example, ±20°). If it is within the error range, select the cluster with the smallest included angle as the category of this point, update the left boundary point of this cluster to this point, and use the new vector obtained this time as the original vector for the next operation; if none of the clusters meet the angle requirement, this point is regarded as a noise point; when searching for points to the right, do the same.

[0109] S5. Output the detection results of the number of rows of corn kernels and the number of kernels per row.

[0110] The optimal number of clusters is the number of rows of corn kernels. After searching for points by the vector-based point searching method, the number of kernels per row and the corresponding number of kernels per row for each row are obtained.

[0111] It can be understood that the point coordinates of all corn kernels in the corn ear image to be detected can be obtained by any existing detection method. To improve the recognition accuracy, it is preferably obtained by the method for counting the number of corn kernels provided in the embodiments of the present invention.

[0112] Preferably, after traversing all the points in the point set, it further includes:

[0113] S6. Determine whether there are unlabeled points in the point set. If so, use the set composed of the unlabeled points as the updated point set, and return to S2; otherwise, enter S5.

[0114] Specifically, considering the property that only one side of the corn can be photographed, there may be a situation where some rows of corn are rotationally curved, so that when searching for points from the middle part to the two ends, there may be a situation where the rows on the back of the corn that cannot be photographed by the camera are rotated to the front of the corn, thus resulting in extra half rows; due to symmetry, there may be extra half rows at both ends of the corn, that is, two half rows; and these half rows cannot be recognized and classified by the vector method because they are not in the rows determined by the previous clustering, so it is necessary to perform additional operations of cutting, clustering, and point searching in S2 - S4 to cluster and recognize the two half rows and then perform point searching; finally, merge the classification results of the half rows with the previous classification results to obtain the overall classification result; if there are still unlabeled points, that is, regarded as error points (noise points), and finally output the detection results of the complete number of rows of corn kernels and the number of kernels per row.

[0115] That is, the remaining unlabeled points are grouped to form a new standardized motor, and then the operations of clipping, K-means clustering, and point searching are performed again to handle the half-row situation caused by rotation, label the remaining points, and merge them with the original classification results. The number of corn kernels in each category can be obtained by counting the points in each category. Finally, the points are rotated back to the original direction to generate the detection results of the number of kernel rows and the number of kernels per row, as Figures 13 - 15 shown.

[0116] Preferably, after step S2 and before step S3, it further includes: compressing the x coordinates of each point in the clipped point set by N times, and keeping the y coordinates unchanged;

[0117] After step S3 and before step S4, it further includes: stretching the abscissa of each point in the clustered point set by N times, and keeping the y coordinates unchanged to restore it to the original ratio.

[0118] Specifically, after step S2 and before step S3, a ratio stretching operation is performed on the clipped point set to eliminate the coordinate scale difference. Based on the fact that the row where the corn kernels are located approximately conforms to a straight line, and after rotating the row to the horizontal previously, the abscissa of the point set is scaled down, for example, by 200 times; the ordinate is kept unchanged, so that the points on the same row are compressed together as much as possible, eliminating the coordinate scale difference and significantly enhancing the subsequent clustering effect. After step S3 and before step S4, the abscissa of the clustered (clipped) point set is stretched by a ratio of, for example, 200 times to restore it to the original ratio.

[0119] An embodiment of the present invention provides an electronic device, including: a computer-readable storage medium and a processor;

[0120] The computer-readable storage medium is used to store executable instructions;

[0121] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the corn kernel number counting method or the corn kernel row number and kernel number per row detection method as described in any of the above embodiments.

[0122] An embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to execute the corn kernel number counting method or the corn kernel row number and kernel number per row detection method as described in any of the above embodiments.

[0123] An embodiment of the present invention provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, they implement the corn kernel number counting method or the corn kernel row number and kernel number per row detection method as described in any of the above embodiments..

[0124] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for counting the number of kernels on a corn ear, characterized in that, Including: Training stage: Construct a corn ear grain number counting model and train it using a dataset to obtain a trained corn ear grain number counting model; wherein, the dataset includes corn ear images collected in living scenarios, field scenarios, and solid color scenarios respectively and their annotation data, and the annotation data includes the number of corn ear grains and the position coordinates of each corn ear grain; the corn ear grain number counting model includes: A feature extraction module for extracting multi-scale feature maps of corn ear images; A Transformer encoder for encoding the multi-scale feature maps; A query point initialization module for processing the corn ear image to evenly distribute the original query points; A quadtree splitter for splitting the original query points in the corn kernel dense area of the corn ear image into four query points according to the encoded feature map, and retaining the original query points in the corn kernel sparse area of the corn ear image; A Transformer decoder for calculating the attention of each target query point based on the encoded multi-scale feature map to obtain the feature representation of each target query point; wherein, the target query points include the split query points and the retained original query points; A prediction head, including a classification branch and a regression branch; the classification branch is used to judge whether the target query point corresponds to a corn ear grain according to the feature representation of the target query point, and the regression branch is used to predict its position offset when the target query point corresponds to a corn ear grain to obtain the position coordinates of the corn ear grain; Application stage: Input the corn ear image to be detected into the trained corn ear grain number counting model to obtain the counting result and position coordinates of the corn ear grains.

2. The method according to claim 1, wherein The Transformer encoder encodes the multi-scale features based on the progressive rectangular window attention mechanism, and the calculation formula is: Among them, x l-1 is the output feature of the (l-1)-th layer of the encoder, is the feature after rectangular window self-attention processing, x l is the final output feature of the l-th layer of the encoder, RectWin-SA represents the rectangular window self-attention mechanism, FFN is the feed-forward network, and LN is the layer normalization.

3. The method according to claim 1 or 2, characterized in that The Transformer decoder calculates the attention of each target query point based on the attention mechanism within the rectangular window, and the calculation formula is: Among them, z l-1 is the output feature of the (l-1)-th layer of the decoder, is the feature after self-attention and cross-attention processing, and z l is the final output feature of the l-th layer of the decoder. RectWin-SA represents the rectangular window self-attention mechanism, RectWin-CA represents the rectangular window cross-attention mechanism, and x N is the final output feature of the encoder.

4. A method for detecting the number of kernel rows and kernels per row of a corn ear, characterized in that, Including: S1, According to the point coordinates of all corn kernels in the corn ear image to be detected, perform principal component analysis on all points to calculate the first main axis direction, and rotate all points so that the first main axis is parallel to the x-axis, and sort them in ascending order of the x coordinate to form a point set; S2, Crop the point set in the x direction according to a preset cropping ratio, and divide the point set into a first point set and a second point set; wherein, the points in the first point set are the points in the middle part of the point set, and the points in the second point set are the points in other positions of the point set; S3, Cluster the first point set, and take the clustering number with the highest silhouette coefficient as the optimal clustering number, and perform clustering analysis on the first point set according to the optimal clustering number to generate clustering labels, and assign the label of its corresponding row to each point in the first point set; S4, Use the vector-based point search method to find the points in the second point set that are in the same row as each row of points in the first point set and assign the label of its corresponding row to them until all points in the second point set are traversed; Among them, the vector-based point searching method is as follows: for each row of points in the first point set, the points with the minimum and maximum x coordinates are respectively used as the left and right boundary points of that row; When searching for points to the left, initial vectors horizontally to the left with the left boundary point of each row as the starting point are respectively set. The points in the first point set that are to the left of the second point set are sequentially used as the points to be processed in the order of decreasing x coordinates and the following operations are performed: the left boundary point of each row and the point to be processed are respectively connected to obtain a new vector, and it is judged whether the angles between the initial vector of each row and the new vector are not within the preset angle range. If so, the point to be processed is regarded as a noise point; otherwise, the point to be processed is assigned the label corresponding to the row of the points in the row with the smallest angle, and the corresponding new vector is used as the initial vector of the points in that row for the next operation. When searching for points to the right, it is the same; S5. Output the detection results of the number of rows of corn cobs and the number of grains per row.

5. The method according to claim 4, wherein After traversing all the points in the second point set, it further includes: S6. Judge whether there are unlabeled points in the second point set. If so, the set composed of the unlabeled points is used as the updated first point set, and return to S3; otherwise, enter S5.

6. The method according to claim 4, wherein The point coordinates of all the corn kernels in the corn cob image to be detected are obtained by the method described in any one of claims 1-3.

7. The method according to claim 4, characterized in that, After step S2 and before S3, it further includes: compressing the x coordinates of the points in the cropped point set by N times and keeping the y coordinates unchanged; After step S3 and before S4, it further includes: stretching the abscissas of the points in the clustered point set by N times and keeping the y coordinates unchanged.

8. An electronic device, characterized in that, It includes: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the corn kernel number counting method described in any one of claims 1-3 or the corn cob row number and number of grains per row detection method described in any one of claims 4-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to make the processor execute the corn kernel number counting method described in any one of claims 1-3 or the corn cob row number and number of grains per row detection method described in any one of claims 4-7.

10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are executed by the processor to execute the corn kernel number counting method described in any one of claims 1-3 or the corn cob row number and number of grains per row detection method described in any one of claims 4-7.

Citation Information

Patent Citations

  • Splicing method for corn ear order images

    CN102982524A

  • Computer vision technique-based corn ear species test method, system and device

    CN103190224A

  • Automatic detection method of corn tassel traits

    CN104573701A

  • Pleiotropic gene for control kernel-number-per-row and kernel-number-per-ear, and applications thereof

    CN104878018A

  • Method for hyperspectral detection of maturity of single corn seed

    CN113777104A

Cited By

  • Corn ear row extraction method

    CN121280736A

  • A method for extracting rows of corn ears

    CN121280736B