Method and device for determining ear number of gramineous plant, electronic equipment and storage medium
By automatically counting the ear count of grass plants in Grapeaceae using density recognition models, the problem of manual counting is solved, and the problem of time-consuming and labor-intensive and difficult to deal with large-scale farmland environmental changes is achieved, and efficient and accurate ear counting is achieved.
Patent Information
- Application Number
- CN202510269688.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The manual counting method of the seeds of the grass family in the prior art is time-consuming and labor-intensive, difficult to cope with rapid changes in large-scale farmland environments, and is prone to artificial errors, affecting the accuracy of yield prediction.
A density recognition model including feature extraction module, feature pyramid, spatial attention module and target prediction head is adopted. The deep learning model is used to train the recognition image and automatically count the number of ears of the grass family plants.
It improves the efficiency and accuracy of ear counting, reduces artificial errors, and can efficiently and accurately count the ear count of grass plants in complex agricultural environments.
Smart Images

Figure CN120198801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular, to a method, device, electronic device and storage medium for determining the number of ears of gramineous plants. Background Art
[0002] In modern agricultural management, accurate crop yield estimation is crucial for optimizing planting strategies, resource allocation, and pest control.
[0003] In the related art, ear counting can only rely on manual surveys, which is not only cumbersome and laborious, has limited sampling areas, is prone to errors, and takes too long, thus severely limiting the accuracy of yield prediction and causing excessive estimation errors. Summary of the Invention
[0004] The present invention provides a method, device, electronic device and storage medium for determining the number of ears of gramineous plants, so as to solve the problems that the manual counting method for ears in the related art is not only time-consuming and laborious, but also difficult to cope with the rapid changes in a large-scale farmland environment.
[0005] According to a method for determining the number of ears of gramineous plants of the present invention, it includes:
[0006] Obtain a to-be-recognized image of a gramineous plant whose number of ears is to be recognized, and determine a density recognition model corresponding to the to-be-recognized image; wherein, the to-be-recognized image is a top-view image of the gramineous plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; the density recognition model is trained by a deep learning model using a sample recognition image and an expected density image;
[0007] Input the to-be-recognized image into the density recognition model to obtain a target density map of the ears of the gramineous plant, and determine the number of ears of the gramineous plant according to the target density map.
[0008] According to another aspect of the present invention, there is provided a device for determining the number of ears of gramineous plants, including:
[0009] A density recognition model determination module, which obtains a to-be-recognized image of a gramineous plant whose number of ears is to be recognized, and determines a density recognition model corresponding to the to-be-recognized image; wherein, the to-be-recognized image is a top-view image of the gramineous plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; the density recognition model is trained by a deep learning model using a sample recognition image and an expected density image;
[0010] The ear number determination module is configured to input the image to be recognized into the density recognition model to obtain a target density map of the ears of gramineous plants, and determine the ear number of the gramineous plants according to the target density map.
[0011] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0012] At least one processor; and
[0013] A memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the ear number of gramineous plants according to any embodiment of the present invention.
[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the method for determining the ear number of gramineous plants according to any embodiment of the present invention when executed.
[0016] The technical solution of the embodiment of the present invention is to obtain an image to be recognized of a gramineous plant with an ear number to be recognized, and determine a density recognition model corresponding to the image to be recognized. Since the image to be recognized is a top view image of the gramineous plant, the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module. The density recognition model is trained by a deep learning model with a sample recognition image and an expected density image, and the ear number of the gramineous plant can be automatically counted through the density recognition model; finally, the image to be recognized is input into the density recognition model to obtain a target density map of the ears of the gramineous plant, and the ear number of the gramineous plant is determined according to the target density map, which can improve the ear counting efficiency and accuracy, solve the problem that the manual counting method of ears in the related art is not only time-consuming and laborious, but also difficult to cope with the rapid changes in a large-scale farmland environment, reduce human errors, and achieve efficient and accurate counting of the ears of gramineous plants in a complex agricultural environment.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 is a flowchart of a method for determining the number of spikes of a gramineous plant according to Embodiment 1 of the present invention;
[0020] Figure 2 is a flowchart of a method for determining the number of spikes of a gramineous plant according to Embodiment 2 of the present invention;
[0021] Figure 3 is a network structure diagram for spike number recognition applicable to the embodiments of the present invention;
[0022] Figure 4 is a method for obtaining the weight diagram corresponding to each hierarchical feature map applicable to the embodiments of the present invention;
[0023] Figure 5 is a schematic diagram of a feature map fusion method based on a spatial attention mechanism applicable to the embodiments of the present invention;
[0024] Figure 6 is a schematic structural diagram of a device for determining the number of spikes of a gramineous plant according to Embodiment 3 of the present invention;
[0025] Figure 7 is a schematic structural diagram of an electronic device for implementing the method for determining the number of spikes of a gramineous plant in the embodiments of the present invention. Detailed Embodiments
[0026] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".
[0029] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and do not limit the scope of these messages or information.
[0030] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to users and user authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0031] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operations of the technical solutions of this disclosure according to the prompt message.
[0032] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manners of this disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of this disclosure.
[0034] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0035] Embodiment 1
[0036] Figure 1 FIG. 1 is a flowchart of a method for determining the number of spikes of a gramineous plant provided by Embodiment 1 of the present invention. This embodiment is applicable to the situation of automatically counting the number of spikes of gramineous plants in large-area farmland. This method can be executed by a device for determining the number of spikes of gramineous plants, and the device for determining the number of spikes of gramineous plants can be implemented in the form of hardware and / or software. Optionally, it is implemented by an electronic device, and the electronic device can be a mobile terminal, a PC or a server, etc.
[0037] As Figure 1 shown, the method may specifically include:
[0038] S110. Obtain a to-be-recognized image of a gramineous plant whose spike number is to be recognized, and determine a density recognition model corresponding to the to-be-recognized image; wherein, the to-be-recognized image is a top-view image of the gramineous plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; the density recognition model is obtained by training a deep learning model with a sample recognition image and an expected density image.
[0039] Among them, the to-be-recognized image can be understood as an image for which the number of spikes of a gramineous plant needs to be determined. The to-be-recognized image can be a top-view image of a gramineous plant taken from above the gramineous plant, which can ensure that all spikes are within the field of view. The to-be-recognized image may include, but is not limited to, images of gramineous plants under at least one condition such as different densities, different angles, and different illuminations. The density recognition model can be understood as a trained deep learning model for the density of spikes of gramineous plants in the to-be-recognized image.
[0040] As Figure 3As shown, the density recognition model consists of multiple modules, each with a specific function. The density recognition model includes, but is not limited to, at least one of a feature extraction module, a feature pyramid, a spatial attention module, and a target prediction head, etc. The feature extraction module can be understood as a functional module for extracting features from the image to be recognized, such as shape, color, or texture, etc. The feature extraction module can be any one of networks such as ResNet50 or VGG16. The feature pyramid can be understood as a functional module that allows the density recognition model to extract features at different levels, thereby capturing the global information and local details in the image to be recognized, helping the density recognition model determine the image features of crops with different sizes and densities, and also helping the density recognition model determine the image features of images taken by the drone at different flight altitudes. The spatial attention module can be understood as helping the density recognition model focus on the area containing the spikes in the image to be recognized, and can ignore the background noise or other irrelevant elements in the spike area, which is beneficial to improving the detection efficiency. The target prediction head can be understood as a functional module that outputs the processed features through the feature extraction module, the feature pyramid, and the spatial attention module, etc. as the target density map of the spikes of gramineous plants. The sample recognition image can be understood as a top-view image containing gramineous plants and their spikes, which is used to train the density recognition model. The expected density image can be understood as an image corresponding to the sample recognition image and annotated with the spike distribution at each position.
[0041] Based on the above solution, optionally, obtaining the image to be recognized of the gramineous plant whose spike number is to be recognized may include: collecting a top-view image of the gramineous plant through a photographing device; or, in response to an image upload operation, obtaining a top-view image of the gramineous plant; or, pulling a top-view image of the gramineous plant from a preset image database, etc., and no specific limitation is made here.
[0042] S120: Input the image to be recognized into the density recognition model to obtain a target density map of the spikes of the gramineous plant, and determine the spike number of the gramineous plant according to the target density map.
[0043] Among them, the target density map can be understood as an image used to characterize the density distribution of the spikes of gramineous plants. The spikes will be marked in the target density map. Each marked point in the target density map corresponds to a position in the image to be recognized and indicates the possible existence of spikes.
[0044] Based on the above solution, optionally, the density recognition model is trained as follows: Obtain multiple sample recognition images and the expected density image corresponding to each sample recognition image, input the sample recognition image into a pre-established deep learning model to obtain a model output image; determine the model output loss according to the model output image, the expected density image, and a preset Bayesian loss function to obtain the density recognition model.
[0045] Among them, the model output image can be understood as the prediction image result output by the deep learning model according to the input sample recognition image, which is used to indicate the ear number distribution of gramineous plants in the sample recognition image, and determine the model output loss in combination with the expected density image and the preset Bayesian loss function.
[0046] By adopting this technical solution, by obtaining multiple sample recognition images and their corresponding expected density images, and training the deep learning model using the preset Bayesian loss function to generate the density recognition model, the density recognition model is more accurate when determining the ear number in a complex agricultural environment.
[0047] Based on the above solution, optionally, obtaining the expected density image corresponding to each sample recognition image includes: For each sample recognition image, label each ear in the sample recognition image to obtain the labeled points corresponding to each ear; for each labeled point, use a two-dimensional Gaussian function to generate the Gaussian heat map corresponding to the labeled point, and perform normalization processing on each Gaussian heat map to obtain the target heat map; accumulate the target heat maps of the labeled points corresponding to multiple ears to obtain the expected density image corresponding to the sample recognition image.
[0048] Among them, the marked points can be understood as the marked positions determined after positioning each ear in each sample recognition image. The marked positions are located in the central part of the ear and can be displayed in two-dimensional coordinates in the sample recognition image. The two-dimensional Gaussian function can be understood as representing the probability distribution of the marked points in the two-dimensional space, used to generate corresponding Gaussian heatmaps according to each marked point, which can simulate the natural variation of the ear distribution in the actual scene, rather than a simple binary presence / absence representation. The Gaussian heatmap can be understood as an image generated by the two-dimensional Gaussian function for the marked points, where each pixel value reflects the possibility or density of the presence of an ear near that position. The Gaussian heatmap can represent the distribution of ears more naturally while retaining the position information and density information. The normalization process can be understood as a process of adjusting the numerical range so that all Gaussian heatmaps have the same scale. The normalization process ensures the comparability between Gaussian heatmaps of different sizes or resolutions, and the total integral of the entire Gaussian heatmap is equal to the total number of ears, thus maintaining the consistency of density determination. The target heatmap can be understood as the Gaussian heatmap after normalization, used to represent the density distribution of ears at their respective positions.
[0049] Based on the above scheme, optionally, the Bayesian loss function can be determined according to the total number of marked points of the ears and the expected value of the marked point count. Among them, the expected value of the marked point count can be determined according to the posterior probability of the pixel point belonging to the marked point and the number of pixels predicted by the deep learning model. The posterior probability of the pixel point belonging to the marked point can be determined according to the likelihood function of the marked points of the ears and the likelihood function of the total number of marked points of the ears; the likelihood function of the marked points of the ears can be determined according to the width parameter of the Gaussian kernel, the coordinates of the marked points, and the pixel points.
[0050] Specifically, the Bayesian loss function is:
[0051]
[0052] Among them, N is the total number of marked points of the ears, E[c n is the expected value of the marked point count of the marked point n, E[c0] is the expected value of the marked point count of the marked point n, M is the total number of pixels of the expected density map, D est (x m ) is the true density heatmap corresponding to the pixel point x n |x m ) is the posterior probability that the pixel point x m belongs to the marked point y n , p(y0|x m ) is the posterior probability that the pixel point x m belongs to the background dummy variable y0, p(x m |y n) represents the likelihood function of the annotation points of the spikes, p(x m |y0) represents the likelihood function of the background dummy variable y0, z n is the coordinate of the annotation point y n . σ is the width parameter of the Gaussian kernel, and d is the distance between the background dummy variable and the nearest spike target point.
[0053] Among them, the annotation expected value can be understood as the count expected value calculated based on the probability estimation of each pixel belonging to a specific annotation point by the model, which reflects the degree to which the model expects there to be an annotation point at this position. The true density heatmap can be understood as a value reflecting the spike density at different positions in the sample recognition image. The width parameter of the Gaussian kernel can be understood as the standard deviation of the Gaussian function, which can be used to smooth the image and reduce noise. The larger σ is, the broader and flatter the Gaussian function is; the smaller σ is, the sharper the function is, and the more concentrated the represented hot spot area is. When choosing a Gaussian function to generate a Gaussian heatmap, if σ is too small, the Gaussian heatmap will be too sparse and lose details; if σ is too large, the Gaussian heatmap may be too blurred to clearly distinguish the hot spot area where the spikes exist. In this embodiment, the optimal width parameter of the Gaussian kernel can be automatically found through cross-validation to ensure that the heatmap has sufficient resolution and is not too sparse. The background dummy variable can be understood as the contribution of the area in the sample recognition image that does not belong to the spikes. In this embodiment, by setting the background dummy variable with an expected value of 0, the influence of background pixels on the "spike" eigenvalue can be eliminated, ensuring that the actual characteristics of the spikes can be accurately evaluated without being interfered by background noise. Adopting this technical solution, by annotating each spike in each sample recognition image, generating corresponding annotation points, creating Gaussian heatmaps for each annotation point using a two-dimensional Gaussian function, and then normalizing these heatmaps to obtain the target heatmap to ensure comparability between images of different scales and resolutions, and accumulating all the target heatmaps to obtain the expected density image, it can not only capture the position information of individual spikes but also reflect the distribution density in the sample recognition image, improving the accuracy and reliability of spike counting in gramineous plants.
[0054] Based on the above solution, optionally, after obtaining multiple sample recognition images, it may further include: dividing the sample recognition images into a training set, a validation set, and a test set according to a preset ratio.
[0055] An optional implementation method is to divide the sample recognition images into a training set, a validation set, and a test set according to a ratio of 8:1:1, and use existing annotation software to annotate each spike in the sample recognition image, and the annotation position is in the central part of the spike.
[0056] Based on the above solution, optionally, after obtaining multiple sample recognition images, it may further include: padding the sample recognition images into squares according to the sizes of the sample recognition images, and scaling the sample recognition images to a preset size.
[0057] An optional implementation manner is to pad the sample recognition images into squares according to the sizes of the sample recognition images, and fill the missing parts with black pixels. At the same time, scale the sample recognition images to a preset size (e.g., 1024×1024), and use a variety of data augmentation techniques to expand the richness of the training data. Randomly rotate the sample recognition images (from 0 degrees to 360 degrees) to simulate different angles during drone shooting. Randomly scale the sample recognition images (from 0.8 times to 1.2 times) to simulate image capture situations at different flight altitudes. In addition, to cope with changes in lighting conditions, randomly adjust at least one of the image parameters such as brightness, contrast, and saturation of the sample recognition images.
[0058] An optional implementation manner is to use Adam as the optimizer, set the initial learning rate to 1e-3, the minimum learning rate to 1e-6, and train for 200 epochs. In the first 20 epochs of training, increase the initial learning rate to the initial value in a linear growth manner to warm up the density recognition model. After the warm-up stage, apply the cosine annealing learning rate scheduling strategy to dynamically adjust the learning rate. Iteratively train the density recognition model on the training set, and evaluate the performance on the validation set in each epoch. When the validation set loss value reaches the historical minimum, save the model weight file of this epoch. After the density recognition model is trained, evaluate the density recognition model on the test set.
[0059] The technical solution of the embodiment of the present invention obtains a to-be-recognized image of a gramineous plant with the number of ears to be recognized, and determines a density recognition model corresponding to the to-be-recognized image. Since the to-be-recognized image is a top-view image of the gramineous plant, the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module. The density recognition model is trained by a deep learning model using sample recognition images and expected density images. The number of ears of the gramineous plant can be automatically counted through the density recognition model. Finally, input the to-be-recognized image into the density recognition model to obtain a target density map of the ears of the gramineous plant, and determine the number of ears of the gramineous plant according to the target density map, which can improve the counting efficiency and accuracy of the ears, solve the problem that in the related art, manual counting of ears is not only time-consuming and laborious, but also difficult to cope with the rapid changes in a large-scale farmland environment, reduce human errors, and achieve efficient and accurate counting of the ears of gramineous plants in a complex agricultural environment.
[0060] Embodiment 2
[0061] Figure 2 This is a flowchart of a method for determining the number of spikes of a gramineous plant provided in the second embodiment of the present invention. Based on the above embodiment, this embodiment further refines how to input the image to be recognized into the density recognition model to obtain the target density map of the spikes of the gramineous plant. Optionally, inputting the image to be recognized into the density recognition model to obtain the target density map of the spikes of the gramineous plant includes: inputting the image to be recognized into the feature extraction module of the density recognition model to obtain a first feature image; inputting the first feature image into the feature pyramid of the density recognition model to obtain second feature images at multiple levels; inputting the second feature images at multiple levels into the spatial attention module of the density recognition model to obtain a spatial feature image; inputting the attention feature map into the target prediction head of the density recognition model to obtain the target density map of the spikes of the gramineous plant. For the specific implementation, reference can be made to the description of this embodiment. Among them, the same or similar technical features as those in the foregoing embodiments will not be elaborated herein.
[0062] As Figure 2 shown, the method may specifically include:
[0063] S210. Obtain an image to be recognized of a gramineous plant for which the number of spikes is to be recognized, and determine a density recognition model corresponding to the image to be recognized; wherein, the image to be recognized is a top-view image of the gramineous plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; the density recognition model is trained by a deep learning model using a sample recognition image and an expected density image.
[0064] S220. Input the image to be recognized into the feature extraction module of the density recognition model to obtain a first feature image.
[0065] Among them, the first feature image can be understood as the feature image obtained after the image to be recognized is processed by the feature extraction module, and still retains the spatial layout information of the image to be recognized. The first feature image may include low-level to mid-level visual features extracted from the image to be recognized, including but not limited to edge features, texture features, or shapes, etc. In this embodiment, the first feature image may include but is not limited to at least one of the outline of the spikes, color changes, or other features that help distinguish the spikes from the background, etc.
[0066] S230. Input the first feature image into the feature pyramid of the density recognition model to obtain second feature images at multiple levels.
[0067] Among them, the feature pyramid network structure can be constructed through lateral connections and upsampling for extracting feature maps at different levels. The second feature image can be understood as the feature image obtained after the first feature image is processed by the feature pyramid. The second feature image contains multiple levels of feature maps, and the feature maps capture the features of the first feature image at different levels, from local details to global details, so that each feature map corresponds to a different spatial resolution or receptive field size. In this embodiment, the feature pyramid allows the density recognition model to consider information at multiple levels within the same framework to ensure that both small spikes and large spike clusters can be accurately captured.
[0068] S240. Input the second feature images at multiple levels into the spatial attention module of the density recognition model to obtain a spatial feature image.
[0069] Among them, the spatial feature image can be understood as the feature map generated after the second feature image is processed by the spatial attention module. Based on the second feature image, the feature representation of the spike region is enhanced through a weighting mechanism, while suppressing the influence of irrelevant or noisy regions. Each pixel value in the spatial feature image not only reflects the feature intensity at that pixel position but also contains the relative importance of that pixel position in the entire spatial feature image. Although weighted processing is performed, the spatial feature image still retains the spatial layout information of the original image, ensuring that subsequent processing can correctly locate and count gramineous plants.
[0070] Based on the above solution, optionally, the spatial attention module includes a size conversion unit, a weight calculation unit, and a feature fusion unit; the step of inputting the second feature images at multiple levels into the spatial attention module of the density recognition model to obtain a spatial feature image includes: through the size conversion unit of the spatial attention module of the density recognition model, process at least some of the second feature images among the second feature images at multiple levels so that the second feature image corresponding to each level presents a preset size; through the weight calculation unit of the spatial attention module of the density recognition model, splice the second feature images of the preset size, calculate the attention weight for each spatial position in the second feature image through a convolution operation, and perform normalization processing on the attention weight through a normalized exponential function to obtain an attention weight map corresponding to each second feature image; through the feature fusion unit of the spatial attention module of the density recognition model, use the attention weight map corresponding to each second feature image to weight the second feature image, and fuse the weighted second feature images by splicing to obtain a spatial feature image.
[0071] Among them, the size conversion unit can be understood as a component responsible for adjusting the size of the second feature image to ensure that the second feature images at all levels have the same preset size, so that the second feature images with different resolutions can be uniformly processed in subsequent steps. The weight calculation unit can be understood as a component for determining the attention weight of each spatial position, and the weight reflects the importance of each position in the second feature image. The feature fusion unit can be understood as a component for fusing the weighted second feature images.
[0072] Adopting this technical solution, by processing the second feature images at multiple levels through the spatial attention module, first, the size conversion unit is used to ensure the consistency of the images at each level to reach the preset size; then, the weight calculation unit splices the second feature images after size adjustment and uses convolution operations to assign attention weights to each spatial position. After being processed by the sigmoid function, an attention weight map is formed, highlighting the important regions in the second feature image; finally, the feature fusion unit uses the attention weight map to perform weighted fusion on the feature images to generate a spatial feature image, effectively improving the accuracy and robustness of the spike number detection of the density recognition model under different scales and complex backgrounds, making the density recognition model more focused on the spikes and reducing the influence of background noise.
[0073] Based on the above solution, optionally, the size conversion unit of the spatial attention module of the density recognition model processes at least some of the second feature images among the second feature images at multiple levels so that the second feature images corresponding to each level present a preset size, including: through the spatial attention module of the density recognition model, respectively upsampling the second feature images output from the first level of the feature pyramid by the bilinear interpolation method so that the second feature images corresponding to each level present a preset size.
[0074] Among them, the preset size can be understood as the target size that the second feature image should reach after processing. The preset size is the image size of the second feature image at the second level, and the first level in the feature pyramid is lower than the second level. The first level of the feature pyramid can be understood as the level with lower resolution and higher semantic level in the feature pyramid. The second feature image at the first level has a smaller spatial size but stronger semantic information. The second level of the feature pyramid can be understood as the level with higher resolution and lower semantic level in the feature pyramid. The second feature image at the second level has a larger spatial size and contains more local details. The bilinear interpolation method can be understood as an image resampling technique used to estimate the value of a new pixel point based on the weighted average of four nearest neighbor pixels. The upsampling can be understood as a method to increase the spatial resolution of an image or feature map, restoring the low-resolution second feature image to a higher resolution to match a specific preset size.
[0075] An optional implementation manner is as Figure 4 shown. Apply a spatial attention module to the second feature image of each level. The second feature image output at the first level is upsampled by the bilinear interpolation method to keep the same size as the second feature image at the second level. Concatenate the feature maps obtained from different levels, use a convolutional operation to calculate the attention weights for each spatial position in the second feature image, and normalize the attention weights at multiple spatial positions through the Softmax function to obtain an attention weight map, and calculate the corresponding attention weight map for each feature map accordingly. As Figure 5 shown, apply the normalized attention weights to the corresponding second feature image, weight the second feature image by element-wise multiplication, and fuse the weighted second feature images by concatenation. This method can optimize the second feature images at different levels through the feature capture capabilities of different levels of the model. For example, the feature map at the first level is sensitive to detailed features, and the feature map at the second level is sensitive to spatial positions. In this way, it can be used to optimize the spatial position of the second feature image at the first level with the second feature image at the second level, and optimize the detailed features of the second feature image at the second level with the second feature image at the first level.
[0076] Adopting this technical solution, through the size conversion unit in the spatial attention module of the density recognition model, the bilinear interpolation method is used to upsample the second feature image output by the first layer of the feature pyramid network, ensuring that feature images of different levels can be adjusted to a preset size. This not only retains the detailed information of the image to be recognized but also enables features from different scales to be effectively compared and fused in the same dimension, enhancing the detection ability of the density recognition model for multi-scale targets. In particular, it is more beneficial for the recognition of the number of ears in the case of small targets or in high-density scenarios, thereby improving the accuracy and efficiency of the density recognition model.
[0077] S250. Input the spatial feature image into the target prediction head of the density recognition model to obtain the target density map of the ears of the gramineous plants, and determine the number of ears of the gramineous plants according to the target density map.
[0078] An optional implementation manner is that the target prediction head includes multiple groups of convolution operations. Except for the last group of convolution, each group of convolution may at least include at least one of a convolution layer, a BatchNorm layer, and a ReLu activation layer, etc. The last group is a 1×1 convolution for changing the number of channels so that the output channel number is a density map of 1.
[0079] The technical solution of the embodiment of the present invention generates a first feature image through the feature extraction module to capture the basic features in the image to be recognized; then, the feature pyramid generates multi-level second feature images to ensure that the density recognition model can process targets of different scales; then, the spatial attention module enhances the attention to the ear region by weighted fusion of the second feature images and reduces the influence of background noise; finally, the target prediction head generates the final target density map, which accurately reflects the distribution of the ears of the gramineous plants. This not only improves the detection accuracy and robustness of the number of ears but also greatly improves the processing efficiency and is applicable to the automatic counting task in a complex agricultural environment.
[0080] Embodiment III
[0081] Figure 6 It is a schematic structural diagram of a device for determining the number of ears of a gramineous plant provided in Embodiment III of the present invention. As Figure 6 shown, the device includes: a density recognition model determination module 610 and an ear number determination module 620.
[0082] Among them, the density recognition model determination module 610 obtains a to-be-recognized image of a gramineous plant with the number of ears to be recognized, and determines a density recognition model corresponding to the to-be-recognized image; wherein, the to-be-recognized image is a top view image of the gramineous plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; the density recognition model is obtained by training a deep learning model with a sample recognition image and an expected density image; the ear number determination module 620 is configured to input the to-be-recognized image into the density recognition model to obtain a target density map of the ears of the gramineous plant, and determine the number of ears of the gramineous plant according to the target density map.
[0083] In the technical solution of the embodiment of the present invention, the density recognition model determination module obtains a to-be-recognized image of a gramineous plant with the number of ears to be recognized, and determines a density recognition model corresponding to the to-be-recognized image. Since the to-be-recognized image is a top view image of the gramineous plant, the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module, and the density recognition model is obtained by training a deep learning model with a sample recognition image and an expected density image, the number of ears of the gramineous plant can be automatically counted through the density recognition model; finally, by inputting the to-be-recognized image into the density recognition model through the density recognition model to obtain a target density map of the ears of the gramineous plant, and determining the number of ears of the gramineous plant according to the target density map, the counting efficiency and accuracy of the ears can be improved, and the problem that in the related art, manual counting of the ears is not only time-consuming and laborious, but also difficult to cope with the rapid changes in a large-scale farmland environment can be solved, the human error can be reduced, and the efficient and accurate counting of the ears of the gramineous plant in a complex agricultural environment can be realized.
[0084] Based on the above method, optionally, the ear number determination module includes: a to-be-recognized image input sub-module, a first feature image input sub-module, a second feature image input sub-module, and an attention feature map input sub-module. Among them, the to-be-recognized image input sub-module is configured to input the to-be-recognized image into the feature extraction module of the density recognition model to obtain a first feature image; the first feature image input sub-module is configured to input the first feature image into the feature pyramid of the density recognition model to obtain second feature images of multiple levels; the second feature image input sub-module is configured to input the second feature images of multiple levels into the spatial attention module of the density recognition model to obtain a spatial feature image; the attention feature map input sub-module is configured to input the spatial feature image into the target prediction head of the density recognition model to obtain a target density map of the ears of gramineous plants.
[0085] Based on the above method, optionally, the spatial attention module includes a size conversion unit, a weight calculation unit, and a feature fusion unit; the second feature image input sub-module may include: a second feature image processing unit, a normalization processing unit, and a fusion unit. Among them, the second feature image processing unit is configured to process at least some of the second feature images among the second feature images of multiple levels through the size conversion unit of the spatial attention module of the density recognition model, so that the second feature image corresponding to each level presents a preset size; the normalization processing unit is configured to splice the second feature images of the preset size through the weight calculation unit of the spatial attention module of the density recognition model, calculate attention weights for each spatial position in the second feature image through a convolution operation, and normalize the attention weights through a normalization exponential function to obtain an attention weight map corresponding to each second feature image; the fusion unit is configured to weight the second feature image by using the attention weight map corresponding to each second feature image through the feature fusion unit of the spatial attention module of the density recognition model, and fuse the weighted second feature images in a splicing manner to obtain a spatial feature image.
[0086] Based on the above method, optionally, the second feature image processing unit is specifically configured to: respectively upsample the second feature image output by the first level of the feature pyramid through the spatial attention module of the density recognition model by means of bilinear interpolation, so that the second feature image corresponding to each level presents a preset size; wherein, the preset size is the image size of the second feature image of the second level; the first level in the feature pyramid is lower than the second level.
[0087] Based on the above method, optionally, the density recognition model is trained as follows: Obtain multiple sample recognition images and the expected density image corresponding to each sample recognition image, input the sample recognition image into a pre-established deep learning model to obtain a model output image; determine the model output loss according to the model output image, the expected density image, and a preset Bayesian loss function, so as to obtain the density recognition model.
[0088] Based on the above method, optionally, the Bayesian loss function is:
[0089]
[0090] where N is the total number of labeled points of the spike, E[c n is the expected count value of the labeled point n, E[c0] is the expected count value of the labeled point n, M is the total number of pixels of the expected density map, D est (x m ) is the true density heat map corresponding to the pixel point x, p(y n |x m ) is the posterior probability that the pixel point x m belongs to the labeled point y N , p(y0|x m ) is the posterior probability that the pixel point x m belongs to the background dummy y0, p(x m |y n ) represents the likelihood function of the labeled points of the spike, p(x m |y0) represents the likelihood function of the background dummy y0, z n is the coordinate of the labeled point y n ; σ is the width parameter of the Gaussian kernel, and d is the distance between the background dummy and the nearest spike target point.
[0091] Based on the above method, optionally, the spike number determination device for gramineous plants further includes: a labeling sub-module, a normalization processing sub-module, and an accumulation sub-module. Among them, the labeling sub-module is used to label each spike in the sample recognition image for each sample recognition image to obtain the labeled points corresponding to each spike; the normalization processing sub-module is used to generate a Gaussian heat map corresponding to the labeled point for each labeled point by using a two-dimensional Gaussian function, and perform normalization processing on each Gaussian heat map to obtain a target heat map; the accumulation sub-module is used to accumulate the target heat maps of the labeled points corresponding to multiple spikes to obtain the expected density image corresponding to the sample recognition image.
[0092] The ear number determination device for gramineous plants provided by the embodiments of the present invention can execute the ear number determination method for gramineous plants provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0093] Embodiment 4
[0094] Figure 7 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0095] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0096] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0097] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a method for determining the number of spikes of a gramineous plant.
[0098] In some embodiments, a method for determining the number of spikes of a gramineous plant may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for determining the number of spikes of a gramineous plant described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute a method for determining the number of spikes of a gramineous plant in any other suitable manner (e.g., by means of firmware).
[0099] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0100] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0101] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0102] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0103] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0104] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0105] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0106] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining the number of ears of a grass plant, characterized in that: include: Obtain an image of a grass plant with a number of ears to be identified, and determine a density recognition model corresponding to the image to be identified; wherein the image to be identified is a top view image of the grass plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; The density recognition model is obtained by training a deep learning model through sample recognition images and expected density images; The image to be identified is input into the density identification model to obtain a target density map of ears of the Gramineae plant, and the number of ears of the Gramineae plant is determined according to the target density map.
2. The method according to claim 1, characterized in that The step of inputting the image to be identified into a density identification model to obtain a target density map of ears of Gramineae plants comprises: Inputting the image to be recognized into the feature extraction module of the density recognition model to obtain a first feature image; Inputting the first feature image into the feature pyramid of the density recognition model to obtain a plurality of levels of second feature images; Inputting the second feature images of multiple levels into the spatial attention module of the density recognition model to obtain spatial feature images; The spatial feature image is input into the target prediction head of the density recognition model to obtain a target density map of ears of Gramineae plants.
3. The method according to claim 2, characterized in that The spatial attention module includes a size conversion unit, a weight calculation unit and a feature fusion unit; the inputting the second feature images of multiple levels into the spatial attention module of the density recognition model to obtain the spatial feature image includes: Processing at least part of the second feature images of multiple levels through a size conversion unit of the spatial attention module of the density recognition model, so that the second feature image corresponding to each level presents a preset size; The plurality of second feature images of the preset size are spliced together by the weight calculation unit of the spatial attention module of the density recognition model, and an attention weight is calculated for each spatial position in the second feature image by a convolution operation, and the attention weight is normalized by a normalized exponential function to obtain an attention weight map corresponding to each second feature image; Through the feature fusion unit of the spatial attention module of the density recognition model, the second feature image is weighted using the attention weight map corresponding to each second feature image, and the weighted second feature images are fused by splicing to obtain a spatial feature image.
4. The method according to claim 3, characterized in that The size conversion unit of the spatial attention module of the density recognition model processes at least part of the second feature images of multiple levels so that the second feature image corresponding to each level presents a preset size, including: The second feature images output by the first level of the feature pyramid are upsampled by a bilinear interpolation method through the spatial attention module of the density recognition model, so that the second feature images corresponding to each level present a preset size; wherein the preset size is the image size of the second feature image of the second level; the first level in the feature pyramid is lower than the second level.
5. The method according to claim 1, characterized in that The density recognition model is trained in the following way: Acquire multiple sample identification images and an expected density image corresponding to each of the sample identification images, input the sample identification images into a pre-established deep learning model, and obtain a model output image; The model output loss is determined according to the model output image, the expected density image and a preset Bayesian loss function to obtain a density recognition model.
6. The method according to claim 5, characterized in that The Bayesian loss function is: Where N is the total number of marked points of the ear, E[c n ] is the expected count value of the labeled point n, E[c0] is the expected count value of the labeled point n, M is the total number of pixels in the expected density map, D est (x m ) is the real density heat map corresponding to the pixel point xm, p(y n |x m ) is the pixel x m Belongs to the marked point y n The posterior probability, p(y0|x m ) is the pixel x m The posterior probability of the background dummy y0, p(x m |y n ) represents the likelihood function of the labeled points of the ear, p(x m |y0) represents the likelihood function of the background dummy variable y0, z n is the marked point y n ; σ is the width parameter of the Gaussian kernel, and d is the distance between the background dummy and the nearest spike target point.
7. The method according to claim 5, characterized in that The obtaining of the expected density image corresponding to each sample identification image includes: For each of the sample identification images, marking each ear in the sample identification image to obtain a marking point corresponding to each ear; For each of the marked points, a two-dimensional Gaussian function is used to generate a Gaussian heat map corresponding to the marked point, and each of the Gaussian heat maps is normalized to obtain a target heat map; The target heat maps of the marked points corresponding to the plurality of ears are accumulated to obtain the expected density image corresponding to the sample identification image.
8. A device for determining the number of ears of grass plants, characterized in that: include: A density recognition model determination module is used to obtain an image to be identified of a Gramineae plant with a number of ears to be identified, and determine a density recognition model corresponding to the image to be identified; wherein the image to be identified is a top view image of the Gramineae plant; the density recognition model includes a feature extraction module, a feature pyramid connected to the feature extraction module, a spatial attention module connected to the feature pyramid, and a target prediction head connected to the spatial attention module; the density recognition model is obtained by training a deep learning model through sample recognition images and expected density images; The ear number determination module is used to input the image to be identified into the density recognition model to obtain a target density map of ears of the Gramineae plant, and determine the number of ears of the Gramineae plant according to the target density map.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for determining the number of ears of Gramineae plants according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for determining the number of ears of Gramineae plants according to any one of claims 1 to 7 when executed.