Lightweight agricultural crop counting method and system
Through the lightweight agricultural crop counting method and the counting model of attention mechanism, the problems of high model complexity and high resource occupancy are solved, and fast and low-cost agricultural crop counting is achieved, adapting to sparse and dense planting environments, and adapting to edge computing equipment with limited resources.
Patent Information
- Application Number
- CN202510333113.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-22
AI Technical Summary
The existing agricultural crop counting method has high complexity, high resource occupancy and long training time, making it difficult to adapt to practical application scenarios in agricultural production.
The lightweight agricultural crop counting method is adopted, and the lightweight agricultural crop intelligent counting model combined with the attention mechanism is used to extract and predict feature and quantity through the coding module, counting module and standardized module, and the convolution operation and attention mechanism are used to reduce the amount of model parameters and calculation complexity.
Optimizing the feature extraction process improves the robustness and adaptability of complex field environments, shortens model training time, reduces dependence on high-performance hardware, and realizes real-time processing of field images and fast crop counting.
Smart Images

Figure CN120356089A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of crop quantity statistics, and particularly relates to a lightweight agricultural crop counting method and system. Background Art
[0002] Agricultural crop counting is a key link in crop yield estimation. By accurately counting the number of crops, the growth status of crops can be comprehensively monitored, and production and management measures can be adjusted based on the data, thereby effectively improving the yield and quality of crops. This process is of great significance in agricultural production and is an important part of precision agriculture and intelligent management. In recent years, with the rapid development of deep learning technology, this technology has been widely applied in the agricultural field. Especially in the task of crop counting, its efficient feature extraction ability and adaptability to complex scenarios have made it the mainstream solution. The object detection method based on deep learning has performed particularly well in agricultural crop counting, providing an effective means to solve the counting problems in sparse planting and dense planting scenarios.
[0003] Currently, the deep learning-based agricultural crop counting methods are mainly divided into the method based on density map regression, the method based on segmentation, and the method based on object detection. The method based on density map regression generates a crop density map and converts the quantity information into pixel intensity distribution, which is suitable for dense planting scenarios but requires high computing resources. The method based on segmentation counts the number of targets through pixel-level segmentation and has high accuracy in sparse planting scenarios, but may have errors when the targets are dense or the background is complex. The method based on object detection directly identifies and locates the crop targets in the image, with good real-time performance and applicability, but the accuracy may be limited in high-density target scenarios.
[0004] For example, Lu et al. proposed a maize tassel counting method based on a local counts regression network in "TasselNet: Counting Maize Tassels in the Wild via Local Counts Regression Network" (Plant Methods, 2017). This method uses the deep learning framework of convolutional neural network, constructs a regression function locally, and follows the standard process of generating a true density map: generating a Gaussian distribution at each annotation point position, constructing a true density map through Gaussian smoothing operation, and then using the local density map for counting integration, and finally outputting the crop counting result. This method effectively improves the accuracy of maize tassel counting, but the model is relatively complex, has high requirements for computing resources, and the density map generation process is more dependent on the position accuracy of manual annotation points.
[0005] In addition, Bai et al. proposed a new deep learning network, RiceNet, for rice plant counting in "Rice Plant Counting, Locating, and Sizing Method Based on High-Throughput UAV RGB Images" (Plant Phenomics, 2023). RiceNet consists of a feature extractor and three designed decoder modules, including a density map estimator (DME), a plant size estimation module (PSE), and a plant location module (PLD). Among them, the first 13 layers of VGG16 are used as the feature extractor to extract multi-level feature maps from the input high-resolution RGB images. Subsequently, the DME module fuses multi-level features through concatenation and upsampling operations to generate a high-quality density map, and introduces a rice plant attention mechanism into the network, significantly enhancing the ability to distinguish crops from the background. RiceNet can use high-throughput RGB images collected by drones to achieve rice plant counting in paddy fields and has achieved good results in density estimation, plant location, and size assessment. However, this method still relies on deep feature fusion and high computing resources, and there is still room for improvement in terms of model lightweight and real-time performance. Summary of the Invention
[0006] The object of the present invention is to solve the problems in existing methods such as high model complexity, large resource occupancy, and long training time. At the same time, while ensuring the high performance of the model, the computing cost and hardware requirements are significantly reduced to adapt to the actual application scenarios in agricultural production.
[0007] To achieve the above object, a lightweight agricultural crop counting method provided by the present application includes the following steps: inputting the RGB image of the crop to be analyzed into a counting model to perform feature extraction and quantity prediction on the image to be analyzed, and outputting the predicted quantity of the crop to be analyzed;
[0008] The counting model includes an encoding module, a counting module, and a normalization module. After the RGB image data undergoes multiple convolutional operations in the encoding module, it is cascaded with an attention mechanism to dynamically enhance the feature representation of the key crop regions in the image;
[0009] The output processed by the encoding module is transmitted to the counting module. In the counting module, global average pooling is first introduced, and then through convolution and normalization operations, the features output by the encoding module are further learned. Finally, after 1× convolution, feature compression and quantity regression are completed, the number of crops in the RGB image is recorded, and a redundant counting map is obtained and reference figure Reference figure Record the number of times each position is calculated;
[0010] The normalization module first maps the low-resolution redundant count map C r back to the resolution of the input image to generate an upsampled count map Then, each pixel value in C u is divided by the value at the corresponding position in P to eliminate redundant counts and generate a count map By aggregating all elements in C n the image-level count c I is calculated as follows:
[0011]
[0012] where is the vectorized form of C r c r (i) is the count of the i-th local region of c r and P i is an r×r region extracted from the reference map P, corresponding to c r (i).
[0013] The processing of the RGB image data in the encoding module includes:
[0014] The RGB image data first undergoes deep feature extraction through the Conv-16 convolutional layer, Conv-32 convolutional layer, Conv-64 convolutional layer, and Conv-128 convolutional layer. The features extracted by the Conv-32 convolutional layer are enhanced by the Ca-32 spatial attention layer, which assigns different weights to each feature channel to highlight key crop regions and allows the model to focus on key areas in the image. The features extracted by the Conv-32 convolutional layer then continue through the Conv-64, and the features extracted by the Conv-64 convolutional layer are enhanced by the Ca-64 spatial attention layer. The fused features are weighted using the weight distribution assigned by the Ca-64 spatial attention layer Figure 1 and then fed into the Conv-128 convolutional layer. The Conv-128 convolutional layer performs deep feature extraction. The Conv-128 convolutional layer uses 256 convolutional kernels to perform convolution operations on the input feature map. The convolutional kernels slide over the input feature map to perform weighted summation on local regions, thereby generating a new feature map.
[0015] The Conv-32 convolutional layer, Conv-64 convolutional layer, and Conv-128 convolutional layer adopt a pyramidal convolutional structure, each including a number of convolutional kernels of different sizes, a normalization layer, and an activation function layer.
[0016] The attention module uses one-dimensional global average pooling to generate a direction-aware feature map, and encodes it into an attention map through convolutional and normalization layers, dynamically enhancing the model's attention and expression ability for key-region features. The specific steps of the global average pooling decomposition are as follows:
[0017] J1. Use one-dimensional global average pooling operations in the vertical and horizontal directions, as follows:
[0018]
[0019] Equation 1 represents the average of the pixel values x c (h, i) over all columns in the h-th row of the feature map of the c-th channel, obtaining the feature value of the h-th row W represents the width of the feature map;
[0020] Equation 2 represents the average of the pixel values x c (j, w) over all rows in the w-th column of the feature map of the c-th channel, obtaining the feature value of the w-th column H represents the height of the feature map;
[0021] This operation performs average pooling on the input feature map of size C×H×W in the x and y directions respectively, generating feature maps of size C×H×1 and C×1×W. This process can capture long-range spatial interactions with precise position information;
[0022] J2. Subsequently, transform the generated feature maps of C×H×1 and C×1×W, and then perform a fusion operation.
[0023] f = δ(c([z h ,z w )), (3)
[0024] Equation 3 represents concatenating z h and z w in the channel dimension or spatial dimension to generate a new feature map. δ is an activation function. After activation, the shape of the feature map becomes C×(H+W)×1 or other concatenation methods (to be determined according to dimension processing).
[0025] J3. Then perform F1 operation using a 1×1 convolutional kernel to reduce the number of channels for dimensionality reduction, generating a feature map Split the feature map f after dimensionality reduction along the spatial dimension into and corresponding to the features in the vertical and horizontal directions respectively, and then perform dimensionality increase operations using 1×1 convolutions on f h and f wThe number of channels is restored to the same dimension as the input feature map to ensure that the attention vector matches the number of channels of the input feature map, and then combined with the sigmoid activation function to obtain the final attention vector and
[0026] g h = σ(F h (f h )),(4)
[0027] g w = σ(F w (f w )),(5)
[0028] J4. Output of the final spatial attention layer:
[0029]
[0030] The output of the spatial attention layer is a weighted value of the original input feature map x c (i, j), and the weight is determined by the attention vectors g h and g w .
[0031] In the counting module, the global average pooling layer is a non-parametric layer. The size of the image area corresponding to the feature points is r×r, the output stride is s, and the kernel resolution size of the global average pooling layer is
[0032] A lightweight agricultural crop counting system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. It is characterized in that when the processor executes the computer program, the steps of the lightweight agricultural crop counting method are implemented.
[0033] The beneficial effects of the present invention are as follows: By adopting a lightweight neural network architecture and combining an attention mechanism, the number of model parameters can be reduced, the computational complexity can be lowered, the feature extraction process can be optimized, and the robustness and adaptability to complex field environments can be improved; The modular design and efficient optimization strategy not only shorten the model training time but also reduce the dependence on high-performance hardware, enabling the model to quickly complete the modeling and training of large-scale crop data. At the same time, through normalization processing, the convergence speed and generalization performance are significantly improved; The attention mechanism dynamically enhances the feature representation of the crop area, suppresses background interference, improves the counting accuracy and the adaptability of the model in sparse and dense planting environments. With the lightweight design, the model can adapt to resource-constrained edge computing devices, realize real-time processing of field images, quickly complete crop counting, and provide an efficient, low-cost, and practical intelligent solution for agricultural production monitoring. Description of the Drawings
[0034] Figure 1 is the model structure diagram of this invention patent;
[0035] Figure 2 is the flow chart of this invention;
[0036] Figure 3 shows the counting performance of CounterNet and other models on the rice seedling and rape flower cluster datasets. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0038] Embodiment 1
[0039] A lightweight agricultural crop counting method provides a lightweight intelligent agricultural crop counting model integrating an attention mechanism. By inputting a dataset to be measured into the counting model, the number of agricultural crops is obtained. The counting model includes an encoding module, a counting module, and a normalization module. The model training steps include:
[0040] S1. Select training samples from the RGB image dataset of field crops;
[0041] S2. Perform feature extraction in the encoding module: Specifically, input the RGB image data into the encoding module to start the feature extraction and encoding process. After the RGB image data undergoes multiple convolutional operations, it is cascaded with the attention mechanism to dynamically enhance the feature representation of the key crop regions in the image. The convolutional layers include Conv-16 convolutional layer, Conv-32 convolutional layer, Conv-64 convolutional layer, and Conv-128 convolutional layer, and the attention layers include Ca-32 and Ca-64.
[0042] Specifically, as Figure 1 shown, the RGB image data first undergoes deep feature extraction through the Conv-16 convolutional layer, Conv-32 convolutional layer, Conv-64 convolutional layer, and Conv-128 convolutional layer. The features extracted by the Conv-32 convolutional layer are enhanced through the Ca-32 spatial attention layer, which assigns different weights to each feature channel to highlight the key crop regions and allows the model to focus on the key regions in the image, such as the main parts of the crops. The features extracted by the Conv-64 convolutional layer are enhanced through the Ca-64 spatial attention layer;
[0043] The features extracted by the Conv-32 convolutional layer continue to be subjected to deep feature extraction through the Conv-64 and Conv-128 convolutional layers. The Conv-128 convolutional layer uses 256 convolutional kernels to perform a convolution operation on the input feature map. The convolutional kernels slide on the input feature map and perform weighted summation on local regions to generate a new feature map.
[0044] The weights distributed by the Ca-32 spatial attention layer are used to weight the features output by the Conv-128 layer. The weighted feature map is fused with the original feature map, usually by element-wise multiplication. In this way, the model can enhance the features in key regions while suppressing the features in unimportant regions to obtain the fused features. Figure 1 。
[0045] The weights distributed by the Ca-64 spatial attention layer are used to weight the fused features Figure 1 and then they are sent to the Conv-128 convolutional layer. This layer further processes the feature map using 128 convolutional kernels to perform feature compression or transformation to meet the requirements of subsequent network layers.
[0046] After multiple convolution operations and two feature enhancement and fusion operations, the dimension of the input data is converted from the original image to a feature map with 128 layers, providing rich feature information for subsequent modules.
[0047] Through this feature fusion design, information loss is minimized to ensure that the detailed features of the crop are retained. After multiple convolution operations and two cascading operations, the dimension of the input data is converted from the original image to a feature map with 128 channels, providing rich feature information for subsequent modules. The core of the encoding module is convolution, which determines the level of feature extraction. Here, the Conv-32, Conv-64, and Conv-128 convolutional layers in the encoding module of this embodiment adopt pyramid-shaped convolutional layers. The Conv-32 convolutional layer uses convolutional kernels of different sizes (such as 3×3, 5×5), the Conv-64 convolutional layer uses convolutional kernels of different sizes (such as 3×3, 5×5, 7×7, 9×9), the Conv-128 convolutional layer uses convolutional kernels of different sizes (such as 3×3, 5×5, 7×7), and the Conv-256 convolutional layer uses convolutional kernels of different sizes (such as 3×3, 5×5). The pyramid convolutional layer can capture features of different scales in the same layer, enabling the convolutional layer to extract local and global features simultaneously, adapt to targets of different sizes, and greatly improve the performance of the network without increasing the computational cost.
[0048] Most existing CNNs use small convolutional kernels (such as 3×3) and expand the receptive field through multi-layer stacking because increasing the convolutional kernel brings huge costs in terms of the number of parameters and computational complexity.
[0049] In the design of the CounterNet deep learning network in the encoding module, an attention module is also introduced through a branch strategy to enhance the expressive ability of image features and the performance of the model. The attention module specifically generates attention maps for height and width through a combination of adaptive average pooling, convolutional layers, batch normalization, and activation functions, further enhancing the channel and spatial information of the feature map. In addition, a calculation method that ensures parameter divisibility is provided to optimize the computational efficiency of the model. This enables the lightweight network to focus on a larger area while avoiding a large amount of computational overhead.
[0050] Preferably, the attention module first uses one-dimensional global average pooling operations in the vertical and horizontal directions in the spatial attention layer to aggregate the input features in the vertical and horizontal directions into two independent direction-aware feature maps. These two feature maps embed information in specific directions and are then encoded into two attention maps. Finally, a convolutional pooling layer and a normalization layer are added to perform the transformation of the feature channels.
[0051] The specific implementation is as follows:
[0052] J1. Use one-dimensional global average pooling operations in the vertical and horizontal directions, specifically as follows:
[0053]
[0054] Equation 1 represents the average of the pixel values x c (h, i) in all columns of the h-th row of the feature map for the c-th channel, obtaining the feature value of the h-th row W represents the width of the feature map.
[0055] Equation 2 represents the average of the pixel values x c (j, w) in all rows of the w-th column of the feature map for the c-th channel, obtaining the feature value of the w-th column H represents the height of the feature map.
[0056] This operation performs average pooling on the input feature map of size C×H×W in the x direction and y direction respectively, generating feature maps of size C×H×1 and C×1×W. This process can capture long-range spatial interactions with precise position information.
[0057] J2. Subsequently, the generated feature maps of C×H×1 and C×1×W are transformed and then fused.
[0058] f = δ(c([z h, z w )) , (3)
[0059] Equation 3 represents concatenating z h and z w in the channel dimension or the spatial dimension to generate a new feature map. δ is an activation function. After activation, the shape of the feature map becomes C×(H+W)×1 or other concatenation methods (to be determined according to dimension processing).
[0060] J3. Then perform F1 operation to reduce the number of channels using a 1×1 convolutional kernel to generate a feature map Split the feature map f after dimensionality reduction along the spatial dimension into and which respectively correspond to the features in the vertical and horizontal directions. Then perform dimensionality increase operations using 1×1 convolutions respectively on f h and f w to restore the number of channels to the same dimension as the input feature map, ensuring that the attention vector matches the number of channels of the input feature map, and then combine with the sigmoid activation function to obtain the final attention vector and
[0061] g h = σ(F h (f h )) , (4)
[0062] g w = σ(F w (f w )) , (5)
[0063] J4. Finally, the output of the spatial attention layer:
[0064]
[0065] The output of the spatial attention layer is a weighting of the original input feature map x c (i, j), and the weights are determined by the attention vectors g h and g w .
[0066] S3. Perform quantity regression in the counting module: Transmit the output processed by the encoding module to the counting module. In the counting module, global average pooling is first introduced, and then through convolution and normalization operations, further learn the features output by the encoding module. Finally, after 1×1 convolution, feature compression and quantity regression are completed to record the number of crops in the RGB image.
[0067] The global average pooling operation performs a comprehensive average calculation on the feature map output by the encoding module in the spatial dimension, calculates the average value of each channel in the feature map over the entire image space, compresses the spatial dimension of the feature map to 1×1, so that each channel only retains one value. Finally, the feature map after convolution and normalization operations is compressed by a 1×1 convolution. The 1×1 convolution kernel performs a weighted sum of the pixels of all input channels at each position to generate the compressed feature map. The feature map after feature compression completes quantity regression through the regression layer. The regression layer is usually a fully connected layer or a convolutional layer, which is used to convert the compressed feature map into an output related to the crop quantity. The regression layer predicts the crop quantity in the image by learning the mapping relationship between the feature map and the crop quantity.
[0068] Global average pooling allows for flexible operation of the basic input size r×r without changing the model complexity because this layer is a non-parametric layer that can flexibly change r according to different object sizes in the image. In CounterNet, it is easy to adjust r. Given the required basic input size r×r and output stride s, simply modify the kernel size of the global average pooling to r / s. Moreover, such modifications do not affect the model complexity. r×r represents the basic size (i.e., resolution) of the input image, r means the input image is divided into r×r blocks, the kernel size of the global average pooling is r / s. That is, if the size of the input image is H×W, the global average pooling layer divides the input image into H / s×W / s blocks, and the size of each block is r / s×r / s. s is the output stride, which is used to adjust the kernel size of the global average pooling layer, thereby flexibly controlling the output resolution of the model, indicating the reduction multiple of the output resolution of the global average pooling layer relative to the input resolution.
[0069] S4. Eliminate redundant counts in the normalization module: Input the prediction result output by the counting module into the normalization module for normalization processing. After processing, redundant counts are eliminated, and the accuracy and generalization ability of the model are improved. The specific steps include:
[0070] S41. Input image where I is the compressed feature map output after normalization by the counting module;
[0071] S42. Generate a redundant counting map through the encoding module and counting module of CounterNet where C r refers to the counting map finally output by the counting module. The counting map C rIt extracts the features of the input image through a network structure and predicts the target object density at each pixel position, representing the density distribution of the target object in the input image. Each pixel value in the counting map represents the target object density at that position. By performing an integration operation (summing over the entire input image or region of interest), the total number of target objects can be obtained; c represents the count value in a local region in the counting map C r where the local region is a non-overlapping small block of the input image, and the size of the small block is s×s.
[0072] S43. Generate a counting map through a normalization operation The normalization operation includes:
[0073] Restore the low-resolution redundant counting map C r to the resolution of the input image to generate an upsampled counting map
[0074] The upsampling operation evenly distributes the count value c∈C in each local region r into an r×r region, and the value of each element is ensuring that the sum of each r×r region is still equal to c. By applying the upsampling operation to the count values of all local regions in C r and rearranging them in the same spatial order and output stride, an upsampled counting map is obtained
[0075] Construct a reference map through CounterNet Record the number of times each position is counted as an indicator of redundancy; the reference map P combines the information of manually labeled points. By using a Gaussian kernel function, each point is expanded into a density distribution, and each labeled point is converted into a Gaussian heatmap. The Gaussian heatmaps of all points are superimposed together to form a continuous density map. The density map reflects the spatial distribution and density information of the target object in the image. Then, based on the density map, further downsampling and normalization operations are performed to obtain the reference map P.
[0076] Normalization processing generates a counting map
[0077]
[0078] where represents the element division operator. In this step, by dividing each pixel value in C u by the value at the corresponding position in P, redundant counts are eliminated, making C n more accurate.
[0079] S44. Image-level counting: Calculate the image-level count c by aggregating all elements in C n ; I ;
[0080]
[0081] where C n (x, y) is the value at position (x, y) in C n , representing the number of crops at that position.
[0082] The normalization process of CounterNet is achieved through Equation (7). To accelerate the normalization operation, the values of C n are reorganized based on the local area count of C r and the contribution of the reference map P, and fast normalization is achieved at the algorithm level. Equation (7) can be equivalent to:
[0083]
[0084] where is the vectorized pattern of C r , c r (i) is the i-th local area count of c r . P is an r×r area extracted from the reference map P, corresponding to c r (i). Equation (8) means decomposing each local count c r (i) into the contributions of the r×r area and eliminating redundancy through normalization.
[0085] In addition, P i can be efficiently constructed using an image processing operator (fold in PyTorch). By defining another variable , Equation (8) can be equivalent to Equation (9), and Equation (9) expresses the calculation of c I in the form of matrix multiplication, facilitating the rapid implementation using the parallel computing power of the GPU:
[0086]
[0087] where Equation (9) can be fully implemented by the GPU, thus greatly reducing the time consumption of the model.
[0088] S5. Iterative optimization: Perform multiple iterative trainings, judge the number of iterations. If the maximum value is not reached, repeat steps S1 - S5; if the maximum value is reached, obtain the best model and save it;
[0089] After training the counting model with a large amount of labeled data, the counting of the crops to be measured can be completed through the counting model. The steps include:
[0090] Step 1. The RGB image of the crop to be analyzed is input into the counting model to perform feature extraction and quantity prediction on the image to be analyzed;
[0091] Step 2. Output the predicted quantity of the crop to be analyzed.
[0092] As Figure 3 shown on the left side in , when the CounterNet model of the present invention is trained by processing a dataset of 397 rice seedling images with point annotation, it only takes 0.22 hours and 2.25 GB of memory. In contrast, TasselNet-v3 requires 1.73 hours and 2.48 GB of memory. TasselNetv3 includes 7 convolutional layers and 3 max pooling layers, specifically defined by C(16)-M-C(32)-M-C(64)-C(64)C(64)-M-C(128)-C(128)-C(1), where C(m) represents a two-dimensional convolutional layer with an m-channel k×k filter, followed by batch normalization (BN) and ReLU, and M is a max pooling operator with a stride of 2 and a kernel size of 2×2. The last C(1) is the prediction layer, excluding BN and ReLU. TasselNetv3 enhances the plant counting performance through guided upsampling and background suppression techniques. The prediction time for each image using TasselNet-v3 is approximately 3 seconds. The prediction time for each image using the CounterNet model is approximately 0.5 seconds.
[0093] As Figure 3 shown on the right side in , when the CounterNet model is trained by processing a dataset of 128 rape flower cluster images with box annotation, it only consumes 0.21 hours and 2.73 GB of memory, while YOLOv5-s requires 0.23 hours and 7.64 GB of memory. The prediction time for each image using YOLOv5-s is approximately 5 seconds. For details, see Table 1.
[0094] Table 1 Comparison of the time and memory consumed by the model when modeling in datasets with different annotation modes.
[0095]
[0096]
[0097] Note: 1) All models are run on a GPU server equipped with an NVIDIA GeForce RTX4090, which has 24 GB of memory and 16384 CUDA (Compute Unified Device Architecture) cores.
[0098] 2) It means that the model is not applicable to this annotation mode, so no results are generated.
[0099] The advantages of the CounterNet of the present invention lie not only in its advantages in terms of time and memory consumption, but more importantly in its adaptability to different annotation modes. In the current agricultural automation technology, CounterNet can handle models with both point annotation and box annotation at the same time. This feature of CounterNet undoubtedly provides more flexibility and choices for researchers and practitioners in the agricultural field.
[0100] A lightweight agricultural crop counting system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the lightweight agricultural crop counting method are implemented.
[0101] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight agricultural crop counting method, characterized in that, It includes the following steps: Input the RGB image of the crop to be analyzed into the counting model for feature extraction and quantity prediction of the image to be analyzed, and output the predicted quantity of the crop to be analyzed; The counting model includes an encoding module, a counting module, and a normalization module. After the RGB image data undergoes multiple convolutional operations in the encoding module, it is cascaded with the attention mechanism to dynamically enhance the feature representation of the key crop regions in the image; The output processed by the encoding module is passed to the counting module. In the counting module, global average pooling is first introduced, and then through convolution and normalization operations, the features output by the encoding module are further learned. Finally, after 1×1 convolution, feature compression and quantity regression are completed, and the number of crops in the RGB image is recorded to obtain a redundant counting map and the reference figure reference figure Record the number of times each position is calculated; The normalization module first maps the low-resolution redundant count map C r back to the resolution of the input image to generate an upsampled count map Then, each pixel value in C u is divided by the value at the corresponding position in P to eliminate redundant counts and generate a count map By aggregating all elements in C n the image-level count c is calculated I : Among them, is the vectorization mode of C r , c r (i) is the count of the i-th local area of c r , and P i is an r×r area extracted from the reference figure P and corresponds to c r (i).
2. The lightweight agricultural crop counting method according to claim 1, wherein, The processing process of the RGB image data in the encoding module includes: The RGB image data first undergoes deep feature extraction through the Conv-16 convolutional layer, Conv-32 convolutional layer, Conv-64 convolutional layer, and Conv-128 convolutional layer. The features extracted by the Conv-32 convolutional layer are enhanced through the Ca-32 spatial attention layer, which assigns different weights to each feature channel to highlight the key crop regions and allows the model to focus on the key regions in the image. The features extracted after the Conv-32 convolutional layer continue to be extracted through the Conv-64, and the features extracted by the Conv-64 convolutional layer are enhanced through the Ca-64 spatial attention layer. The fused feature map one is weighted using the weight distribution assigned by the Ca-64 spatial attention layer and then fed into the Conv-128 convolutional layer. The Conv-128 convolutional layer performs deep feature extraction. The Conv-128 convolutional layer uses 256 convolutional kernels to perform convolutional operations on the input feature map. The convolutional kernels slide on the input feature map to perform weighted summation on local regions, thereby generating a new feature map.
3. The lightweight agricultural crop counting method according to claim 1, wherein The Conv-32 convolutional layer, Conv-64 convolutional layer, and Conv-128 convolutional layer adopt pyramid-shaped convolutional layers, each of which includes a number of convolutional kernels of different sizes, a normalization layer, and an activation function layer.
4. A lightweight agricultural crop counting method according to claim 1, characterized in that, The attention module uses one-dimensional global average pooling to generate a direction-aware feature map and encodes it into an attention map through convolutional and normalization layers, dynamically enhancing the model's attention and expression ability for key region features. The specific steps of the global average pooling decomposition are as follows: J1. Use one-dimensional global average pooling operations in the vertical and horizontal directions to perform average pooling on the input feature map with dimensions C×H×W in the x direction and y direction respectively, generating feature maps with dimensions C×H×1 and C×1×W respectively, as follows: Equation 1 represents the pixel values x of all columns in the h-th row of the feature map of the c-th channel c are averaged to obtain the feature value of the h-th row W represents the width of the feature map; Equation 2 represents the pixel value x of all rows in the w-th column of the feature map of the c-th channel c is averaged for (j, w) to obtain the feature value of the w-th column H represents the height of the feature map; J2. Subsequently, transform the generated feature maps with dimensions C×H×1 and C×1×W, and then perform a fusion operation; f = δ(c([z h , z w )), (3) Equation 3 represents concatenating z h and z w in the channel dimension or the spatial dimension to generate a new feature map. δ is an activation function. After activation, the shape of the feature map becomes C×(H+W)×1 or other concatenation methods; J3. Then perform the F1 operation to reduce the dimension using a 1×1 convolutional kernel to reduce the number of channels and generate a feature map Perform a split operation on the dimension-reduced feature map f along the spatial dimension and divide it into and which respectively correspond to the features in the vertical and horizontal directions. Then, use 1×1 convolutions to perform dimension-increasing operations on f h and f w to restore the number of channels to the same dimension as the input feature map, ensure that the attention vector matches the number of channels of the input feature map, and then combine with the sigmoid activation function to obtain the final attention vector and g h = σ(F h (f h )), (4) g w = σ(F w (f w )), (5) J4. Finally, the output of the spatial attention layer: The output of the spatial attention layer is a weighting of the original input feature map x c at (i,j), with the weights determined by the attention vectors g h and g w respectively.
5. A lightweight agricultural crop counting method according to claim 1, characterized in that, The global average pooling layer in the counting module is a non-parametric layer. The size of the image region corresponding to the feature point is r×r, the output stride is s, and the kernel resolution size of the global average pooling layer is 6. A lightweight agricultural crop counting system, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5.