Small target detection network and rice panicle detection method in complex environment
Patent Information
- Application Number
- CN202610920553.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]现有技术中,通常以神经网络来对稻穗、麦穗等小目标进行检测,但是,现有技术中存在如下缺陷:由于检测目标的尺寸较小,而且在稻田、麦田等复杂的背景下(比如稻苗、麦苗以及杂草等叶片遮挡,或者土壤、枯叶等环境因素的影响),导致在检测过程中存在较为严重的背景噪声,现有的神经网络难以准确获取稻穗等小目标检测结果,为了提升检测精度,比如网络将轻量级卷积模块与空间金字塔空洞卷积相结合等手段,虽然在检测精度上得到一定提升,但是,其计算量庞大,提升了设备成本,而且由于检测目标的尺寸较小,在特征提取过程中丢失了关键空间信息,仍然存在较大的漏检率或者误检率
[0033]本发明的有益效果:通过本发明,采用SPD模块,能够对空间维度进行压缩并保留细粒度信息,在不牺牲特征质量的前提下降低计算量,增强对水稻穗等小目标边缘纹理等细节特征的表征能力,而且在特征处理模块的作用下,能够实现复杂环境下小目标的细节的重构与特征恢复,从而能够有效提升最终检测结果的精度,为后续生产措施的制定提供准确的数据支持。
Smart Images

Figure CN122821086A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an image processing network and method, and more particularly to a small target detection network and a rice ear detection method in complex environments. Background Technology
[0002] The testing is required for grain crops such as rice and wheat to predict yields and provide data guidance for subsequent production operations.
[0003] In existing technologies, neural networks are typically used to detect small targets such as rice ears and wheat ears. However, these technologies have the following drawbacks: due to the small size of the targets and the complex backgrounds in rice paddies and wheat fields (such as the obstruction by leaves of rice seedlings, wheat seedlings, and weeds, or the influence of environmental factors such as soil and dead leaves), there is significant background noise during the detection process. Existing neural networks struggle to accurately obtain detection results for small targets such as rice ears. To improve detection accuracy, methods such as combining lightweight convolutional modules with spatial pyramid dilated convolutions have been employed. While these methods have improved detection accuracy to some extent, they involve a large computational load, increasing equipment costs. Furthermore, due to the small size of the targets, key spatial information is lost during feature extraction, resulting in a significant false negative or false positive rate.
[0004] Therefore, in order to solve the above-mentioned technical problems, it is urgent to propose a new technical approach. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a small target detection network and a rice ear detection method in complex environments, which can accurately detect small targets in complex environments such as rice ears, while maintaining the lightweight nature of the entire network and reducing equipment costs.
[0006] The present invention provides a small target detection network in complex environments, including a feature extraction module, a feature processing module, and an output module;
[0007] The feature extraction module is used to extract features from small target images acquired in complex environments and output four different extracted features.
[0008] The feature processing module is used to receive four extracted features output by the feature extraction module, process the four extracted features to generate three features to be identified, and then output them.
[0009] The output module is used to receive three identification features and output the small target detection result.
[0010] Furthermore, the feature extraction module includes SPD module I, SPD module II, convolutional unit I, convolutional unit II, benchmark module I, benchmark module II, benchmark module III, benchmark module IV, and convolutional module I;
[0011] The SPD module I takes a small target image as input, the output features of the SPD module I are input into convolutional unit I, the output features of convolutional unit I are input into convolutional unit II, the output features of convolutional unit II are input into SPD module II, the output features of SPD module II are input into benchmark module I, and benchmark module I outputs the extracted feature F1.
[0012] The extracted feature F1 is input into the baseline module II, and the output of the baseline module II is used to extract feature F2; the extracted feature F2 is input into the baseline module III, and the output of the baseline module III is used to extract feature F3; the extracted feature F3 is input into the baseline module IV, and the output feature of the baseline module IV is input into the convolution module I, and the output of the convolution module I is used to extract feature F4; wherein, the scale of the convolution module I is 1×1.
[0013] Furthermore, the structures of convolutional unit I and convolutional unit II are identical;
[0014] The convolutional unit I includes a two-dimensional convolutional module I, a batch normalization module I, and a ReLU activation function module I;
[0015] The input of the two-dimensional convolution module I serves as the input of the convolution unit I. The output features of the two-dimensional convolution module I are input into the batch normalization module I. The output features of the batch normalization module I are input into the ReLU activation function module I. The output features of the ReLU activation function module I serve as the output features of the convolution unit I.
[0016] Furthermore, the reference module I, reference module II, reference module III, and reference module IV have the same structure;
[0017] The baseline module I includes an SPD module, a two-dimensional convolution module II, a batch normalization module II, and a ReLU activation function module II;
[0018] The input of the SPD module serves as the input of the baseline module I. The output of the SPD module is connected to the input of the two-dimensional convolution module II. The output features of the two-dimensional convolution module II are input into the batch normalization module II. The output features of the batch normalization module II are input into the ReLU activation function module II. The output features of the ReLU activation function module II are added element-wise to the input features of the SPD module, and the resulting feature is used as the output feature of the baseline module I.
[0019] Furthermore, the feature processing includes SPD module III, SPD module IV, SPD module V, AIFI module, convolution module II, convolution module III, convolution module IV, convolution module V, multi-scale feature fusion module I, multi-scale feature fusion module II, multi-scale feature fusion module III, multi-scale feature fusion module IV, multi-scale feature fusion module V, CARAFE module I, CARAFE module II, RepC3 module I, RepC3 module II, RepC3 module III, RepC3 module IV, and RepC3 module V;
[0020] The SPD module III receives the extracted feature F1, and the output feature of the SPD module III is input to the multi-scale feature fusion module I;
[0021] Convolutional module II receives the extracted feature F2. The output features of convolutional module II are input into multi-scale feature fusion module II and multi-scale feature fusion module I. The output features of multi-scale feature fusion module II are input into RepC3 module II. The output features of RepC3 module II are input into multi-scale feature fusion module I.
[0022] Convolutional module III receives extracted feature F3. The output feature of convolutional module III is input into multi-scale feature fusion module III. The output feature of multi-scale feature fusion module III is input into RepC3 module I. The output feature of RepC3 module I is input into CARAFE module II. The output feature of CARAFE module II is input into multi-scale feature fusion module II.
[0023] The AIFI module receives the extracted feature F4. The output feature of the AIFI module is input into convolution module IV. The output feature of convolution module IV is input into convolution module V. The output feature of convolution module V is input into CARAFE module I. The output feature of CARAFE module I is input into multi-scale feature fusion module III.
[0024] The output features of multi-scale feature fusion module I are input into RepC3 module III. RepC3 module III outputs recognition feature DF1. The output features of RepC3 module III are also input into SPD module IV. The output features of SPD module IV are input into multi-scale feature fusion module IV. Multi-scale feature fusion module IV also receives the output features of RepC3 module I and convolution module III. The output features of multi-scale feature fusion module IV are input into RepC3 module IV. RepC3 module IV outputs recognition feature DF2. The output features of RepC3 module IV are also input into SPD module V. The output features of SPD module V are input into multi-scale feature fusion module V. Multi-scale feature fusion module V also receives the output features of convolution module V. Multi-scale feature fusion module V is input into RepC3 module V. RepC3 module V outputs recognition feature DF3.
[0025] Among them, the scale of convolutional module II, convolutional module III, convolutional module IV and convolutional module V are all 1×1.
[0026] Furthermore, the RepC3 module I, RepC3 module II, RepC3 module III, RepC3 module IV, and RepC3 module V have the same structure;
[0027] RepC3 module I includes a 1×1 convolution module, a 3×3 convolution module, a batch normalization module III, a batch normalization module IV, and a ReLU activation function module III;
[0028] Both the 1×1 convolution module and the 3×3 convolution module input the same features. The output features of the 1×1 convolution module are input into the batch normalization module III, and the output features of the 3×3 convolution module are input into the batch normalization module IV. The output features of the batch normalization module III and the batch normalization module IV are summed element by element and then input into the ReLU activation function module III. The output features of the ReLU activation function module III are used as the output features of the RepC3 module I.
[0029] Accordingly, the present invention also provides a method for detecting rice ears based on the above-mentioned small target detection network, comprising the following steps:
[0030] S1. Obtain sample images of rice ears and preprocess the sample images;
[0031] S2. Input the preprocessed rice ear sample image into the small target detection network and train the small target detection network;
[0032] S3. Acquire real-time images of rice ears, preprocess the images, and input the preprocessed real-time images of rice ears into the trained small target detection network to obtain the rice ear detection results.
[0033] The beneficial effects of this invention are as follows: By using the SPD module, the spatial dimension can be compressed while retaining fine-grained information. The computational load is reduced without sacrificing feature quality, and the ability to represent detailed features such as the edge texture of small targets like rice ears is enhanced. Moreover, under the action of the feature processing module, the details of small targets in complex environments can be reconstructed and features restored, thereby effectively improving the accuracy of the final detection results and providing accurate data support for the formulation of subsequent production measures. Attached Figure Description
[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0035] Figure 1 This is a schematic diagram of the structure of the present invention.
[0036] Figure 2This is a schematic diagram of the feature extraction module structure of the present invention.
[0037] Figure 3 This is a schematic diagram of the feature processing module structure of the present invention.
[0038] Figure 4 This is a schematic diagram of the basic module structure of the present invention.
[0039] Figure 5 This is a schematic diagram of the convolutional unit structure of the present invention.
[0040] Figure 6 This is a schematic diagram of the RepC3 module structure of the present invention.
[0041] Figure 7 This is a schematic flowchart of the detection method of the present invention. Detailed Implementation
[0042] The present invention will be further described in detail below:
[0043] The present invention provides a small target detection network in complex environments, including a feature extraction module, a feature processing module, and an output module;
[0044] The feature extraction module is used to extract features from small target images acquired in complex environments and output four different extracted features.
[0045] The feature processing module is used to receive four extracted features output by the feature extraction module, process the four extracted features to generate three features to be identified, and then output them.
[0046] The output module is used to receive three features with recognition and output the small target detection result; wherein, the output module adopts an existing detection head, such as the YOLO series (v3-v11) detection head, which can process features of three different scales to output the detection result.
[0047] Specifically: the feature extraction module includes SPD module I, SPD module II, convolution unit I, convolution unit II, benchmark module I, benchmark module II, benchmark module III, benchmark module IV, and convolution module I;
[0048] The SPD module I takes a small target image as input, the output features of the SPD module I are input into convolutional unit I, the output features of convolutional unit I are input into convolutional unit II, the output features of convolutional unit II are input into SPD module II, the output features of SPD module II are input into benchmark module I, and benchmark module I outputs the extracted feature F1.
[0049] The extracted feature F1 is input into the baseline module II, and the output of the baseline module II is used to extract feature F2; the extracted feature F2 is input into the baseline module III, and the output of the baseline module III is used to extract feature F3; the extracted feature F3 is input into the baseline module IV, and the output feature of the baseline module IV is input into the convolution module I, and the output of the convolution module I is used to extract feature F4; wherein, the scale of the convolution module I is 1×1.
[0050] The convolutional unit I and convolutional unit II have the same structure;
[0051] The convolutional unit I includes a two-dimensional convolutional module I, a batch normalization module I, and a ReLU activation function module I;
[0052] The input of the two-dimensional convolution module I serves as the input of the convolution unit I. The output features of the two-dimensional convolution module I are input into the batch normalization module I. The output features of the batch normalization module I are input into the ReLU activation function module I. The output features of the ReLU activation function module I serve as the output features of the convolution unit I.
[0053] The reference modules I, II, III, and IV have the same structure.
[0054] The baseline module I includes an SPD module, a two-dimensional convolution module II, a batch normalization module II, and a ReLU activation function module II;
[0055] The input of the SPD module serves as the input of the baseline module I, and its output is connected to the input of the 2D convolution module II. The output features of the 2D convolution module II are input into the batch normalization module II, and then into the ReLU activation function module II. The output features of the ReLU activation function module II are element-wise added to the input features of the SPD module, and this sum is used as the output features of the baseline module I. The SPD module, short for Space-to-Depth + Conv, combines a space-to-depth transformation layer (SPD layer) with a non-stiff convolutional layer (stiff by 1) to achieve lossless downsampling and feature reconstruction. This preserves detailed information about small targets, enhances feature representation capabilities, and ensures the accuracy of the final detection results. The feature extraction module, with the above structure, can output extracted features in four different dimensions, thus ensuring the accuracy of the final detection results.
[0056] In this embodiment, the feature processing includes SPD module III, SPD module IV, SPD module V, AIFI module, convolution module II, convolution module III, convolution module IV, convolution module V, multi-scale feature fusion module I, multi-scale feature fusion module II, multi-scale feature fusion module III, multi-scale feature fusion module IV, multi-scale feature fusion module V, CARAFE module I, CARAFE module II, RepC3 module I, RepC3 module II, RepC3 module III, RepC3 module IV, and RepC3 module V;
[0057] The SPD module III receives the extracted feature F1, and the output feature of the SPD module III is input to the multi-scale feature fusion module I;
[0058] Convolutional module II receives the extracted feature F2. The output features of convolutional module II are input into multi-scale feature fusion module II and multi-scale feature fusion module I. The output features of multi-scale feature fusion module II are input into RepC3 module II. The output features of RepC3 module II are input into multi-scale feature fusion module I.
[0059] Convolutional module III receives extracted feature F3. The output feature of convolutional module III is input into multi-scale feature fusion module III. The output feature of multi-scale feature fusion module III is input into RepC3 module I. The output feature of RepC3 module I is input into CARAFE module II. The output feature of CARAFE module II is input into multi-scale feature fusion module II.
[0060] The AIFI module receives the extracted feature F4. The output feature of the AIFI module is input into convolution module IV. The output feature of convolution module IV is input into convolution module V. The output feature of convolution module V is input into CARAFE module I. The output feature of CARAFE module I is input into multi-scale feature fusion module III.
[0061] The output features of multi-scale feature fusion module I are input into RepC3 module III. RepC3 module III outputs recognition feature DF1. The output features of RepC3 module III are also input into SPD module IV. The output features of SPD module IV are input into multi-scale feature fusion module IV. Multi-scale feature fusion module IV also receives the output features of RepC3 module I and convolution module III. The output features of multi-scale feature fusion module IV are input into RepC3 module IV. RepC3 module IV outputs recognition feature DF2. The output features of RepC3 module IV are also input into SPD module V. The output features of SPD module V are input into multi-scale feature fusion module V. Multi-scale feature fusion module V also receives the output features of convolution module V. Multi-scale feature fusion module V is input into RepC3 module V. RepC3 module V outputs recognition feature DF3.
[0062] The scale of convolutional modules II, III, IV, and V is 1×1. The CARAFE module, short for Content-Aware ReAssembly of Features, upsamples feature maps by predicting convolutional kernels. It generates position-related reassembly kernels through a content-adaptive mechanism and uses these kernels to perform weighted aggregation on the input feature maps. This module enables large receptive field semantic aggregation, supports large-scale kernels, surpasses the sub-pixel neighborhood limitations of bilinear interpolation, and captures cross-regional contextual relationships. It features dynamic content adaptation, achieving adaptive allocation of local feature weights and overcoming the global kernel rigidity problem of deconvolution. It boasts high computational efficiency, introducing only 1 / 30th the number of parameters of deconvolution. In small target detection tasks such as rice ears, it improves mAP by 0.6% without a significant decrease in push speed and alleviates semantic misalignment in multi-scale feature fusion. This structure allows for the fusion of four extracted features to form the final recognition features, effectively ensuring the accuracy of the final result.
[0063] Among them, RepC3 module I, RepC3 module II, RepC3 module III, RepC3 module IV and RepC3 module V have the same structure;
[0064] RepC3 module I includes a 1×1 convolution module, a 3×3 convolution module, a batch normalization module III, a batch normalization module IV, and a ReLU activation function module III;
[0065] Both the 1×1 convolution module and the 3×3 convolution module input the same features. The output features of the 1×1 convolution module are input into the batch normalization module III, and the output features of the 3×3 convolution module are input into the batch normalization module IV. The output features of the batch normalization module III and the batch normalization module IV are summed element by element and then input into the ReLU activation function module III. The output features of the ReLU activation function module III are used as the output features of the RepC3 module I.
[0066] Accordingly, the present invention also provides a method for detecting rice ears based on the above-mentioned small target detection network, comprising the following steps:
[0067] S1. Obtain sample images of rice ears and preprocess the sample images;
[0068] S2. Input the preprocessed rice ear sample image into the small target detection network and train the small target detection network;
[0069] S3. Acquire real-time images of rice ears and preprocess them. Input the preprocessed real-time images of rice ears into the trained small object detection network to obtain the rice ear detection results. The image preprocessing includes denoising and geometric correction.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A small target detection network in a complex environment, characterized in that: It includes a feature extraction module, a feature processing module, and an output module; The feature extraction module is used to extract features from small target images acquired in complex environments and output four different extracted features. The feature processing module is used to receive four extracted features output by the feature extraction module, process the four extracted features to generate three features to be identified, and then output them. The output module is used to receive three identification features and output the small target detection result.
2. The small target detection network in complex environments according to claim 1, characterized in that: The feature extraction module includes SPD module I, SPD module II, convolutional unit I, convolutional unit II, benchmark module I, benchmark module II, benchmark module III, benchmark module IV, and convolutional module I; The SPD module I takes a small target image as input, the output features of the SPD module I are input into convolutional unit I, the output features of convolutional unit I are input into convolutional unit II, the output features of convolutional unit II are input into SPD module II, the output features of SPD module II are input into benchmark module I, and benchmark module I outputs the extracted feature F1. The extracted feature F1 is input into the benchmark module II, and the output of the benchmark module II is used to extract feature F2. The extracted feature F2 is input into the baseline module III, the baseline module III outputs the extracted feature F3, the extracted feature F3 is input into the baseline module IV, the output feature of the baseline module IV is input into the convolution module I, and the convolution module I outputs the extracted feature F4; wherein, the scale of the convolution module I is 1×1.
3. The small target detection network in complex environments according to claim 2, characterized in that: The convolutional unit I and convolutional unit II have the same structure; The convolutional unit I includes a two-dimensional convolutional module I, a batch normalization module I, and a ReLU activation function module I; The input of the two-dimensional convolution module I serves as the input of the convolution unit I. The output features of the two-dimensional convolution module I are input into the batch normalization module I. The output features of the batch normalization module I are input into the ReLU activation function module I. The output features of the ReLU activation function module I serve as the output features of the convolution unit I.
4. The small target detection network in complex environments according to claim 2, characterized in that: The reference modules I, II, III, and IV have the same structure. The baseline module I includes an SPD module, a two-dimensional convolution module II, a batch normalization module II, and a ReLU activation function module II; The input of the SPD module serves as the input of the baseline module I. The output of the SPD module is connected to the input of the two-dimensional convolution module II. The output features of the two-dimensional convolution module II are input into the batch normalization module II. The output features of the batch normalization module II are input into the ReLU activation function module II. The output features of the ReLU activation function module II are added element-wise to the input features of the SPD module, and the resulting feature is used as the output feature of the baseline module I.
5. The small target detection network in complex environments according to claim 2, characterized in that: The feature processing includes SPD module III, SPD module IV, SPD module V, AIFI module, convolution module II, convolution module III, convolution module IV, convolution module V, multi-scale feature fusion module I, multi-scale feature fusion module II, multi-scale feature fusion module III, multi-scale feature fusion module IV, multi-scale feature fusion module V, CARAFE module I, CARAFE module II, RepC3 module I, RepC3 module II, RepC3 module III, RepC3 module IV, and RepC3 module V; The SPD module III receives the extracted feature F1, and the output feature of the SPD module III is input to the multi-scale feature fusion module I; Convolutional module II receives the extracted feature F2. The output features of convolutional module II are input into multi-scale feature fusion module II and multi-scale feature fusion module I. The output features of multi-scale feature fusion module II are input into RepC3 module II. The output features of RepC3 module II are input into multi-scale feature fusion module I. Convolutional module III receives extracted feature F3. The output feature of convolutional module III is input into multi-scale feature fusion module III. The output feature of multi-scale feature fusion module III is input into RepC3 module I. The output feature of RepC3 module I is input into CARAFE module II. The output feature of CARAFE module II is input into multi-scale feature fusion module II. The AIFI module receives the extracted feature F4. The output feature of the AIFI module is input into convolution module IV. The output feature of convolution module IV is input into convolution module V. The output feature of convolution module V is input into CARAFE module I. The output feature of CARAFE module I is input into multi-scale feature fusion module III. The output features of multi-scale feature fusion module I are input into RepC3 module III. RepC3 module III outputs recognition feature DF1. The output features of RepC3 module III are also input into SPD module IV. The output features of SPD module IV are input into multi-scale feature fusion module IV. Multi-scale feature fusion module IV also receives the output features of RepC3 module I and convolution module III. The output features of multi-scale feature fusion module IV are input into RepC3 module IV. RepC3 module IV outputs recognition feature DF2. The output features of RepC3 module IV are also input into SPD module V. The output features of SPD module V are input into multi-scale feature fusion module V. Multi-scale feature fusion module V also receives the output features of convolution module V. Multi-scale feature fusion module V is input into RepC3 module V. RepC3 module V outputs recognition feature DF3. Among them, the scale of convolutional module II, convolutional module III, convolutional module IV and convolutional module V are all 1×1.
6. The small target detection network in complex environments according to claim 5, characterized in that: The RepC3 module I, RepC3 module II, RepC3 module III, RepC3 module IV, and RepC3 module V have the same structure; RepC3 module I includes a 1×1 convolution module, a 3×3 convolution module, a batch normalization module III, a batch normalization module IV, and a ReLU activation function module III; Both the 1×1 convolution module and the 3×3 convolution module input the same features. The output features of the 1×1 convolution module are input into the batch normalization module III, and the output features of the 3×3 convolution module are input into the batch normalization module IV. The output features of the batch normalization module III and the batch normalization module IV are summed element by element and then input into the ReLU activation function module III. The output features of the ReLU activation function module III are used as the output features of the RepC3 module I.
7. A method for detecting rice ears based on the small target detection network described in any one of claims 1-6, characterized in that: Includes the following steps: S1. Obtain sample images of rice ears and preprocess the sample images; S2. Input the preprocessed rice ear sample image into the small target detection network and train the small target detection network; S3. Acquire real-time images of rice ears, preprocess the images, and input the preprocessed real-time images of rice ears into the trained small target detection network to obtain the rice ear detection results.