Wheat Detection Method and Device Based on RGB-D Adaptive Fusion Information of Lightweight Model

Through the RGB-D adaptive fusion information detection method based on lightweight model, a binocular camera and lightweight YSNv2 network are used, combined with the cross entropy loss function and the light adaptive weight factor, the efficient and accurate detection problem of wheat ear detection in complex environments is solved, and efficient detection under light changes is achieved.

CN116778476BActive Publication Date: 2025-07-22ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310719551.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-07-22
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

The existing wheat image detection methods are difficult to achieve efficient and accurate wheat ear detection in complex field environments, especially in the face of light changes, occlusion and background interference. The traditional method has a large amount of calculation and requires high-end hardware. The deep learning-based method has high computing resources and is difficult to meet the actual application needs.

Method used

Using the RGB-D adaptive fusion information detection method based on the lightweight model, the RGB-D images of wheat ears were obtained using a binocular camera, and the detection was performed through a lightweight YSNv2 network. Combining the cross entropy loss function and the light adaptive weight factor, RGB and depth image information were fused to build a wheat ear detection model to improve the detection accuracy.

Benefits of technology

Efficient and accurate wheat ear detection is achieved in complex field environments, reducing the impact of light changes, improving detection speed and accuracy, and meeting the computing power and memory requirements of practical application scenarios, which is of great research significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778476B_ABST
    Figure CN116778476B_ABST
Patent Text Reader

Abstract

The present invention discloses a wheat detection method and device based on RGB-D adaptive fusion information of a lightweight model. The steps of the method include: 1. Using a binocular camera to obtain the RGB-D images of wheat ears in the target wheat field, and performing processing such as alignment, background removal, image enhancement, and annotation on the RGB image and the depth pseudo-color image to construct a wheat RGB-D dataset; 2. Using the YSNv2 network to train the training set to obtain the weights of the wheat RGB and depth images; 3. Using the light adaptation machine weight factor to change the contribution of the RGB detection model and output the final fusion detection result. The present invention uses RGB-D images and the lightweight YSNv2 network to detect wheat ears, which can reduce the influence of light changes on wheat detection, thereby improving the detection accuracy of wheat ears.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of field crop image detection, and specifically relates to a wheat detection method and device based on RGB-D adaptive fusion information of a lightweight model. This method can monitor and analyze the growth status of wheat, and improve the yield estimation and quality monitoring of wheat. Background Art

[0002] Wheat is one of the most important crops in the world. Accurately and timely detecting wheat ears helps improve the yield and quality of wheat. However, traditional single-modal wheat image detection methods are difficult to cope with complex field environments, such as light changes, occlusion, background interference, etc. RGB-D image detection is a technology that uses RGB and depth images for object detection. Compared with traditional RGB image detection, RGB-D image detection can obtain a real depth map, and the RGB-D fusion detection technology with little influence from the object's own color provides a data basis for detection in agricultural application scenarios.

[0003] Currently, existing crop image detection methods mainly include detection methods based on traditional machine learning algorithms and detection methods based on deep learning algorithms. Among them, the detection methods based on traditional machine learning algorithms mainly perform matching based on features such as the brightness, contrast, and color of images. This method is simple and easy to implement, but the detection accuracy is relatively low. The detection methods based on deep learning algorithms have high detection accuracy, but require a large amount of computing resources and time for training, and are difficult to meet the detection requirements.

[0004] The wheat image detection method based on fusion information has received great attention, which can reduce the false alarm rate and missed detection rate, thus reducing the need for manual intervention. However, the detection methods based on traditional methods are limited in their practical applications in smart agricultural equipment due to problems such as large computational amounts, the need for high-end hardware, and the lack of multi-source wheat data. Summary of the Invention

[0005] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes a wheat detection method and device based on RGB-D adaptive fusion information of a lightweight model, in order to efficiently and accurately detect wheat ears using RGB-D images and the lightweight YSNv2 network, reduce the influence of light changes on wheat detection, thereby improving the detection accuracy of wheat ears, and improving agricultural production efficiency and quality.

[0006] To achieve the above invention purpose, the present invention adopts the following technical solutions:

[0007] The wheat detection method based on RGB-D adaptive fusion information of a lightweight model according to the present invention is characterized by including the following steps:

[0008] Step 1: Use a binocular camera to obtain the RGB-D images of wheat ears in the target wheat field ( , ), where represents the RGB image of wheat ears, represents the depth pseudo-color image of wheat ears;

[0009] After aligning the RGB image of wheat ears and the depth pseudo-color image of wheat ears , perform image enhancement and background removal to obtain the processed RGB-D image ( , ) and label the positions of the wheat ears, thereby constructing an RGB-D dataset of wheat ears; where represents the processed RGB image of wheat ears, represents the processed depth pseudo-color image of wheat ears; let the detection box labeled at the position of the wheat ear in be denoted as the detection box labeled at the position of the wheat ear in be denoted as

[0010] Step 2: Construct the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network based on lightweight models, and process , respectively, and obtain the prediction results of the RGB image of wheat ears and the prediction results of the depth pseudo-color image of wheat ears ; where represents several detection boxes of the RGB image of wheat ears , represents several confidence levels corresponding to the detection box of the RGB wheat ear image, represents several detection boxes of the depth pseudo-color image of wheat ears , represents several confidence levels corresponding to the detection box of the wheat ear depth image;

[0011] Step 3: Construct a cross-entropy loss function:

[0012] According to the prediction results of the RGB image of wheat ears and the labeled detection box , construct the cross-entropy loss function of the YSNv2_RGB wheat ear detection network;

[0013] According to the predicted results of the wheat ear depth image and the labeled detection boxes ,construct the cross-entropy loss function of the YSNv2_D wheat ear detection network ; ;

[0014] Based on the wheat ear RGB-D dataset, use the gradient descent method to train the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network respectively, and calculate and to update the model parameters until and converge, so as to obtain the optimal wheat ear RGB image detection model and the optimal wheat ear depth image detection model ;

[0015] Step 4. Construct the dynamic weight factor of the wheat ear RGB image detection model :

[0016] (1)

[0017] In formula (1), and respectively represent the upper and lower limits of the light intensity; represents the light intensity of the wheat ear RGB image ;

[0018] Update the confidence of the detection box of the wheat ear RGB image to ;

[0019] Update the detection result of the wheat ear RGB image to ;

[0020] Step 5. RGB-D wheat ear fusion detection:

[0021] Step 5.1. Merge the detection box of the wheat ear RGB image and the detection box of the wheat ear depth image into a temporary detection box ;Set the confidence of the temporary detection box ;

[0022] Step 5.2. According to the confidence Sort all the detection boxes in descending order, and mark the detection box with the highest confidence as ; ;

[0023] Step 5.3: Calculate the intersection over union (IoU) between the detection box and each of the remaining detection boxes, and filter out all the detection boxes whose IoU exceeds the threshold;

[0024] Step 5.4: Calculate the confidence of the filtered detection boxes, and take the detection box corresponding to the maximum confidence as the final fused detection box , and take the corresponding confidence together as the fused detection result of the RGB-D wheat ear , ). .

[0025] Another feature of the wheat detection method based on a lightweight model for RGB-D adaptive fusion information according to the present invention is that the lightweight model in step 2 includes: a backbone feature extraction module, a feature pyramid fusion module based on weighted bidirectionality, and a detection box generation module;

[0026] Step 2.1: The backbone feature extraction module is successively composed of N feature extraction sub-blocks, and each feature extraction sub-block is composed of 2 Shuffle_Block modules followed by a CABlock module; wherein, each Shuffle_Block module is successively composed of a channel separation module, a common convolutional layer, a depth convolutional layer, a channel connection layer, and a ReLU activation layer; each CABlock module is successively composed of two average pooling layer modules, a two-dimensional convolutional channel fusion module, a regularization non-linear activation layer module, and a Sigmoid activation layer module;

[0027] Input the processed RGB image of the wheat ear and the depth pseudo-color image of the wheat ear into the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network respectively, and after being successively processed by several feature extraction blocks of the corresponding backbone feature extraction network, generate N-scale RGB wheat ear feature maps , , …, …, and N-scale depth wheat ear feature maps , , …, , …, ; wherein, Indicate the RGB feature map of the wheat ear at the nth scale, Indicate the depth feature map of the wheat ear at the nth scale;

[0028] Step 2.2: The weighted bidirectional pyramid fusion module includes N columns of convolutional neural network paths in parallel; among them, each column of convolutional neural network path is successively composed of a normal convolution module, an upsampling convolution module, a bidirectional pyramid module, and a depth convolution module;

[0029] The RGB feature maps of wheat ears at N scales , ,…, …, And the depth ear feature maps of N scales , ,…, ,…, Are respectively input into the corresponding pyramid fusion module, and the nth column of convolutional neural network path processes And To obtain the nth wheat ear RGB feature map And the nth wheat ear depth feature map ; Thus, the RGB feature maps of wheat ears are obtained , ,…, …, And the depth feature maps of wheat ears , ,…, ,…,

[0030] Step 2.3: The detection box generation module successively includes: a feature map fusion module and an anchor box prediction module; among them, the feature map fusion module is successively composed of a normal convolutional layer and a Sigmoid activation layer;

[0031] Will And Are respectively input into the corresponding detection box generation module, and after being processed by the normal convolutional layer in the corresponding feature map fusion module, the nth fused wheat ear RGB feature fusion map with the shape of And the nth fused wheat ear depth feature fusion map And the nth fused wheat ear depth feature fusion map ; Where C represents the number of channels, and H and W respectively represent the height and width of the feature map;

[0032] Will And The shape of is adjusted to Subsequently, it is sent into the corresponding Sigmoid activation layer for processing to obtain the nth cross-layer RGB feature map and the nth cross-layer depth feature map ;

[0033] The and are respectively input into the corresponding anchor prediction module to obtain the prediction results of the wheat ear in the wheat ear RGB image and the prediction results of the wheat ear in the wheat ear depth image .

[0034] A wheat detection device based on RGB-D adaptive fusion information of a lightweight model according to the present invention is characterized by including: a wheat ear RGB-D image acquisition module, a model construction module, a wheat ear RGB image adaptive module, and a wheat ear RGB-D detection frame fusion module;

[0035] The wheat ear RGB-D image acquisition module uses a binocular camera to acquire the wheat ear RGB-D image of the target wheat field ( , ), after aligning and , then performing image enhancement and removing the background to obtain the processed RGB-D image ( , ) and marking the position of the wheat ear, thereby constructing a wheat ear RGB-D data set;

[0036] The model construction module is used to construct the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network based on the lightweight model, and optimize the model using the cross-entropy loss function to obtain the optimal wheat ear RGB image detection model and the optimal wheat ear depth image detection model ;

[0037] The wheat ear RGB image adaptive module adjusts the contribution degree of the prediction result of the YSNv2_RGB wheat ear detection network by establishing the dynamic weight factor of the wheat ear RGB image detection model ;

[0038] The wheat ear RGB-D detection frame fusion module is used to fuse the detection results of the adjusted YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. The present invention uses a lightweight model to achieve wheat ear detection, meeting the requirements for computing power, memory, and running speed in practical application scenarios. This method is an efficient and accurate method for detecting field crop images, and has important research significance for realizing wheat ear detection in complex field environments.

[0041] 2. RGB-D wheat ear data is adopted. RGB-D image detection can obtain the depth map of the wheat ear, and is less affected by the color of the wheat ear itself, increasing the depth information of the wheat ear image. At the same time, the RGB image and the depth map come from independent sensors, having a large fault tolerance.

[0042] 3. The illumination adaptive weight is used to fuse the RGB-D detection results. When the illumination intensity is too high, the image may be overexposed, resulting in the loss of detailed information of the object; when the illumination intensity is too low, the image may be too dark, resulting in unclear outlines of the object. Using the illumination adaptive weight to adjust the fusion ratio of the RGB and depth channel detection results can effectively perform wheat ear detection under different illumination conditions, improving the detection accuracy and robustness of the model.

[0043] The present invention uses a binocular camera to obtain the color image and depth image of wheat, and adopts a lightweight model and adaptive fusion information, which can improve the detection speed and accuracy of wheat ears. The present invention can be applied to the field of wheat field growth monitoring, and has important significance for improving wheat yield estimation and production management. At the same time, the present invention can also be applied to RGB-D image detection in other fields, having a wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flow chart of a wheat detection method based on a lightweight model with RGB-D adaptive fusion information according to the present invention;

[0045] Figure 2 is a flow chart of a wheat detection device based on a lightweight model with RGB-D adaptive fusion information according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0046] In this example, a wheat detection method based on a lightweight model with RGB-D adaptive fusion information, as Figure 1 shown, includes the following steps:

[0047] Step 1. Use a binocular camera to obtain the RGB-D image of the wheat ears in the target wheat field ( , ), where represents the RGB image of the wheat ear, and represents the depth pseudo-color image.

[0048] For the wheat RGB image and the depth pseudo-color image After alignment, the image is enhanced and the background is removed to obtain the processed RGB-D image ( , ), and the positions of the wheat ears are marked, thereby constructing a wheat ear RGB-D dataset; among them, represents the processed RGB image of the wheat ear, represents the processed depth pseudo-color image; let The detection frame marked at the position of the wheat ear in is denoted as The detection frame marked at the position of the wheat ear in ;

[0049] In a specific implementation, 500 RGB and depth images of wheat ears with a resolution of 640×640 are respectively subjected to pixel alignment, image enhancement, and background removal.

[0050] In a specific implementation, the pixel alignment process includes: using the internal and external parameter matrix information, running the external calibration for the binocular camera, and combining the external parameter parameters of the color image and the depth image into a final transformation matrix, which will be applied to the depth image to project it to the position overlapping with the color image, so as to obtain the internal parameter matrix of the camera using equations (1)-(4) , , the external parameter matrix , :

[0051] (1)

[0052] (2)

[0053] (3)

[0054] (4)

[0055] Different from the RGB wheat ear dataset, the original depth information directly obtained by the binocular camera cannot be directly used for training. The directly obtained depth image expresses the distance between each pixel in the image and the detection through black and white changes, that is, the so-called grayscale image.

[0056] In a specific implementation, the process of image enhancement includes steps such as performing pseudo-coloring, applying filtering, and hole filling on the grayscale image. The specific operation is as follows: Calculate the depth value of each pixel in the depth image, map the depth value of each pixel to a relative color, use different colors to represent different depth values, and apply the changed colors to draw the color of each pixel in the depth image. Specifically, code is written using the OpenCV framework to perform pseudo-color rendering on all depth images.

[0057] In a specific implementation, the background removal is as follows: Let the depth value of a certain point in the depth image be d, then there is:

[0058] (5)

[0059] In this way, the effective wheat ear information at close range can be retained, background interference can be reduced, and the effect of removing the background is achieved. Subsequently, the annotation of the detection box is carried out.

[0060] Step 2: Construct the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network based on lightweight models, and process respectively, to obtain the prediction results of the wheat ear RGB image and the prediction results of the wheat ear depth pseudo-color image ; among them, represents several detection boxes of the wheat ear RGB image ; represents several confidence levels corresponding to the detection boxes of the RGB wheat ear image ; represents several detection boxes of the wheat ear depth pseudo-color image ; represents several confidence levels corresponding to the detection boxes of the wheat ear depth image ; represents several confidence levels corresponding to the detection boxes of the wheat ear depth image ;

[0061] Step 2.1: The backbone feature extraction module is successively composed of N feature extraction sub-blocks, and each feature extraction sub-block is composed of 2 Shuffle_Block modules followed by a CABlock module; among them, each Shuffle_Block module is successively composed of a channel separation module, a common convolutional layer, a depth convolutional layer, a channel connection layer, and a ReLU activation layer; each CABlock module is successively composed of two average pooling layer modules, a two-dimensional convolutional channel fusion module, a regularization non-linear activation layer module, and a Sigmoid activation layer module;

[0062] The processed wheat ear RGB image And the pseudo-color image of wheat ear depth They are respectively input into the small YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network, and after being processed by several feature extraction blocks of the corresponding backbone feature extraction network in sequence, N kinds of scale RGB wheat ear feature maps of wheat ears are generated respectively , ,…, …, And N kinds of scale depth wheat ear feature maps of wheat ears , ,…, ,…, ; Among them, represents the RGB feature map of the nth scale of wheat ears, represents the depth feature map of the nth scale of wheat ears;

[0063] In a specific implementation, the feature extraction sub-block is composed of 6 Shuffle_Block modules and 3 CABlock modules connected in series in sequence, and a CABlock is connected after every 2 Shuffle_Block;

[0064] The processed RGB image and the processed depth image are respectively input into the RGB-D wheat ear detection network YSNv2, and after being processed by the Shuffle_Block module of the wheat ear backbone feature extraction network in sequence, the RGB image feature of the wheat ear is output and the depth image feature of the wheat ear , and then after being processed by the CABlock module respectively, the RGB wheat ear feature and the depth wheat ear feature are obtained; After being processed by 6 Shuffle_Block modules and 3 CABlock modules, 3 kinds of scale RGB wheat ear feature maps are generated respectively , , and the depth wheat ear feature map , , . Among them, the feature map of the first scale is processed according to formula (6):

[0065] (6)

[0066] When entering the backbone part of the YSNv2 network, the image will be processed by multiple convolutional layers, Shuffle_Block modules and CABlock modules.

[0067] First, the image will pass through a convolutional layer for preliminary feature extraction of the image. Assume the size of the convolutional kernel is 2, the stride is 2, and the number is 32, then we have:

[0068] (7)

[0069] Among them, and respectively represent the feature maps obtained after the RGB image and the depth image pass through the convolutional layer. Their sizes are 320×320×32 respectively.

[0070] Then, the image will enter the Shuffle_Block module. Assume the size of the depth convolutional kernel in the Shuffle_Block module is 3, the stride is 1, and the number is 128, then we have:

[0071] (8)

[0072] Among them, and respectively represent the feature maps obtained after the RGB image and the depth image pass through the Shuffle_Block module. Their sizes are 318×318×128 respectively.

[0073] After being processed by the Shuffle_Block module, assume the output dimension of the multi-layer perceptron in the CA module is 512, then we have:

[0074] (9)

[0075] Among them, and respectively represent the feature maps obtained after the RGB image and the depth image pass through the CA module. Their sizes are 318×318×512 respectively.

[0076] Step 2.2: The weighted bidirectional pyramid fusion module, including N columns of convolutional neural network paths in parallel; among them, each column of convolutional neural network path is successively composed of a common convolutional module, an upsampling convolutional module, a bidirectional pyramid module, and a depth convolutional module;

[0077] The RGB feature maps of wheat ears at N scales , , …, …, and the depth feature maps of wheat ears at N scales , , …, , …, are respectively input into the corresponding pyramid fusion modules, and the nth column of convolutional neural network path respectively processes and Process to obtain the RGB feature map of the nth wheat ear and the depth feature map of the nth wheat ear ; thus obtaining the RGB feature map of the wheat ear , ,…, …, and the depth feature map of the wheat ear , ,…, ,…,

[0078] In a specific implementation, the RGB ear feature maps of 3 scales , , and the depth ear feature map , , are respectively input into the pyramid fusion module, and are processed by 3 columns of convolutional neural network paths respectively for , , and and , , to obtain the RGB feature map of the wheat ear , , and the depth feature map of the wheat ear , , ; among them, The sizes are 318×318×128 respectively; The size is 318×318×256; The size is 318×318×512; The sizes are 318×318×128 respectively; The size is 318×318×256; The size is 318×318×512;

[0079] Step 2.3: The detection box generation module, which successively includes: a feature map fusion module and an anchor box prediction module; among them, the feature map fusion module is successively composed of a common convolutional layer and a Sigmoid activation layer;

[0080] Input and respectively into the corresponding detection box generation modules, and after being processed by the common convolutional layer in the corresponding feature map fusion module, the nth fused RGB feature fusion map of the wheat ear with the shape of and the nth fused depth feature fusion map of the wheat ear are obtained respectively where C represents the number of channels, and H and W represent the height and width of the feature map respectively;

[0081] Adjust the shapes of and to and then send them into the corresponding Sigmoid activation layer for processing to obtain the nth cross-layer RGB feature map and the nth cross-layer depth feature map ;

[0082] Input and into the corresponding anchor prediction modules respectively to obtain the prediction results of the wheat ear in the RGB image of the wheat ear and the prediction result of the wheat ear in the depth image of the wheat ear .

[0083] In a specific implementation, input 3 RGB wheat ear features , , and 3 depth wheat ear features , , into the detection box generation module respectively, and after being processed by the feature map fusion module, obtain 1 fused RGB wheat ear feature fusion map and depth feature fusion map .

[0084] and The feature map of the first scale is 318×318×128, the feature map of the second scale is 318×318×256, and the feature map of the third scale is 318×318×512.

[0085] Adjust the shapes of and of the first scale to and then send them into the Sigmoid activation layer for processing to obtain the cross-layer RGB feature map of the wheat ear and the cross-layer depth feature map of the wheat ear ; The other two scales have the same processing process.

[0086] Input and into the anchor prediction modules respectively to obtain the prediction results of the wheat ear in the RGB image of the wheat ear and the prediction result of the wheat ear in the depth image of the wheat ear .

[0087] Prediction result of the RGB image of the wheat ear And the prediction results of deep wheat ear image The first 4 values are the coordinates of the detection box , the fifth value is the confidence of the detection box .

[0088] Step 3: Construct the cross entropy loss function:

[0089] Based on the wheat ear RGB image The prediction results And the marked detection box , construct the cross entropy loss function of the YSNv2_RGB wheat ear detection network ;

[0090] Based on the depth image of wheat ears The prediction results And the marked detection box , construct the cross entropy loss function of the YSNv2_D wheat ear detection network ;

[0091] Based on the wheat ear RGB-D dataset, the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network are trained using the gradient descent method, and the corresponding and To update the model parameters until and Until convergence, the optimal wheat ear RGB image detection model after training is obtained. and the optimal wheat ear depth image detection model ;

[0092] Step 4: Use formula (10) to build a wheat ear RGB image detection model Dynamic weight factor :

[0093] (10)

[0094] In formula (10), and Respectively represent the upper and lower limits of light intensity; Represents an RGB image of wheat ears Light intensity;

[0095] Wheat ears RGB image The detection box The confidence of is updated to ;

[0096] Wheat ears RGB image Test results Updated to ;

[0097] In a specific implementation, it is measured that in the field under direct sunlight on a sunny day, the illuminance can reach 10 wLx, and it is lower than 1000 Lx at night. Set 100000, . .

[0098] Prediction results of the YSNv2 (RGB) model for RGB images of wheat ears will be affected by the environmental light intensity Set and respectively represent the maximum light intensity and the minimum light intensity when is empty, specifically according to equations (11) and (12):

[0099] (11)

[0100] (12)

[0101] In this embodiment, the current environmental illuminance is measured .

[0102] RGB wheat ear image detection results , ≈0.1353, where . The weighted confidence =0.10.

[0103] In the actual process of wheat ear detection in the field, when the light intensity changes, the YSNv2_RGB wheat ear detection network will dynamically adjust the prediction weights according to the above formula . When the light intensity changes from increases to , the prediction weight will gradually increase from 0 to the maximum value, and then gradually decrease to 0. In this example, too high environmental light causes a significant decrease in the contribution degree of the YSNv2_RGB wheat ear detection network model, and the fusion result will rely more on the prediction results of the YSNv2_D wheat ear detection network model. This dynamic adjustment can effectively adapt to different lighting conditions and provide strong support for wheat ear detection.

[0104] Step 5, RGB-D wheat ear fusion detection:

[0105] Step 5.1, Combine the detection frame of the RGB image of the wheat ear and the detection frame of the depth image of the wheat ear into a temporary detection frame ; Set up a temporary detection box confidence ;

[0106] Step 5.2, According to the confidence for all the detection boxes in are sorted in descending order, and the detection box with the highest confidence is denoted as ;

[0107] Step 5.3, Calculate the intersection over union (IoU) between the detection box and each of the remaining detection boxes, and filter out all the detection boxes whose IoU exceeds the threshold;

[0108] Step 5.4, Calculate the confidence of the filtered detection boxes, and take the detection box corresponding to the maximum confidence as the final fused detection box , and take the corresponding confidence together as the fused detection result of RGB-D wheat ears , ). .

[0109] In this embodiment, the detection result of the RGB wheat ear image , ≈0.1353, where . The weighted confidence =0.10.

[0110] In this embodiment, the prediction result of the wheat ear depth image is , and it can be calculated that the IoU between the two is 0.81. According to , , filter out the detection boxes with an IoU higher than 0.8 and a smaller confidence. Obviously, . Add the prediction result of the wheat ear depth image to the final RGB-D wheat ear fused detection result .

[0111] In this embodiment, as Figure 2 shown, a wheat detection device based on a lightweight model for RGB-D adaptive fusion information includes: a wheat ear RGB-D image acquisition module, a model construction module, a wheat ear RGB image adaptive module, and a wheat ear RGB-D detection box fusion module;

[0112] The wheat ear RGB-D image acquisition module uses a binocular camera to acquire the wheat ear RGB-D image of the target wheat field( , ), for and After alignment, perform image enhancement and remove the background to obtain the processed RGB-D image ( , ) and mark the positions of the wheat ears, thereby constructing a wheat ear RGB-D dataset;

[0113] Model construction module, used to construct the above-mentioned YSNv2_RGB wheat ear detection network and YSNv2_D wheat ear detection network based on lightweight models, optimize the model using the cross-entropy loss function, and obtain the optimal wheat ear RGB image detection model after training and the optimal wheat ear depth image detection model ;

[0114] Wheat ear RGB image adaptive module, by establishing the dynamic weight factor of the wheat ear RGB image detection model , to adjust the contribution degree of the prediction result of the YSNv2_RGB wheat ear detection network;

[0115] Wheat ear RGB-D detection box fusion module, used to fuse the detection results of the adjusted YSNv2_RGB wheat ear detection network and YSNv2_D wheat ear detection network.

Claims

1. A wheat detection method based on RGB-D adaptive fusion information of a lightweight model, characterized in that It includes the following steps: Step 1: Use a binocular camera to obtain the RGB-D images of wheat ears in the target wheat field ( , ), where represents the RGB image of wheat ears, and represents the depth pseudo-color image of wheat ears; For the RGB image of the wheat ear and the depth pseudo-color image of the wheat ear After alignment, the image is enhanced and the background is removed to obtain the processed RGB-D image ( , ), and the position of the ear is marked, thereby constructing a wheat ear RGB-D dataset; where represents the processed RGB image of the wheat ear, represents the processed depth pseudo-color image of the wheat ear; let The detection frame marked at the position of the ear in is denoted as , The detection frame marked at the position of the ear in is denoted as ; Step 2: Build the YSNv2_RGB wheat ear detection network and YSNv2_D wheat ear detection network based on the lightweight model, and , After processing, the corresponding RGB image of wheat ears is obtained The prediction results and wheat ears deep pseudo-color image The prediction results ;in, Represents an RGB image of wheat ears Several detection boxes of Represents the detection frame of the RGB wheat ear image The corresponding confidence levels are: Represents a pseudo-color image of wheat ears in depth Several detection boxes of Represents the wheat ear depth image detection frame The corresponding confidence levels are as follows; wherein the lightweight model includes: a backbone feature extraction module, a feature pyramid fusion module based on weighted bidirectionality, and a detection box generation module; Step 3: Construct a cross-entropy loss function: According to the RGB images of wheat ears prediction results and the marked detection boxes , construct the cross-entropy loss function of the YSNv2_RGB wheat ear detection network ; According to the depth image of wheat ears prediction results and the marked detection boxes , construct the cross-entropy loss function of the YSNv2_D wheat ear detection network ; Based on the wheat ear RGB-D dataset, the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network are respectively trained using the gradient descent method, and and are calculated accordingly to update the model parameters until and converge, so as to obtain the optimal trained wheat ear RGB image detection model and the optimal wheat ear depth image detection model ; Step 4. Use Equation (1) to construct a wheat ear RGB image detection model of the dynamic weight factor : (1) In formula (1), and respectively represent the upper and lower limits of the light intensity; represents the light intensity of the RGB image of the wheat ear ; The detection box of the RGB image of wheat ears is updated to the confidence level of ; Update the detection result of the RGB image of wheat ears to ; ​ Step 5: RGB-D wheat ear fusion detection: Step 5.1: Combine the detection box of the RGB image of the wheat ear and the detection box of the depth image of the wheat ear into a temporary detection box ; Set the confidence level of the temporary detection box ; ; Step 5.

2. According to the confidence level sort all the detection boxes in descending order, and mark the detection box with the highest confidence level as ; Step 5.3, calculate the detection box Calculate the intersection over union (IoU) between the detection box and each of the remaining detection boxes, and filter out all detection boxes with an IoU exceeding the threshold; Step 5.4: Calculate the confidence of the filtered detection boxes, and use the detection box corresponding to the maximum confidence as the final fused detection box , and use the corresponding confidence together as the fused detection result of the RGB-D wheat ear , ). .

2. The wheat detection method based on RGB-D adaptive fusion information of a lightweight model according to claim 1, characterized in that, The lightweight model in Step 2 includes: a backbone feature extraction module, a weighted bidirectional feature pyramid fusion module, and a detection box generation module; Step 2.1: The backbone feature extraction module is sequentially composed of N feature extraction sub-blocks, and each feature extraction sub-block is composed of 2 Shuffle_Block modules followed by a CABlock module; among them, each Shuffle_Block module is sequentially composed of a channel separation module, a common convolutional layer, a depth convolutional layer, a channel connection layer, and a ReLU activation layer; each CABlock module is sequentially composed of two average pooling layer modules, a two-dimensional convolutional channel fusion module, a regularization non-linear activation layer module, and a Sigmoid activation layer module; Input the processed RGB image of wheat ears and the depth pseudo-color image of wheat ears into the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network respectively. After being processed by several feature extraction blocks of the corresponding backbone feature extraction network in sequence, N kinds of scale RGB wheat ear feature maps , ,…, …, and N kinds of scale depth wheat ear feature maps , ,…, ,…, are generated respectively; where represents the RGB wheat ear feature map of the nth scale, represents the depth wheat ear feature map of the nth scale; Step 2.2: The weighted bidirectional pyramid fusion module includes N columns of convolutional neural network paths in parallel; among them, each column of convolutional neural network paths is sequentially composed of a common convolutional module, an upsampling convolutional module, a bidirectional pyramid module, and a depth convolutional module; RGB feature maps of wheat ears at N scales , ,…, …, and depth feature maps of wheat ears at N scales , ,…, ,…, are respectively input into the corresponding pyramid fusion modules, and the and are processed by the nth column convolutional neural network path to obtain the nth RGB feature map of wheat ears and the nth depth feature map of wheat ears ; thus, RGB feature maps of wheat ears , ,…, …, and depth feature maps of wheat ears , ,…, ,…, ; Step 2.3: The detection box generation module sequentially includes: a feature map fusion module and an anchor box prediction module; among them, the feature map fusion module is sequentially composed of a common convolutional layer and a Sigmoid activation layer; Will and They are input into the corresponding detection frame generation module respectively, and after being processed by the ordinary convolution layer in the corresponding feature map fusion module, the corresponding shape is The nth fused wheat ear RGB feature fusion image The deep feature fusion image of wheat ears after the nth fusion ; Where C represents the number of channels, H and W represent the height and width of the feature map respectively; Adjust the shapes of and to and then send them into the corresponding Sigmoid activation layers for processing to obtain the nth cross-layer RGB feature map and the nth cross-layer depth feature map ; Input and into the corresponding anchor prediction modules respectively, and obtain the prediction results of wheat ear in the RGB image of wheat ear and the prediction results of wheat ear in the depth image of wheat ear .

3. A wheat detection device based on an RGB-D adaptive fusion information of a lightweight model, characterized in that, It includes: A wheat ear RGB-D image acquisition module, a model construction module, a wheat ear RGB image adaptive module, and a wheat ear RGB-D detection box fusion module; The wheat ear RGB-D image acquisition module uses a binocular camera to acquire the RGB-D images of wheat ears in the target wheat field ( , ). After aligning and , the image is enhanced and the background is removed to obtain the processed RGB-D image ( , ), and the positions of the wheat ears are marked, thereby constructing a wheat ear RGB-D dataset; The model construction module is used to construct the YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network based on the lightweight model in the wheat detection method for adaptively fusing RGB-D information based on the lightweight model described in claim 2, and optimize the model using the cross-entropy loss function to obtain the optimal trained wheat ear RGB image detection model and the optimal wheat ear depth image detection model ; The wheat ear RGB image adaptive module adjusts the contribution degree of the prediction result of the YSNv2_RGB wheat ear detection network by establishing a dynamic weight factor for the detection model of wheat ear RGB images. ; The wheat ear RGB-D detection box fusion module is used to fuse the detection results of the adjusted YSNv2_RGB wheat ear detection network and the YSNv2_D wheat ear detection network.

Citation Information

Patent Citations

  • RGB-D saliency target detection method based on deep learning

    CN113159068A

  • Wheat ear detection and counting method based on YOLOv5 improvement

    CN116152684A