Tea Canopy Depth Distribution Detection Method Based on Computer Vision and Deep Learning
Through the tea puff depth distribution detection method based on computer vision and deep learning, the problem of inaccurate three-dimensional distribution recognition of tea dragons in tea garden image recognition is solved, and efficient perception of tea garden information is achieved, and reliable information support is provided for the automatic navigation of tea pickers.
Patent Information
- Application Number
- CN202211107167.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-13
AI Technical Summary
The prior art has not fully utilized deep learning in tea garden image recognition, and cannot accurately identify the three-dimensional distribution of tea leaves, resulting in inaccurate navigation of tea pickers and lack of a unified applicable model.
The tea-wool depth distribution detection method based on computer vision and deep learning is adopted, and images are collected and marked by RGB cameras. The single-input and double-output tea-wool image segmentation model is used, and the backbone feature extraction network and feature fusion network are combined to extract tea-wool centerline and tea garden area information, and the bipartite graph matching algorithm is used to calculate the tea-wool depth distribution.
It realizes high-accuracy detection of the depth distribution of tea puppet area in Tea Long, which is suitable for intelligent agricultural equipment navigation in tea gardens, and provides reliable tea garden information to support automatic navigation of tea pickers.
Smart Images

Figure CN115512106B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically designs a method for detecting the depth distribution of tea canopies based on computer vision and deep learning. Background Art
[0002] China is the origin and the largest producer of tea. Tea picking is the most labor-consuming field operation in the tea industry. The labor for tea picking generally accounts for more than 50% of the annual management labor. With the continuous aging of China's population, the realization of automated tea picking is an inevitable trend for future development. Vigorously promoting the mechanization, automation, and intelligentization of the tea industry is also a development requirement for further improving the tea yield and quality. Intelligent picking can not only reduce labor costs and labor intensity, but also reduce the phenomena of missed picking and random picking, and improve the quality and yield of tea.
[0003] Existing technologies mostly focus on the research of the tea picking components of tea picking machines. To achieve the automation and intelligentization of the tea industry, the automatic navigation of tea picking machines is essential. Using computer vision technology can accurately and quickly identify the tea ridge areas in the tea garden, which provides very crucial information for the automatic navigation of tea picking machines, the adjustment of the cutting knife height and bending angle of tea picking machines, etc.
[0004] Compared with traditional computer vision technologies (threshold method, region growing algorithm, etc.), the computer vision technology based on deep learning can greatly improve the recognition accuracy of tea garden areas, which has important research value and wide practical significance for realizing the automatic navigation of tea picking machines. However, this technology has not been deeply studied in the existing tea garden image recognition research. Most research only focuses on the pixel information (color, etc.) in the image, and rarely pays attention to the semantic information of tea ridges (morphological features, depth distribution features, etc. of tea canopies), which cannot provide reliable tea garden information for tea picking machines during operation. In addition, there are also problems such as difficulty in annotating tea garden images and lack of a unified applicable model for tea garden images. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology, and propose a method for detecting the depth distribution of tea canopies based on computer vision and deep learning. This method can accurately identify the three-dimensional distribution of tea ridges in different environmental conditions and different tea gardens, and provide a benchmark model for the detection of tea ridges in tea gardens. This method has good applicability and robustness, and lays a foundation for the information perception of the automatic navigation technology of tea picking machines.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A method for detecting the depth distribution of tea canopies based on computer vision and deep learning, comprising the following steps:
[0008] S1. Use an RGB camera to collect field tea garden images, mark the tea area and tea clump centerline positions in the images, and create a tea garden image dataset;
[0009] S2. Using a deep learning-based tea ridge image segmentation model, the tea garden image dataset described in step S1 is used for training until the tea ridge image segmentation model converges; according to the tea ridge image segmentation model, an RGB camera is input to capture the original image to be detected, and the tea ridge centerline and tea garden area segmentation image contained in the image are output;
[0010] The tea shed image segmentation model adopts a single-input, dual-output forward propagation flow, including a backbone feature extraction network, a feature fusion network, and two pre-designed detection branches. The first branch of the detection branch is based on an attention and position encoding mechanism, and is composed of a row-column attention module, a learnable position encoding module, and a multi-layer perceptron module. It is used to extract semantic information related to the centerline of the tea shed in the image and output the centerline of the tea shed in the image. The second branch is based on an encoding-decoding mechanism, and is composed of a multi-head self-attention module, an upsampling module, and a fully connected layer with an output channel of 2. It is used to extract foreground information of the tea shed area in the image and output a semantic segmentation image of the tea shed area.
[0011] S3. Based on the segmented image of the tea garden area in step S2, a connected domain of the segmented image of the tea garden area is extracted, and the connected domain is matched with the center line of the tea head in step S2 using a bipartite graph matching algorithm to obtain a matching relationship between the center line of the tea head and the tea garden area. Based on the matching relationship between the center line of the tea head and the tea garden area, the depth distribution characteristics of the tea head are calculated and mapped into the segmented image of the tea garden area to complete the detection of the depth distribution of the tea head.
[0012] The tea garden image dataset in step S1 consists of three parts: original images, tea garden area image labels, and tea head centerline data labels, wherein the tea garden area image labels and tea head centerline data labels are drawn independently of each other; during the acquisition process, the camera is at an angle of 45° to 90° with the horizontal direction, and the vertical height between the camera and the highest point of the tea head is maintained within 0.5m to 1m to ensure that the camera field of view includes the tea head extending outward along the camera's viewing angle.
[0013] The tea vine image segmentation model in step S2:
[0014] (1) The backbone feature extraction network and feature fusion network extract features from the original image data and generate a high-dimensional feature map of the image data;
[0015] (2) The first branch of the detection branch consists of a row-column attention module, a learnable position encoding module, and a multi-layer perceptron module. This branch extracts the row-column features from the high-dimensional feature map, uses the learnable position encoding to map the row-column features to the tea bush centerline features, and finally generates the tea bush centerline through the multi-layer perceptron.
[0016] (3) The second branch of the detection branch consists of a multi-head self-attention module, an upsampling module, and a fully connected layer with an output channel of 2. This branch first uses the multi-head self-attention module to enhance the importance of foreground information in the high-dimensional feature map, then uses the upsampling module to decode and restore the high-dimensional feature map to the original image size, and finally classifies the foreground and background of all pixel points in the image through the fully connected layer to obtain the semantic segmentation image of the tea ridge area.
[0017] Calculate the tea bush depth distribution feature in step S3:
[0018] S31. Extract the connected components from the segmented image of the tea ridge area in the step, obtain the category, area, center point position, circumscribed rectangle, and edge contour of all connected components in the tea ridge area, and remove the connected components with an area smaller than S.
[0019] S32. Use the bipartite graph matching algorithm to match the tea bush centerline in step S2 with all connected components, and remove the unmatched tea bush centerlines to obtain the successfully matched tea ridge area.
[0020] S33. Traverse all the successfully matched tea ridge areas in step S32, create a tea bush depth mask. The gray value of the pixel points at the tea bush centerline position in the mask is the maximum value, and the gray value of the pixel points at the left and right edge positions of the tea ridge connected component is the minimum value. The size of the gray value in the mask represents the relative height of the tea bush from the ground. The calculation steps for the left and right edges of the tea ridge connected component are as follows:
[0021] S331. Establish a coordinate system with the upper left corner of the image as the origin, the horizontal right direction as the x-axis, and the vertical downward direction as the y-axis. Take out the width w, height h, and the upper left vertex P(x, y) of the circumscribed rectangle of the tea ridge, and calculate that the upper edge of the tea ridge is y and the lower edge is y + h.
[0022] S332. Traverse all the edge points of the tea ridge connected component, calculate the distance d between the ordinate and the upper and lower edges of the tea ridge. If d is greater than h / 100, then determine that this edge point is a left and right edge point.
[0023] S333. Aggregate all the left and right edge points to obtain the left and right edge positions of the tea ridge connected component.
[0024] S34. Check whether there is an un-matched tea ridge connected region. If so, judge the relative position relationship between the center point position of the connected region and the vertical center line of the image. If the center point position is on the left side of the vertical center line of the image, construct a tea bush depth mask with the right edge of the connected region as the minimum gray value and the left edge as the maximum gray value; if the center point position is on the right side of the vertical center line of the image, construct a tea bush depth mask with the right edge of the connected region as the maximum gray value and the left edge as the minimum gray value.
[0025] S35. After splicing the tea bush depth masks constructed in step S33 and step S34, it is the depth distribution feature of the tea bush in the input image.
[0026] Advantages of the present invention:
[0027] The tea bush depth distribution detection method constructed by the present invention can identify the depth distribution of the tea bush area in the tea ridge, has high detection accuracy and good field applicability, and can be applied to the navigation of intelligent agricultural equipment in tea gardens. Description of the drawings
[0028] Figure 1 It is a flow chart of the implementation steps of the technical solution of the present invention.
[0029] Figure 2 It is the original image, tea ridge area label map and tea bush center line label map of the tea garden dataset of the present invention.
[0030] Figure 3 It is the overall structure diagram of the tea ridge image segmentation model based on deep learning of the present invention.
[0031] Figure 4 It is the output result predicted by the deep learning model of the present invention according to the original image, where the white line is the predicted result of the tea bush center line and the green area is the predicted area of the tea ridge.
[0032] Figure 5 It is the tea ridge depth distribution image obtained by the tea ridge depth detection method of the present invention based on the predicted tea ridge and tea bush center line, where the higher the brightness of the green area, the higher its height from the ground. Detailed implementation manners
[0033] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the implementation manners of the present invention are not limited thereto.
[0034] As Figure 1 shown, the tea bush depth distribution detection method based on computer vision and deep learning includes the following steps:
[0035] S1. Use an RGB camera to collect images of the tea garden in the field and label the tea ridges in the images to make a tea garden image dataset.
[0036] In this example, the specific ones are:
[0037] S11. Use an RGB camera to collect tea garden videos of 5 tea gardens at different time periods and in different weather conditions. The camera frame rate is set to 25fps / s, and an image is captured every 5 frames, resulting in a total of 2,000 original images of the tea garden. During the acquisition process, the camera is at a 60° angle to the horizontal direction, and the vertical height between the camera and the highest point of the tea bush is maintained at 0.5m; S12. The image size is uniformly adjusted to 960×540 pixels, and the collected tea garden images are annotated with the tea bush centerline and tea ridge area using LabelMe annotation software. The tea bush centerline is marked with 72 coordinate points (x, y), and the tea ridge area is marked with a binary image mask of the same size as the original image; an example of annotated data is shown below. Figure 2 As shown;
[0038] S13. Randomly divide the labeled data into a training set (1400 images), a validation set (400 images), and a test set (200 images) in a ratio of 7:2:1 to obtain the tea garden image dataset required for this example.
[0039] S2. Use a deep learning-based tea ridge image segmentation model and train it on a NVIDIA A6000 graphics processor using the tea garden dataset described in step S1 until the tea ridge image segmentation model converges. According to the tea ridge image segmentation model, the RGB camera is input to capture the original image, and the tea ridge centerline and tea garden area segmentation image contained in the image are output; Figure 4 As shown in the output of this example, the four white lines are the predicted results of the tea bush center line, and the five gray areas are the predicted tea ridge areas;
[0040] In this example, if Figure 3 As shown in the figure, the tea grove image segmentation model based on deep learning is as follows:
[0041] S21. ResNet50 is used as the backbone network for feature extraction. Image data is fed into ResNet50 to extract high-dimensional features. The feature pyramid module is then connected to fuse multi-scale feature information to obtain a three-dimensional high-dimensional feature map of the image.
[0042] S22. Branch A, used to output the tea-bush centerline, comprises a row-column attention module, a position encoding module, and two multi-layer perceptrons. The row-column attention module in branch A consists of a six-layer multi-head attention module, the position encoding module is a learnable two-dimensional position encoding module, and the multi-layer perceptron comprises a fully connected layer with an output dimension of 2 for category prediction and a fully connected layer with an output dimension of 144 for position prediction. The high-dimensional feature map enters branch A, which outputs the 72 coordinates of the tea-bush centerline and the confidence score for each tea-bush centerline.
[0043] Branch B for outputting the segmented image of the tea garden area. The multi-head self-attention module in Branch B is composed of 6 layers of transformer encoder layer and transformer decoder layer. Each transformer module includes 8 heads of attention. The upsampling module includes an anti-pooling layer and two convolutional layers. The anti-pooling layer uses bilinear interpolation method for upsampling. The convolutional kernel of the convolutional layer is 3×3, the stride is 1, and the dilation is 0. The high-dimensional feature map enters Branch B and outputs a tea ridge segmentation mask with the same size as the original image. Each pixel value represents the confidence that the pixel point at that position is in the tea ridge area.
[0044] S3. According to the tea bush center line and the segmented image of the tea garden area described in step S2, calculate the tea bush depth distribution characteristics and map them to the segmented image of the tea garden area to complete the detection of the tea bush depth distribution.
[0045] In this example, the specific steps for calculating the tea bush depth distribution characteristics are as follows:
[0046] S31. Extract the connected components of the tea ridge area to obtain the category, area, center point position, and edge contour of all connected components of the tea ridge area. Remove the connected components with an area smaller than S. In this example, the image resolution is 1920*1080, and the value of S is set to 300;
[0047] S32. Use the bipartite graph matching algorithm to match the tea bush center line described in step S2 with all connected components, and remove the unmatched tea bush center lines to obtain the successfully matched tea ridge areas. In this example, the bipartite graph matching algorithm uses the Hungarian algorithm to calculate the matching matrix, and the value in the matching matrix M is the distance from the center point P of the connected component to the tea bush center line;
[0048] S33. Traverse all the successfully matched tea ridge areas obtained in step S32, create a tea bush depth mask, where the pixel value at the position of the tea bush center line in the mask is the maximum value, and the pixel values at the left and right edge positions of the tea ridge connected component are the minimum values. The size of the gray value in the mask represents the relative height of the tea bush from the ground. The calculation steps for the left and right edges of the tea ridge connected component are as follows:
[0049] S331. Establish a coordinate system with the upper left corner of the image as the origin, the horizontal right direction as the x-axis, and the vertical downward direction as the y-axis. Take out the width w, height h, and the upper left vertex P(x, y) of the circumscribed rectangle of the tea ridge, and calculate that the upper edge of the tea ridge is y, and the lower edge is y+h;
[0050] S332. Traverse all the edge points of the tea ridge connected domain, calculate the distances d between their vertical coordinates and the upper and lower edges of the tea ridge. If d is greater than h / 100, then count this edge point as a left or right edge point;
[0051] S333. Aggregate all the left and right edge points to obtain the left and right edge positions of the tea ridge connected domain;
[0052] S34. Check whether there is an un-matched tea ridge connected domain. If so, judge the relative position relationship between the center point position of the connected domain and the vertical center line l of the image. If the center point position is on the left side of the vertical center line of the image, construct a tea bush depth mask with the right edge of the connected domain as the minimum gray value and the left edge as the maximum gray value; if the center point position is on the right side of the vertical center line of the image, construct a tea bush depth mask with the right edge of the connected domain as the maximum gray value and the left edge as the minimum gray value. In this example, the maximum gray value is set to 255 and the minimum value is set to 20;
[0053] S35. After splicing the tea bush depth masks constructed in step S33 and step S34, it is the depth distribution feature of the tea bush in the input image, as Figure 5 shown.
[0054] The dataset proposed by the present invention is established for the ridge features of tea ridges in the tea garden. The training dataset can highlight the semantic features of the tea ridges, so the design, training and testing of the tea ridge image segmentation model can be carried out.
[0055] The tea ridge image segmentation model proposed by the present invention adopts a single-input and dual-output workflow, and outputs the center line information of the tea bush and the tea bush area segmentation image through two different branches respectively, which can effectively extract the semantic information of the tea ridges in the tea garden, and this is crucial for the visual navigation process of the tea picking machine.
[0056] The tea bush depth detection method proposed by the present invention can convert the planar tea ridge information in the image into the three-dimensional depth distribution information of the tea bush surface, which can facilitate the selection of the height and angle of the cutting tool in the picking operation of the tea picking machine.
[0057] It should also be noted that in this specification, terms such as "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0058] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.
Claims
1. A method for detecting the depth distribution of tea canopies based on computer vision and deep learning, characterized in that, It includes the following steps: S1. Use an RGB camera to collect images of the field tea garden, annotate the tea ridge area and the position of the tea bush center line in the images, and make a tea garden image dataset; S2. Use a deep learning-based tea ridge image segmentation model, and use the tea garden image dataset described in step S1 for training until the tea ridge image segmentation model converges; according to the tea ridge image segmentation model, input the original image collected by the RGB camera to be detected, and output the semantic segmentation images of the tea bush center line and the tea ridge area contained in the image; The tea ridge image segmentation model adopts a forward propagation flow with single input and dual output, including a backbone feature extraction network, a feature fusion network, and two pre-designed detection branches; the first branch in the detection branch is based on the attention and position encoding mechanism, and is composed of a row-column attention module, a learnable position encoding module, and a multi-layer perceptron module, which is used to extract the semantic information related to the tea bush center line in the image and output the center line of the tea bush in the image; the second branch is based on the encoding-decoding mechanism, and is composed of a multi-head self-attention module, an upsampling module, and a fully connected layer with 2 output channels, which is used to extract the foreground information of the tea ridge area in the image and output the semantic segmentation image of the tea ridge area; S3. According to the tea ridge area semantic segmentation image described in step S2, extract the connected components of the tea ridge area segmentation image, use the bipartite graph matching algorithm to match the connected components with the tea bush center line described in step S2, obtain the matching relationship between the tea bush center line and the tea ridge area, and calculate the depth distribution characteristics of the tea bush according to the matching relationship between the tea bush center line and the tea ridge area, and map it to the tea garden area segmentation image to complete the detection of the tea bush depth distribution.
2. The tea canopy depth distribution detection method based on computer vision and deep learning according to claim 1, characterized in that The tea garden image dataset in step S1 consists of three parts: the original image, the tea ridge area image label, and the tea bush center line data label, where the tea ridge area image label and the tea bush center line data label are independently drawn; during the collection process, the camera forms an angle of 45° to 90° with the horizontal direction, and the vertical height between the camera and the highest point of the tea bush is kept within 0.5m to 1m, ensuring that the camera's field of view includes the tea ridge extending outward along the camera's perspective.
3. The method for detecting the depth distribution of tea canopies based on computer vision and deep learning according to claim 1, wherein The tea ridge image segmentation model in step S2: (1) The backbone feature extraction network and the feature fusion network extract features from the original image data to generate a high-dimensional feature map of the image data; (2) The first branch of the detection branch is composed of a row-column attention module, a learnable position encoding module, and a multi-layer perceptron module. The first branch of the detection branch extracts the row-column features in the high-dimensional feature map, and uses the learnable position encoding to map the row-column features to the tea bush center line features, and finally generates the tea bush center line through a multi-layer perceptron; (3) The second branch of the detection branch is composed of a multi-head self-attention module, an upsampling module and a fully connected layer with an output channel of 2. The second branch of the detection branch first uses the multi-head self-attention module to enhance the importance of foreground information in the high-dimensional feature map, and then uses the upsampling module to decode the high-dimensional feature map and restore it to the original image size. Finally, the foreground and background of all pixels in the image are classified through the fully connected layer to obtain a semantic segmentation image of the tea area.
4. The tea bush depth distribution detection method based on computer vision and deep learning according to claim 1, characterized in that In step S3, the tea bush depth distribution characteristics are calculated: S3 1. Perform connected domain extraction on the segmented image of the tea grove area in step S2 to obtain the category, area, center point position, circumscribed rectangle, and edge contour of all connected domains in the tea grove area, and eliminate connected domains with an area smaller than S; S32. Using a bipartite graph matching algorithm, the tea bush centerline in step S2 is matched with all connected domains, and the unmatched tea bush centerline is eliminated to obtain a successfully matched tea area. S33. Traverse all successfully matched tea vine areas in step S32 and create a tea vine depth mask. The grayscale value of the pixels at the center of the tea vine in the mask is the maximum, and the grayscale value of the pixels at the left and right edges of the tea vine connected domain is the minimum. The grayscale value in the mask represents the relative height of the tea vine from the ground. The left and right edges of the tea vine connected domain are calculated as follows: S331. Establish a coordinate system with the upper left corner of the image as the origin, the x-axis extending horizontally to the right, and the y-axis extending vertically downward. Determine the width w, height h, and top-left vertex P(x, y) of the circumscribed rectangle of the tea field. Calculate the upper edge of the tea field as y and the lower edge as y+h. S332. Traverse all the edge points of the connected domain of the tea ridge, calculate the distance d between the vertical coordinate and the upper and lower edges of the tea ridge, and if d is greater than h / 100, determine that the edge point is a left or right edge point; S333. Collect all left and right edge points to obtain the left and right edge positions of the connected domain of Chalong; S34. Check whether there is a connected domain of tea clumps that has not been successfully matched. If so, determine the relative position of its center point to the vertical centerline of the image. If the center point is to the left of the vertical centerline of the image, construct a tea clump depth mask with the right edge of the connected domain as the minimum grayscale value and the left edge as the maximum grayscale value. If the center point is on the right side of the vertical center line of the image, the tea clover depth mask is constructed with the right edge of the connected domain as the maximum grayscale value and the left edge as the minimum grayscale value; S35. The tea tuft depth masks constructed in step S33 and step S34 are spliced together to obtain the depth distribution features of the tea tuft in the input image.
Citation Information
Patent Citations
Orchard complex road segmentation method based on lightweight semantic segmentation algorithm
CN114463542A
Cotton row center line image extraction method for agricultural machine embedded equipment
CN114782455A