Intelligent robot tea tender shoot picking point identification method based on multi-source data fusion

Through the intelligent robot tea tender shoot picking point recognition method with multi-source data fusion and deep data optimization, the problem of tea tender shoot picking point recognition in night environments is solved, high-precision picking is achieved, and tea quality and picking efficiency are improved.

CN120014460AActive Publication Date: 2025-05-16CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510092295.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the picking points of tea tender shoots in night environments, and traditional mechanical tea picking has problems of small target missed detection and occlusion, resulting in low picking efficiency and poor tea quality.

Method used

The intelligent robot tea tender shoot picking point recognition method based on multi-source data fusion is adopted, and the tea two-dimensional detection frame is obtained through the DFOCC-YOLOv8 model, and the tea tender shoot segmentation model with adaptive occlusion loss function and multi-source feature adaptive fusion is improved to improve the recognition accuracy. At the same time, a lidar is introduced to obtain depth information, and a picking point positioning system with deep data optimization and wind speed correction is accurately positioned.

Benefits of technology

It realizes high-precision identification of tea tender shoot picking points in night environments, improves picking efficiency and tea quality, reduces labor costs, and adapts to the segmentation accuracy and robustness of complex tea images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014460A_ABST
    Figure CN120014460A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent robot tea tender shoot picking point identification method based on multi-source data fusion, and belongs to the technical field of intelligent tea picking. Firstly, a tea night image is obtained, and a night image enhancement model based on a decomposition model is proposed to solve the problems that the image has a large number of noise points and is insufficient in brightness in a night environment. And then a DFOCC-YOLOv8 model is proposed to obtain a two-dimensional detection frame of the tea leaves, channels and spatial features are mutually compensated, and the problem of leak detection of small tea targets is solved. And a TeaReLU activation function and a self-adaptive shielding loss function are provided, so that the problem that the tender shoots of the tea leaves cannot be accurately identified due to shielding is solved. Then, constructing a tea leaf tender shoot segmentation model based on heterogenous multi-feature weighted fusion, and removing backgrounds irrelevant to tea leaves; and finally, introducing point cloud data of a laser radar to obtain depth information, constructing a picking point positioning system based on depth data optimization and wind speed correction, and determining the position of a tea tender shoot picking point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of intelligent tea picking, and relates to an intelligent robot tea tender shoot picking point recognition method based on multi-source data fusion. Background Art

[0002] At present, high-quality tea relies heavily on manual picking, with a low mechanization rate. The mechanization rate of tea garden management operations nationwide is less than 10%, which is still a huge gap compared with field crops. In addition, the picking of high-quality tea is highly seasonal and the picking time is very short, which puts forward a great demand for labor.

[0003] At present, the positioning methods of famous and high-quality tea are mostly based on two-dimensional images. The premise of successful positioning is that the picking points of tender shoots are clearly visible. In the actual tea garden environment in the field, the tea stems of famous and high-quality tea are mostly occluded. The goal of occlusion task research is to recover image space data to obtain visual inertia, rather than to estimate shape and position. Such methods require a lot of time and computing resources. Although they can improve the robustness of the system, such methods obviously cannot meet the needs of agricultural picking tasks. This urgently requires the extraction of target features for specific task objects and the study of shape recovery and position estimation. In addition, due to the problems of short tea picking cycle, high labor costs and tea occlusion, it is of great significance to propose a method that can automatically pick tea by identifying picking points at night. This method can enable the machine to achieve fully automatic tea picking at night, which can save costs, improve tea quality, and contribute to the rapid development of the tea industry. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a method for identifying the picking point of tea shoots by an intelligent robot based on multi-source data fusion. To solve the problem of a large number of noise points and insufficient brightness in images under nighttime environments. Then, the DFOCC-YOLOv8 model is proposed to obtain the two-dimensional detection frame of tea leaves, and the channel and spatial features are compensated for each other to solve the problem of missed detection of small tea targets and obtain more accurate tea location information. And an adaptive occlusion loss function is introduced to improve the problem of inaccurate recognition due to occlusion. Then, a tea shoot segmentation model with multi-source feature adaptive fusion is constructed, and a feature extraction method of color gradient analysis, multi-scale texture response, dynamic shape symmetry analysis, and multi-directional tea growth characteristics is proposed. The two-dimensional frame of tea shoots obtained by S2 is processed, and the segmentation accuracy and robustness of the model in complex tea images are improved through an adaptive fusion mechanism. Finally, the point cloud data of the laser radar is introduced to obtain depth information, and a picking point positioning system based on depth data optimization and wind speed correction is constructed. By modeling the growth characteristics of tea leaves, combining the deep data optimization technology, and utilizing the fusion of leaf growth angle and point cloud data, the depth of the picking point of tea shoots is inferred. In addition, considering the impact of wind speed, wind direction and environmental factors on depth measurement, wind error correction technology was introduced. By adjusting the depth data in real time, the error caused by wind speed fluctuations was eliminated, thereby significantly improving the picking accuracy and stability.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] An intelligent robot tea shoot picking point recognition method based on multi-source data fusion, the method comprises the following steps:

[0007] S1: Obtain tea leaves at night and enhance the images based on the image decomposition model to solve the problem of a large number of noise points and insufficient brightness in the images under night conditions;

[0008] S2: Use the DFOCC-YOLOv8 model to detect tea leaves in the enhanced image, obtain the two-dimensional detection frame of the tea leaves, and filter the detection frame to remove tea leaves that do not meet the requirements;

[0009] S3: constructing a tea shoot segmentation model based on weighted fusion of heterogeneous multi-features, processing the tea 2D detection frame obtained in step S2, removing the background irrelevant to the tea leaves, and obtaining a tea shoot segmentation image;

[0010] S4: Introduce the point cloud data of the lidar to obtain depth information, and combine it with the segmented image of the young tea shoots for data synchronization and registration; build a picking point positioning system based on depth data optimization and wind speed correction, optimize and correct the acquired depth data, and determine the three-dimensional coordinates of the picking point of the young tea shoots.

[0011] Further, the S1 specifically includes the following steps:

[0012] S11: The camera is installed on the arm of the tea-picking robot, and the camera is tilted downward from the tea field, 30 to 50 cm away from the tea leaves, and the angle between the camera and the tea leaves is [30°, 60°], so as to obtain images of the tea leaves in a night environment; the laser radar is installed beside the camera to obtain point cloud data of the tea leaves in a night environment;

[0013] S12: A night image enhancement method based on an image decomposition model is proposed. The collected night image is decomposed into a structure layer and a texture layer. A global brightness factor is constructed to enhance the brightness of the structure map, as shown in formula (1):

[0014]

[0015] Among them, L g (x, y) represents the global brightness factor, I'(x, y) represents the brightness value of the input image, μ image represents the global brightness mean of the image, σ image represents the brightness standard deviation of the image, ν represents a small constant to avoid division by zero, η controls the intensity of brightness enhancement, γ is a nonlinear exponential adjustment parameter, and I' max The maximum brightness value of the image is 255, β is the scaling factor, λ is the parameter that adjusts the enhancement intensity of low-light areas, σ max Indicates the maximum value of the brightness standard deviation of the image;

[0016] Normalize the pixel values ​​of the input image and calculate the local contrast by combining the global mean and standard deviation;

[0017] Gaussian distribution function is used to assign weights according to the distance between pixel value and maximum pixel value, thus highlighting the target brightness range;

[0018] Combined with local texture or brightness gradient information, the dynamic range is compressed through logarithmic transformation to optimize image details and edge performance;

[0019] S13: Construct a local brightness factor to enhance the details of the texture map, as shown in formula (2):

[0020]

[0021] Among them, L local (x, y) represents the local brightness factor, I(x, y) represents the pixel value of the input image at position (x, y), and N k (x, y) represents the k-th filter of the Gaussian pyramid, k is a positive integer, μ local (x,y),σ local(x, y) represent the mean and standard deviation of the local area of ​​the image at position (x, y), respectively. The local area is a circular area with a radius of 2 and centered at (x, y). represents the gradient of the position (x, y), and K represents the total number of filters;

[0022] The brightness difference between a pixel and its neighborhood is used to measure local changes, while gradient information is introduced to suppress mutation areas to reduce noise interference.

[0023] By locally normalizing pixel values, detail changes are highlighted and local contrast is enhanced;

[0024] Use adjustable parameters to balance local detail enhancement and overall smoothing effect;

[0025] S14: Fusing the structure layer and the texture layer processed by the night image enhancement method based on the image decomposition model for subsequent processing.

[0026] Further, the S2 specifically includes:

[0027] S21: Use YOLOv8 to extract features from the image and obtain the location information of the tea leaves; build a DFOCC-YOLOv8 model to obtain the two-dimensional detection frame of the tea leaves, and build a dual-feature winding model based on YOLOv8 to intertwine and compensate the spatial features and channel features; first perform data enhancement on the data set to improve the generalization ability of the model; then send the three feature maps of different scales output by YOLOv8 to the dual-feature winding module;

[0028] S22: Perform channel feature extraction, as shown in formula (3):

[0029]

[0030] Among them, Conv represents one-dimensional convolution, Conv k The subscript k in the convolution kernel represents the size of the convolution kernel, F c (e,u) represents the feature map of the e-th row, u-th column, and c-th channel, H represents the height of the feature map, W represents the width of the feature map, σ represents the Sigmoid activation function, F1 represents the feature map after channel feature recalibration, Conv1 represents a one-dimensional convolution with a convolution kernel of 1×1, and Conv3 represents a one-dimensional convolution with a convolution kernel of 3×3;

[0031] Perform global average pooling on the input feature map to obtain a feature map of scale 1×1×c, and use cross-channel interaction to improve the feature correlation between different channels; use 1×1 convolution to extract features again, and use the Sigmoid function to generate the corresponding weight coefficients for channel-weighted fusion of the input feature map;

[0032] S23: spatial feature extraction, as shown in formula (4);

[0033] F2=σ(Conv21(ξ(Conv23(Tea ReLU(Conv1(F c (e,u)))))) (4)

[0034] Perform convolution operation on the input feature map to halve the number of channels; use 3×3 convolution to extract spatial information, and activate with ReLU function to enhance the nonlinearity of the model; then use 1×1 convolution to reduce the number of channels, and construct TeaReLU activation function to filter features, improve speed and reduce parameters, as shown in formula (5):

[0035]

[0036] Among them, Conv2 represents two-dimensional convolution, Conv k The subscript k in the figure indicates the size of the convolution kernel, σ indicates the Sigmoid activation function, F2 indicates the feature map after spatial feature recalibration, F indicates the original feature map, ξ indicates the ReLU function, Conv21 indicates a two-dimensional convolution with a convolution kernel of 1×1, Conv23 indicates a two-dimensional convolution with a convolution kernel of 3×3, and x indicates the input value of the neuron. is the threshold parameter, controlling the nonlinear turning point,

[0037] Finally, the Sigmoid function is used to generate the corresponding weight coefficients for spatial weighted fusion of the input feature maps;

[0038] S24: The feature map obtained after the channel feature and the spatial feature are entangled is sent to the DFOCC-YOLOv8 model to detect the tea shoot frame and calculate the length of the tea shoot frame. If it does not meet the required length requirement, it will be removed. To improve the accuracy, the DFOCC-YOLOv8 model constructs an adaptive occlusion loss function, as shown in formula (6), where the smaller the IoU (a, b), the more serious the occlusion, the greater the difficulty of sample learning, and the greater the corresponding weight;

[0039]

[0040] Among them, 1-IoU(a,b) is the occlusion weight, L occ represents the occlusion loss function, a represents the predicted box, b represents the real box, χ pos is the position loss weight coefficient, κ cls is the weight coefficient of the classification loss function, represents the true label of the nth class (0 or 1), logc n Represents the predicted probability c of the nth class nThe logarithm of S is used to measure the uncertainty between the predicted probability and the true label. a Represents the area of ​​the prediction box, S b Represents the area of ​​the true box, and N represents the total number of categories.

[0041] Further, the S3 specifically includes the following steps:

[0042] S31: Extract the color feature information of tea leaves; firstly, introduce the gradient information by calculating the change rate of the intensity values ​​of the three color channels of red, green and blue in the image in the horizontal and vertical directions, and detect the area with drastic color change; then, combine the local pixel difference and statistical characteristics, highlight the details and enhance the contrast through the local normalization operation, and use the adjustable parameters to balance the local detail enhancement and the overall smoothing effect, so as to optimize the image details and texture, and adapt to the visual tasks in complex environments, as shown in formula (7):

[0043]

[0044] Among them, I R (x,y),I G (x,y),I B (x, y) are the color values ​​of the red, green, and blue channels of a pixel point (x, y) in the image, respectively. Respectively represent the color difference between the pixel and the adjacent pixel in the horizontal and vertical directions, f color Indicates the weight of the color characteristics of tea leaves in image segmentation;

[0045] S32: Extracting the texture feature information of tea leaves; First, the image texture related information is obtained by summing up different directions and scales; Then, the sum of local pixel intensities is calculated in the denominator, and a specific value is subtracted from it and the absolute value is taken to measure the local difference; Then, this difference is processed by exponential functions and parameters to suppress the mutation area and reduce noise; Finally, the various parts of the formula and the adjustable parameters are used to balance the local detail enhancement and the overall smoothing effect to optimize the image texture, as shown in formula (8):

[0046]

[0047] Among them, R d (x, y, s) represents the texture response, L s(x, y, d) represents the local contrast or local texture intensity, n' represents the size of the complexity metric area, ε represents the learning rate of the weighting factor, I(x+i', y+j') represents the pixel value at the position (x+i', y+j') in the image, d represents the number of directions, which refers to how many different directions are considered to capture the texture change information, s represents the number of scales, which refers to the number of different sized areas or resolutions we choose when extracting texture features, i' represents the offset in the x direction relative to the current position, j' represents the offset in the y direction relative to the current position, and f texture represents the weight of the texture characteristics of tea leaves in image segmentation, D represents the total number of different directions selected when extracting texture features, and S represents the total number of different region sizes selected when extracting texture features;

[0048] S33: In order to reflect the morphological changes of tea leaves, a new method based on shape manifold and local symmetry analysis is proposed, as shown in formula (9):

[0049]

[0050] in, is the second-order derivative, indicating the curvature of the edge, y(q) represents the vertical coordinate value of the qth point of the shape contour, L is the total length of the contour, and f shape Indicates the weight of the shape characteristics of tea leaves in image segmentation;

[0051] First, the second-order derivative part is similar to the “gradient information”, which measures the rate of shape change, including the brightness difference between a pixel and its neighbors to measure local changes;

[0052] Next, the summation part combines “local pixel differences” and “statistical characteristics” to highlight the local changes of the function;

[0053] Finally, the formula combines the two, balancing the shape change rate and local differences to achieve the evaluation of the overall characteristics of the shape;

[0054] S34: A new feature extraction method for tea growth law is proposed, as shown in formula (10), focusing on the growth direction difference, local spatial consistency and growth pattern correlation of tea leaves; by quantifying the growth direction difference between adjacent pixels, measuring the growth consistency of local areas, and introducing angle difference measurement, the spatial structure changes of tea leaves at different growth stages are captured;

[0055]

[0056] Among them, x and y represent the position of the current pixel in the image, i and j represent the coordinates of the adjacent pixels, and Δθ ijrepresents the difference in growth direction between pixel i and pixel j, g represents the standard deviation of Gaussian weighting, which is used to control the influence range of growth direction difference in local area, ΔG ij represents the difference in growth pattern between pixel i and pixel j, g L represents the standard deviation of the local spatial weighting, which is used to control the weighting range, θ ij represents the growth direction angle between adjacent pixels i and j, describing the difference in the growth angle of tea leaves, g P represents the standard deviation of the growth mode weights, which is used to control the weighted range of the growth mode correlation. Represents the standard deviation of the fusion weight, which is used to control the influence range of each feature in the fusion process. grow Indicates the weight of tea growth characteristics in image segmentation;

[0057] S35: Construct a tea shoot segmentation model based on adaptive fusion of multi-source features, as shown in formula (11), and design a learning mechanism to predict the importance of each pixel in different feature dimensions;

[0058] Specifically: by calculating the features of the local image area and combining the context information to predict the weight of each feature;

[0059] First, the mapping function learned through the convolutional neural network can determine the feature weight based on the local features, and the Sigmoid activation function suppresses the interference caused by local differences;

[0060] Next, the different feature maps are processed and multiplied with the parameters to highlight the changes in different features and enhance the local contrast;

[0061] Finally, by comprehensively processing different features, the segmentation algorithm is used to balance the local detail enhancement and the overall fusion effect, so as to achieve the optimization processing of the comprehensive features of the image;

[0062] f fusion =UNet(σ(φ color (f color ))·P+σ(φ texture (f texture ))·P+σ(φ shape (f shape ))·P+σ(φ grow (f grow ))·P)(11)

[0063] Among them, f fusion (x, y) is the feature map after segmentation, σ represents the Sigmoid activation function, φ color ,φ texture ,φ shape ,φ growIt is a mapping function learned through the convolutional neural network CNN, which predicts the weight of each feature based on the local features of the input image. UNet represents the segmentation algorithm, and P represents the tea shoot detection frame obtained by S24.

[0064] Further, the S4 specifically includes:

[0065] S41: extract key points from the tea segmentation image obtained in S35; use the key point detection method HRNet to extract the b(b∈N + ) The coordinate information of the intersection of the leaves of the young tea shoots, and then take 1-2mm below the intersection to determine the two-dimensional coordinates of the picking point of the young tea shoots Then the segmented image and the point cloud data of the LiDAR are fused to achieve data synchronization and registration; and the LiDAR data is projected onto the segmented image to obtain the depth value of the tea leaf part;

[0066] S42: Using the depth data optimization technology, combining the tea growth characteristics with the tea shoot leaf point cloud depth data for inference, as shown in formula (12); specifically, by modeling the tea leaf growth angle, adjusting the leaf depth value, overcoming the depth deviation caused by tea growth, and inferring the depth of the picking point; at the same time, considering that the leaf points closer to the leaf centroid are in the same plane as the picking point, these points are given higher weights to avoid contradictions when the picking point position is unknown;

[0067]

[0068] in, represents the depth value of the i-th point on the b-th tea shoot leaf. Indicates the depth value of the bth tea shoot picking point, represents the growth angle of the bth tea shoot leaf, It means to establish a model for the growth characteristics of the b-th tea shoot leaf to obtain the depth of the tea shoot picking point. c is a constant that controls the influence of distance on the weight. n represents the total number of point clouds on the b-th tea shoot leaf.

[0069] S43: In the tea picking task, the correction of wind speed and wind direction to depth is considered; as shown in formula (13), by dynamically considering the real-time changes of wind speed, wind direction and environmental factors, the depth data is accurately adjusted in combination with the correction function, the wind error is eliminated, and the picking accuracy and stability are improved;

[0070]

[0071] Among them, ΔZ wind(t) represents the wind error correction value, which represents the correction amount of wind speed and wind direction to depth, v(t) represents the wind speed at time t, θ(t) represents the wind direction angle at time t, w1 represents the initial influence coefficient of wind speed and wind direction on depth error, w2 is an exponential attenuation coefficient, which is used to represent the weakening of wind speed influence over time, and w3 is a time-related correction function, which is used to control the influence of wind speed on depth over time. represents the time constant of the wind error, and f(E(t)) represents the environmental correction function; considering the influence of environmental factors such as humidity and temperature on the correction of wind speed and wind direction, the environmental correction function f(E(t)) is used to adjust the wind error, and E(t) represents the environmental variable at time t;

[0072] S44: After completing the depth correction, determine the picking point of the tea leaves

[0073] Intelligent robot tea shoot picking point recognition system based on multi-source data fusion, including:

[0074] Image acquisition module, used to obtain nighttime images of tea leaves;

[0075] An image enhancement module, used to enhance the image based on an image decomposition model;

[0076] The tea leaf shoot detection module is used to detect tea leaf shoots in the enhanced image and obtain a two-dimensional tea leaf detection frame;

[0077] The tea shoot segmentation module is used to process the tea 2D detection frame, remove the background irrelevant to the tea leaves, and obtain the tea shoot segmentation image;

[0078] The depth information acquisition module is used to obtain depth information by introducing the point cloud data of the laser radar, and to synchronize and align the data with the segmented image of the young tea leaves;

[0079] The picking point positioning module is used to build a picking point positioning system based on depth data optimization and wind speed correction, optimize and correct the depth data, and determine the three-dimensional coordinates of the picking point of the young tea shoots.

[0080] The beneficial effects of the present invention are:

[0081] (1) The present invention proposes a night image enhancement model based on a decomposition model to solve the problem of a large number of noise points and insufficient brightness in images under night conditions, so that the tea picking robot can work at night, reduce labor costs, and avoid missing the best time for tea picking.

[0082] (2) Compared with the traditional mechanical tea picking, which has the problem of damaging the tender buds and the presence of a large amount of impurities in the tender buds, the present invention proposes a DFOCC-YOLOv8 model to obtain a two-dimensional detection frame of tea leaves. The size of the two-dimensional detection frame can be calculated to filter out tea leaves that are too small or too old. A dual-channel winding model is constructed based on YOLOv8 to complement the spatial and channel features to solve the problem of missing small tea targets. The occlusion loss function is introduced to improve the problem of inaccurate recognition due to occlusion.

[0083] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0085] Figure 1 This is a main flow chart of the intelligent robot tea shoot identification method based on multi-source data fusion according to an embodiment of the present invention;

[0086] Figure 2 This is a structural diagram of a dual-channel winding module described in an embodiment of the invention. DETAILED DESCRIPTION

[0087] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0088] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0089] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0090] In this embodiment, a 32-line Velodyne laser radar is selected as a depth sensor, a binocular depth camera d435i is used as an image data sensor, and an implementation algorithm is written in the Ubuntu system.

[0091] An intelligent robot tea shoot picking point recognition method based on multi-source data fusion, the method comprises the following steps:

[0092] S1: A night image enhancement method based on image decomposition is proposed. By decomposing night low-light images into structure layers and texture layers, large-scale object contours and small-scale details are optimized respectively. The global brightness factor optimizes the brightness of the structure layer and highlights the target brightness range; the local brightness factor enhances the details of the texture layer, suppresses noise and strengthens local contrast.

[0093] S2: A tea shoot detection method based on the DFOCC-YOLOv8 model is proposed. This model aims at the problem of tea shoot small target detection, constructs a dual feature winding module, compensates spatial features with channel features, and proposes the TeaReLU activation function to smooth low-value inputs, retain high-value inputs, and enhance detail feature extraction. Subsequently, an adaptive occlusion loss function is proposed to dynamically adjust the weights of occluded samples.

[0094] S3: A tea shoot segmentation model based on adaptive fusion of multi-source features is proposed. The model proposes a color feature extraction method based on gradient information and local normalization, introduces shape manifold and local symmetry analysis methods, and designs a feature extraction method focusing on the differences in tea growth laws to capture the spatial structure changes of tea leaves at different growth stages.

[0095] S4: A picking point positioning system based on depth data optimization and wind speed correction is proposed. This system solves the problem that the laser radar cannot scan the stalk area. It uses depth data optimization technology combined with tea growth characteristics to infer the depth of the picking point. Then, considering the dynamic changes of wind speed and wind direction, a correction function is constructed to adjust the depth data and eliminate wind error.

[0096] like Figure 1 As shown, the S1 specifically includes the following steps:

[0097] S11: The camera is installed on the arm of the tea-picking robot, and the camera is tilted downward from the tea field, 30 to 50 cm away from the tea leaves, with an angle of [30°, 60°] to the tea leaves, to obtain images of the tea leaves in a night environment; the laser radar can be installed next to the camera to obtain point cloud data of the tea leaves in a night environment.

[0098] S12: A night image enhancement method based on an image decomposition model is proposed, which decomposes the collected night images into a structure layer and a texture layer. Since the night images usually contain a lot of noise and insufficient brightness, a global brightness factor is constructed, as shown in formula (1), to enhance the brightness of the structure image. First, the pixel values ​​of the input image are standardized, and the local contrast is calculated by combining the global mean and standard deviation; then, the Gaussian distribution function is used to assign weights according to the distance between the pixel value and the maximum pixel value to highlight the target brightness range; finally, combined with the local texture or brightness gradient information, the dynamic range is compressed through logarithmic transformation to optimize the details and edge performance of the image.

[0099]

[0100] Among them, L g (x, y) represents the global brightness factor, I'(x, y) represents the brightness value of the input image, μ image represents the global brightness mean of the image, σ image represents the brightness standard deviation of the image, ν represents a small constant to avoid division by zero, η controls the intensity of brightness enhancement, γ is a nonlinear exponential adjustment parameter, and I' max The maximum brightness value of the image is 255, β is the scaling factor, λ is the parameter that adjusts the enhancement intensity of low-light areas, σ max Indicates the maximum standard deviation of the brightness of the image.

[0101] S13: Construct a local brightness factor, as shown in formula (2). To enhance the details of the texture map, first use the brightness difference between the pixel and its neighborhood to measure the local change, and introduce gradient information to suppress the mutation area to reduce noise interference; then, by locally normalizing the pixel values, highlight the detail changes and enhance the local contrast; finally, use adjustable parameters to balance the local detail enhancement and overall smoothing effect.

[0102]

[0103] Among them, L local (x, y) represents the local brightness factor, I(x, y) represents the pixel value of the input image at position (x, y), and N k(x, y) represents the k-th filter of the Gaussian pyramid, k is a positive integer, μ local (x,y),σ local (x, y) represent the mean and standard deviation of the local area of ​​the image at position (x, y), respectively. The local area is a circular area with a radius of 2 and centered at (x, y). represents the gradient of the position (x, y), and K represents the total number of filters.

[0104] S14: Effectively fuse the structure layer and the texture layer processed by the night image enhancement method based on the image decomposition model for subsequent processing.

[0105] The S2 specifically includes:

[0106] S21: Use YOLOv8 to extract features from the image and obtain the location information of the tea leaves. Since tea leaves are small targets, directly using YOLOv8 for detection will lead to missed detection and false detection. Based on this, the DFOCC-YOLOv8 model is built to obtain the two-dimensional detection frame of the tea leaves. Based on YOLOv8, a dual-feature entanglement model is built to entangle and compensate the spatial features and channel features. First, data enhancement is performed on the data set to improve the generalization ability of the model. Then the three feature maps of different scales output by YOLOv8 are sent to the dual-feature entanglement module, whose structure is as follows: Figure 2 As shown:

[0107] S22: Perform channel feature extraction, as shown in formula (3). First, perform global average pooling on the input feature map to obtain a feature map of scale 1×1×c, and then use cross-channel interaction to improve the feature correlation between different channels. Use 1×1 convolution to extract features again, and finally use the Sigmoid function to generate the corresponding weight coefficient for channel weighted fusion of the input feature map.

[0108]

[0109] Among them, Conv represents one-dimensional convolution, Conv k The subscript k in the convolution kernel represents the size of the convolution kernel, F c (e,u) represents the feature map of the e-th row, u-th column, and c-th channel, H represents the height of the feature map, W represents the width of the feature map, σ represents the Sigmoid activation function, F1 represents the feature map after channel feature recalibration, Conv1 represents a one-dimensional convolution with a convolution kernel of 1×1, and Conv3 represents a one-dimensional convolution with a convolution kernel of 3×3.

[0110] S23: Spatial feature extraction, as shown in formula (4). First, perform a convolution operation on the input feature map to halve the number of channels; then use 3×3 convolution to extract spatial information, and activate the ReLU function to enhance the nonlinearity of the model; then use 1×1 convolution to reduce the number of channels, and construct the TeaReLU activation function to filter features, as shown in formula (5), which improves the speed and reduces the number of parameters. Finally, the Sigmoid function is used to generate the corresponding weight coefficients for spatial weighted fusion of the input feature map.

[0111] F2=σ(Conv21(ξ(Conv23(Tea ReLU(Conv1(F c (e,u)))))) (4)

[0112]

[0113] Among them, Conv2 represents two-dimensional convolution, Conv k The subscript k in the figure indicates the size of the convolution kernel, σ indicates the Sigmoid activation function, F2 indicates the feature map after spatial feature recalibration, F indicates the original feature map, ξ indicates the ReLU function, Conv21 indicates a two-dimensional convolution with a convolution kernel of 1×1, Conv23 indicates a two-dimensional convolution with a convolution kernel of 3×3, and x indicates the input value of the neuron. is the threshold parameter, controlling the nonlinear turning point,

[0114] S24: The feature map obtained after the channel feature and the spatial feature are entangled is sent to the DFOCC-YOLOv8 model to detect the tea shoot frame and calculate the length of the tea shoot frame. If it does not meet the required length requirement, it will be removed. In order to further improve the accuracy, the model also constructs an adaptive occlusion loss function, as shown in formula (6), where the smaller the IoU (a, b), the more serious the occlusion, the greater the difficulty of sample learning, and the larger the corresponding weight.

[0115]

[0116] Among them, 1-IoU(a,b) is the occlusion weight, L occ represents the occlusion loss function, a represents the predicted box, b represents the real box, χ pos is the position loss weight coefficient, κ cls is the weight coefficient of the classification loss function, represents the true label of the nth class (0 or 1), logc n Represents the predicted probability c of the nth class n The logarithm of is used to measure the uncertainty between the predicted probability and the true label.

[0117] Sa Represents the area of ​​the prediction box, S b Represents the area of ​​the true box, and N represents the total number of categories.

[0118] The S3 specifically includes the following steps:

[0119] S31: Extract the color feature information of tea leaves. First, the gradient information is introduced by calculating the change rate of the intensity values ​​of the three color channels of red, green and blue in the image in the horizontal and vertical directions to detect the area with drastic color changes. Then, the local pixel difference and statistical characteristics are combined, and the details and contrast are highlighted through local normalization and other operations. The adjustable parameters are used to balance the local detail enhancement and the overall smoothing effect to optimize the image details and textures and adapt to the visual tasks in complex environments, as shown in formula (7).

[0120]

[0121] Among them, I R (x,y),I G (x,y),I B (x, y) are the color values ​​of the red, green, and blue channels of a pixel point (x, y) in the image, respectively. Respectively represent the color difference between the pixel and the adjacent pixel in the horizontal and vertical directions, f color Represents the weight of the color characteristics of tea leaves in image segmentation.

[0122] S32: Extract the texture feature information of tea leaves. First, the image texture related information is obtained by summing different directions and scales. Then, the sum of local pixel intensities is calculated in the denominator, and the specific value is subtracted from it and the absolute value is taken to measure the local difference. Then, this difference is processed by exponential functions and parameters to suppress the mutation area and reduce noise. Finally, the various parts of the formula and adjustable parameters are used to balance the local detail enhancement and the overall smoothing effect to optimize the image texture, as shown in formula (8).

[0123]

[0124] Among them, R d (x, y, s) represents the texture response, L s(x, y, d) represents the local contrast or local texture intensity, n' represents the size of the complexity metric area, ε represents the learning rate of the weighting factor, I(x+i', y+j') represents the pixel value at the position (x+i', y+j') in the image, d represents the number of directions, which refers to how many different directions are considered to capture the texture change information, s represents the number of scales, which refers to the number of regions of different sizes (or resolutions) selected when extracting texture features, i' represents the offset in the x direction relative to the current position, j' represents the offset in the y direction relative to the current position, and f texture It represents the weight of the texture characteristics of tea leaves in image segmentation, D represents the total number of different directions selected when extracting texture features, and S represents the total number of different region sizes selected when extracting texture features.

[0125] S33: In order to more accurately reflect the morphological changes of tea leaves, a new method based on shape manifold and local symmetry analysis is proposed, as shown in formula (9). First, the second-order derivative part is similar to "gradient information" to measure the shape change rate, such as the brightness difference between a pixel and its neighborhood to measure the local change. Then, the summation part combines "local pixel difference" and "statistical characteristics" to highlight the local changes of the function. Finally, the formula combines the two to balance the shape change rate and local difference, and realizes the evaluation of the overall shape characteristics.

[0126]

[0127] in, is the second-order derivative, indicating the curvature of the edge, y(q) represents the vertical coordinate value of the qth point of the shape contour, L is the total length of the contour, and f shape Represents the weight of the shape characteristics of tea leaves in image segmentation.

[0128] S34: A new feature extraction method for tea growth law is proposed, as shown in formula (10), focusing on the growth direction difference, local spatial consistency and growth pattern correlation of tea leaves. By quantifying the growth direction difference between adjacent pixels, measuring the growth consistency of local areas, and introducing the angle difference metric, the spatial structure changes of tea leaves at different growth stages are captured.

[0129]

[0130] Among them, x and y represent the position of the current pixel in the image, i and j represent the coordinates of the adjacent pixels, and Δθ ij represents the difference in growth direction between pixel i and pixel j, g represents the standard deviation of Gaussian weighting, which is used to control the influence range of growth direction difference in local area, ΔG ij represents the difference in growth pattern between pixel i and pixel j, g Lrepresents the standard deviation of the local spatial weighting, which is used to control the weighting range, θ ij represents the growth direction angle between adjacent pixels i and j, describing the difference in the growth angle of tea leaves, g P represents the standard deviation of the growth mode weights, which is used to control the weighted range of the growth mode correlation. Represents the standard deviation of the fusion weight, which is used to control the influence range of each feature in the fusion process. grow Represents the weight of tea growth characteristics in image segmentation.

[0131] S35: A tea shoot segmentation model based on adaptive fusion of multi-source features is constructed. As shown in formula (11), a learning mechanism is designed to predict the importance of each pixel in different feature dimensions. Specifically, the weight of each feature is predicted by calculating the features of the local image area and combining the context information. First, the mapping function learned by the convolutional neural network can determine the feature weight based on the local features, and the Sigmoid activation function can suppress the interference caused by excessive local differences, just like introducing gradient information. Then, different feature maps are processed and multiplied with parameters to highlight the changes in different features and enhance the local contrast. Finally, different features are processed comprehensively, and the segmentation algorithm is used to balance the local detail enhancement and the overall fusion effect to achieve the optimization of the comprehensive features of the image.

[0132] f fusion =UNet(σ(φ color (f color ))·P+σ(φ texture (f texture ))·P+σ(φ shape (f shape ))·P+σ(φ grow (f grow ))·P) (11)

[0133] Among them, f fusion (x, y) is the feature map after segmentation, σ represents the Sigmoid activation function, φ color ,φ texture ,φ shape ,φ grow It is a mapping function learned through a convolutional neural network (CNN). The weight of each feature is predicted based on the local features of the input image. UNet represents the segmentation algorithm, and P represents the tea shoot detection frame obtained by S24.

[0134] The S4 specifically includes:

[0135] S41: Extract key points from the tea segmentation image obtained in S35. Use the key point detection method HRNet to extract the b(b∈N +) The coordinate information of the intersection of the leaves of the young tea shoots, and then take 1-2mm below the intersection to determine the two-dimensional coordinates of the picking point of the young tea shoots Then the segmented image and the point cloud data of the LiDAR are fused to achieve data synchronization and registration. The LiDAR data is projected onto the segmented image to obtain the depth value of the tea leaf part.

[0136] S42: Since the picking point of the young tea shoots is located at the stem position and is small in size, the laser radar cannot effectively scan this area, resulting in the inability to directly obtain the depth value of the picking point. Therefore, the depth data optimization technology is used to combine the tea growth characteristics with the depth data of the tea shoot leaf point cloud for inference, as shown in formula (12). Specifically, by modeling the growth angle of the tea leaves, the depth value of the leaves is adjusted to overcome the depth deviation caused by the growth of the tea leaves, and the depth of the picking point is accurately inferred. At the same time, considering that the leaf points closer to the centroid of the leaf are more likely to be in the same plane as the picking point, these points are given higher weights. This strategy effectively avoids the depth contradiction that may arise when the position of the picking point is unknown.

[0137]

[0138] in, represents the depth value of the i-th point on the b-th tea shoot leaf, Indicates the depth value of the bth tea shoot picking point, represents the growth angle of the bth tea shoot leaf, It means that a model is established for the growth characteristics (angle) of the b-th tea shoot leaves to obtain the depth of the tea shoot picking point. c is a constant that controls the influence of distance on the weight. n represents the total number of point clouds on the b-th tea shoot leaves.

[0139] S43: In the tea picking task, wind speed and wind direction may affect the depth measurement results, especially because the dynamic change of wind may cause fluctuations in sensor readings and changes in the position of the tea shoot picking point, resulting in errors (wind-driven errors). Therefore, the correction of wind speed and wind direction to depth must be considered. As shown in formula (13), by dynamically considering the real-time changes of wind speed, wind direction and environmental factors, the depth data is accurately adjusted in combination with the correction function, thereby eliminating wind-driven errors and improving picking accuracy and stability.

[0140]

[0141] Among them, ΔZ wind(t) represents the wind error correction value, which represents the correction amount of wind speed and wind direction to depth, v(t) represents the wind speed at time t, θ(t) represents the wind direction angle at time t, w1 represents the initial influence coefficient of wind speed and wind direction on depth error, w2 is an exponential decay coefficient, which is used to represent the weakening of wind speed influence over time, usually due to air resistance or other environmental factors, w3 is a time-related correction function, which is used to control the influence of wind speed on depth over time. The time constant representing the wind error is usually the response time of wind speed fluctuation or environmental factors. f(E(t)) represents the environmental correction function. Considering the influence of environmental factors such as humidity and temperature on wind speed and wind direction correction, the environmental correction function f(E(t)) is used to further adjust the wind error to ensure the accuracy of depth. E(t) represents the environmental variables at time t, such as humidity, temperature, etc.

[0142] S44: After completing the depth correction, determine the picking point of the tea leaves

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. An intelligent robot tea shoot picking point recognition method based on multi-source data fusion, characterized by: The method comprises the following steps: S1: Obtain tea leaves at night and enhance the images based on the image decomposition model to solve the problem of a large number of noise points and insufficient brightness in the images under night conditions; S2: Use the DFOCC-YOLOv8 model to detect tea leaves in the enhanced image, obtain the two-dimensional detection frame of the tea leaves, and filter the detection frame to remove tea leaves that do not meet the requirements; S3: constructing a tea shoot segmentation model based on weighted fusion of heterogeneous multi-features, processing the tea 2D detection frame obtained in step S2, removing the background irrelevant to the tea leaves, and obtaining a tea shoot segmentation image; S4: Introduce the point cloud data of the lidar to obtain depth information, and combine it with the segmented image of the young tea shoots for data synchronization and registration; build a picking point positioning system based on depth data optimization and wind speed correction, optimize and correct the acquired depth data, and determine the three-dimensional coordinates of the picking point of the young tea shoots.

2. The method for identifying tea shoot picking points by an intelligent robot based on multi-source data fusion according to claim 1 is characterized in that: The S1 specifically includes the following steps: S11: The camera is installed on the arm of the tea picking robot, and the camera is tilted downward from the tea field, 30 to 50 cm away from the tea leaves, and the angle between the camera and the tea leaves is [300, 600], so as to obtain images of the tea leaves in a night environment; the laser radar is installed beside the camera to obtain point cloud data of the tea leaves in a night environment; S12: A night image enhancement method based on an image decomposition model is proposed. The collected night image is decomposed into a structure layer and a texture layer. A global brightness factor is constructed to enhance the brightness of the structure map, as shown in formula (1): Among them, L g (x, y) represents the global brightness factor, I'(x, y) represents the brightness value of the input image, μ image represents the global brightness mean of the image, σ image represents the brightness standard deviation of the image, ν represents a small constant to avoid division by zero, η controls the intensity of brightness enhancement, γ is a nonlinear exponential adjustment parameter, and I' max The maximum brightness value of the image is 255, β is the scaling factor, λ is the parameter that adjusts the enhancement intensity of low-light areas, σ max Indicates the maximum value of the brightness standard deviation of the image; Normalize the pixel values ​​of the input image and calculate the local contrast by combining the global mean and standard deviation; Gaussian distribution function is used to assign weights according to the distance between pixel value and maximum pixel value, thus highlighting the target brightness range; Combined with local texture or brightness gradient information, the dynamic range is compressed through logarithmic transformation to optimize image details and edge performance; S13: Construct a local brightness factor to enhance the details of the texture map, as shown in formula (2): Among them, L local (x, y) represents the local brightness factor, I(x, y) represents the pixel value of the input image at position (x, y), and N k (x, y) represents the k-th filter of the Gaussian pyramid, k is a positive integer, μ local (x,y),σ local (x, y) represent the mean and standard deviation of the local area of ​​the image at position (x, y), respectively. The local area is a circular area with a radius of 2 and centered at (x, y). represents the gradient of the position (x, y), and K represents the total number of filters; The brightness difference between a pixel and its neighborhood is used to measure local changes, while gradient information is introduced to suppress mutation areas to reduce noise interference. By locally normalizing pixel values, detail changes are highlighted and local contrast is enhanced; Use adjustable parameters to balance local detail enhancement and overall smoothing effect; S14: Fusing the structure layer and the texture layer processed by the night image enhancement method based on the image decomposition model for subsequent processing.

3. The method for identifying tea shoot picking points by an intelligent robot based on multi-source data fusion according to claim 2 is characterized in that: The S2 specifically includes: S21: Use YOLOv8 to extract features from the image and obtain the location information of the tea leaves; build a DFOCC-YOLOv8 model to obtain the two-dimensional detection frame of the tea leaves, and build a dual-feature winding model based on YOLOv8 to intertwine and compensate the spatial features and channel features; first perform data enhancement on the data set to improve the generalization ability of the model; then send the three feature maps of different scales output by YOLOv8 to the dual-feature winding module; S22: Perform channel feature extraction, as shown in formula (3): Among them, Conv represents one-dimensional convolution, Conv k The subscript k in the convolution kernel represents the size of the convolution kernel, F c (e,u) represents the feature map of the e-th row, u-th column, and c-th channel, H represents the height of the feature map, W represents the width of the feature map, σ represents the Sigmoid activation function, F1 represents the feature map after channel feature recalibration, Conv1 represents a one-dimensional convolution with a convolution kernel of 1×1, and Conv3 represents a one-dimensional convolution with a convolution kernel of 3×3; Perform global average pooling on the input feature map to obtain a feature map of scale 1×1×c, and use cross-channel interaction to improve the feature correlation between different channels; use 1×1 convolution to extract features again, and use the Sigmoid function to generate the corresponding weight coefficients for channel-weighted fusion of the input feature map; S23: spatial feature extraction, as shown in formula (4); F2=σ(Conv21(ξ(Conv23(Tea ReLU(Conv1(F c (e,u)))))) (4) Perform convolution operation on the input feature map to halve the number of channels; use 3×3 convolution to extract spatial information, and activate with ReLU function to enhance the nonlinearity of the model; then use 1×1 convolution to reduce the number of channels, and construct TeaReLU activation function to filter features, improve speed and reduce parameters, as shown in formula (5): Among them, Conv2 represents two-dimensional convolution, Conv k The subscript k in the figure indicates the size of the convolution kernel, σ indicates the Sigmoid activation function, F2 indicates the feature map after spatial feature recalibration, F indicates the original feature map, ξ indicates the ReLU function, Conv21 indicates a two-dimensional convolution with a convolution kernel of 1×1, Conv23 indicates a two-dimensional convolution with a convolution kernel of 3×3, and x indicates the input value of the neuron. is the threshold parameter, controlling the nonlinear turning point, Finally, the Sigmoid function is used to generate the corresponding weight coefficients for spatial weighted fusion of the input feature maps; S24: The feature map obtained after the channel feature and the spatial feature are entangled is sent to the DFOCC-YOLOv8 model to detect the tea shoot frame and calculate the length of the tea shoot frame. If it does not meet the required length requirement, it will be removed. To improve the accuracy, the DFOCC-YOLOv8 model constructs an adaptive occlusion loss function, as shown in formula (6), where the smaller the IoU (a, b), the more serious the occlusion, the greater the difficulty of sample learning, and the greater the corresponding weight; Among them, 1-IoU(a,b) is the occlusion weight, L occ represents the occlusion loss function, a represents the predicted box, b represents the real box, χ pos is the position loss weight coefficient, κ cls is the weight coefficient of the classification loss function, represents the true label of the nth class (0 or 1), logc n Represents the predicted probability c of the nth class n The logarithm of S is used to measure the uncertainty between the predicted probability and the true label. a Represents the area of ​​the prediction box, S b Represents the area of ​​the true box, and N represents the total number of categories.

4. The method for identifying tea shoot picking points by an intelligent robot based on multi-source data fusion according to claim 3 is characterized in that: The S3 specifically includes the following steps: S31: Extract the color feature information of tea leaves; firstly, introduce the gradient information by calculating the change rate of the intensity values ​​of the three color channels of red, green and blue in the image in the horizontal and vertical directions, and detect the area with drastic color change; then, combine the local pixel difference and statistical characteristics, highlight the details and enhance the contrast through the local normalization operation, and use the adjustable parameters to balance the local detail enhancement and the overall smoothing effect, so as to optimize the image details and texture, and adapt to the visual tasks in complex environments, as shown in formula (7): Among them, I R (x,y),I G (x,y),I B (x, y) are the color values ​​of the red, green, and blue channels of a pixel point (x, y) in the image, respectively. Respectively represent the color difference between the pixel and the adjacent pixel in the horizontal and vertical directions, f color Indicates the weight of the color characteristics of tea leaves in image segmentation; S32: Extracting the texture feature information of tea leaves; First, the image texture related information is obtained by summing up different directions and scales; Then, the sum of local pixel intensities is calculated in the denominator, and a specific value is subtracted from it and the absolute value is taken to measure the local difference; Then, this difference is processed by exponential functions and parameters to suppress the mutation area and reduce noise; Finally, the various parts of the formula and the adjustable parameters are used to balance the local detail enhancement and the overall smoothing effect to optimize the image texture, as shown in formula (8): Among them, R d (x, y, s) represents the texture response, L s (x, y, d) represents the local contrast or local texture intensity, n' represents the size of the complexity metric area, ε represents the learning rate of the weighting factor, I(x+i', y+j') represents the pixel value at the position (x+i', y+j') in the image, d represents the number of directions, which refers to how many different directions are considered to capture the texture change information, s represents the number of scales, which refers to the number of different sized areas or resolutions we choose when extracting texture features, i' represents the offset in the x direction relative to the current position, j' represents the offset in the y direction relative to the current position, and f texture represents the weight of the texture characteristics of tea leaves in image segmentation, D represents the total number of different directions selected when extracting texture features, and S represents the total number of different region sizes selected when extracting texture features; S33: In order to reflect the morphological changes of tea leaves, a new method based on shape manifold and local symmetry analysis is proposed, as shown in formula (9): in, is the second-order derivative, indicating the curvature of the edge, y(q) represents the vertical coordinate value of the qth point of the shape contour, L is the total length of the contour, and f shape Indicates the weight of the shape characteristics of tea leaves in image segmentation; First, the second-order derivative part is similar to "gradient information", which measures the rate of shape change, including the difference in brightness between a pixel and its neighbors to measure local changes; Next, the summation part combines "local pixel differences" and "statistical characteristics" to highlight the local changes of the function; Finally, the formula combines the two, balancing the shape change rate and local differences to achieve the evaluation of the overall shape characteristics; S34: A new feature extraction method for tea growth law is proposed, as shown in formula (10), focusing on the growth direction difference, local spatial consistency and growth pattern correlation of tea leaves; by quantifying the growth direction difference between adjacent pixels, measuring the growth consistency of local areas, and introducing angle difference measurement, the spatial structure changes of tea leaves at different growth stages are captured; Among them, x and y represent the position of the current pixel in the image, i and j represent the coordinates of the adjacent pixels, and Δθ ij represents the difference in growth direction between pixel i and pixel j, g represents the standard deviation of Gaussian weighting, which is used to control the influence range of growth direction difference in local area, ΔG ij represents the difference in growth pattern between pixel i and pixel j, g L represents the standard deviation of the local spatial weighting, which is used to control the weighting range, θ ij represents the growth direction angle between adjacent pixels i and j, describing the difference in the growth angle of tea leaves, g P represents the standard deviation of the growth mode weights, which is used to control the weighted range of the growth mode correlation. Represents the standard deviation of the fusion weight, which is used to control the influence range of each feature in the fusion process. grow Indicates the weight of tea growth characteristics in image segmentation; S35: Construct a tea shoot segmentation model based on adaptive fusion of multi-source features, as shown in formula (11), and design a learning mechanism to predict the importance of each pixel in different feature dimensions; Specifically: by calculating the features of the local image area and combining the context information to predict the weight of each feature; First, the mapping function learned through the convolutional neural network can determine the feature weight based on the local features, and the Sigmoid activation function suppresses the interference caused by local differences; Next, the different feature maps are processed and multiplied with the parameters to highlight the changes in different features and enhance the local contrast; Finally, by comprehensively processing different features, the segmentation algorithm is used to balance the local detail enhancement and the overall fusion effect, so as to achieve the optimization processing of the comprehensive features of the image; f fusion =UNet(σ(φ color (f color ))·P+σ(φ texture (f texture ))·P+σ(φ shape (f shape ))·P+σ(φ grow (f grow ))·P)(11) Among them, f fusion (x, y) is the feature map after segmentation, σ represents the Sigmoid activation function, φ color ,φ texture ,φ shape ,φ grow It is a mapping function learned through the convolutional neural network CNN, which predicts the weight of each feature based on the local features of the input image. UNet represents the segmentation algorithm, and P represents the tea shoot detection frame obtained by S24.

5. The method for identifying tea shoot picking points by an intelligent robot based on multi-source data fusion according to claim 4 is characterized in that: The S4 specifically includes: S41: extract key points from the tea segmentation image obtained in S35; use the key point detection method HRNet to extract the b(b∈N + ) The coordinate information of the intersection of the leaves of the young tea shoots, and then take 1-2mm below the intersection to determine the two-dimensional coordinates of the picking point of the young tea shoots Then the segmented image and the point cloud data of the LiDAR are fused to achieve data synchronization and registration; and the LiDAR data is projected onto the segmented image to obtain the depth value of the tea leaf part; S42: Using the depth data optimization technology, combining the tea growth characteristics with the tea shoot leaf point cloud depth data for inference, as shown in formula (12); specifically, by modeling the tea leaf growth angle, adjusting the leaf depth value, overcoming the depth deviation caused by tea growth, and inferring the depth of the picking point; at the same time, considering that the leaf points closer to the leaf centroid are in the same plane as the picking point, these points are given higher weights to avoid contradictions when the picking point position is unknown; in, represents the depth value of the i-th point on the b-th tea shoot leaf. Indicates the depth value of the bth tea shoot picking point, represents the growth angle of the bth tea shoot leaf, It means to establish a model for the growth characteristics of the b-th tea shoot leaf to obtain the depth of the tea shoot picking point. c is a constant that controls the influence of distance on the weight. n represents the total number of point clouds on the b-th tea shoot leaf. S43: In the tea picking task, the correction of wind speed and wind direction to depth is considered; as shown in formula (13), by dynamically considering the real-time changes of wind speed, wind direction and environmental factors, the depth data is accurately adjusted in combination with the correction function, the wind error is eliminated, and the picking accuracy and stability are improved; Among them, ΔZ wind (t) represents the wind error correction value, which represents the correction amount of wind speed and wind direction to depth, v(t) represents the wind speed at time t, θ(t) represents the wind direction angle at time t, w1 represents the initial influence coefficient of wind speed and wind direction on depth error, w2 is an exponential decay coefficient, which is used to represent the weakening of wind speed influence over time, and w3 is a time-related correction function, which is used to control the influence of wind speed on depth over time. represents the time constant of the wind error, and f(E(t)) represents the environmental correction function; considering the influence of environmental factors such as humidity and temperature on the correction of wind speed and wind direction, the environmental correction function f(E(t)) is used to adjust the wind error, and E(t) represents the environmental variable at time t; S44: After completing the depth correction, determine the picking point of the tea leaves 6. Intelligent robot tea shoot picking point recognition system based on multi-source data fusion, characterized by: include: Image acquisition module, used to obtain nighttime images of tea leaves; An image enhancement module, used to enhance the image based on an image decomposition model; The tea leaf shoot detection module is used to detect tea leaf shoots in the enhanced image and obtain a two-dimensional tea leaf detection frame; The tea shoot segmentation module is used to process the tea 2D detection frame, remove the background irrelevant to the tea leaves, and obtain the tea shoot segmentation image; The depth information acquisition module is used to obtain depth information by introducing the point cloud data of the laser radar, and to synchronize and align the data with the segmented image of the young tea leaves; The picking point positioning module is used to build a picking point positioning system based on depth data optimization and wind speed correction, optimize and correct the depth data, and determine the three-dimensional coordinates of the picking point of the young tea shoots.

Citation Information

Patent Citations

  • Precise identification method for tea tender shoot grade in complex environment

    CN115810106A

  • Tea dynamic target identification method fusing three-dimensional wind speed and vision

    CN116619368A

  • Tea leaf tender shoot identification and picking point positioning method

    CN116958823A

  • Tea bud and leaf pose estimation method and system based on picking robot

    CN118096891A

Cited By

  • Tea picking method, system and equipment based on depth point cloud analysis and medium

    CN120572522A

  • Double-arm intelligent collaborative tea picking method and device based on multi-strategy dynamic scheduling

    CN120753093A

  • Method for recognizing and positioning flower buds and flowers for picking robot to carry out partitioned picking

    CN121305366A

  • Petrochemical engineering operation monitoring method based on image enhancement and identification joint optimization

    CN121661409A