Tea tender shoot identification and positioning method based on double-dictionary learning and sparse selection memory

Through the method based on double dictionary learning and sparse selection memory, combined with improved YOLOV11 and random forest algorithm, the precise identification and positioning of tea tender shoots in complex environments is achieved, and the problem of inaccurate recognition and positioning of tea picking robots in the existing technology is solved, and the picking efficiency and quality are improved.

CN120182369APending Publication Date: 2025-06-20JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242459.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing tea picking robots are difficult to accurately identify and locate tea tender shoots in complex unstructured environments, resulting in low picking efficiency and low quality.

Method used

The method based on double dictionary learning and sparse selection memory is adopted, combined with the improved YOLOV11 object detection algorithm and random forest algorithm, the global and local posture characteristics of multi-pose overlapping tea tender shoots are extracted, and the precise identification and positioning of tea tender shoots is achieved through double dictionary learning and sparse encoding.

Benefits of technology

It improves the accurate identification and positioning performance of tea tender shoots in complex backgrounds, enhances the robustness of multi-pose overlapping situations, and supports the rapid and accurate perception and positioning of tea picking robots in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005294572730000033
    Figure BDA0005294572730000033
  • Figure BDA0005294572730000041
    Figure BDA0005294572730000041
  • Figure BDA0005294572730000061
    Figure BDA0005294572730000061
Patent Text Reader

Abstract

The invention discloses a tea tender tip recognition and positioning method based on double dictionary learning and sparse selection memory, and the method comprises the steps: extracting the features of a multi-pose overlapping tea tender tip depth map through an improved YOLOV11 target detection algorithm, and constructing global pose feature vectors and local pose feature vectors under different scales; modeling the attitude features of the tea tender shoots from global and local dimensions by using a double-dictionary learning method, and training global attitude feature vectors and local attitude feature vectors of a global dictionary and a local dictionary based on sparse coding; training and classifying the feature vectors by using a random forest algorithm; constructing a decision tree model by utilizing the classified feature vectors, and effectively judging the postures and positions of the multi-posture overlapping tea tender shoots by integrating and learning a plurality of decision trees and combining the results of all decision trees; according to the poses of the tea tender shoots judged according to the training result of the decision tree, the picking positions of the tea tender shoots are judged in combination with the tea tender shoot depth map; and accurate identification and positioning of the tea tender tips under the complex background and multi-posture conditions are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural automation and intelligent picking, and particularly to a method for identifying and positioning tea shoots based on double dictionary learning and sparse selection memory. Background Art

[0002] This special economic crop, tea, mainly grows in hilly mountainous areas and other places. The orchard terrain is uneven and the planting standards are inconsistent. Special economic crops refer to crops with special uses or relatively high economic value in addition to the main food crops. The tea picking environment is mostly a complex unstructured environment such as overlapping picking objects and leaf occlusion. At present, manual picking or semi-automatic picking is adopted, with high labor costs and low efficiency. According to the growth characteristics of tea, the state of tender buds or one bud with one leaf is usually selected during picking to ensure the picking quality of tea and its subsequent edible value.

[0003] Currently, domestic and foreign enterprises and institutions represented by Abundant in New Zealand, Denso in Japan, KAWASAKI, and Nanjing Institute of Agricultural Machinery have all carried out research on key technologies such as precise perception and positioning of tea picking robots. The key perception and positioning technologies adopted by the fruit picking robots of Abundant in New Zealand and Denso in Japan have problems such as low recognition accuracy, long recognition time, and especially low positioning accuracy. It belongs to drag-type picking, with low picking quality and is not conducive to tea picking. The perception and positioning systems adopted by the tea picking robots of KAWASAKI and Nanjing Institute of Agricultural Machinery are only used for global recognition and navigation during picking, belonging to razor-type picking and cannot be used for picking tea shoots. Summary of the Invention

[0004] In order to solve the deficiencies in the prior art, the present application proposes a method for identifying and positioning tea shoots based on double dictionary learning and sparse selection memory, aiming to solve the problems of multi-pose evaluation of overlapping tea shoots, sparse selection memory, and pose-robust sparse dictionary construction method. By combining the random forest algorithm with the double dictionary learning method for multi-pose evaluation of tea shoots, precise identification and positioning of tea shoots in complex backgrounds and multi-pose situations can be achieved.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for identifying and positioning tea shoots based on double dictionary learning and sparse selection memory, comprising the following steps:

[0007] Step 1: Collect depth images of tea shoots in a complex scene with multi-poses and overlapping, and use the improved YOLOV11 object detection algorithm to extract multi-scale features of the depth map of multi-pose overlapping tea shoots, and construct global pose feature vectors and local pose feature vectors at different scales;

[0008] Step 2: Based on the global pose feature vector and the local pose feature vector, use the double dictionary learning method to model the pose features of the tea shoot from both the global and local dimensions simultaneously, train the global dictionary and the local dictionary, and apply the trained dictionaries to the sparse coding process of the pose feature vector;

[0009] Step 3: Based on the globally and locally pose feature vectors after sparse coding, use the random forest algorithm to train and classify the above feature vectors; construct a decision tree model using the classified feature vectors, and through the ensemble learning of multiple decision trees, combine the results of all decision trees to effectively judge the pose and position of the multi-pose overlapping tea shoots;

[0010] Step 4: According to the pose of the tea shoot determined by the decision tree training result, combined with the depth map of the tea shoot, determine the picking part of the tea shoot.

[0011] Furthermore, the process of Step 1 includes:

[0012] Step 1.1: Preprocess the collected depth image of the tea shoot,

[0013] Step 1.2: Construct an improved YOLOV11 object detection algorithm; the improved YOLOV11 object detection algorithm introduces a bidirectional feature pyramid network BiFPN module and a multi-scale attention aggregation module MSAA in the neck; introduce an EIoU loss function and a Shape-IoU loss function in the loss calculation module;

[0014] Step 1.3: When using the above improved YOLOV11 object detection algorithm to identify the depth map of multi-pose overlapping tea shoots, it can extract global pose features and local pose features at different scales, and construct a global pose feature vector and a local pose feature vector respectively.

[0015] Furthermore, the image preprocessing in Step 1.1 is preprocessed using a Gaussian filter to remove noise and isolated points in the depth map and smooth the depth map image.

[0016] Furthermore, the network structure of the improved YOLOV11 object detection algorithm includes a Backbone module, a Neck module, and a detection head; the Backbone module includes a first Conv unit, a second Conv unit, a first C3k2 unit, a third Conv unit, a second C3k2 unit, a fourth Conv unit, a third C3k2 unit, a fifth Conv unit, a fourth C3k2 unit, an SPPF unit, and a C2PSA unit connected in sequence; the Backbone module extracts features from the input depth map of the tea shoot;

[0017] The Neck module is divided into three processing routes. The first processing route includes a sixth Conv unit, an MSAA module, a first Concat unit, and a fifth C3k2 unit connected in sequence. The sixth Conv unit is connected to the second C3k2 unit;

[0018] The second processing route includes an eighth Conv unit, a second Concat unit, a sixth C3k2 unit, a ninth Conv unit, and a second Concat unit connected in sequence. The ninth Conv unit is connected to the first Concat unit through a first Upsample unit; the second Concat unit is connected to the fifth C3k2 unit through a seventh Conv unit. The eighth Conv unit is connected to the third C3k2 unit;

[0019] The third processing route includes an eleventh Conv unit, a BiFPN module, a second Upsample unit, a third Concat unit, and an eighth C3k2 unit connected in sequence. Among them, the third Concat unit is connected to the seventh Conv unit through a tenth Conv unit. The eleventh Conv unit is connected to the C2PSA unit.

[0020] Furthermore, the double dictionary learning method is as follows: Use the online dictionary learning method to train the local dictionary in the local pose feature vector; Use the K-SVD algorithm to train the global dictionary on the global pose feature vector; Obtain the globally and locally pose feature vectors after sparse coding.

[0021] Furthermore, the process of step 3 includes:

[0022] Step 3.1: Assume that the size of the dataset is N. Bagging will generate B training subsets, and each subset contains a dataset D of N samples = {(x i , y i )| i = 1, 2, 3......N}, where x i is the pose feature vector, and y i is the corresponding label, corresponding to the pose category or position coordinates of the tea shoot tip; Divide the original dataset into a training set and a test set to ensure that the training set contains multi-pose and overlapping tea shoot tip poses and position samples, improving the generalization ability of the model;

[0023] Step 3.2: On each subset D i of the decision tree, construct a decision tree h i , until the stopping condition is met;

[0024] Step 3.3: For a new input sample x, each decision tree h i will give a prediction result

[0025] Furthermore, for the classification task, the majority voting method is adopted to determine the final prediction result:

[0026]

[0027] For the regression task, the average value of all prediction results is calculated:

[0028]

[0029] where m is the number of decision trees.

[0030] Advantages of the present invention:

[0031] (1) The objective of the present invention is to propose a tea shoot recognition network with high precision and high efficiency, while enhancing the robustness to multi-pose overlapping situations. By specifically learning and sparsely representing the global and local pose feature vectors at different scales, the model can solve the problems of accurate recognition and positioning of special economic crops in complex unstructured environments such as overlapping and occlusion. The global dictionary represents the general tea shoot pose patterns, and the local dictionary identifies the detailed features. During the training process, sparse coding is used to represent the pose features as a linear combination of the dictionaries, and the dictionaries are continuously optimized through the K-SVD algorithm, enabling the model to effectively handle the multi-pose overlapping problem in complex backgrounds and improving the accurate recognition and positioning performance of tea shoots.

[0032] (2) The advantage of the present invention is to provide services for the picking of special economic crops by the latest intelligent precision picking robot, and solve the problems of accurate recognition and positioning of special economic crops in complex unstructured environments such as overlapping and occlusion. In a complex environment, the picking robot for special economic crops can quickly and accurately sense and position, meeting the requirements of the system for the pose data of the picking robot for special economic crops. Description of the Drawings

[0033] Figure 1 is the overall operation logic diagram of the method of the present invention.

[0034] Figure 2 is the network structure diagram of the improved YOLOv11 algorithm described in the method of the present invention.

[0035] Figure 3 is the network structure of the improved YOLOv11 method of the present invention, where a is the network structure diagram of the MSAA network, b(1) is the network structure diagram of the FPN module, and b(2) is the network structure diagram of the BiFPN module.

[0036] Figure 4 is the double dictionary learning method used in the present invention, where a is the local dictionary trained by the online learning algorithm, and b is the global dictionary trained by the K-SVD algorithm.

[0037] Figure 5The network structure diagram of the improved random forest algorithm in the present invention. Detailed implementation manners

[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0039] Refer to Figures 1-5 , the present invention designs a tea shoot recognition and positioning method based on double dictionary learning and sparse selection memory, and the specific implementation steps are as follows:

[0040] Step 1: Use an RGB-D450 depth camera to collect depth images of tea shoots in complex scenarios such as multi-poses and overlaps, and preprocess the depth images of tea shoots. Construct an improved YOLOV11 object detection algorithm, and use the improved YOLOV11 object detection algorithm to extract multi-scale features of multi-posed and overlapped tea shoot depth images, and construct pose feature vectors of global features (overall contour, size, proportion of tea shoots; average color of tea shoots; texture pattern, roughness, etc. on the surface of tea shoots) and local features (local texture, color, shape) at different scales. The specific process is as follows:

[0041] Step 1.1: Use a Gaussian filter to preprocess the depth image of tea shoots, remove noise and isolated points in the depth image, smooth the depth image, improve the quality of the depth image and reduce the influence of random noise.

[0042] The formula of the Gaussian filter is:

[0043]

[0044] Among them, G(x, y) is the value of the Gaussian filter, (x, y) is the coordinate distance from the center point, and σ is the standard deviation, which controls the width of the filter.

[0045] Step 1.2: Construct an improved YOLOV11 object detection algorithm, the network structure of which is as Figure 2 shown in FIGS. 2 and 3, and includes a Backbone module, a Neck module, and a detection head; the Backbone module includes a first Conv unit, a second Conv unit, a first C3k2 unit, a third Conv unit, a second C3k2 unit, a fourth Conv unit, a third C3k2 unit, a fifth Conv unit, a fourth C3k2 unit, an SPPF unit, and a C2PSA unit connected in sequence; the Backbone module extracts features from the input depth image of tea shoots.

[0046] The Neck module is divided into three processing routes. The first processing route includes a sixth Conv unit, an MSAA module, a first Concat unit, and a fifth C3k2 unit connected in sequence. The sixth Conv unit is connected to the second C3k2 unit.

[0047] The second processing route includes an eighth Conv unit, a second Concat unit, a sixth C3k2 unit, a ninth Conv unit, and a second Concat unit connected in sequence. The ninth Conv unit is connected to the first Concat unit through a first Upsample unit; the second Concat unit is connected to the fifth C3k2 unit through a seventh Conv unit. The eighth Conv unit is connected to the third C3k2 unit.

[0048] The third processing route includes an eleventh Conv unit, a BiFPN module, a second Upsample unit, a third Concat unit, and an eighth C3k2 unit connected in sequence. Among them, the third Concat unit is connected to the seventh Conv unit through a tenth Conv unit. The eleventh Conv unit is connected to the C2PSA unit.

[0049] The detection head is used to output features of different sizes.

[0050] In the present invention, a bidirectional feature pyramid network (BiFPN) module and a multi-scale attention aggregation (MSAA) module are introduced into the neck of the YOLOV11 object detection algorithm to improve the model's detection ability for multi-scale objects. An EIoU loss function and a Shape-IoU loss function are introduced into the loss calculation module, which takes into account the shape and scale of the bounding box while improving the detection accuracy of the model for small objects, further enhancing the model's detection ability.

[0051] Step 1.3: When using the improved YOLOV11 object detection algorithm to identify the depth map of multi-pose overlapping tea shoots, global pose features and local pose features at different scales can be extracted, and a global pose feature vector and a local pose feature vector are respectively constructed.

[0052] Step 2: Based on the global pose feature vector and the local pose feature vector, the double dictionary learning method is used to model the pose features of tea shoots from both the global and local dimensions simultaneously, training the global dictionary and the local dictionary to achieve accurate representation of multi-pose overlapping tea shoots. The trained dictionaries are applied to the sparse coding process of the pose feature vectors, and each pose feature vector is transformed into a sparse code through sparse representation, enhancing the efficiency and robustness of the accurate recognition and positioning model optimized by double dictionary learning and sparse coding.

[0053] Step 2.1: Use the online dictionary learning method to train a local dictionary from the local pose feature vectors extracted from the depth map of tea shoots, so that the feature vectors are linearly combined sparsely through the dictionary matrix.

[0054] The process of training the local dictionary is as follows: Take the local pose feature vectors as the input data of the online dictionary learning algorithm, process one small batch or a single data point each time and update the model, and optimize the dictionary through repeated iterations so that it can effectively represent the data, and constrain the representation sparsity through sparse coding.

[0055] First, construct the training dictionary matrix D. Assume that there are k categories in the training dictionary D, and each category contains n samples, that is, the training dictionary D contains k individuals, and each individual contains n special economic crop images. Assume the resolution is w×h, and pull it into a column, v∈R m*N (m = w*h). The entire training dictionary D can be expressed as:

[0056] D = [A1, A2, …, A k = [v 1,1 , v 1,2 , …, v k,n

[0057] where D∈R m*N (N = k*n) represents the training dictionary, A i ∈R m*n represents all the training special economic crop images of the i-th category, and v i,j represents the j-th special economic crop image in A i .

[0058] Suppose the test special economic crop image sample belongs to the i-th category, with a resolution of w×h, and pull it into a column to get y∈R m (m = w*h). According to the theory of sparse representation, the training samples of the same category as y can be represented by it in the form of a linear combination:

[0059] y = a i,1 v i,1 + a i,2 v i,2 + … + a i,n v i,n

[0060] where a i,j represents the corresponding coefficient in v i,j . The above formula only describes the linear combination of the i-th category and must be extended to the entire training dictionary D, that is, the linear combination under all samples:

[0061]

[0062] ​Among them, is a sparse vector. In the theory of sparse representation, the components unrelated to the i-th category are zero, that is, the non-i-th special economic crop image cannot represent the i-th special economic crop. The sparse vector can be written as:

[0063]

[0064] To establish a multi-object recognition model for special economic crops such as tea shoots and perform the final recognition and determination, the linear equation y = Dx must be solved, that is, the sparse coding of the test sample y in the training dictionary D needs to be solved. When solving y = Dx for the recognition problem, a very high requirement is placed on the sparsity of the solution. Problem analysis:

[0065]

[0066] In the process of converting various norms, due to the inherent advantages of the l1 norm, the l0 norm problem is converted into an l1 norm problem for solution:

[0067]

[0068] In the actual process, it is necessary to introduce a noise feature term to enhance the robustness of the system to noise, and the column vectors in the training dictionary D can more accurately represent the test sample y, satisfying:

[0069]

[0070] In online dictionary learning, data samples arrive one by one or in batches. Therefore, sparse coding is usually incremental. For each new sample y t , it is necessary to perform sparse coding with the current dictionary Dt:

[0071]

[0072] In online dictionary learning, every time a new data sample y t arrives, the current dictionary Dt is used to calculate the sparse coefficient x t , and then the dictionary is updated through these systems. When each column in the dictionary changes very little, the dictionary stops iterating. The change in dictionary update can be measured by calculating the difference before and after dictionary update:

[0073] ΔD = ||D t+1 - D t || F

[0074] Among them, ∥·∥F is the Frobenius norm, indicating the overall change of the dictionary matrix. If the dictionary change amount is less than a certain threshold ∈2, the iteration can be stopped:

[0075] ΔD < ∈2

[0076] Step 2.2: Use the K-SVD algorithm to train the global dictionary on the global pose feature vectors. Among them, the K-SVD algorithm uses sparse coding and singular value decomposition (SVD) to iteratively update the dictionary, learning dictionary atoms from a large number of samples, enabling the global pose feature vectors of tea shoots to be compactly encoded through sparse representation, and helping the optimized model to accurately identify the poses and positions of common tea shoots.

[0077] First, construct the training dictionary matrix D. Assume that the training dictionary matrix D contains k categories, and each category contains n samples, that is, the training dictionary D contains k individuals, and each individual contains n special economic crop images. Assume the resolution is w×h, and it is pulled into a column, v∈R m*N (m = w*h).

[0078] For each sample y t ∈R m , calculate its sparse coefficient x t ∈R k such that each sample can be represented by a linear combination of dictionary atoms while maintaining the sparsity of the coefficient x t . The optimization problem is as follows:

[0079]

[0080] where λ is the sparsity regularization parameter that controls the sparsity of the coefficient. Usually, the coordinate descent method is used to solve this optimization problem to obtain the sparse coefficient x t of each sample.

[0081] After calculating the sparse coefficients of each sample, use singular value decomposition (SVD) to update each column of the dictionary. Assume that X = [x1, x2,..., x n is the sparse coefficient matrix of all samples, where each column x t corresponds to the sample y t . Update the dictionary atoms by minimizing the following objective function:

[0082]

[0083] where Y j is the part of all samples corresponding to the dictionary atoms D j (that is, all samples corresponding to the atom Dj).

[0084] For each dictionary atom D j , extract the samples Y j related to D j from the sparse coefficient matrix X. These samples are the samples corresponding to the dictionary atom D jThe sample set when activated.

[0085] For each dictionary atom D j is updated by optimizing the dictionary atom through singular value decomposition (SVD) for the matrix Y j = D j x j perform SVD decomposition:

[0086]

[0087] where U j and V j are orthogonal matrices, and Σ j is a diagonal matrix. According to the SVD decomposition result, the dictionary atom D j is updated. Usually, the constraint norm of the dictionary is maintained during the update, i.e., ||D j ||2 = 1.

[0088] After multiple iterations, we will obtain an optimal dictionary D, which can be effectively used to sparsely represent the global pose features of new tea shoots. The following steps can be used:

[0089]

[0090] where y is the sample, D is the learned dictionary, λ is the sparsity regularization parameter that controls the sparsity of the coefficients. The sparse coefficients x are used to predict the pose and position.

[0091] In the K - SVD algorithm, when the reconstruction error E t of the dictionary changes very little, it is considered that the dictionary has learned enough features and the training stops. The formula for calculating the reconstruction error is:

[0092]

[0093] Stopping condition:

[0094] |E t+1 - E t | < ∈1

[0095] where E t and E t+1 are the reconstruction errors of the current and the previous iteration respectively, and ∈1 is the set threshold, indicating that when the change in the reconstruction error is less than a very small value, the iteration stops.

[0096] The dual dictionary learning method combines the feature learning results of the global and local dictionaries, which can more accurately represent the pose features of multi-pose overlapping tea tender shoots. The global dictionary characterizes the general pose patterns of tea tender shoots, and the local dictionary identifies the detailed features. During the training process, sparse coding is used to represent the pose features as a linear combination of the dictionaries, and the dictionary is continuously optimized by the K-SVD algorithm, enabling the model to effectively handle the multi-pose overlapping problem in complex backgrounds and improving the accurate recognition and positioning performance of tea tender shoots.

[0097] Step 3: Based on the globally pose feature vector and locally pose feature vector after sparse coding obtained by the above dual dictionary learning method; use the random forest algorithm to train and classify the above feature vectors. Based on the classified feature vectors, construct a decision tree model, and through the ensemble learning of multiple decision trees, finally combine the results of all decision trees to effectively judge the pose and position of multi-pose overlapping tea tender shoots.

[0098] Step 3.1: Assume the size of the dataset is N, and Bagging will generate B training subsets, each subset containing a dataset D of N samples = {(x i , y i )| i = 1, 2, 3......N}, where x i is the pose feature vector and y i is the corresponding label (the pose category or position coordinates corresponding to the tea tender shoot). Divide the original dataset into a training set and a test set to ensure that the training set contains samples of multi-pose and overlapping tea tender shoot poses and positions, improving the generalization ability of the model.

[0099] Step 3.2: On each subset D i of the decision tree, construct a decision tree h i , until the stopping condition is met.

[0100] Step 3.3: For a new input sample x, each decision tree h i will give a prediction result For the classification task, use the majority voting method to determine the final prediction result:

[0101]

[0102] For the regression task, calculate the average value of all prediction results:

[0103]

[0104] where m is the number of decision trees.

[0105] Step 4: According to the pose of the tea tender shoot determined by the decision tree training result, combined with the depth map of the tea tender shoot, determine the picking part of the tea tender shoot.

[0106] The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A method for identifying and locating young tea shoots based on dual dictionary learning and sparse selection memory, characterized in that: The steps include: Step 1: Collect the depth image of tea shoots in complex scenes with multiple poses and overlaps, use the improved YOLOV11 target detection algorithm to extract the depth map features of multi-pose overlapping tea shoots using multi-scale features, and construct the global pose feature vectors and local pose feature vectors at different scales; Step 2: Based on the global posture feature vector and the local posture feature vector, the posture features of the tea shoots are modeled from both global and local dimensions using the dual dictionary learning method, and the global dictionary and the local dictionary are trained. The trained dictionary is applied to the sparse coding process of the posture feature vector; Step 3: Based on the global posture feature vector and the local posture feature vector after sparse coding; the above feature vectors are trained and classified using the random forest algorithm; a decision tree model is constructed using the classified feature vectors, and the posture and position of the tea shoots with multiple postures overlapping are effectively judged by integrating the learning of multiple decision trees and combining the results of all decision trees; Step 4: Determine the position of the tea shoots based on the decision tree training results and the tea shoot depth map to determine the picking position of the tea shoots.

2. The method for identifying and locating young tea shoots based on dual dictionary learning and sparse selective memory according to claim 1, characterized in that: The process for Step 1 includes: Step 1.1: Preprocess the collected tea shoot depth images. Step 1.2: Construct an improved YOLOV11 target detection algorithm; the improved YOLOV11 target detection algorithm introduces a bidirectional feature pyramid network BiFPN module and a multi-scale attention aggregation module MSAA in the neck; and introduces an EIoU loss function and a Shape-IoU loss function in the loss calculation module; Step 1.3: When using the above-mentioned improved YOLOV11 target detection algorithm to identify the multi-pose overlapping tea shoot depth map, the global pose features and local pose features at different scales can be extracted, and the global pose feature vector and the local pose feature vector can be constructed respectively.

3. The method for identifying and locating young tea shoots based on dual dictionary learning and sparse selective memory according to claim 1, characterized in that: The image preprocessing in step 1.1 uses a Gaussian filter to remove noise and isolated points in the depth map and smooth the depth map image.

4. The method for identifying and locating young tea shoots based on dual dictionary learning and sparse selective memory according to claim 1, characterized in that: The network structure of the improved YOLOV11 target detection algorithm includes a Backbone module, a Neck module, and a detection head; the Backbone module includes a first Conv unit, a second Conv unit, a first C3k2 unit, a third Conv unit, a second C3k2 unit, a fourth Conv unit, a third C3k2 unit, a fifth Conv unit, a fourth C3k2 unit, an SPPF unit, and a C2PSA unit connected in sequence; the Backbone module extracts features from an input tea shoot depth map; The Neck module is divided into three processing routes. The first processing route includes the sixth Conv unit, the MSAA module, the first Concat unit, and the fifth C3k2 unit connected in sequence. The sixth Conv unit is connected to the second C3k2 unit; The second processing route includes the eighth Conv unit, the second Concat unit, the sixth C3k2 unit, the ninth Conv unit, and the second Concat unit connected in sequence, and the ninth Conv unit is connected to the first Concat unit through the first Upsample unit; the second Concat unit is connected to the fifth C3k2 unit through the seventh Conv unit. The eighth Conv unit is connected to the third C3k2 unit; The third processing route includes the eleventh Conv unit, the BiFPN module, the second Upsample unit, the third Concat unit, and the eighth C3k2 unit connected in sequence, wherein the third Concat unit is connected to the seventh Conv unit through the tenth Conv unit. The eleventh Conv unit is connected to the C2PSA unit.

5. The method for identifying and locating young tea shoots based on dual dictionary learning and sparse selective memory according to claim 1, characterized in that: The dual dictionary learning method is as follows: using an online dictionary learning method to train a local dictionary in a local posture feature vector; using a K-SVD algorithm to train a global dictionary on a global posture feature vector; and obtaining a sparsely coded global posture feature vector and a local posture feature vector.

6. The method for identifying and locating young tea shoots based on dual dictionary learning and sparse selective memory according to claim 1, characterized in that: The process for step 3 includes: Step 3.1: Assuming the size of the dataset is N, Bagging will generate B training subsets, each of which contains a dataset D = {(x i ,y i )|i=1,2,3......N}, where x i is the posture feature vector, y i is the corresponding label, corresponding to the posture category or position coordinate of the tea shoot; the original data set is divided into a training set and a test set, ensuring that the training set contains multi-posture, overlapping tea shoot posture and position samples to improve the generalization ability of the model; Step 3.2: In each subset D of the decision tree i On the top, build a decision tree h i , until the stopping condition is met; Step 3.3: For a new input sample x, each decision tree h i A prediction result will be given 7. The method for identifying and locating young tea shoots based on dual dictionary learning and sparse selective memory according to claim 6, characterized in that: For classification tasks, the majority voting method is used to determine the final prediction result: For regression tasks, calculate the average of all predictions: Where m is the number of decision trees.