Dense tea tree bud leaf picking point positioning method, device and equipment and medium

By adopting the trained detection optimization model and picking point positioning model in tea tree bud and leaf picking technology, combined with Res2Net and improved HRNet network model, the technical bottlenecks of dense tea bud detection and occluded picking point recognition are solved, and high-precision tea tree bud and leaf detection and picking point positioning are achieved.

CN120198796APending Publication Date: 2025-06-24SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510259488.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has technical bottlenecks in intensive tea bud detection and occluded picking point identification, resulting in low detection accuracy and difficulty in positioning the picking point.

Method used

A dense tea tree bud and leaf picking point positioning method is adopted, and the target detection and picking point positioning model is optimized through the trained tea tree bud and leaf detection and the obstructed tea tree bud and leaf picking point positioning model is combined with the Res2Net module, the path aggregate feature pyramid network and the improved HRNet network model to perform target detection and picking point positioning.

Benefits of technology

It significantly improves the detection accuracy of tea tree buds and leaves and the positioning accuracy of occluded picking points, enhances the robustness and adaptability of the network, and can effectively identify and locate tea tree buds and leaves in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198796A_ABST
    Figure CN120198796A_ABST
Patent Text Reader

Abstract

The invention relates to a dense tea tree bud leaf picking point positioning method and device, equipment and a medium, and the method comprises the steps: extracting unshielded tea tree bud leaves and the convolution features of the shielded tea tree bud leaves in a convolution backbone network, and mapping the convolution features to graph nodes for representation in a node embedding module, embedding graph node information through the learned adjacent matrix to capture a neighborhood relationship; the relation between different graph nodes is dynamically adjusted in a dynamic adjacency matrix weighting module, so that adaptive learning is carried out in a graph structure, and graph convolution operation is adopted in a graph relation layer to allow information interaction and propagation between different picking points of unshielded tea tree buds and leaves so as to capture global and local context information; and optimizing the relative position relationship between the picking points of the unshielded tea tree bud leaves and the shielded tea tree bud leaves by adopting the relative position loss so as to deduce the picking points of the shielded tea tree bud leaves. According to the invention, the picking point positioning of the shielded tea tree bud leaves can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of agricultural production, and particularly to a method for locating picking points of intensive tea tree buds and leaves, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Tea is one of the important cash crops in China. Tea picking is a key link in tea production, and the quality of picking directly affects the final quality of tea. At present, tea picking mainly relies on manual operation. However, manual picking has problems such as low efficiency, high cost, and being greatly affected by the proficiency of picking workers. Therefore, it is particularly important to develop a high-precision automated tea picking technology to replace manual picking and traditional mechanical picking, and the core of achieving this goal lies in the precise detection of tea greens and the precise positioning of picking points.

[0003] In recent years, automatic picking robots based on vision technology have gradually been applied to the picking process of famous and high-quality teas, greatly improving the efficiency and intelligent level of tea picking. However, the natural growth environment of tea leaves in tea gardens is complex and changeable, and there are still many technical bottlenecks and practical application challenges in the detection of intensive tea buds and the recognition of occluded picking points in the existing technology.

[0004] However, the growth environment of tea leaves in tea gardens is complex, and there are still many deficiencies in the detection of intensive tea buds and the recognition of occluded picking points in the existing technology. On the one hand, traditional detection models are easily interfered when dealing with intensive tea buds, resulting in low detection accuracy; on the other hand, when the picking point is occluded by leaves or other tea buds, the existing methods often cannot accurately locate it. In addition, although the single-stage detection model has a certain efficiency, its robustness in complex backgrounds is poor and it cannot meet the actual needs of tea bud picking.

[0005] In summary, in view of the fact that there are still many technical bottlenecks in the detection of intensive tea buds and the recognition of occluded picking points in the existing technology, and traditional detection models are easily interfered when dealing with intensive tea buds, resulting in problems such as low detection accuracy, the applicant has made corresponding explorations to solve this problem. Summary of the Invention

[0006] The purpose of the present application is to solve the above problems and provide a method for locating picking points of intensive tea tree buds and leaves, a corresponding device, an electronic device, and a computer-readable storage medium.

[0007] To meet the various purposes of the present application, the following technical solutions are adopted:

[0008] A method for locating picking points of intensive tea tree buds and leaves proposed for one of the purposes of the present application includes:

[0009] Respond to the instruction for locating the picking points of dense tea tree buds and leaves, obtain a tea leaf image containing the tea tree buds and leaves to be picked, and use the optimized tea tree bud and leaf detection model trained to the convergence state to perform object detection on the tea leaf image to determine the detection box information corresponding to the tea tree buds and leaves to be picked in the tea leaf image. Based on the detection box information corresponding to the tea tree buds and leaves to be picked, divide the tea tree buds and leaves to be picked into unobstructed tea tree buds and leaves and occluded tea tree buds and leaves;

[0010] Input the detection box information corresponding to the unobstructed tea tree buds and leaves and the occluded tea tree buds and leaves into the pre-trained picking point positioning model for occluded tea tree buds and leaves. Extract the convolutional features of the unobstructed tea tree buds and leaves and the occluded tea tree buds and leaves in the convolutional backbone network of the picking point positioning model for occluded tea tree buds and leaves. Map the convolutional features to graph node representations in the node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships;

[0011] Dynamically adjust the relationships between different graph nodes in the dynamic adjacency matrix weighting module for adaptive learning in the graph structure. In the graph relationship layer, use graph convolution operations to allow information interaction and propagation between different picking points of the unobstructed tea tree buds and leaves to capture global and local context information;

[0012] Optimize the relative position relationship between the picking points of the unobstructed tea tree buds and leaves and the occluded tea tree buds and leaves using the relative position loss to infer the picking points of the occluded tea tree buds and leaves to complete the picking point positioning of dense tea tree buds and leaves.

[0013] Optionally, the optimized tea tree bud and leaf detection model includes a Res2Net module, a path aggregation feature pyramid network, and a detection head. Among them, the detection head includes an IoU-aware class score and a bounding box fine-tuning module.

[0014] Optionally, the step of using the optimized tea tree bud and leaf detection model trained to the convergence state to perform object detection on the tea leaf image to determine the detection box information corresponding to the tea tree buds and leaves to be picked in the tea leaf image includes:

[0015] In the Res2Net module, by introducing a multi-scale convolution module to process feature information of multiple scales in one residual block, effectively extract features in complex and dense tea leaf images to capture details and context information;

[0016] In the path aggregation feature pyramid network, use the semantic information of high-level features and the spatial information of low-level features to generate fine multi-scale feature maps;

[0017] In the detection head, the IoU perception class score accurately identifies the target by introducing IoU information and fusing the confidence and positioning accuracy of each tea shoot and leaf to be picked. In the bounding box fine-tuning module, the position and size of the bounding box are further optimized to accurately match the detection result with the actual target.

[0018] Optionally, in the graph relationship layer, through the mechanism of the graph neural network, the adjacency matrix and feature propagation are used to explicitly model the spatial and semantic relationships between the picking points of the unoccluded tea shoots and leaves, so as to enhance the model's understanding of the dependency relationships between the picking points of the tea shoots and leaves to be picked. By integrating the information of other picking points associated with the unoccluded tea shoots and leaves, the picking points of the occluded tea shoots and leaves are inferred.

[0019] Optionally, the basic network architecture of the occluded tea shoot and leaf picking point positioning model is an improved HRNet network model, and the improved HRNet network model includes a convolutional backbone network, a node embedding module, a dynamic adjacency matrix weighting module, and a graph relationship layer.

[0020] Optionally, the graph relationship layer includes a GraphConv layer, a BatchNorm layer, and a ReLU layer.

[0021] Optionally, the tea shoots and leaves to be picked include single buds, one bud with one leaf, and one bud with two leaves.

[0022] A dense tea shoot and leaf picking point positioning device provided to meet another object of the present application includes:

[0023] A tea shoot and leaf occlusion detection module, configured to respond to an instruction to locate picking points for dense tea shoots and leaves, obtain a tea leaf image containing the tea shoots and leaves to be picked, perform target detection on the tea leaf image using a tea shoot and leaf detection optimization model trained to a convergent state, so as to determine the detection box information corresponding to the tea shoots and leaves to be picked in the tea leaf image, and divide the tea shoots and leaves to be picked into unoccluded tea shoots and leaves and occluded tea shoots and leaves based on the detection box information corresponding to the tea shoots and leaves to be picked;

[0024] A neighborhood relationship capturing module, configured to input the detection box information corresponding to the unoccluded tea shoots and leaves and the occluded tea shoots and leaves into a pre-trained occluded tea shoot and leaf picking point positioning model, extract the convolutional features of the unoccluded tea shoots and leaves and the occluded tea shoots and leaves in the convolutional backbone network of the occluded tea shoot and leaf picking point positioning model, map the convolutional features to graph node representations in the node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships;

[0025] The context information capture module is set to dynamically adjust the relationships between different graph nodes in the dynamic adjacency matrix weighting module for adaptive learning in the graph structure. Graph convolution operations are employed in the graph relationship layer to allow information interaction and propagation among different picking points of the unoccluded tea tree buds and leaves, so as to capture global and local context information;

[0026] The picking point positioning module is set to optimize the relative position relationship between the picking points of the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves by using relative position loss, so as to infer the picking points of the occluded tea tree buds and leaves and complete the picking point positioning of dense tea tree buds and leaves.

[0027] An electronic device provided to meet another objective of the present application includes a central processing unit and a memory. The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method for positioning picking points of dense tea tree buds and leaves according to the present application.

[0028] A computer-readable storage medium provided to meet another objective of the present application stores, in the form of computer-readable instructions, a computer program implemented according to the method for positioning picking points of dense tea tree buds and leaves. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.

[0029] Compared with the prior art, in view of the many technical bottlenecks still existing in the prior art in the aspects of dense tea bud detection and occluded picking point recognition, and the problem that traditional detection models are vulnerable to interference when dealing with dense tea buds, resulting in low detection accuracy, the present application includes but is not limited to the following beneficial effects:

[0030] First, the present application can significantly improve the detection accuracy of tea tree buds and leaves. The tea tree bud and leaf detection optimization model in the present application combines the Res2Net module with the path aggregation feature pyramid network (PAFPN), enhancing the multi-scale expression ability of the feature map. This combination helps to provide finer-grained multi-scale features, enabling the network to more accurately perform target detection when facing dense tea buds. This means that even when tea buds are in a complex background with dense overlap, the detection model can still identify the targets with high accuracy, avoiding misjudgment and missed judgment.

[0031] Second, this application can greatly improve the positioning accuracy of the picking points of occluded tea tree buds and leaves. In the background of dense tea buds, many tea tree buds and leaves may be occluded by other leaves, and traditional detection methods are prone to inaccurate identification of these occluded tender buds. By combining an improved HRNet network model, especially by introducing graph convolution operations in the graph relationship layer, it is possible to more effectively capture neighborhood relationships and global context information. In this way, the model can adaptively adjust the relationships between graph nodes in the graph structure, promote the interaction of information between different picking points, and thus infer the positions of occluded picking points, greatly improving the positioning ability of occluded tea tree buds and leaves.

[0032] Third, the robustness and adaptability of the network are enhanced. The optimized model for tea tree bud and leaf detection in this application makes the detection head better fuse the confidence of the presence of tea buds and the positioning accuracy by introducing the IoU-aware class score and the bounding box fine-tuning module, and optimizes the detection of dense tea buds in complex backgrounds. Especially in the tea garden environment with complex backgrounds, the model can effectively cope with various interference factors, improving the robustness of the network in practical applications. The bounding box fine-tuning module further optimizes the size and position of the bounding box to ensure a more accurate match with the actual target.

[0033] Fourth, this application solves the problem of grading the picking points in the tea green area. This application combines dense object detection and occluded key point detection. First, it identifies the tea green area through an improved dense tea green detection network, and then performs precise picking point positioning in these areas through the improved HRNet. This hierarchical picking point positioning method can implement various picking strategies such as single-bud picking, one-bud-one-leaf picking, and one-bud-two-leaf picking, meeting the picking requirements of different tea varieties and tea green grades, and providing an efficient solution for the graded picking of tea leaves.

[0034] Fifth, this application solves the technical bottleneck in the complex tea garden environment. The growth environment of the tea garden is complex and changeable, and there are often occlusions and overlaps between tea buds. The detection accuracy of existing technologies is often greatly affected in such an environment. By combining the optimized model for tea tree bud and leaf detection and the model for positioning the picking points of occluded tea tree buds and leaves, this application can still maintain a high detection accuracy in complex backgrounds, especially in scenes of dense and occluded tea buds, breaking through the limitations of traditional methods.

[0035] Furthermore, by combining vision technology and advanced deep learning models, this application can achieve precise detection and positioning of tea tree buds and leaves. Automated picking robots can efficiently perform picking tasks based on the detection results of the model. This greatly improves the efficiency of tea leaf picking, reduces the cost of manual picking, and enhances the intelligent level of the tea garden, promoting the development of agricultural automation technology. Brief Description of the Drawings

[0036] The above and / or additional aspects and advantages of the present application will become apparent and understandable from the following description of embodiments in conjunction with the accompanying drawings, where:

[0037] Figure 1 is a schematic flow chart of a method for locating picking points of dense tea tree buds and leaves in an embodiment of the present application;

[0038] Figure 2 is an exemplary network architecture of a Res2Net module in an embodiment of the present application;

[0039] Figure 3 is an exemplary network architecture of an optimized model for detecting tea tree buds and leaves in an embodiment of the present application;

[0040] Figure 4 is an exemplary network architecture of a bounding box fine-tuning module in an embodiment of the present application;

[0041] Figure 5 is an exemplary network architecture of an improved HRNet in an embodiment of the present application;

[0042] Figure 6 is an exemplary network architecture of a graph relationship layer in HRNet in an embodiment of the present application;;

[0043] Figure 7 is a principle block diagram of a device for locating picking points of dense tea tree buds and leaves in an embodiment of the present application;

[0044] Figure 8 is a schematic structural diagram of a computer device in an embodiment of the present application. Detailed Description of the Embodiment

[0045] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0046] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0047] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present application pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0048] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive without transmitting, and devices with receiving and transmitting hardware that have the receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm-top computers or other devices, which are conventional laptop and / or palm-top computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed form at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be devices such as a smart TV, a set-top box, etc.

[0049] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0050] It should be noted that the concept of "server" in this application can similarly be extended to the case of server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be called through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this when implementing the network deployment method of this application.

[0051] One or several technical features of this application, unless expressly specified, can either be deployed on the server and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.

[0052] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client with sufficient device capabilities for direct invocation. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.

[0053] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0054] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for each of the embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0055] For each of the embodiments to be disclosed in this application, unless expressly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0056] Please refer to Figure 1 , in one embodiment of the dense tea shoot and leaf picking point positioning method of this application, it includes:

[0057] Step S10: In response to the instruction for locating the picking points of dense tea tree buds and leaves, obtain a tea leaf image containing the tea tree buds and leaves to be picked, and use the optimized tea tree bud and leaf detection model trained to the convergence state to perform object detection on the tea leaf image to determine the detection box information corresponding to the tea tree buds and leaves to be picked in the tea leaf image. Based on the detection box information corresponding to the tea tree buds and leaves to be picked, divide the tea tree buds and leaves to be picked into unoccluded tea tree buds and leaves and occluded tea tree buds and leaves;

[0058] The tea tree bud and leaf picking point positioning system in the terminal device can respond to the instruction for locating the picking points of dense tea tree buds and leaves, obtain a tea leaf image containing the tea tree buds and leaves to be picked, and use the optimized tea tree bud and leaf detection model trained to the convergence state to perform object detection on the tea leaf image to determine the detection box information corresponding to the tea tree buds and leaves to be picked in the tea leaf image. Based on the detection box information corresponding to the tea tree buds and leaves to be picked, divide the tea tree buds and leaves to be picked into unoccluded tea tree buds and leaves and occluded tea tree buds and leaves; wherein, the tea tree buds and leaves to be picked include single buds, one bud with one leaf, and one bud with two leaves; the optimized tea tree bud and leaf detection model includes a Res2Net module, a path aggregation feature pyramid network, and a detection head, and the detection head includes an IoU-aware class score and a bounding box fine-tuning module.

[0059] In some embodiments, use an image acquisition device to obtain a tea leaf image in a vertical downward shooting, a 45° or 30° oblique shooting with the ground, or a vertical side shooting with the ground. The steps include: data preparation, using the image acquisition device to obtain a tea leaf image containing the tea tree buds and leaves to be picked in a 45° or 30° oblique shooting with the ground. Perform rectangular box annotation on the tea leaf image containing the tea tree buds and leaves to be picked to construct a tea tree bud and leaf data set. Divide the annotated tea leaf image containing the tea tree buds and leaves to be picked into a training set and a validation set in a ratio of 9:1.

[0060] In a further embodiment, the step of using the optimized tea tree bud and leaf detection model trained to the convergence state to perform object detection on the tea leaf image to determine the detection box information corresponding to the tea tree buds and leaves to be picked in the tea leaf image includes:

[0061] Step S101: In the Res2Net module, by introducing a multi-scale convolution module to process feature information of multiple scales in one residual block, effectively extract features in complex and dense tea leaf images to capture details and context information;

[0062] Step S102: In the path aggregation feature pyramid network, adopt the semantic information of high-level features and the spatial information of low-level features to generate a fine multi-scale feature map;

[0063] Step S103: In the detection head, the IoU-aware class score accurately identifies the target by introducing IoU information and fusing the confidence and localization accuracy of each tea shoot and leaf to be picked. In the bounding box fine-tuning module, the position and size of the bounding box are further optimized to precisely match the detection result with the actual target.

[0064] Specifically, an optimized model for tea shoot and leaf detection is constructed. The detection performance of dense tea shoots and leaves is improved through the IoU-aware class score and bounding box fine-tuning. The convolutional backbone network is responsible for extracting the feature representation of the input image.

[0065] Please refer to Figure 2 , in the optimized model for tea shoot and leaf detection, the Res2Net module is used as the convolutional backbone network of the optimized model for tea shoot and leaf detection in this application. The Res2Net module extracts features from the input tea leaf image. As the backbone network, Res2Net can process feature information at multiple scales simultaneously in a single residual block by introducing a multi-scale convolutional module. This multi-scale feature processing method enables Res2Net to more effectively extract features in complex and dense tea leaf images, capturing the details and context information in the images.

[0066] Furthermore, the features extracted in the Res2Net module are passed to the Path Aggregation Feature Pyramid Network (PAFPN) in the neck network. The Path Aggregation Feature Pyramid Network (PAFPN) establishes rich connections between feature maps at various scales, making full use of the semantic information of high-level features and the spatial information of low-level features to generate more refined multi-scale feature maps. The detection head consists of two parts: the IoU-aware class score and bounding box fine-tuning, which are used to generate the bounding box, confidence, and class information of tea shoots and leaves on these multi-scale feature maps. The IoU-aware class score introduces IoU information during the classification process and simultaneously fuses the confidence of the existence of tea shoots and leaves and the localization accuracy. The bounding box fine-tuning module further optimizes the position and size of the bounding box to ensure a more accurate match between the bounding box and the actual target.

[0067] As Figure 3 shown, the detection head consists of two parts: the IoU-aware class score and the bounding box fine-tuning module, which are used to generate the bounding box, confidence, and class information of tea shoots and leaves on these multi-scale feature maps. The IoU-aware class score introduces IoU information during the classification process and simultaneously fuses the confidence of the existence of tea shoots and leaves and the localization accuracy. The bounding box fine-tuning module further optimizes the position and size of the bounding box to ensure a more accurate match between the bounding box and the actual target.

[0068] Please refer to Figure 4, the role of the bounding box fine-tuning module is to further optimize the position and size of the initially generated bounding boxes during the object detection process, which helps improve the accuracy of the object detection network for the position and size of the object, especially suitable for complex tasks in dense object detection and occlusion scenarios. As Figure 2 shown, PAFPN further enhances the fusion ability of multi-scale features through path aggregation and top-down, bidirectional information flow, generates more refined multi-scale feature maps, and improves the network's detection ability for dense objects.

[0069] In some embodiments, the above-mentioned optimized tea shoot and leaf detection model is trained, and training parameters are set to train the optimized tea shoot and leaf detection model, which specifically includes:

[0070] Step S1001: During training, use the SGD optimizer, use Var i Foca l Loss as the classification loss function, use G I oU as the loss function for the object detection box, set the initial learning rate to 0.01, set the weight decay to 0.0001, set epochs to 24, and set batch s i ze to 16;

[0071] Step S1002: Input the training set of the tea shoot and leaf dataset into the model for training, and train according to the set parameters. Update the network parameters once per epoch iteration. Input the validation set data into the model for prediction, evaluate the model accuracy, save the model file parameters once every 1 eopchs of training, and compare and select the model parameters with the highest validation accuracy as the optimal solution. Finally, obtain the optimized tea shoot and leaf detection model;

[0072] Step S1003: Use Pytorch to load the model file parameters and construct the network model, and input the tea leaf image containing the tea shoots and leaves to be picked into the trained optimized tea shoot and leaf detection model. The model outputs the object detection box information corresponding to the unoccluded tea shoots and leaves and the occluded tea shoots and leaves in the tea leaf image.

[0073] As can be seen from the above steps, after the optimized tea shoot and leaf detection model is trained to the convergence state, it can be used to identify the object detection box information corresponding to the unoccluded tea shoots and leaves and the occluded tea shoots and leaves in the tea leaf image, and determine the picking points of the unoccluded tea shoots and leaves. For the picking points of the occluded tea shoots and leaves, subsequent steps are required to further determine.

[0074] Step S20: Input the detection box information corresponding to the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves into a pre-trained picking point localization model for occluded tea tree buds and leaves. Extract the convolutional features of the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves in the convolutional backbone network of the picking point localization model for occluded tea tree buds and leaves. Map the convolutional features to graph node representations in the node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships;

[0075] Obtain a tea leaf image containing tea tree buds and leaves to be picked. Use a tea tree bud and leaf detection optimization model that has been trained to a convergent state to perform object detection on the tea leaf image to determine the detection box information corresponding to the tea tree buds and leaves to be picked in the tea leaf image. After dividing the tea tree buds and leaves to be picked into unoccluded tea tree buds and leaves and occluded tea tree buds and leaves based on the detection box information corresponding to the tea tree buds and leaves to be picked, input the detection box information corresponding to the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves into a pre-trained picking point localization model for occluded tea tree buds and leaves. Extract the convolutional features of the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves in the convolutional backbone network of the picking point localization model for occluded tea tree buds and leaves. Map the convolutional features to graph node representations in the node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships;

[0076] In some embodiments, the basic network architecture of the picking point localization model for occluded tea tree buds and leaves is an improved HRNet network model. The improved HRNet network model includes a convolutional backbone network, a node embedding module, a dynamic adjacency matrix weighting module, and a graph relationship layer; among them, the graph relationship layer includes a GraphConv layer, a BatchNorm layer, and a ReLU layer.

[0077] Step S30: Dynamically adjust the relationships between different graph nodes in the dynamic adjacency matrix weighting module for adaptive learning in the graph structure. Use graph convolution operations in the graph relationship layer to allow information interaction and propagation between different picking points of the unoccluded tea tree buds and leaves to capture global and local context information;

[0078] Step S40: Optimize the relative position relationship between the picking points of the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves using a relative position loss to infer the picking points of the occluded tea tree buds and leaves to complete the picking point localization of dense tea tree buds and leaves.

[0079] Input the detection box information corresponding to the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves into a pre-trained positioning model for the picking points of obstructed tea tree buds and leaves. Extract the convolutional features of the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves in the convolutional backbone network of the positioning model for the picking points of obstructed tea tree buds and leaves. Map the convolutional features to graph node representations in the node embedding module, embed the graph node information through the learned adjacency matrix to capture neighborhood relationships, and then dynamically adjust the relationships between different graph nodes in the dynamic adjacency matrix weighting module for adaptive learning in the graph structure. In the graph relationship layer, adopt graph convolution operations to allow information interaction and propagation among different picking points of the unobstructed tea tree buds and leaves to capture global and local context information; optimize the relative position relationship between the picking points of the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves using relative position loss to infer the picking points of the obstructed tea tree buds and leaves, so as to complete the positioning of the picking points of dense tea tree buds and leaves.

[0080] In some embodiments, in the graph relationship layer, through the mechanism of a graph neural network, use an adjacency matrix and feature propagation to explicitly model the spatial and semantic relationships between the picking points of the unobstructed tea tree buds and leaves, so as to enhance the model's understanding of the dependency relationships between the picking points of the tea tree buds and leaves to be picked. By integrating the information of other picking points associated with the unobstructed tea tree buds and leaves, infer the picking points of the obstructed tea tree buds and leaves.

[0081] Specifically, please refer to Figure 5 , the improved HRNet network model mainly consists of four key parts: a convolutional backbone network, a node embedding module, a dynamic adjacency matrix weighting module, and a graph relationship layer. HRNet calculates feature maps from the input tea bud images, processes multi-scale features in parallel throughout the network, through a multi-resolution parallel sub-network structure, and conducts frequent information exchange and fusion between sub-networks of different resolutions, so as to better capture multi-scale features and context information. The node embedding module maps the convolutional features into graph node representations, combines the dynamic graph adjacency matrix learned from the dataset, and inputs it into the graph relationship layer to infer the positions of the obstructed picking points. The graph relationship layer allows information interaction and propagation among most relevant picking points through graph convolution. Finally, optimize the relative position relationship of the picking points through relative position loss, enabling the network to more accurately predict the positions of the obstructed picking points.

[0082] Please refer to Figure 5, HRNet calculates the feature map from the input tea bud image and processes multi-scale features in parallel throughout the network. Through a multi-resolution parallel sub-network structure and frequent information exchange and fusion between sub-networks of different resolutions, it can better capture multi-scale features and context information, which helps to retain more detailed information and infer the occluded picking points. The dynamic adjacency matrix weighting module effectively enhances the model's structure modeling ability, inference ability in occluded scenarios, and adaptability to diverse scenarios by adaptively constructing and updating the relationship weights between key points. The relative position loss enables the model to learn the relative geometric relationship of key points by defining the relative distance or angle between points. This relationship usually conforms to the natural structure of the tea bud, thus constraining the picking points generated by the model to conform to the real physical structure.

[0083] Please refer to Figure 6 , The graph relation layer consists of a GraphConv layer, a BatchNorm layer, and a ReLU layer. There is usually a certain structural association between the picking points and the auxiliary points. The graph relation layer explicitly models the spatial and semantic relationships between these points through the mechanism of a graph neural network (GNN), using the adjacency matrix and feature propagation, thereby enhancing the model's understanding of the dependence relationship between points. The graph relation layer infers the reasonable position of the occluded picking points by integrating the information of other associated points.

[0084] In some embodiments, the improved HRNet network model is trained, and training parameters are set to train the key point model of tea tree buds and leaves. Specifically, it includes:

[0085] Step S301: During training, the Adam optimizer is used, the initial learning rate is set to 0.0005, the weight decay is set to 0.0001, and the number of iterations of epochs is set to 210.

[0086] Step S302: Update the network parameters once every epoch iteration. Input the validation set data into the model for prediction, evaluate the model accuracy, save the model file parameters once every 10 epochs of training, and compare and select the model parameters with the highest validation accuracy as the optimal solution.

[0087] Step S303: Use Pytorch to load the model file parameters and construct the improved HRNet network model. The detection box information corresponding to the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves is input into the improved HRNet network model, and finally the picking point coordinates of the occluded tea tree buds and leaves are obtained.

[0088] In some embodiments, tea green images are obtained by taking pictures obliquely at an angle of 45° or 30° between the image acquisition device and the ground. Key points are marked on the tea green images, including the picking points of single tea buds, one bud with one leaf, and one bud with two leaves, to construct a dataset for the optimized model of tea tree bud and leaf detection. The marked tea green images are divided into a training set and a validation set at a ratio of 9:1.

[0089] In summary, first, the optimized model for tea tree bud and leaf detection is used to identify the target areas of unobstructed and occluded tea tree buds and leaves in the tea leaf images. Then, the occluded tea tree bud and leaf picking point positioning model that has been trained to convergence is utilized to determine the picking point positions of single buds, one bud with one leaf, and one bud with two leaves according to the growth state of the tea tree buds and leaves, thereby realizing the graded picking of tea tree buds and leaves. This method effectively improves the integrity of picking tea tree buds and leaves, reduces the interference of the surrounding environment on positioning, and significantly enhances the accuracy and efficiency of positioning the picking points of tea tree buds and leaves.

[0090] As can be seen from the above embodiments, compared with the prior art, in view of the many technical bottlenecks still existing in the prior art in the detection of dense tea buds and the recognition of occluded picking points, and the problem that traditional detection models are easily interfered when dealing with dense tea buds, resulting in low detection accuracy, etc., the present application includes but is not limited to the following beneficial effects:

[0091] First, the present application can significantly improve the detection accuracy of tea tree buds and leaves. The optimized model for tea tree bud and leaf detection in the present application combines the Res2Net module with the path aggregation feature pyramid network (PAFPN), enhancing the multi-scale expression ability of the feature map. This combination helps to provide more fine-grained multi-scale features, enabling the network to more accurately detect targets when facing dense tea buds. This means that even in the case of complex backgrounds and dense overlaps of tea buds, the detection model can still identify the targets with relatively high accuracy, avoiding misjudgment and missed judgment.

[0092] Second, the present application can greatly improve the positioning accuracy of the picking points of occluded tea tree buds and leaves. In the background of dense tea buds, many tea tree buds and leaves may be occluded by other leaves, and traditional detection methods are prone to inaccurate identification of these occluded tender buds. By combining the improved HRNet network model, especially introducing graph convolution operations in the graph relationship layer, it is possible to more effectively capture neighborhood relationships and global context information. In this way, the model can adaptively adjust the relationships between graph nodes in the graph structure, promote the interaction of information between different picking points, and thus infer the positions of occluded picking points, greatly improving the positioning ability of occluded tea tree buds and leaves.

[0093] Thirdly, the robustness and adaptability of the network are enhanced. In the optimized model for tea shoot and leaf detection in this application, by introducing the IoU-aware class score and the bounding box fine-tuning module, the detection head can better fuse the confidence of the existence of tea shoots and the positioning accuracy, and optimize the detection of dense tea shoots in complex backgrounds. Especially in the tea garden environment with complex backgrounds, the model can effectively cope with various interference factors, improving the robustness of the network in practical applications. The bounding box fine-tuning module further optimizes the size and position of the bounding box to ensure a more accurate match with the actual target.

[0094] Fourthly, this application solves the problem of grading the picking points in the tea green area. This application combines dense object detection and occluded keypoint detection. First, the improved dense tea green detection network is used to identify the tea green area, and then the improved HRNet is used to accurately locate the picking points in these areas. This method of grading picking point positioning can implement various picking strategies such as single-bud picking, one-bud-one-leaf picking, and one-bud-two-leaf picking, meeting the picking requirements of different tea varieties and tea green grades, and providing an efficient solution for the graded picking of tea leaves.

[0095] Fifthly, this application solves the technical bottleneck in the complex tea garden environment. The growth environment of the tea garden is complex and changeable, and there are often occlusion and overlap phenomena among tea shoots. The detection accuracy of existing technologies is often greatly affected in such an environment. However, through the combination of the optimized model for tea shoot and leaf detection and the model for locating the picking points of occluded tea shoots and leaves, a high detection accuracy can still be maintained in complex backgrounds, especially in scenes of dense and occluded tea shoots, breaking through the limitations of traditional methods.

[0096] Furthermore, by combining vision technology and advanced deep learning models, this application can achieve accurate detection and positioning of tea shoots and leaves. The automated picking robot can efficiently execute the picking task according to the detection results of the model. This greatly improves the efficiency of tea leaf picking, reduces the cost of manual picking, and enhances the intelligent level of the tea garden, promoting the development of agricultural automation technology.

[0097] Please refer to Figure 7, A positioning device for picking points of dense tea tree buds and leaves provided to meet one of the purposes of this application, including a tea tree bud and leaf occlusion detection module 1100, a neighborhood relationship capture module 1200, a context information capture module 1300, and a picking point positioning module 1400. Among them, the tea tree bud and leaf occlusion detection module 1100 is configured to respond to an instruction for positioning picking points of dense tea tree buds and leaves, obtain a tea leaf image containing the tea tree buds and leaves to be picked, and use a trained-to-converge tea tree bud and leaf detection optimization model to perform object detection on the tea leaf image to determine the detection frame information corresponding to the tea tree buds and leaves to be picked in the tea leaf image, and divide the tea tree buds and leaves to be picked into unoccluded tea tree buds and leaves and occluded tea tree buds and leaves based on the detection frame information corresponding to the tea tree buds and leaves to be picked; the neighborhood relationship capture module 1200 is configured to input the detection frame information corresponding to the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves into a pre-trained picking point positioning model for occluded tea tree buds and leaves, extract the convolutional features of the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves in the convolutional backbone network in the picking point positioning model for occluded tea tree buds and leaves, map the convolutional features to graph node representations in the node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships; the context information capture module 1300 is configured to dynamically adjust the relationships between different graph nodes in the dynamic adjacency matrix weighting module for adaptive learning in the graph structure, and use graph convolution operations in the graph relationship layer to allow information interaction and propagation between different picking points of the unoccluded tea tree buds and leaves to capture global and local context information; the picking point positioning module 1400 is configured to optimize the relative position relationship between the picking points of the unoccluded tea tree buds and leaves and the occluded tea tree buds and leaves using a relative position loss to infer the picking points of the occluded tea tree buds and leaves to complete the positioning of the picking points of dense tea tree buds and leaves.

[0098] Based on any embodiment of this application, please refer to Figure 8 , Another embodiment of this application further provides an electronic device, which can be implemented by a computer device, such as Figure 8As shown, it is a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a method for locating picking points of dense tea tree buds and leaves. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for locating picking points of dense tea tree buds and leaves of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand, Figure 8 The structure shown in [figure reference] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0099] In this embodiment, the processor is used to execute Figure 7 the specific functions of each module in [module reference]. The memory stores the program code and various types of data required to execute the above modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program code and data required to execute all modules in the device for locating picking points of dense tea tree buds and leaves of the present application. The server can call the program code and data of the server to execute the functions of all modules.

[0100] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for locating picking points of dense tea tree buds and leaves according to any embodiment of the present application.

[0101] The present application also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by one or more processors, the steps of the method for locating picking points of dense tea tree buds and leaves according to any embodiment of the present application are implemented.

[0102] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiments of the present application can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described embodiments of each method. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0103] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the technical field, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for locating dense tea tree bud and leaf picking points, characterized in that: include: In response to an instruction to locate a picking point for dense tea tree buds and leaves, a tea image containing tea tree buds and leaves to be picked is obtained, and a tea tree bud and leaf detection optimization model that has been trained to a convergent state is used to perform target detection on the tea image to determine detection frame information corresponding to the tea tree buds and leaves to be picked in the tea image, and based on the detection frame information corresponding to the tea tree buds and leaves to be picked, the tea tree buds and leaves to be picked are divided into unobstructed tea tree buds and leaves and obstructed tea tree buds and leaves; Input the detection frame information corresponding to the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves into the pre-trained obstructed tea tree bud and leaf picking point positioning model, extract the convolution features of the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves in the convolution backbone network in the obstructed tea tree bud and leaf picking point positioning model, map the convolution features to graph node representations in the node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships; Dynamically adjusting the relationship between different graph nodes in a dynamic adjacency matrix weighting module to perform adaptive learning in a graph structure, and using a graph convolution operation in a graph relationship layer to allow information interaction and propagation between different picking points of the unobstructed tea tree buds and leaves to capture global and local context information; The relative position loss is used to optimize the relative position relationship between the picking point of the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves, so as to infer the picking point of the obstructed tea tree buds and leaves, so as to complete the picking point positioning of the dense tea tree buds and leaves.

2. The method for locating dense tea tree bud and leaf picking points according to claim 1, characterized in that: The tea tree bud and leaf detection optimization model includes a Res2Net module, a path aggregation feature pyramid network and a detection head, wherein the detection head includes an IoU-aware class score and a bounding box fine-tuning module.

3. The method for locating dense tea tree bud and leaf picking points according to claim 2, characterized in that: The step of using the tea bud and leaf detection optimization model that has been trained to a convergent state to perform target detection on the tea image to determine the detection frame information corresponding to the tea bud and leaf to be picked in the tea image includes: In the Res2Net module, a multi-scale convolution module is introduced to process feature information of multiple scales in one residual block, so as to effectively extract features in complex and dense tea images to capture details and context information; In the path aggregation feature pyramid network, the semantic information of high-level features and the spatial information of low-level features are used to generate refined multi-scale feature maps; In the detection head, the IoU-aware class score introduces IoU information, integrates the confidence and positioning accuracy of each tea bud and leaf to be picked to accurately identify the target, and further optimizes the position and size of the bounding box in the bounding box fine-tuning module to accurately match the detection result with the actual target.

4. The method for locating dense tea tree bud and leaf picking points according to claim 1, characterized in that: In the graph relationship layer, the adjacency matrix and feature propagation are used to explicitly model the spatial and semantic relationships between the picking points of the unobstructed tea buds and leaves through the mechanism of graph neural network, so as to enhance the model's understanding of the dependency relationship between the various picking points of the tea buds and leaves to be picked, and to infer the picking points of the obstructed tea buds and leaves by integrating other information related to the picking points of the unobstructed tea buds and leaves.

5. The method for locating dense tea tree bud and leaf picking points according to claim 1, characterized in that: The basic network architecture of the obstructed tea tree bud and leaf picking point positioning model is an improved HRNet network model, which includes a convolutional backbone network, a node embedding module, a dynamic adjacency matrix weighting module and a graph relationship layer.

6. The method for locating dense tea tree bud and leaf picking points according to claim 5, characterized in that: The graph relationship layer includes a GraphConv layer, a BatchNorm layer and a ReLU layer.

7. The method for locating dense tea tree bud and leaf picking points according to any one of claims 1 to 6, characterized in that: The tea tree buds and leaves to be picked include single buds, one bud and one leaf, and one bud and two leaves.

8. A device for locating dense tea bud and leaf picking points, characterized in that: include: A tea tree bud and leaf occlusion detection module is configured to respond to an instruction to locate a picking point for dense tea tree buds and leaves, obtain a tea image containing tea tree buds and leaves to be picked, use a tea tree bud and leaf detection optimization model that has been trained to a convergent state to perform target detection on the tea image to determine detection frame information corresponding to the tea tree buds and leaves to be picked in the tea image, and divide the tea tree buds and leaves to be picked into unobstructed tea tree buds and leaves and obstructed tea tree buds and leaves based on the detection frame information corresponding to the tea tree buds and leaves to be picked; A neighborhood relationship capture module is configured to input the detection frame information corresponding to the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves into a pre-trained obstructed tea tree bud and leaf picking point positioning model, extract the convolution features of the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves in the convolution backbone network in the obstructed tea tree bud and leaf picking point positioning model, map the convolution features to graph node representations in a node embedding module, and embed the graph node information through the learned adjacency matrix to capture neighborhood relationships; A context information capture module is configured to dynamically adjust the relationship between different graph nodes in the dynamic adjacency matrix weighting module so as to perform adaptive learning in the graph structure, and adopt a graph convolution operation in the graph relationship layer to allow information interaction and propagation between different picking points of the unobstructed tea tree buds and leaves to capture global and local context information; The picking point positioning module is configured to use relative position loss to optimize the relative position relationship between the picking point of the unobstructed tea tree buds and leaves and the obstructed tea tree buds and leaves, so as to infer the picking point of the obstructed tea tree buds and leaves, so as to complete the picking point positioning of the dense tea tree buds and leaves.

9. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

Citation Information

Cited By

  • Tea bud and leaf three-dimensional picking point prediction method for distinguishing shielding conditions

    CN120747236A

  • Face recognition detection method and device, computer equipment and program product

    CN121096006A