Method, device and equipment for planning picking path of famous high-quality tea and medium

Through the improved Transformer model, the inefficiency problem caused by randomness of tea green picking paths is solved, and more efficient and accurate picking path planning is achieved.

CN120063268APending Publication Date: 2025-05-30SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510132018.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the tea green picking path is random, resulting in a long motion path and significantly reducing the picking efficiency.

Method used

The improved Transformer model is used to plan the tea green picking path, and the complex dependencies between the tea green space coordinates are extracted through the encoding module, and the shortest picking path is generated in the decoding module, and the self-supervised learning is used to optimize the path generation.

Benefits of technology

It significantly reduces the movement distance of the robotic arm, improves the picking efficiency, enhances the accuracy and flexibility of path planning, and has good adaptability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120063268A_ABST
    Figure CN120063268A_ABST
Patent Text Reader

Abstract

The invention relates to a famous high-quality tea picking path planning method, device and equipment and a medium, and the method comprises the steps: coding a tea leaf space coordinate corresponding to each to-be-picked tea leaf in a coding module of a picking path planning model, processing is carried out through a first position coding module, a multi-head self-attention layer, a first feedforward network module and a first multi-layer stacking structure in the coding module in sequence; in a decoding module of the picking path planning model, taking a high-dimensional vector representation corresponding to each tea leaf space coordinate as data input, and processing the data through a second position coding module, an autoregressive attention layer, a second feedforward network module, a final attention layer, an index sampling layer and a second multi-layer stacking structure in the decoding module in sequence; and taking the path length corresponding to each candidate picking path as a reward signal to carry out self-supervised learning so as to determine the shortest picking path corresponding to the to-be-picked tea leaves. The picking efficiency is remarkably improved, and energy consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of agricultural production, and particularly relates to a method for planning picking paths of famous and high-quality tea, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Compared with other fruits and vegetables, famous and high-quality tea grows more densely, there are many tea green targets per unit area, and they are unevenly distributed. Considering the actual picking efficiency of tea greens, it is necessary to complete the detection and positioning tasks of all tea green targets in the current field of view and then control the robotic arm to pick tea greens one by one. If the tea greens are picked disorderly, the movement path of the end effector shows randomness, resulting in a relatively long overall path and significantly reducing the picking efficiency.

[0003] In summary, in order to adapt to the problems in the prior art that if the tea greens are picked disorderly, the movement path of the end effector shows randomness, resulting in a relatively long overall path and significantly reducing the picking efficiency, etc., the applicant has made corresponding explorations in consideration of solving this problem. Summary of the Invention

[0004] The purpose of the present application is to solve the above problems and provide a method for planning picking paths of famous and high-quality tea, a corresponding device, an electronic device, and a computer-readable storage medium.

[0005] To meet the various purposes of the present application, the present application adopts the following technical solutions:

[0006] A method for planning picking paths of famous and high-quality tea proposed to meet one of the purposes of the present application includes:

[0007] Respond to the picking path planning instruction of famous and high-quality tea, and obtain the tea green space coordinates corresponding to each tea green to be picked;

[0008] Call a preset picking path planning model, encode the tea green space coordinates corresponding to each tea green to be picked in the encoding module of the picking path planning model, and sequentially process them through the first position encoding module, multi-head self-attention layer, first feed-forward network module, and first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea green space coordinate, so as to extract the complex dependence relationship between each tea green space coordinate;

[0009] In the decoding module of the picking path planning model, use the high-dimensional vector representation corresponding to each tea green space coordinate as the data input, and sequentially process it through the second position encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacking structure in the decoding module to calculate and determine the selected probability distribution corresponding to the tea green space coordinate, so as to gradually generate each candidate picking path;

[0010] Use the path lengths corresponding to the respective candidate picking paths as reward signals for self-supervised learning to determine the shortest picking path corresponding to the tea shoots to be picked, so as to complete the picking path planning of famous and high-quality tea.

[0011] Optionally, the basic network architecture of the picking path planning model is an improved Transformer model. Among them, the improved Transformer model includes an encoding module and a decoding module. The encoding module is constructed by an input embedding layer, a first positional encoding module, a multi-head self-attention layer, a first feed-forward network module, and a first multi-layer stacking structure. The decoding module is constructed by a second positional encoding module, an autoregressive attention layer, a second feed-forward network module, a final attention layer, an index sampling layer, and a second multi-layer stacking structure.

[0012] Optionally, in the encoding module of the picking path planning model, the tea shoot spatial coordinates corresponding to each tea shoot to be picked are encoded, and are sequentially processed through the first positional encoding module, the multi-head self-attention layer, the first feed-forward network module, and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea shoot spatial coordinate, so as to extract the complex dependence relationship between each tea shoot spatial coordinate. The steps include:

[0013] Input the tea shoot spatial coordinates corresponding to each tea shoot to be picked into a preset picking path planning model. In the input embedding layer, convert the tea shoot spatial coordinates into input embedding vectors;

[0014] In the first positional encoding module, add positional encoding to the input embedding vectors, and add a positional encoding to each tea shoot spatial coordinate. Among them, the positional encoding represents the relative position of the tea shoot spatial coordinate in the entire sequence;

[0015] In the multi-head self-attention layer, for each tea shoot spatial coordinate, the model calculates the dependence relationship with all other tea shoot spatial coordinates, and adjusts the representation of each tea shoot spatial coordinate according to these dependence relationships to determine the self-attention processing result. The first feed-forward network module performs further non-linear transformation on the self-attention processing result;

[0016] Stack in multiple Transformer encoder layers in the first multi-layer stacking structure. Each layer of Transformer encoder layer further extracts and fuses global information, so that the model can capture the complex dependence relationship between different tea shoot spatial coordinates.

[0017] Optionally, in the decoding module of the picking path planning model, taking the high-dimensional vector representation corresponding to each tea shoot spatial coordinate as data input, and successively passing through the second position encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacking structure in the decoding module for processing, and calculating and determining the selected probability distribution corresponding to the tea shoot spatial coordinate, so as to gradually generate each candidate picking path, the steps include:

[0018] In the autoregressive attention layer, calculate the attention values between the selected tea shoot spatial coordinates in the generated picking path and the unvisited tea shoot spatial coordinates respectively to determine the first output result;

[0019] In the final attention layer, perform attention calculation on the first output result in the autoregressive attention layer and the unvisited tea shoot spatial coordinates again;

[0020] Based on the softmax function, calculate the probability distribution of each unvisited tea shoot spatial coordinate being selected to gradually generate each picking path.

[0021] Optionally, the steps of training the picking path planning model include:

[0022] Taking the improved Transformer model as the basic network architecture of the picking path planning model, randomly initialize and generate two network models, one as the baseline network and the other as the training model;

[0023] Input the tea shoot spatial coordinates corresponding to each tea shoot to be picked into the training model. The encoder in the picking path planning model encodes the tea shoot spatial coordinates corresponding to each tea shoot to be picked to generate its corresponding encoded node;

[0024] The decoder then uses these encoded nodes to generate an initial picking path and calculates the total length of this path, which is used as the basis for the reward signal. If the length of the current picking path is shorter than the baseline path, update the length of the baseline path. The shorter the path, the higher the reward signal;

[0025] Through the backpropagation algorithm, update the model parameters according to the reward signal, thereby optimizing the path generation ability;

[0026] Check whether the preset number of training rounds or convergence conditions are reached. If not, continue to train using the dataset; otherwise, output the finally optimized picking path planning model.

[0027] Optionally, before the step of obtaining the tea shoot spatial coordinates corresponding to each tea shoot to be picked, it includes:

[0028] Obtain the RGB image corresponding to the tea shoots to be picked, and use the preset LabelMe annotation software to annotate the RGB image corresponding to the tea shoots to be picked, so as to annotate the pixel coordinates corresponding to the picking points of each tea shoot to be picked;

[0029] Map the pixel coordinates into the depth map to obtain the tea shoot space coordinates corresponding to each tea shoot to be picked.

[0030] Optionally, the step of obtaining the tea shoot space coordinates corresponding to each tea shoot to be picked includes:

[0031] Obtain the pixel point coordinates corresponding to each pixel point of each tea shoot to be picked in the RGB image, the principal point coordinates in the camera intrinsics, the focal lengths in the camera intrinsics, and the depth value of the pixel point in the camera coordinate system. Among them, the pixel point coordinates include the pixel coordinates in the horizontal direction and the pixel coordinates in the vertical direction. The principal point coordinates in the camera intrinsics include the principal point coordinates in the horizontal direction of the image and the principal point coordinates in the vertical direction of the image. The focal lengths in the camera intrinsics include the focal length in the horizontal direction of the camera intrinsics and the focal length in the vertical direction of the camera intrinsics;

[0032] Calculate and determine the first difference between the pixel coordinates in the horizontal direction and the principal point coordinates in the horizontal direction of the image, calculate and determine the first ratio between the first difference and the focal length in the horizontal direction of the camera intrinsics, and determine the horizontal direction spatial coordinate of the pixel point in the camera coordinate system according to the first product between the first ratio and the depth value of the pixel point in the camera coordinate system;

[0033] Calculate and determine the second difference between the pixel coordinates in the vertical direction and the principal point coordinates in the vertical direction of the image, calculate and determine the second ratio between the second difference and the focal length in the vertical direction of the camera intrinsics, and determine the vertical direction spatial coordinate of the pixel point in the camera coordinate system according to the second product between the second ratio and the depth value of the pixel point in the camera coordinate system;

[0034] Determine the tea shoot space coordinates corresponding to the tea shoot to be picked according to the horizontal direction spatial coordinate of the pixel point in the camera coordinate system, the vertical direction spatial coordinate of the pixel point in the camera coordinate system, and the depth value of the pixel point in the camera coordinate system.

[0035] A famous and high-quality tea picking path planning device provided for another purpose of this application includes:

[0036] A space coordinate acquisition module, configured to respond to a famous and high-quality tea picking path planning instruction and acquire the tea shoot space coordinates corresponding to each tea shoot to be picked;

[0037] The encoding module is configured to call a preset picking path planning model, encode the tea green spatial coordinates corresponding to each tea green to be picked in the encoding module of the picking path planning model, and sequentially process them through the first position encoding module, multi-head self-attention layer, first feed-forward network module, and first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea green spatial coordinate, so as to extract the complex dependency relationships between the tea green spatial coordinates;

[0038] The decoding module is configured to use the high-dimensional vector representation corresponding to each tea green spatial coordinate as input data in the decoding module of the picking path planning model, and sequentially process them through the second position encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacking structure in the decoding module to calculate and determine the selected probability distribution corresponding to the tea green spatial coordinates, so as to gradually generate each candidate picking path;

[0039] The picking path planning module is configured to perform self-supervised learning using the path lengths corresponding to the candidate picking paths as reward signals to determine the shortest picking path corresponding to the tea green to be picked, so as to complete the picking path planning of famous and high-quality tea.

[0040] An electronic device provided to meet another object of the present application includes a central processing unit and a memory. The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the famous and high-quality tea picking path planning method of the present application.

[0041] A computer-readable storage medium provided to meet another object of the present application stores a computer program implemented according to the famous and high-quality tea picking path planning method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.

[0042] Compared with the prior art, in view of the problems in the prior art that if the tea greens are picked disorderly, the movement path of the end effector is random, resulting in a long overall path and significantly reducing the picking efficiency, the present application includes but is not limited to the following beneficial effects:

[0043] First, in the traditional tea green picking process, the picking path is often random, resulting in the movement trajectory of the robotic arm not being optimized, causing a long path and low picking efficiency. Through the picking path planning model, the picking path can be accurately planned to ensure that the movement of the robotic arm is more efficient and direct, which can greatly reduce the movement distance of the robotic arm and thus improve the overall picking efficiency;

[0044] Second, the picking path planning model of the present application is based on an improved Transformer architecture. In the picking path planning model, the encoding module processes the input data through the first positional encoding, the multi-head self-attention layer, and the feed-forward network module, effectively capturing the complex relationships between the spatial coordinates of the tea shoots. The decoding module gradually generates candidate paths and optimizes path selection through components such as the autoregressive attention layer and the feed-forward network module. This efficient encoding and decoding structure can fully learn the spatial distribution characteristics of the tea shoots, improving the accuracy and flexibility of path planning;

[0045] Third, the picking path planning model of the present application uses self-supervised learning. By taking the path length of the candidate path as the reward signal, the system can self-adjust during the picking process to continuously shorten the path length and improve the picking efficiency. This path optimization method can not only self-adjust according to the current picking task but also adaptively change in different environments, having strong robustness;

[0046] Fourth, through the multi-head self-attention mechanism, the model of the present application can handle the complex spatial dependence relationships between the tea shoots. Compared with traditional methods, the model can effectively identify which tea shoot targets have priority during the picking process, reducing unnecessary repeated movements and path overlaps, making the picking process more orderly and efficient.

[0047] Fifth, the picking path planning model of the present application uses an improved Transformer architecture, having good adaptability and scalability, and being able to handle different types of tea garden environments and tea shoot distributions. Even when facing a highly variable environment (such as uneven distribution of tea shoots in the tea garden), the model can adjust the path planning strategy according to the data to ensure that the picking efficiency is not affected. With the assistance of the model, the robotic arm can move precisely along the planned path, avoiding misoperations that may occur under human intervention. The detection and positioning of tea shoot targets are more accurate, further improving the automation level and precision of the picking process.

[0048] Furthermore, by integrating the improved Transformer model, the system can intelligently and automatically perform tea shoot picking path planning. The effective path planning, complex relationship modeling, application of self-supervised learning, and improvement in the precise control of the robotic arm significantly improve the picking efficiency, reduce energy consumption, and provide strong technical support for the development of future intelligent agriculture. This technology is not only applicable to tea picking but can also be extended to automated picking tasks for other fruits and vegetables. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0050] Figure 1 It is a schematic flowchart of the famous and high-quality tea picking path planning method in the embodiment of the present application;

[0051] Figure 2 It is an exemplary network structure of the improved Transformer model in the embodiment of the present application;

[0052] Figure 3 It is an exemplary network structure of the autoregressive attention module in the embodiment of the present application;

[0053] Figure 4 It is a flowchart of the process for training the improved Transformer model in the embodiment of the present application;

[0054] Figure 5 It is a principle block diagram of the famous and high-quality tea picking path planning device in the embodiment of the present application;

[0055] Figure 6 It is a schematic diagram of the structure of the computer device in the embodiment of the present application. Detailed implementation manners

[0056] The embodiments of the present application are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation of the present application.

[0057] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0058] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0059] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no transmitting ability, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers, tablets, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to operate locally, and / or operate in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback functions, or can also be devices such as a smart TV, a set-top box, etc.

[0060] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, which is a hardware device with the necessary components revealed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0061] It should be noted that the concept of "server" in this application can similarly be extended to the case of server clusters. According to the network deployment principles understood by those skilled in the art, the various servers should be logically divided. Physically, these servers can either be independent of each other but can be called through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by it in the implementation of the network deployment method of this application.

[0062] One or several technical features of this application, unless clearly specified, can either be deployed on the server and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.

[0063] The neural network models cited or potentially cited in this application, unless clearly specified, can either be deployed on a remote server and remotely invoked by the client, or deployed on a client with sufficient device capabilities for direct invocation. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.

[0064] All kinds of data involved in this application, unless clearly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0065] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0066] For each of the embodiments to be disclosed in this application, unless explicitly stated to be mutually exclusive, the relevant technical features involved in each embodiment can be cross - combined to flexibly construct new embodiments, as long as such combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0067] Please refer to Figure 1 , in one embodiment of the famous tea picking path planning method of this application, it includes:

[0068] Step S10: Respond to the famous tea picking path planning instruction, and obtain the tea - leaf spatial coordinates corresponding to each tea - leaf to be picked;

[0069] The famous tea picking path planning system in the terminal device can respond to the famous tea picking path planning instruction and obtain the tea - leaf spatial coordinates corresponding to each tea - leaf to be picked;

[0070] In some embodiments, before the step of obtaining the tea - leaf spatial coordinates corresponding to each tea - leaf to be picked, it includes:

[0071] Step S101: Obtain the RGB image corresponding to the tea - leaf to be picked, and use the preset LabelMe annotation software to annotate the RGB image corresponding to the tea - leaf to be picked, so as to annotate the pixel coordinates corresponding to the picking points of each tea - leaf to be picked;

[0072] Step S102: Map the pixel coordinates to the depth map to obtain the tea - leaf spatial coordinates corresponding to each tea - leaf to be picked.

[0073] In a further embodiment, the step of obtaining the tea - leaf spatial coordinates corresponding to each tea - leaf to be picked includes:

[0074] Step S101: Obtain the pixel coordinates corresponding to each pixel point of each tea - leaf to be picked in the RGB image, the principal point coordinates in the camera internal parameters, the focal lengths in the camera internal parameters, and the depth value of the pixel point in the camera coordinate system. Among them, the pixel coordinates include the pixel coordinates in the horizontal direction and the pixel coordinates in the vertical direction, the principal point coordinates in the camera internal parameters include the principal point coordinates in the horizontal direction of the image and the principal point coordinates in the vertical direction of the image, and the focal lengths in the camera internal parameters include the focal length in the horizontal direction of the camera internal parameters and the focal length in the vertical direction of the camera internal parameters;

[0075] Step S102: Calculate and determine the first difference between the pixel coordinates in the horizontal direction and the principal point coordinates in the horizontal direction of the image, calculate and determine the first ratio between the first difference and the focal length in the horizontal direction in the camera internal parameters, and determine the spatial coordinate in the horizontal direction of the pixel point in the camera coordinate system according to the first product between the first ratio and the depth value of the pixel point in the camera coordinate system;

[0076] Step S103: Calculate and determine the second difference between the pixel coordinates in the vertical direction and the principal point coordinates in the vertical direction of the image, calculate and determine the second ratio between the second difference and the focal length in the vertical direction in the camera internal parameters, and determine the spatial coordinate in the vertical direction of the pixel point in the camera coordinate system according to the second product between the second ratio and the depth value of the pixel point in the camera coordinate system;

[0077] Step S104: Determine the tea green spatial coordinates corresponding to the to-be-picked tea green according to the spatial coordinate in the horizontal direction of the pixel point in the camera coordinate system, the spatial coordinate in the vertical direction of the pixel point in the camera coordinate system, and the depth value of the pixel point in the camera coordinate system.

[0078] Specifically, in order to obtain the actual spatial coordinates of the tea green, an RGBD camera can be used for data acquisition. The camera is tilted at an angle of 15° to 75° with respect to the ground for shooting, and this angle range can better capture the actual spatial coordinates of the tea green. The data acquisition work is carried out at the experimental base of a tea research institute of a certain agricultural academy of sciences, and the acquisition time is from 9:00 am to 5:00 pm to obtain images under different lighting conditions. The collected RGB images are labeled by the LabelMe annotation software, and the picking points of all tea greens are labeled to obtain their pixel coordinates; then, these labeled pixel coordinates are mapped to the depth map to obtain the actual spatial coordinates of the to-be-picked tea green. All the tea green spatial coordinates obtained in each frame of the image constitute a data set; finally, the collected data set is divided into a training set and a validation set according to a ratio of 9:1.

[0079] More specifically, the RGB-D camera is installed on a hydraulic pan-tilt head and fixed to the frame of the famous tea picking device. By adjusting the hydraulic pan-tilt head, the optical axis of the camera forms an angle of 15° to 75° with the ground. The RGB-D camera is equipped with multiple lenses such as RGB and color infrared, and can synchronously output RGB images and depth images within its field of view. The picking device is controlled to operate slowly at a constant speed between the tea ridges, and the video is continuously recorded, so as to simultaneously obtain synchronous RGB images and depth images. Subsequently, the images are extracted frame by frame from the video to form an image dataset of the tea shoots in the field. This image dataset covers the images of the tea shoots in the field at different times, under different lighting conditions and different weather conditions. The images collected at this stage are still two-dimensional data and do not directly meet the requirements of the famous tea picking path planning for the actual spatial coordinates of the tea shoots. Therefore, the collected field images need to be further processed. In this application, the LabelMe annotation tool is used to annotate the tea shoot areas in the RGB images to obtain the pixel coordinates of the tea shoots. Then, the annotated pixel coordinates are mapped to the corresponding depth images to extract the depth values of the tea shoots. However, the depth values are only the components of the axis in the camera coordinate system, and it is necessary to further solve the actual coordinates of the tea shoots in the three-dimensional space through conversion. The conversion method is as follows:

[0080] It can be assumed that the coordinates of a certain pixel point in the industrial camera image are (u, v), and its corresponding three-dimensional space coordinates in the camera coordinate system are (X_c, Y_c, Z_c), where Z_c is the depth value of this pixel point in the camera coordinate system, which is directly provided by the depth image, and the depth map is aligned and synchronously output with the RGB image, ensuring the real-time consistency of the pixel coordinates and the depth values. According to the inverse transformation formula of the camera pinhole imaging model, the spatial coordinate X_c of this pixel point in the horizontal direction in the camera coordinate system and the spatial coordinate Y_c of this pixel point in the vertical direction in the camera coordinate system can be calculated through the following calculation formula, which is expressed as:

[0081]

[0082] Among them, X_c represents the spatial coordinate of this pixel point in the horizontal direction in the camera coordinate system, u represents the pixel coordinate in the horizontal direction, cx represents the principal point coordinate in the horizontal direction of the image, fx represents the focal length in the horizontal direction of the camera internal parameters, and Z_c represents the depth value of this pixel point in the camera coordinate system;

[0083]

[0084] Among them, Y_c represents the spatial coordinate of this pixel point in the vertical direction in the camera coordinate system, v represents the pixel coordinate in the vertical direction, cy represents the principal point coordinate in the vertical direction of the image, fy represents the focal length in the vertical direction of the camera internal parameters, and Z_c represents the depth value of this pixel point in the camera coordinate system.

[0085] Through the conversion of the above calculation formula, the three-dimensional spatial coordinates of the tea leaves to be picked can be obtained. In each frame of the image, the spatial coordinates of all the tea leaves form a data set. Finally, all the collected data is divided into a training set and a validation set according to a ratio of 9:1.

[0086] Step S20: Invoke a preset picking path planning model, and encode the tea leaf spatial coordinates corresponding to each tea leaf to be picked in the encoding module of the picking path planning model. The encoding process sequentially passes through the first position encoding module, the multi-head self-attention layer, the first feed-forward network module, and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependency relationships between the spatial coordinates of each tea leaf.

[0087] After obtaining the tea leaf spatial coordinates corresponding to each tea leaf to be picked, invoke a preset picking path planning model, and encode the tea leaf spatial coordinates corresponding to each tea leaf to be picked in the encoding module of the picking path planning model. The encoding process sequentially passes through the first position encoding module, the multi-head self-attention layer, the first feed-forward network module, and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependency relationships between the spatial coordinates of each tea leaf.

[0088] In some embodiments, the basic network architecture of the picking path planning model is an improved Transformer model. Among them, the improved Transformer model includes an encoding module and a decoding module. The encoding module is constructed by an input embedding layer, a first position encoding module, a multi-head self-attention layer, a first feed-forward network module, and a first multi-layer stacking structure. The decoding module is constructed by a second position encoding module, an autoregressive attention layer, a second feed-forward network module, a final attention layer, an index sampling layer, and a second multi-layer stacking structure.

[0089] In a further embodiment, the step of encoding the tea leaf spatial coordinates corresponding to each tea leaf to be picked in the encoding module of the picking path planning model, and sequentially passing through the first position encoding module, the multi-head self-attention layer, the first feed-forward network module, and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependency relationships between the spatial coordinates of each tea leaf, includes:

[0090] Step S201: Input the tea leaf spatial coordinates corresponding to each tea leaf to be picked into a preset picking path planning model. In the input embedding layer, convert the tea leaf spatial coordinates into input embedding vectors.

[0091] Step S202: In the first position encoding module, add position encoding to the input embedding vector, and add a position encoding to each tea shoot spatial coordinate, where the position encoding represents the relative position of the tea shoot spatial coordinate in the entire sequence;

[0092] Step S203: In the multi-head self-attention layer, for each tea shoot spatial coordinate, the model calculates the dependencies with all other tea shoot spatial coordinates, and adjusts the representation of each tea shoot spatial coordinate according to these dependencies to determine the self-attention processing result, and further performs a non-linear transformation on the self-attention processing result in the first feed-forward network module;

[0093] Step S204: Stack in multiple Transformer encoder layers in the first multi-layer stacked structure, and each layer of Transformer encoder layer further extracts and fuses global information so that the model can capture the complex dependencies between different tea shoot spatial coordinates.

[0094] In some embodiments, the steps of training the picking path planning model include:

[0095] Step S2001: Use the improved Transformer model as the basic network architecture of the picking path planning model, and randomly initialize to generate two network models, one as the baseline network and the other as the training model;

[0096] Step S2002: Input the tea shoot spatial coordinates corresponding to each tea shoot to be picked into the training model. The encoder in the picking path planning model encodes the tea shoot spatial coordinates corresponding to each tea shoot to be picked to generate its corresponding encoded node;

[0097] Step S2003: The decoder then uses these encoded nodes to generate an initial picking path, calculates the total length of this path, and uses it as the basis for the reward signal. If the length of the current picking path is shorter than the baseline path, update the length of the baseline path. The shorter the path, the higher the reward signal;

[0098] Step S2004: Through the backpropagation algorithm, update the model parameters according to the reward signal, thereby optimizing the path generation ability;

[0099] Step S2005: Check whether the preset number of training rounds or convergence conditions are reached. If not, continue to train using the dataset; otherwise, output the finally optimized picking path planning model.

[0100] Specifically, the basic network architecture of the picking path planning model of this application is an improved Transformer model, which mainly consists of an encoding module and a decoding module. The encoding module is composed of a position encoding module, a multi-head self-attention layer, a feed-forward network module, and a multi-layer stacking structure, aiming to encode the spatial coordinates of the tea greens and their mutual relationships into high-dimensional vector representations, so as to provide context-aware input for the decoder and ultimately be used to generate the optimal picking path; the decoding module is composed of a position encoding module, an autoregressive attention layer, a feed-forward network module, a final attention layer, an index sampling layer, and a multi-layer stacking structure, aiming to generate the optimal picking path by means of iterative decoding, and perform interactive calculations using the previously selected tea green spatial coordinates and the unvisited tea green coordinates at each step.

[0101] Furthermore, the improved Transformer model is trained. This application uses a reinforcement learning method for self-supervised training, introduces a baseline network to stabilize the training and reduce the gradient variance, and at the same time uses the path length as a reward signal to encourage the model to generate shorter paths. Specifically, it includes:

[0102] First, two network models are randomly initialized: one as the baseline network and the other as the training model. Then, the tea green spatial coordinates collected in the above embodiment are input into the training model. Next, the encoder encodes these tea green spatial coordinates to generate their representation vectors. The decoder then uses these encoded nodes to generate an initial path and calculates the total length of this path, which will serve as the basis for the reward signal. If the length of the current path is shorter than the baseline path, the length of the baseline path is updated. The shorter the path, the higher the reward signal. Through the backpropagation algorithm, the model parameters are updated according to the reward signal, thereby optimizing the path generation ability. Finally, it is checked whether the preset number of training epochs or convergence conditions are reached. If not, continue to train using the dataset; otherwise, output the finally optimized model.

[0103] Even further, the PyTorch framework model is used. The PyTorch framework model provides flexible deep learning tools and can conveniently load and construct complex network structures. The tea green spatial coordinates to be planned are input into the network model trained with a large amount of labeled data, and this model can quickly and accurately output the optimal picking path, that is, the shortest picking path, providing a reliable basis for subsequent automated picking.

[0104] Step S30: In the decoding module of the picking path planning model, using the high-dimensional vector representations corresponding to each tea shoot spatial coordinate as data inputs, and successively processing them through the second position encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacking structure in the decoding module to calculate and determine the selected probability distribution corresponding to the tea shoot spatial coordinate, so as to gradually generate each candidate picking path;

[0105] Call the preset picking path planning model. In the encoding module of the picking path planning model, encode the tea shoot spatial coordinates corresponding to each tea shoot to be picked, and successively process them through the first position encoding module, multi-head self-attention layer, first feed-forward network module, and first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representations corresponding to each tea shoot spatial coordinate, so as to extract the complex dependencies between each tea shoot spatial coordinate. Then, in the decoding module of the picking path planning model, use the high-dimensional vector representations corresponding to each tea shoot spatial coordinate as data inputs, and successively process them through the second position encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacking structure in the decoding module to calculate and determine the selected probability distribution corresponding to the tea shoot spatial coordinate, so as to gradually generate each candidate picking path.

[0106] In some embodiments, the step of using the high-dimensional vector representations corresponding to each tea shoot spatial coordinate as data inputs in the decoding module of the picking path planning model, and successively processing them through the second position encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacking structure in the decoding module to calculate and determine the selected probability distribution corresponding to the tea shoot spatial coordinate, so as to gradually generate each candidate picking path, includes:

[0107] Step S301: In the autoregressive attention layer, calculate the attention values between the selected tea shoot spatial coordinates in the generated picking path and the unvisited tea shoot spatial coordinates respectively to determine the first output result;

[0108] Step S302: In the final attention layer, perform attention calculation again on the first output result in the autoregressive attention layer and the unvisited tea shoot spatial coordinates;

[0109] Step S303: Based on the softmax function, calculate the probability distribution of each unvisited tea shoot spatial coordinate being selected, so as to gradually generate each picking path.

[0110] Step S40: Use the path lengths corresponding to the respective candidate picking paths as reward signals for self-supervised learning to determine the shortest picking path corresponding to the tea leaves to be picked, so as to complete the picking path planning of famous and high-quality tea.

[0111] In the decoding module of the picking path planning model, use the high-dimensional vector representations corresponding to each tea leaf spatial coordinate as data inputs, and sequentially process them through the second positional encoding module, autoregressive attention layer, second feed-forward network module, final attention layer, index sampling layer, and second multi-layer stacked structure in the decoding module to calculate and determine the selected probability distribution corresponding to the tea leaf spatial coordinate. After gradually generating each candidate picking path, use the path lengths corresponding to the respective candidate picking paths as reward signals for self-supervised learning to determine the shortest picking path corresponding to the tea leaves to be picked, so as to complete the picking path planning of famous and high-quality tea.

[0112] Specifically, the basic network architecture of the picking path planning model proposed in this application is an improved Transformer model. The Transformer model is a deep learning model based on the self-attention mechanism. Through multi-head self-attention and feed-forward neural networks, it can effectively capture the dependencies between different positions in the sequence. This model is mainly composed of an encoding module and a decoding module, as Figure 2 shown.

[0113] In the encoding module, first convert the input data set into an embedding representation through the input embedding layer, and then add it to the positional encoding. The positional encoding is used to provide position information for each input vector to ensure that the model can perceive the order of elements in the input sequence. The embedding representation after fusing the position information is input into multiple layers of Transformer encoders. Each layer of encoder includes a multi-head self-attention mechanism, a feed-forward neural network, a residual connection, and batch normalization. Through multi-layer processing, the encoder gradually extracts and fuses global information to capture complex patterns and dependencies in the input data.

[0114] The process of this decoding module is carried out step by step. First, the generated coordinate sequence is converted into an embedding representation through the Output Embedding layer, and then added to the Positional Encoding. The fused embedding representation is input into the Transformer Decoder Layers. Each layer of the decoder contains components such as the Autoregressive Attention Module, the Final Attention Module, the Feed-Forward Neural Network, the Residual Connection, and Batch Normalization. Among them, the Autoregressive Attention Module and the Final Attention Module are the most crucial parts. The Autoregressive Attention Module calculates the attention for the generated coordinate sequence and the unvisited coordinates respectively, and its calculation schematic diagram is as shown in Figure 3 shown. It enhances the global view of the decoder, enabling it to simultaneously focus on the correlation between the coordinates in the current path and the unvisited coordinates. The Final Attention Module then calculates the attention again for the calculation result of the Autoregressive Attention Module and the unvisited coordinates, and finally calculates the probability distribution of each coordinate being selected through the softmax function.

[0115] This application uses the reinforcement learning method to perform self-supervised training on the improved model, and the training flow chart is as shown in Figure 4 shown. The computer hardware configuration for training includes an 11th Gen Intel(R) Core(TM) i5-11500@2.70GHz CPU and an NVIDIA GeForce RTX 3060 GPU. The training environment is the Ubuntu22.04 operating system based on Linux, the Python version is 3.9, and the model training parameters are set as follows: the number of training epochs is 5000, the batch size is 128. The optimizer used for training is Adam, the learning rate is 0.0001, and the training tolerance (used to determine whether to update the baseline model) is 20. The specific training process is as follows:

[0116] Start with randomly initializing two network models: one is the Baseline Network and the other is the Training Network, which will be used for path generation and performance comparison respectively. The tea green coordinate dataset collected in Step 1 is input into the Training Network, and the Training Network generates an initial path based on these inputs, representing the order of tea green picking. Subsequently, the system calculates the total length of this path and uses it as the basis for the reward signal. The design principle of the reward signal is that "the shorter the path, the better the model performance". To evaluate the performance of the Training Network, the Baseline Network uses the same inputs to generate another path, and the length of its path serves as a control benchmark. If the path length generated by the Training Network is shorter than that of the Baseline Network, the weights of the Baseline Network will be updated to those of the Training Network to improve the performance of the Baseline Network in subsequent evaluations. Next, the system uses the reward signal to optimize the parameters of the Training Network. The goal of the Training Network is to generate shorter paths, so the reward signal is negatively correlated with the path length. Through the backpropagation algorithm, the model adjusts its parameters according to the reward signal to optimize its ability to generate paths. Specifically, the policy gradient method is used to guide the model to learn how to obtain higher rewards through path optimization. After each round of training, the system checks whether the preset training termination conditions are met, such as reaching the set number of training rounds or the model performance converging. If the conditions are not met, the system will repeat the above process based on new training data, continuously generating new paths and updating the model parameters; if the conditions are met, the training stops and the finally optimized model is output. Finally, after multiple rounds of training and parameter optimization, the Training Network can generate the optimal path based on the input tea green space coordinates. At the same time, the Baseline Network also records the optimized performance for verifying the model effect. This process design ensures that the model can not only learn effective path generation but also has good generalization ability and is applicable to different scenarios.

[0117] Finally, this application deploys the trained model using the PyTorch framework, which is widely used in the field of deep learning for its flexibility and efficiency. The model completes the inference task with high precision by loading a deep network trained based on a large amount of labeled data. In practical applications, the tea green space coordinates to be planned can be used as inputs and passed to the trained model. Based on the input coordinates, the model quickly calculates the optimal picking path using deep learning algorithms, thus realizing the picking path planning of famous and high-quality teas and significantly improving the picking efficiency and accuracy.

[0118] As can be seen from the above embodiments, compared with the prior art, in view of the problems in the prior art that if the tea greens are picked disorderly, the movement path of the end effector is random, resulting in a relatively long overall path and significantly reducing the picking efficiency, etc., this application includes but is not limited to the following beneficial effects:

[0119] First, in the traditional tea shoot picking process, the picking path is often random, resulting in an unoptimized movement trajectory of the robotic arm, causing a long path and low picking efficiency. Through the picking path planning model, the picking path can be accurately planned to ensure that the movement of the robotic arm is more efficient and direct, which can greatly reduce the movement distance of the robotic arm and thus improve the overall picking efficiency;

[0120] Second, the picking path planning model of the present application is based on an improved Transformer architecture. The encoding module in the picking path planning model processes the input data through the first positional encoding, multi-head self-attention layer, and feed-forward network module, effectively capturing the complex relationships between the spatial coordinates of tea shoots. The decoding module consists of components such as the autoregressive attention layer and feed-forward network module, gradually generating candidate paths and optimizing path selection. This efficient encoding and decoding structure can fully learn the spatial distribution characteristics of tea shoots, improving the accuracy and flexibility of path planning;

[0121] Third, the picking path planning model of the present application uses self-supervised learning. By taking the path length of the candidate path as the reward signal, the system can self-adjust during the picking process to continuously shorten the path length and improve the picking efficiency. This path optimization method can not only self-adjust according to the current picking task but also adaptively change in different environments, having strong robustness;

[0122] Fourth, through the multi-head self-attention mechanism in the present application, the model can handle the complex spatial dependence relationships between tea shoots. Compared with traditional methods, the model can effectively identify which tea shoot targets have priority during the picking process, reducing unnecessary repeated movements and path overlaps, making the picking process more orderly and efficient.

[0123] Fifth, the picking path planning model of the present application uses an improved Transformer architecture, having good adaptability and scalability, and being able to handle different types of tea garden environments and tea shoot distributions. Even in the face of a large change in the environment (such as uneven distribution of tea shoots in the tea garden), the model can adjust the path planning strategy according to the data to ensure that the picking efficiency is not affected. With the assistance of the model, the robotic arm can accurately move along the planned path, avoiding misoperations that may occur under manual intervention. The detection and positioning of tea shoot targets are more accurate, further improving the automation level and accuracy of the picking process.

[0124] Furthermore, by integrating an improved Transformer model, the system can intelligently and automatically plan the picking paths of fresh tea leaves. Effective path planning, complex relationship modeling, the application of self-supervised learning, and the improvement of the precise control of the robotic arm have significantly improved the picking efficiency, reduced energy consumption, and provided strong technical support for the future development of intelligent agriculture. This technology is not only applicable to tea leaf picking but can also be extended to the automated picking tasks of other fruits and vegetables.

[0125] Please refer to Figure 5 , a picking path planning device for famous and high-quality tea provided to meet one of the purposes of this application, includes a spatial coordinate acquisition module 1100, an encoding module 1200, a decoding module 1300, and a picking path planning module 1400. Among them, the spatial coordinate acquisition module 1100 is set to respond to the picking path planning instruction for famous and high-quality tea and acquire the tea leaf spatial coordinates corresponding to each fresh tea leaf to be picked; the encoding module 1200 is set to call a preset picking path planning model and encode the tea leaf spatial coordinates corresponding to each fresh tea leaf to be picked in the encoding module of the picking path planning model, and successively process them through the first position encoding module, the multi-head self-attention layer, the first feed-forward network module, and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependence relationships between the tea leaf spatial coordinates; the decoding module 1300 is set to use the high-dimensional vector representation corresponding to each tea leaf spatial coordinate as the data input in the decoding module of the picking path planning model, and successively process them through the second position encoding module, the autoregressive attention layer, the second feed-forward network module, the final attention layer, the index sampling layer, and the second multi-layer stacking structure to calculate and determine the selected probability distribution corresponding to the tea leaf spatial coordinates, so as to gradually generate each candidate picking path; the picking path planning module 1400 is set to use the path lengths corresponding to each candidate picking path as the reward signal for self-supervised learning to determine the shortest picking path corresponding to the fresh tea leaf to be picked, so as to complete the picking path planning of famous and high-quality tea.

[0126] Based on any embodiment of this application, please refer to Figure 6 , another embodiment of this application further provides an electronic device, which can be implemented by a computer device, such as Figure 6As shown, it is a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a method for planning the picking path of high-quality tea. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for planning the picking path of high-quality tea of this application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand, Figure 6 The structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0127] In this embodiment, the processor is used to execute Figure 5 the specific functions of each module in [the figure]. The memory stores the program codes and various types of data required to execute the above-mentioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules in the high-quality tea picking path planning device of this application. The server can call the program codes and data of the server to execute the functions of all modules.

[0128] This application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for planning the picking path of high-quality tea described in any embodiment of this application.

[0129] This application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method for planning the picking path of high-quality tea described in any embodiment of this application are implemented.

[0130] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiments of the method of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-described methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0131] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

[0132] In summary, by integrating the improved Transformer model, the system of the present application can intelligently and automatically plan the tea green picking path. The effective path planning, complex relationship modeling, application of self-supervised learning, and improvement of the precise control of the robotic arm significantly improve the picking efficiency, reduce energy consumption, and provide strong technical support for the development of future intelligent agriculture. This technology is not only applicable to tea picking but can also be extended to the automated picking tasks of other fruits and vegetables.

Claims

1. A method for planning a path for picking high-quality tea, characterized in that: include: Respond to the famous tea picking path planning instruction and obtain the spatial coordinates of each tea leaf to be picked; Calling a preset picking path planning model, encoding the tea leaf spatial coordinates corresponding to each tea leaf to be picked in the encoding module of the picking path planning model, and sequentially processing through the first position encoding module, the multi-head self-attention layer, the first feedforward network module and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependency relationship between each tea leaf spatial coordinate; In the decoding module of the picking path planning model, the high-dimensional vector representation corresponding to each tea leaf spatial coordinate is used as data input, and is processed in sequence by the second position encoding module, the autoregressive attention layer, the second feedforward network module, the final attention layer, the index sampling layer and the second multi-layer stacking structure in the decoding module to calculate and determine the selection probability distribution corresponding to the tea leaf spatial coordinate, so as to gradually generate various candidate picking paths; The path lengths corresponding to the candidate picking paths are used as reward signals for self-supervised learning to determine the shortest picking path corresponding to the green tea leaves to be picked, so as to complete the picking path planning for high-quality tea.

2. The method for planning a path for picking high-quality tea according to claim 1, characterized in that: The basic network architecture of the picking path planning model is an improved Transformer model, wherein the improved Transformer model includes an encoding module and a decoding module. The encoding module is constructed by an input embedding layer, a first position encoding module, a multi-head self-attention layer, a first feedforward network module and a first multi-layer stacking structure, and the decoding module is constructed by a second position encoding module, an autoregressive attention layer, a second feedforward network module, a final attention layer, an index sampling layer and a second multi-layer stacking structure.

3. The method for planning a path for picking high-quality tea according to claim 2, characterized in that: The steps of encoding the tea leaf spatial coordinates corresponding to each tea leaf to be picked in the encoding module of the picking path planning model, and sequentially processing the tea leaf spatial coordinates through the first position encoding module, the multi-head self-attention layer, the first feedforward network module and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependency relationship between each tea leaf spatial coordinate, include: Inputting the tea leaf spatial coordinates corresponding to each tea leaf to be picked into a preset picking path planning model, and converting the tea leaf spatial coordinates into an input embedding vector in the input embedding layer; In the first position encoding module, the position encoding is added to the input embedding vector, and a position encoding is added to each tea leaf spatial coordinate, wherein the position encoding represents the relative position of the tea leaf spatial coordinate in the entire sequence; In the multi-head self-attention layer, for each tea green space coordinate, the model calculates the dependency relationship with all other tea green space coordinates, and adjusts the representation of each tea green space coordinate according to these dependencies to determine the self-attention processing result, and further nonlinearly transforms the self-attention processing result in the first feedforward network module; In the stacking of multiple Transformer encoder layers in the first multi-layer stacking structure, each Transformer encoder layer further extracts and fuses global information so that the model captures the complex dependencies between different tea leaf spatial coordinates.

4. The method for planning a path for picking high-quality tea according to claim 2, characterized in that: In the decoding module of the picking path planning model, the high-dimensional vector representation corresponding to each tea leaf spatial coordinate is used as data input, and is processed in sequence by the second position encoding module, the autoregressive attention layer, the second feedforward network module, the final attention layer, the index sampling layer and the second multi-layer stacking structure in the decoding module, and the selection probability distribution corresponding to the tea leaf spatial coordinate is calculated and determined to gradually generate each candidate picking path, including: In the autoregressive attention layer, the attention values ​​between the selected tea leaves spatial coordinates and the unvisited tea leaves spatial coordinates in the generated picking path are respectively calculated to determine a first output result; In the final attention layer, the first output result in the autoregressive attention layer and the unvisited tea green space coordinates are again subjected to attention calculation; Based on the softmax function, the probability distribution of each unvisited tea leaf spatial coordinate being selected is calculated to gradually generate each picking path.

5. The method for planning a path for picking high-quality tea according to claim 2, characterized in that: The steps to train the picking path planning model include: The improved Transformer model is used as the basic network architecture of the picking path planning model, and two network models are randomly initialized to generate one as a baseline network and the other as a training model; The spatial coordinates of each tea leaf corresponding to the tea leaves to be picked are input into the training model, and the encoder in the picking path planning model encodes the spatial coordinates of each tea leaf corresponding to the tea leaves to be picked to generate corresponding encoding nodes; The decoder then uses these encoding nodes to generate an initial picking path and calculates the total length of the path as the basis for the reward signal. If the length of the current picking path is shorter than the baseline path, the length of the baseline path is updated. The shorter the path, the higher the reward signal. By means of a back-propagation algorithm, model parameters are updated according to the reward signal, thereby optimizing the path generation capability; Check whether the preset number of training rounds or convergence conditions are reached. If not, continue to use the dataset for training; otherwise, output the final optimized picking path planning model.

6. The method for planning a path for picking high-quality tea according to claim 1, characterized in that: Before the step of obtaining the spatial coordinates of each tea leaf corresponding to the tea leaves to be picked, the following steps are included: Obtaining an RGB image corresponding to the green tea leaves to be picked, and using a preset LabelMe annotation software to annotate the RGB image corresponding to the green tea leaves to be picked, so as to annotate the pixel coordinates corresponding to the picking point of each green tea leaves to be picked; The pixel coordinates are mapped into a depth map to obtain the spatial coordinates of each tea leaf corresponding to the tea leaves to be picked.

7. The method for planning a path for picking high-quality tea according to claim 1, characterized in that: The steps of obtaining the spatial coordinates of each tea leaf corresponding to the tea leaves to be picked include: Obtain the pixel coordinates corresponding to each pixel of each to-be-picked green tea in the RGB image, the principal point coordinates in the camera intrinsic parameters, the focal length in the camera intrinsic parameters, and the depth value of the pixel in the camera coordinate system, wherein the pixel coordinates include the pixel coordinates in the horizontal direction and the pixel coordinates in the vertical direction, the principal point coordinates in the camera intrinsic parameters include the principal point coordinates in the horizontal direction of the image and the principal point coordinates in the vertical direction of the image, and the focal length in the camera intrinsic parameters includes the focal length in the horizontal direction of the camera intrinsic parameters and the focal length in the vertical direction of the camera intrinsic parameters; Calculate and determine a first difference between the pixel coordinates in the horizontal direction and the principal point coordinates in the horizontal direction of the image, calculate and determine a first ratio between the first difference and the focal length in the horizontal direction in the camera intrinsic parameter, and determine the spatial coordinates of the pixel point in the horizontal direction in the camera coordinate system according to a first product between the first ratio and the depth value of the pixel point in the camera coordinate system; Calculate and determine a second difference between the pixel coordinates in the vertical direction and the principal point coordinates in the vertical direction of the image, calculate and determine a second ratio between the second difference and the focal length in the vertical direction in the camera intrinsic parameter, and determine the spatial coordinates of the pixel in the vertical direction in the camera coordinate system according to a second product between the second ratio and the depth value of the pixel in the camera coordinate system; The spatial coordinates of the tea leaves corresponding to the tea leaves to be picked are determined according to the spatial coordinates of the pixel point in the horizontal direction in the camera coordinate system, the spatial coordinates of the pixel point in the vertical direction in the camera coordinate system and the depth value of the pixel point in the camera coordinate system.

8. A device for planning a path for picking high-quality tea, characterized in that: include: A spatial coordinate acquisition module is configured to respond to the famous tea picking path planning instruction and obtain the spatial coordinates of each tea leaf corresponding to the tea leaves to be picked; The encoding module is configured to call a preset picking path planning model, encode the tea leaf spatial coordinates corresponding to each tea leaf to be picked in the encoding module of the picking path planning model, and sequentially process the first position encoding module, the multi-head self-attention layer, the first feedforward network module and the first multi-layer stacking structure in the encoding module to determine the high-dimensional vector representation corresponding to each tea leaf spatial coordinate, so as to extract the complex dependency relationship between each tea leaf spatial coordinate; A decoding module is configured to input the high-dimensional vector representation corresponding to each spatial coordinate of the green tea leaves as data in the decoding module of the picking path planning model, and sequentially process the high-dimensional vector representation through the second position encoding module, the autoregressive attention layer, the second feedforward network module, the final attention layer, the index sampling layer and the second multi-layer stacking structure in the decoding module to calculate and determine the selection probability distribution corresponding to the spatial coordinate of the green tea leaves, so as to gradually generate each candidate picking path; The picking path planning module is configured to use the path lengths corresponding to the candidate picking paths as reward signals for self-supervised learning to determine the shortest picking path corresponding to the green tea leaves to be picked, so as to complete the picking path planning for the famous and high-quality tea.

9. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

Citation Information

Patent Citations

  • Intelligent coherent picking system for famous high-quality tea

    CN115271200A

  • Unmanned aerial vehicle search path planning method and system suitable for emergency rescue

    CN116301005A

  • Seedling missing detection and positioning method and device and storage medium

    CN118096690A

  • Tea leaf grading picking method, device, equipment and medium

    CN118366027A