Intelligent picking method, agricultural picking robot, chip system and medium
By combining visual perception and LiDAR, using a neural network model to enhance fruit features and LiDAR to accurately scan fruit locations, the problem of low efficiency in traditional harvesting methods and low accuracy in existing robot recognition is solved, achieving efficient and low-damage fruit harvesting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANTONG INST OF TECH
- Filing Date
- 2025-01-06
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional crop harvesting methods are inefficient and labor-intensive. Existing agricultural harvesting robots have low accuracy in identifying fruits, are prone to damaging them, and are affected by environmental factors.
By combining visual perception and LiDAR, images are processed through a pre-trained neural network model to enhance fruit features and identify the approximate area of the fruit. LiDAR is used to precisely scan the location of the fruit, and the fruit stalk is identified as the picking point for harvesting.
It improves the precision and efficiency of fruit harvesting, reduces damage to the fruit, lowers labor costs, and is unaffected by environmental factors.
Smart Images

Figure CN119817329B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of smart agriculture, and in particular relates to an intelligent harvesting method, an agricultural harvesting robot, a chip system, and a medium. Background Technology
[0002] With the development of agricultural modernization, traditional crop harvesting methods can no longer meet the growing demand for agricultural products. For example, more and more plantations are starting to grow apple trees on a large scale. When the fruit ripens, in order to harvest the ripe apples as quickly as possible within the optimal harvesting window, it is usually necessary to hire a large number of harvesters, which is not only inefficient but also incurs high labor costs.
[0003] Therefore, how to achieve fully automated and intelligent harvesting of crops is a problem that needs to be solved. Summary of the Invention
[0004] This application provides an intelligent harvesting method, an agricultural harvesting robot, a chip system, and a medium, which can improve the accuracy and efficiency of fruit harvesting.
[0005] Firstly, a smart harvesting method is provided, which is applied to agricultural harvesting robots. This method includes:
[0006] Acquire a first image, the image content of which includes the target fruit tree, and the target fruit tree includes the target fruit;
[0007] The first image is processed based on the first neural network model to obtain the second image. The first neural network model is trained on multiple training datasets. Each training dataset includes multiple images augmented from an original image. The original image contains fruits of the same type as the target fruit. The first neural network model is trained with the aim of reducing the values of the first loss function, the second loss function, and the third loss function. The first loss function is used to characterize the feature difference between the input image and the output image of the first neural network model. The second loss function is the adversarial loss function value of the fourth loss function, which is used to characterize the feature difference between the multiple output images corresponding to each training dataset. The third loss function is used to characterize the difference between the fruit recognition result and the target recognition result for the training dataset.
[0008] Identify a first region in the second image, the second image is composed of the first region and the second region, the first image is composed of a third region and a fourth region, the first region corresponds to the third region, the second region corresponds to the fourth region, the feature intensity deviation of the corresponding pixels of the first region and the third region is less than or equal to a first threshold, and the feature intensity deviation of the corresponding pixels of the second region and the fourth region is greater than the first threshold.
[0009] Based on the first region, a target detection area is determined, and a laser pulse is emitted into the target detection area via a laser emitter.
[0010] The echo signal corresponding to the laser pulse is received by the laser receiving array to obtain the measurement information of multiple pixels corresponding to the target detection area. The number of receiving units in the laser receiving array is less than or equal to a preset threshold.
[0011] The upper boundary point of the target fruit on the target fruit tree is determined from the multiple pixels, and the picking point is determined based on the upper boundary point.
[0012] The target fruits are harvested from this picking point.
[0013] Optionally, determining the upper boundary point of the target fruit on the target fruit tree from the plurality of pixels includes:
[0014] Each pixel in the candidate boundary points is detected, and the candidate boundary point is some or all of the multiple boundary points;
[0015] If the i-th pixel among the candidate boundary points meets the preset conditions, the i-th pixel is determined as the upper boundary point of the target fruit. The preset conditions include: the difference between the measured distance of the x pixels above the i-th pixel and the measured distance of the i-th pixel is greater than a first threshold, and the difference between the measured distance of the y pixels below the i-th pixel and the measured distance of the i-th pixel is less than a second threshold. x, y, and i are all positive integers, and the values of x and y are determined based on the type of the target fruit tree.
[0016] Optionally, the method further includes:
[0017] Select p rows of pixels from the plurality of pixels as the candidate boundary point. Among the p rows of pixels, the measurement distances corresponding to more than m pixels are greater than the third threshold, and the measurement distances corresponding to more than n pixels are less than the fourth threshold.
[0018] Optionally, the number of receiving units in the receiving array is determined based on the type of the target fruit tree and the model of the agricultural harvesting robot.
[0019] Optionally, different receiving units in the receiving array have different measurement ranges, and the measurement range of each receiving unit is determined based on the historical maximum measurement range of that receiving unit.
[0020] Optionally, channels with dependency coefficients less than or equal to a fifth threshold in the first neural network model are pre-deleted, and the parameters of the first neural network model are pre-adjusted based on a loss amount, which is used to characterize the degree of structural change of the first neural network model, and the dependency coefficient is used to characterize the dependency relationship between the channel and the model.
[0021] In a second aspect, a terminal device is provided, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the terminal device performs the steps of the method as described in any of the first aspects above.
[0022] Thirdly, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the steps of the method as described in any of the first aspects above.
[0023] Fourthly, a computer program product is provided that, when run on a terminal device, causes the terminal device to perform the method described in any one of the first aspects.
[0024] Fifthly, a chip system is provided, the chip system including a processor coupled to a memory, the processor executing a computer program stored in the memory to implement the method described in any of the first aspects above.
[0025] The chip system can be a single chip or a chip module composed of multiple chips.
[0026] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0027] Figure 1 An exemplary flowchart of an intelligent harvesting method for fruit trees provided in an embodiment of this application;
[0028] Figure 2 A schematic diagram of one training method for the first neural network model;
[0029] Figure 3 This is an example image of the first image in this application;
[0030] Figure 4 This is an example of a second image in this application;
[0031] Figure 5 This is a schematic diagram of the first and second regions of the second image in this application;
[0032] Figure 6This is a schematic diagram of the third and fourth regions of the first image in this application;
[0033] Figure 7 This is an exemplary structure for a lidar system;
[0034] Figure 8 A schematic diagram of the target fruit;
[0035] Figure 9 This is a schematic diagram of the results of lidar detection. Detailed Implementation
[0036] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0037] With the development of agricultural modernization, traditional crop harvesting methods can no longer meet the growing demand for agricultural products. For example, more and more plantations are starting to grow apple trees on a large scale. When the fruit ripens, in order to harvest the ripe apples as quickly as possible within the optimal harvesting window, it is usually necessary to hire a large number of harvesters, which is not only inefficient but also incurs high labor costs.
[0038] Therefore, many researchers have begun to study agricultural harvesting robots that can achieve intelligent harvesting. However, current agricultural harvesting robots still have many problems. For example, the accuracy of current harvesting robots in identifying fruits is not high enough, and they often damage the fruits when harvesting with robotic arms, affecting the value of the fruits. Also, the accuracy of current harvesting robots in identifying fruits is affected by many factors, such as changes in light, which may greatly reduce the accuracy of identification.
[0039] In view of this, embodiments of this application provide an agricultural harvesting robot capable of accurately identifying target fruits (such as apples, pears, peaches, etc.). Specifically, in the solution provided in this application, after processing the captured images using a pre-trained artificial intelligence model, the approximate area where the target fruit is located is determined, and then this area is scanned using LiDAR. Identifying the location of the target fruit using point cloud data obtained from LiDAR scanning is not only highly accurate but also unaffected by environmental factors. In addition, this application also creatively proposes a method for determining the harvesting point based on the upper boundary point of the target fruit, taking into account the characteristics of the target fruit. Since the fruit stalk is usually located on its upper boundary, harvesting by grasping the stalk can avoid damaging the fruit surface and improve the commercial value of the fruit.
[0040] The following is combined with Figure 1Method 100 in this application provides an illustrative description of the intelligent harvesting method for fruit trees provided in the embodiments of this application. Method 100 can be executed by an agricultural harvesting robot, or by an internal module or component of the agricultural harvesting robot; this application does not limit this, and the following description uses the execution of method 100 by an agricultural harvesting robot as an example. The agricultural harvesting robot described in the embodiments of this application is a fully automatic or semi-automatic robot capable of intelligently harvesting fruits from fruit trees. This agricultural harvesting robot can be a dedicated robot, i.e., specifically used for harvesting, or it can be a multifunctional robot with functions other than harvesting; this application does not limit this. This application does not limit the type of fruit tree to which the agricultural harvesting robot is applicable. Theoretically, this agricultural harvesting robot can be used to harvest fruits from various types of fruit trees, such as apples, pears, and lychees. Furthermore, this agricultural harvesting robot can also be used to harvest fruits from other types of crops, such as cucumbers and tomatoes.
[0041] S110, Acquire the first image.
[0042] For example, the agricultural harvesting robot in this application includes a visual recognition system, which may specifically include one or more cameras. These cameras may be RGB cameras, ToF cameras, hyperspectral cameras, or combinations of different types of cameras. This application does not limit the scope of the application.
[0043] Agricultural harvesting robots can control cameras to take pictures of orchards through a control system, and obtain a first image. The content of the first image includes the target fruit tree, which includes one or more target fruits. In other words, the first image is an image obtained by taking pictures of the target fruit tree.
[0044] In this application, the term "target fruit tree" is used to refer to the fruit tree to be harvested. The term "target fruit" is used to refer to the fruit to be harvested from the target fruit tree. In this application, the target fruit tree can be an apple tree, a pear tree, a lychee tree, etc., and the corresponding target fruit can be an apple, a pear, or a lychee.
[0045] It should be understood that the target fruit tree in the first image refers to one of the fruit trees in the orchard.
[0046] It should also be understood that the first image in this application may be one or more images.
[0047] Figure 3 An exemplary first image is shown. From Figure 3 As can be seen, the first image includes the target fruit tree and the target fruit hanging from the branches of the tree. In addition to the target fruit tree and the target fruit, the first image may also include other interfering features such as weeds and leaves.
[0048] It is understandable that the purpose of capturing the first image in this application is not to identify the location of the target fruit based on the first image. Existing technologies primarily rely on images for fruit identification, but the results are often unsatisfactory. This is mainly because leaves and fruits have visual similarities, and factors such as lighting conditions, image clarity, and dew can create visual interference. Therefore, image-based fruit identification frequently results in errors or omissions. Even if the fruit's location is identified based on the captured image, inaccurate identification can lead to damage to the fruit's skin during harvesting, shortening its shelf life and impacting its commercial value.
[0049] Based on the above analysis, this application does not use visual images to identify the location of the target fruit, but uses LiDAR to identify the location of the target fruit (the specific process can be referred to in the description of S140 to S160, which will not be repeated here). However, it is actually quite difficult to identify the location of the fruit using point cloud data obtained by LiDAR scanning, and it also consumes a lot of energy. Therefore, there is currently no relevant technology in the prior art that applies LiDAR to the scenario of fruit identification. The main reason is that LiDAR can obtain information such as the measurement distance and reflectivity of each pixel in the target detection area by scanning the laser array. If a large area is scanned, the scanning result will be very complex, because the leaves, trunks, fruits, etc. of the fruit tree are intertwined, and the measurement distances corresponding to different positions are alternating and changing. To detect the location of a relatively small fruit in such a result, a very complex algorithm is required, and the recognition accuracy is difficult to guarantee.
[0050] Based on the above analysis, this application employs a combination of visual perception and lidar to detect target fruits. The processing approach is as follows: first, the approximate location of the fruit is identified using a visual method; then, the lidar emits laser pulses purposefully, scanning only the fruit and a small area around it. Identifying the target fruit's location from this scan result becomes much easier. Therefore, the main inventive points of this application include two parts. The first part determines how to detect the approximate area of the target fruit from an image. This application refers to the approximate area of the target fruit as the target detection area. The size of the target detection area should be smaller than a preset value, because if the target detection area is too large, subsequent detection using lidar will be very difficult; conversely, the smaller the target detection area, the higher the accuracy of subsequent fruit identification, but the greater the difficulty in determining the target detection area. The second part determines how to use lidar to accurately identify the picking point from the target detection area. It should be understood that identifying the picking point is not as simple as identifying the location of the target fruit. Numerous real-world examples have shown that picking the fruit by identifying its location and using a robotic arm to grasp it easily damages the surface of the fruit. Therefore, this application aims to accurately identify the fruit stalk and use it as the picking point, thereby minimizing damage during harvesting. The following describes one possible implementation of Scheme 1 in conjunction with S120-S130; and another possible implementation of Scheme 2 in conjunction with S140-S160.
[0051] S120. The first image is processed based on the first neural network model to obtain the second image.
[0052] For example, after capturing the first image, the fruit is not directly used for detection. Instead, a first neural network model is used to process the first image to obtain a second image. The first neural network model enhances the first image, strengthening fruit-related features while weakening fruit-irrelevant features (such as leaf features, trunk features, light features, and weed features). This makes it easier to locate the specific position of the target fruit. Because the first neural network model can weaken fruit-irrelevant features, these features can interfere with the accuracy of fruit recognition. For example, fruit-related features like leaves and trunks can interfere with accuracy, and light falling on crops at different times can create shadows or highlights, also interfering with accuracy. The first neural network model involved in this application can weaken the influence of these features, enabling more precise location of the target fruit.
[0053] In other words, the second image is obtained by enhancing the first image. Therefore, the feature intensity of fruit-related features in the second image is higher than that of fruit-related features in the first image, and the feature intensity of non-fruit-related features in the second image is lower than that of non-fruit-related features in the first image.
[0054] The first neural network model in this application is a creatively proposed fruit feature enhancement model tailored to the scenario of fruit recognition. The following section will combine... Figure 2 An example of a training method for the first neural network model is provided.
[0055] The first neural network model is trained on multiple training datasets. Each training dataset is augmented from a single original image, and each dataset contains multiple images. The original image contains the target fruit tree and the target type of fruit. The type of crop in the image is the same as the type of crop to be detected, and the target type is also the same as the type of fruit to be detected. For example, if we want to detect apples on an apple tree, we need to prepare several images of apple trees, each containing a number of apples. Then, we augment these apple tree images to obtain more images. Augmentation here refers to processing the original image, such as rotating the image, adding shadows, adding different angles of light, scaling, adding occlusion, applying blur effects, etc. In other words, for each image in the training set, they all come from the same original image, or are augmented from the same original image.
[0056] It should be understood that establishing different training datasets instead of mixing all images together into one large dataset is to facilitate subsequent characterization of the differences between images obtained by augmenting the same image. For details, please refer to the following description, which will not be elaborated here.
[0057] During model training, images from the training dataset are input into the first neural network model. The first neural network model extracts feature vectors, which are used to describe the features of different contents in the image. Then, the extracted features are processed, and finally, the processed image is output.
[0058] The training process for the first neural network model is essentially a process of training and adjusting the parameters of the first neural network model. The following is an example illustrating one method for training model parameters:
[0059] The first neural network model was trained with the aim of reducing the values of the first, second, and third loss functions.
[0060] The first loss function value characterizes the feature difference between the input and output images of the first neural network model. Specifically, it represents the difference between the input image and the corresponding output image of the first neural network model. A smaller first loss function value indicates a smaller difference between the input and output images. In other words, during model parameter training, we aim to minimize image changes, ideally maintaining the original appearance. However, this doesn't meet our expectations, as we want fruit-related features to remain unchanged while other features are weakened or even eliminated. Therefore, we introduce a third and fourth loss function value. We hope that these two loss function values will cause significant changes in features other than the fruit in the image, resulting in a significant improvement in fruit recognition accuracy. The implementation process is explained below.
[0061] The fourth loss function value is used to characterize the feature differences between multiple output images corresponding to each training dataset. The smaller the fourth loss function value, the smaller the difference between different output images corresponding to different input images in a training dataset; conversely, the larger the fourth loss function value, the greater the difference between different output images corresponding to different input images in a training dataset, and correspondingly, the smaller the adversarial loss function value corresponding to the fourth loss function value. This application uses the second loss function value to characterize the adversarial loss function of the fourth loss function value.
[0062] It should be understood that the fourth loss function value in this application represents the difference between multiple output images corresponding to a training dataset. Therefore, theoretically, each dataset will have a corresponding fourth loss function value, and correspondingly, a corresponding second loss function value. This is done because, for a training dataset, the images are all augmented from a single original image; therefore, the poses of the fruits in the images within the dataset are generally consistent. However, for a training image, most of the content consists of non-fruit features, as fruits are typically small. For a training dataset, amplifying the differences between output images primarily amplifies non-fruit features, since these constitute the vast majority of features in the image. If these features remain unchanged, the requirement to reduce the second loss function value cannot be met. Of course, if only the second loss function value is reduced, the fruit features in the image will also change significantly; therefore, it is necessary to combine the first and third loss function values.
[0063] The third loss function value is used to characterize the difference between the fruit recognition result and the target recognition result for the training dataset. Therefore, the smaller the third loss function value, the higher the recognition accuracy. Here, the fruit recognition result refers to the actual recognition result of the input image in the training dataset after feature extraction, using the second neural network model for fruit recognition. This second neural network model can be any existing fruit recognition model. The target recognition result refers to the actual fruit recognition result corresponding to the input image, which can be manually calibrated by humans.
[0064] Therefore, the second loss function value forms an adversarial relationship with the first and third loss function values. This adversarial relationship ensures that when the image changes, the reduction of the second loss function value is not random, but rather it modifies non-fruit features. This modification improves the accuracy of fruit recognition. Thus, this modification gradually weakens or even eliminates non-fruit features in the image, while the remaining features (i.e., fruit features) are naturally preserved or even enhanced. Through this processing, the third image obtained after processing by the first neural network model becomes more conducive to fruit recognition, thereby improving the fruit recognition results.
[0065] It is understood that the first image may consist of only one image or multiple images (e.g., m images, where m is an integer greater than 1), and this application does not impose any limitation. If the first image consists of multiple images, step S120 can be performed sequentially on these multiple images, and then the resulting multiple second images can be fused to improve the effect of the final enhanced image. Alternatively, the image with the best effect (e.g., the image with the most uniform lighting and the highest clarity) can be selected from the multiple images to perform step S120, and the other images can be discarded. Or, the multiple images can be fused before performing step S120, and step S120 can be performed on the fused image. It is understood that performing the fusion processing of multiple images can refer to the fusion processing of image features to achieve feature equalization and improve the effect of subsequent image processing.
[0066] Optionally, the size of the first neural network model can be reduced by channel pruning to improve the lightweight nature of the harvesting robot and increase the efficiency of model processing. For example, channels with dependency coefficients less than or equal to a fifth threshold in the first neural network model are pre-deleted, and the parameters of the first neural network model are pre-adjusted based on a loss value, which characterizes the degree of structural change in the first neural network model, and the dependency coefficient characterizes the dependency relationship between the channel and the model.
[0067] S130, Identify the first region in the second image.
[0068] For example, after the processing flow of S120, a second image is finally obtained. In the second image, most non-fruit features have been weakened or even filtered out, while fruit features are basically preserved or even enhanced. Figure 3 and Figure 4 For example: Suppose Figure 3 The first image includes the target fruit tree and the target fruit. It is understood that... Figure 3 In the example shown, the first image includes the complete target fruit. However, in real-world applications, the first image may only include a portion of the target fruit tree, and this application does not limit this. It is also understood that... Figure 3 In the example shown, the first image includes one target fruit. However, in practical applications, the first image may include two or more target fruits, and this application does not limit this. In addition to the target fruit tree and the target fruit, the first image also includes other elements such as weeds, soil, and light spots. Figure 3 (Not all shown). Since the purpose of this application is to harvest the target fruit, all features in the first image other than the target fruit are interference features, which may cause interference when identifying the location of the target fruit. By processing the first image through the first neural network model, these interference features can be weakened or even eliminated. Specific effects include... Figure 4 As shown. It is understandable that... Figure 4 yes Figure 3 The corresponding processed result image, but Figure 4 The effect diagram in the diagram is merely an illustration, representing a relatively ideal state. It is mainly used to understand the processing effect of the first neural network model. The actual processing effect of the first neural network model may be affected by many factors.
[0069] Since the interference features in the second image are weakened, using the second image to identify the location of the target fruit can improve the recognition effect. In this embodiment, a first region is used to characterize the approximate area where the target fruit is located. The first region can be a regular shape or an irregular shape; this application does not limit this. However, it should be understood that the first region at least includes the complete target fruit.
[0070] In this embodiment, it is not necessary to use a dedicated neural network model for fruit recognition to process the second image to determine the first region. Instead, the first region is determined directly based on the first and second images. This is because the solution in this application does not intend to directly use images and artificial intelligence models to determine the location of the target fruit and then harvest it. Fruit harvesting requires very precise identification of the picking point, preferably on the fruit stem. Otherwise, the target fruit may be damaged during harvesting. Using images and artificial intelligence models to identify the picking point is highly susceptible to environmental factors in terms of accuracy, and the power consumption is relatively high.
[0071] The following is an example of a possible implementation: A first region is determined based on a first image and a second image, wherein the second image is composed of the first region and the second region, and the first image is composed of a third region and a fourth region. The first region and the third region correspond to each other, and the second region and the fourth region correspond to each other. The feature intensity deviation of corresponding pixels in the first region and the third region is less than or equal to a first threshold, and the feature intensity deviation of corresponding pixels in the second region and the fourth region is greater than the first threshold. This scheme will be explained in detail below.
[0072] Because the fruit is relatively small, it occupies a small area in the entire image, meaning the second image contains a lot of interfering information. Therefore, a method can be used to crop the second image, retaining the image content corresponding to the fruit, and the resulting image size is much smaller than the second image, allowing for precise scanning using LiDAR. However, the key issue is how to preserve the fruit in the cropped image (i.e., the aforementioned first region). To address this, we consider that in step S120, the first neural network model processes the first image to obtain the second image, and this model can preserve or even enhance the feature intensity of fruit-related features while weakening other features. Therefore, this characteristic can be used to determine the aforementioned first region. For example, comparing the first and second images reveals that since the second image is obtained by processing the first image, each pixel in the second image corresponds one-to-one with each pixel in the first image; the only difference is that the pixels in the second image are obtained after processing the pixels in the first image. Comparing each pixel sequentially reveals that the variation in feature intensity differs between pixels. This is because the first neural network model can weaken non-fruit features while preserving or even strengthening fruit features. Therefore, a numerical value can be used to measure the change in feature intensity of corresponding pixels in the first and second images. For convenience, this value is called the feature change amount. If the feature change amount of a pixel in the second image exceeds a preset threshold (e.g., greater than or equal to a first threshold), it is assigned to the second region; if the feature change amount of a pixel is less than the preset threshold, it is assigned to the first region. The second image consists of the first and second regions. In other words, we divide the second image into two parts based on the feature change amount: one part (the first region) includes pixels with relatively small feature changes, and the other part (the second region) includes pixels with relatively large feature changes. Based on the characteristics of the first neural network model, the image content corresponding to the fruit is highly likely to be contained in the first region. This allows us to discard the second region and reduce information interference. Figure 5 and Figure 6 For example ( Figure 5 and Figure 6 The examples in the text are respectively with Figure 4 and Figure 3 (The example in the text corresponds to this): Figure 5 A schematic diagram of the distribution of the first and second regions in the second image is shown, wherein the first region encompasses the target fruit; Figure 6 A schematic diagram of the distribution of the third and fourth regions in the first image is shown. The position of the third region corresponds to the position of the first region, and the position of the fourth region corresponds to the position of the second region.
[0073] Understandably, since the first neural network model is only used for feature processing, it is difficult to accurately determine the boundary of the target crop using the methods described above. Therefore, we want the first region to at least cover the target fruit, even if it includes some interfering features. In practice, the first threshold can be set slightly larger so that the first region is large enough to cover the target fruit.
[0074] S140. Determine the target detection area based on the first area, and emit laser pulses to the target detection area through a laser emitter.
[0075] S150. Receive the echo signal corresponding to the laser pulse through the laser receiving array to obtain measurement information of multiple pixels corresponding to the target detection area.
[0076] For example, the main purpose of steps S110 to S130 is to determine the approximate area where the target fruit is located, referred to as the first area. The size of the first area is typically smaller than a preset value (the size of the first area can be controlled by setting a first threshold). Therefore, a lidar can be used to precisely scan the first area without needing to perform a large-scale scan of the entire visible area. For instance, after determining the first area, the three-dimensional space corresponding to the first area is used as the target detection area of the lidar, and then a laser pulse is emitted from the lidar's laser emitter towards the target detection area.
[0077] To facilitate understanding of the solution provided in this application, a brief introduction to the working principle of lidar is given below: LiDAR is an active remote sensing device that uses a laser as the emission source and employs photoelectric detection technology. It can be used to detect the position, velocity, and other characteristics of objects within a target scene. Its working principle involves emitting laser pulses (detection signals) towards the target object within the scene, comparing the received echo signal reflected back from the target object with the emitted laser pulse, and outputting the corresponding electrical signal. The signal processing unit then processes the electrical signal appropriately to form a point cloud. By processing the point cloud, parameters such as the distance, orientation, height, velocity, attitude, and shape of the target object can be obtained, thus realizing the laser detection function. This allows for applications in navigation and avoidance, obstacle recognition, ranging, speed measurement, and autonomous driving scenarios for products such as automobiles, robots, logistics vehicles, and inspection vehicles. The following section combines... Figure 7 An example of a lidar structure is provided below. Figure 7 As shown, a lidar system includes a transmitting array (i.e., an array of transmitting devices) and a receiving array (i.e., an array of receiving devices). The transmitting array consists of multiple transmitting units. Figure 7 Each small square in the transmitting array represents a transmitting unit, and the receiving array consists of multiple receiving units. Figure 7Each small square in the receiver array represents a receiver unit. In one scenario, the number of transmitting units can be set to be the same as the number of receiving units, i.e., one transmitting unit corresponds to one receiving unit. In another scenario, it can be designed so that one transmitting unit corresponds to multiple receiving units, or multiple transmitting units correspond to one receiving unit. It is understandable that... Figure 1 This is merely a schematic diagram illustrating one possible correspondence between the transmitting and receiving units of a lidar system. The specific correspondence between the transmitting and receiving units in this application is not limited to this. Figure 7 Optionally, the lidar in this application may also include rotating components, such as a rotation drive platform, in addition to the transmitting and receiving channels; or scanning components, such as a rotating mirror, a galvanometer, or a combination of both. It is understood that this application does not limit the type or number of scanning components. It is also understood that this application does not limit the specific architecture of the lidar system.
[0078] Based on the above description, a lidar system can emit laser pulses towards the target detection area via a transmitting array, and then receive the corresponding echo signals via a receiving array. Since the target detection area is defined by a first region, and this first region includes the complete target fruit, the point cloud data generated based on the echo signals includes the point cloud data of the target fruit.
[0079] It is understandable that the size of the first region is smaller than a preset value. In other words, to achieve accurate measurement, the area of the target detection region is generally not very large. Therefore, the number of receiving units in the laser receiving array of the lidar can be set to be less than or equal to a preset threshold. It is also understandable that the more receiving units there are, the more echo signals can be received, the larger the measurement range, and the higher the measurement accuracy. However, correspondingly, the power consumption will be higher, and the data processing time will be longer. Given that this application has already narrowed the target detection region to a relatively small size through the aforementioned steps, it is unnecessary to use a high-precision, high-energy-consumption lidar. Therefore, a lidar with fewer than a preset threshold of receiving units can be used to achieve the measurement work required by this application, which can meet the measurement requirements, maintain relatively low power consumption, and improve measurement efficiency.
[0080] In one possible implementation, the specific number of receiving units in the LiDAR's receiving array can be determined based on the type of target fruit tree and the model of the agricultural harvesting robot. Because different fruit trees vary in size, the corresponding target fruits also differ in size, shape, and color. Furthermore, the processing effect of the first neural network model on the first image varies in different scenarios. Therefore, the final determined size of the first region will differ in different application scenarios. Thus, a uniform standard cannot be used to determine the number of receiving units in the LiDAR; instead, it must be determined based on the actual scenario.
[0081] When determining the number of receiving units, the influence of the agricultural harvesting robot model can also be considered. Different models of agricultural harvesting robots vary in their detection range, robotic arm length, and lidar installation location. These factors all affect the size of the first measured area and the specific measurement method. Considering the agricultural harvesting robot model when determining the number of receiving units can minimize the number of receiving units while meeting the detection requirements of the current scenario.
[0082] Optionally, in one possible implementation, different measurement ranges can be configured for the receiving units in the receiving array. It is understood that for a receiving unit in the receiving array, its measurement range refers to the laser detection range, which is determined by the transmission power of the corresponding transmitting unit. All other things being equal, a higher transmission power results in a larger measurement range, and a lower transmission power results in a smaller measurement range. However, increasing the transmission power consumes more resources. In the scenario corresponding to the embodiments of this application, the lidar is used to measure the target detection area, and the positions corresponding to different receiving units are also different. Since the first area is larger than the target fruit, for the receiving units in the receiving array, some receiving units receive echo signals reflected back from the target fruit, some from leaves or tree trunks, and others from objects behind the target fruit tree. In this application, the agricultural harvesting robot typically detects the target fruit tree at a preset distance. Therefore, when detecting different fruit trees, the measurement range of the receiving units for different fruit trees is actually within a certain range. For example, the measurement range of the receiving units in the middle position is relatively small because they are mostly reflected back from the target fruit, while the measurement range of the receiving units at the edge position is relatively large because they are mostly reflected back from other objects behind the target fruit tree. If all receiving units (and their corresponding transmitting units) adopt a relatively large measurement range, this can certainly handle the measurement work described in this application well, but it will also lead to a waste of detection resources. Therefore, we can set the current measurement range of a receiving unit based on its historical maximum measurement range. In this way, we can reduce the waste of detection resources.
[0083] S160. Determine the upper boundary point of the target fruit on the target fruit tree from multiple pixels, and determine the picking point based on the upper boundary point.
[0084] For example, step S150 can obtain measurement information of multiple pixels corresponding to the target detection area, i.e., point cloud data corresponding to the target detection area. Then, the picking point needs to be determined based on this measurement information.
[0085] In this embodiment, the fruit stalk of the target fruit is used as the picking point. Picking through the fruit stalk can minimize the damage to the fruit during picking.
[0086] In one possible implementation, the upper boundary point is determined from multiple measured pixels, and then the picking point is determined based on the upper boundary point. It is understood that the receiving array includes multiple receiving units, each corresponding to the measurement information of one pixel, including at least the measured distance and reflectivity. These multiple pixels are obtained by the lidar through repeated backscan measurements, and therefore these pixels also form an array. Assuming the array corresponding to these multiple pixels is as follows... Figure 9 As shown in the grid.
[0087] For ease of explanation, Figure 9 In the process, the pixels corresponding to the boundary of the target fruit are filled with shadows. These pixels are the boundary points corresponding to the target fruit, including the upper boundary point.
[0088] Since the fruit stalk connects to the upper boundary point of the fruit, accurately identifying the location of the upper boundary point is crucial for identifying the fruit stalk. The following is an example of one possible implementation.
[0089] First, candidate boundary points are determined. These candidate boundary points include the upper boundary point and are a subset of all pixels. The reason for determining candidate boundary points is that detecting the upper boundary point requires checking each pixel individually. Checking all pixels would be too time-consuming and power-intensive. Therefore, we aim to select a subset of pixels (candidate boundary points) first, and then determine the upper boundary point from this subset, thus reducing the number of checks and lowering power consumption.
[0090] In one possible implementation, p rows of pixels from the plurality of pixels are selected as candidate boundary points. More than m pixels in these p rows have a measurement distance greater than a third threshold, and more than n pixels have a measurement distance less than a fourth threshold. It should be understood that the measurement distance corresponding to the target fruit is necessarily less than the measurement distances of other surrounding areas; that is, the measurement distance corresponding to the upper boundary point is less than the measurement distance of the area above the upper boundary. If more than m pixels in the p rows have a measurement distance greater than the third threshold, it means that more than m pixels do not correspond to the target fruit; if more than n pixels have a measurement distance less than the fourth threshold, it means that more than n pixels correspond to the target fruit. As long as the values of m and n are set correctly, it can be determined that the p rows of pixels include the upper boundary point. Candidate boundary points can be easily determined in this way.
[0091] After determining the candidate boundary points, each pixel within these candidate boundary points is inspected. Taking the i-th pixel as an example: if the i-th pixel in the candidate boundary points meets preset conditions, it is determined as the upper boundary point of the target fruit. These preset conditions include: the difference between the measured distances of the x pixels above the i-th pixel and the measured distance of the i-th pixel itself is greater than a first threshold, and the difference between the measured distances of the y pixels below the i-th pixel and the measured distance of the i-th pixel itself is less than a second threshold. x, y, and i are all positive integers, and the values of x and y are determined based on the type of the target fruit tree. As the previous analysis shows, the measured distance of the upper boundary point is less than the measured distance of the area above the upper boundary. Therefore, by analyzing whether the measured distances of the corresponding pixels above and below a pixel meet the threshold conditions, the upper boundary point can be determined.
[0092] After determining the upper boundary point, the location of the picking point can be directly estimated based on the upper boundary point. Alternatively, the pixels located above the boundary point can be analyzed one by one to determine the specific location of the picking point. For example, the location of the picking point can be determined by analyzing the measured distance. This will not be explained in detail here.
[0093] S170. Harvest the target fruit based on the harvesting point.
[0094] Once the picking point is determined, the target fruit is picked based on that point. The specific picking method is not limited here, as it depends on the type of robotic arm of the picking robot. It can be picked by gripping with a claw-type arm, or by pruning the fruit stem with a scissor-like arm, causing the target fruit to fall onto a soft mat or net laid on the ground. Details will not be provided here.
[0095] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0096] This application provides a computer program product that, when run on a device, enables the device to perform the steps described in the various method embodiments above.
[0097] This application provides a chip for executing instructions. When the chip is running, it executes the technical solutions described in the above embodiments. Its implementation principle and technical effects are similar and will not be repeated here.
[0098] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0099] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0100] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0101] It should be understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0102] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
[0103] Furthermore, it should be noted that the various numerical designations used in this application (such as the terms "first," "second," "third," "fourth," and other terminology used in the specification, claims, and accompanying drawings, if any) are merely for descriptive convenience and are not intended to limit the scope of this application. The order of the process numbers does not imply the sequence of execution; the execution order of each process should be determined by its function and internal logic.
Claims
1. A smart harvesting method, characterized in that, The method is applied to agricultural harvesting robots, and the method includes: A first image is acquired, the image content of the first image includes a target fruit tree, and the target fruit tree includes target fruit; The first image is processed based on a first neural network model to obtain a second image. The first neural network model is trained on multiple training datasets. Each training dataset includes multiple images augmented from an original image. The original image contains fruits of the same type as the target fruit. The first neural network model is trained with the aim of reducing the values of a first loss function, a second loss function, and a third loss function. The first loss function value is used to characterize the feature differences between the input image and the output image of the first neural network model. The second loss function value is an adversarial loss function value of a fourth loss function value. The fourth loss function value is used to characterize the feature differences between multiple output images corresponding to each training dataset. The third loss function value is used to characterize the difference between the fruit recognition result and the target recognition result for the training dataset. Identify a first region in the second image, the second image is composed of the first region and the second region, the first image is composed of the third region and the fourth region, the first region corresponds to the third region, the second region corresponds to the fourth region, the feature intensity deviation of the corresponding pixels of the first region and the third region is less than or equal to a first threshold, and the feature intensity deviation of the corresponding pixels of the second region and the fourth region is greater than the first threshold. Based on the first region, a target detection area is determined, and a laser pulse is emitted toward the target detection area via a laser emitter; The laser receiving array receives the echo signal corresponding to the laser pulse to obtain the measurement information of multiple pixels corresponding to the target detection area. The number of receiving units in the laser receiving array is less than or equal to a preset threshold. The upper boundary point of the target fruit on the target fruit tree is determined from the plurality of pixels, and the picking point is determined based on the upper boundary point; The target fruit is harvested from the designated harvesting point.
2. The method according to claim 1, characterized in that, Determining the upper boundary point of the target fruit on the target fruit tree from the plurality of pixels includes: Each pixel in the candidate boundary points is detected, where the candidate boundary points are some or all of the plurality of boundary points; If the i-th pixel among the candidate boundary points meets the preset conditions, the i-th pixel is determined as the upper boundary point of the target fruit. The preset conditions include: the difference between the measured distances of the x pixels above the i-th pixel and the measured distance of the i-th pixel is greater than a first threshold, and the difference between the measured distances of the y pixels below the i-th pixel and the measured distance of the i-th pixel is less than a second threshold. The x, y, and i are all positive integers, and the values of x and y are determined based on the type of the target fruit tree.
3. The method according to claim 2, characterized in that, The method further includes: The p rows of pixels among the plurality of pixels are selected as the candidate boundary points. The measured distances of more than m pixels in the p rows are greater than the third threshold, and the measured distances of more than n pixels are less than the fourth threshold.
4. The method according to claim 1, characterized in that, The number of receiving units in the receiving array is determined based on the type of the target fruit tree and the model of the agricultural harvesting robot.
5. The method according to claim 4, characterized in that, The measurement ranges corresponding to different receiving units in the receiving array are different, and the measurement range corresponding to each receiving unit is determined based on the historical maximum measurement range of the receiving unit.
6. The method according to claim 1, characterized in that, In the first neural network model, channels with dependency coefficients less than or equal to the fifth threshold are pre-deleted, and the parameters of the first neural network model are pre-adjusted based on the loss amount, which is used to characterize the degree of structural change of the first neural network model, and the dependency coefficient is used to characterize the dependency relationship between the channel and the model.
7. An agricultural harvesting robot, characterized in that, The agricultural harvesting robot includes: The image acquisition module is used to acquire a first image, the image content of which includes a target fruit tree, and the target fruit tree includes target fruit; The model processing module is used to process the first image based on a first neural network model to obtain a second image. The first neural network model is trained on multiple training datasets, each training dataset including multiple images augmented from an original image. The original image contains fruits of the same type as the target fruit. The first neural network model is trained to reduce the values of a first loss function, a second loss function, and a third loss function. The first loss function value characterizes the feature differences between the input and output images of the first neural network model. The second loss function value is an adversarial loss function value, representing the feature differences between multiple output images corresponding to each training dataset. The third loss function value characterizes the difference between the fruit recognition result and the target recognition result for the training dataset. The recognition module is used to recognize a first region in the second image, the second image is composed of the first region and the second region, the first image is composed of the third region and the fourth region, the first region corresponds to the third region, the second region corresponds to the fourth region, the feature intensity deviation of the corresponding pixels of the first region and the third region is less than or equal to a first threshold, and the feature intensity deviation of the corresponding pixels of the second region and the fourth region is greater than the first threshold. The detection module is used to determine a target detection area based on the first area and to emit a laser pulse to the target detection area via a laser emitter; and to receive the echo signal corresponding to the laser pulse via a laser receiving array to obtain measurement information of multiple pixels corresponding to the target detection area, wherein the number of receiving units in the laser receiving array is less than or equal to a preset threshold. The detection module is used to determine the upper boundary point of the target fruit on the target fruit tree from the plurality of pixels, and to determine the picking point based on the upper boundary point; The harvesting module is used to control the mechanical device of the agricultural harvesting robot to harvest the target fruit based on the harvesting point.
8. An agricultural harvesting robot, comprising one or more processors and a memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the agricultural harvesting robot to perform the method as described in any one of claims 1 to 6.
9. A chip system, characterized in that, The chip system is applied to an agricultural harvesting robot, and the chip system includes one or more processors, which are used to invoke computer instructions to cause the agricultural harvesting robot to perform the method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on an agricultural harvesting robot, cause the agricultural harvesting robot to perform the method as described in any one of claims 1 to 6.