Crop position detection device and harvesting device

The crop position detection device uses a depth camera and advanced image processing techniques to accurately determine crop positions, addressing the challenge of complex environments and crop shapes by estimating target points from non-target portions.

JP2025073845APending Publication Date: 2025-05-13NAT UNIV CORP HOKKAIDO NAT UNIV ORG +1

Patent Information

Application Number
JP2023184954
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing crop position detection devices struggle to accurately locate specific parts of crops, such as the cut point, due to complexities in the surrounding environment and the shape of the crop itself.

Method used

A crop position detection device equipped with a depth camera, an object position acquisition unit, and a calculation unit that performs partial recognition, partial point extraction, and target position estimation processes to accurately determine the position of a target point on a crop by utilizing point clouds and image data.

Benefits of technology

The device enables precise detection of crop positions, even in complex environments, by estimating the target point's position based on partial point clouds from non-target portions, thereby improving harvesting accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073845000001_ABST
    Figure 2025073845000001_ABST
Patent Text Reader

Abstract

To make it easy to accurately obtain a position of a specific part of a crop.SOLUTION: Three-dimensional positions of a point group corresponding to objects existing in a space including a grape are acquired on the basis of an imaging result of the grape by a hand camera 140 as a depth camera. A hand camera recognition processing unit 152 executes segmentation processing of dividing image data indicating the imaging result of the grape by the hand camera 140 into regions relating to respective portions of the grape. The hand camera recognition processing unit 152 extracts the point group corresponding to the regions relating to a bunch of the grape from the point group where the three-dimensional positions have been acquired. The hand camera recognition processing unit 152 estimates a cut point position of a cob of the grape on the basis of the three-dimensional positions of the extracted point group relating to the bunch of the grape.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a crop position detection device and a harvesting device. [Background technology]

[0002] Conventionally, there are devices that acquire the position of crops based on the results of photographing the crops by a camera. Patent Document 1 relates to an example of such a device, which identifies the position of the crops to be harvested based on the results of two-stage photographing by a first imaging unit and a second imaging unit, and then harvests the crops. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2023-40516 A Summary of the Invention [Problem to be solved by the invention]

[0004] There are cases where it is required to accurately obtain the position of a specific part of a crop. For example, in order to properly harvest a crop, it is necessary to accurately obtain the position of a cut point in a specific part of the crop. However, depending on the surrounding environment of the crop to be detected and the shape characteristics of the crop itself, it may not be easy to accurately obtain the position of such a specific part.

[0005] An object of the present invention is to provide a crop position detection device and a harvesting device that can easily obtain the position of a specific part of a crop accurately. [Means for solving the problem]

[0006] The crop position detection device of the present invention comprises a depth camera, an object position acquisition means for acquiring the three-dimensional position of a point cloud corresponding to an object in a space including the crop based on the result of photographing the crop by the depth camera, and a calculation unit, wherein the calculation unit executes a partial recognition process for recognizing each part of the crop based on image data showing the result of photographing the crop by the depth camera, a partial point extraction process for extracting a partial point cloud, which is a point cloud corresponding to each part recognized by the partial recognition process, from the point cloud whose three-dimensional positions have been acquired by the object position acquisition means, and a target position estimation process for estimating the position of the target point based on the partial point cloud extracted by the partial point extraction process for non-target parts, which are parts of the crop that do not include the target point to be detected.

[0007] As a result of intensive research, the inventors have discovered that when it is not easy to accurately obtain the position of a specific part of a crop (i.e., a target part), it is easier to obtain the accurate position by basing it on parts other than the specific part (i.e., non-target parts). According to the crop position detection device of the present invention, the position of a target point to be detected is estimated based on a group of partial points extracted for non-target parts that do not include the target point. In this way, when it is not easy to accurately detect the position of a target part that includes the target point, it is easier to obtain the accurate position by estimating the position based on the other parts.

[0008] In addition, in the present invention, it is preferable that the calculation unit further executes a target position determination process in which either a first candidate position, which is the position of the target point obtained based on the partial point group extracted by the partial point extraction process for a target part that is a part of the crop including the target point, or the second candidate position, which is the position of the target point estimated by the target position estimation process, is determined as the position of the target point, or a result of calculating both of them is determined as a detection result of the position of the target point. According to this, either a first candidate position based on a point group extracted for the target part or a second candidate position based on a point group extracted for a non-target part can be selected. Therefore, since it is possible to select either an appropriate method, a more accurate position is likely to be obtained. Note that in the present invention, the "result of calculating both" refers to, for example, a result of weighted averaging the coordinates of the first candidate position and the coordinates of the second candidate position.

[0009] In addition, in the present invention, it is preferable that the calculation unit, in the target position estimation process, determines either the first candidate position or the second candidate position as the position of the target point based on at least one of the number of points in the partial point group extracted by the partial point extraction process for the target part and the positional relationship between the partial point group extracted by the partial point extraction process for the target part and the partial point group extracted by the partial point extraction process for the non-target part. According to this, the number of points and the positional relationship are indexes representing the reliability of the first candidate position. Therefore, the position of the target point is appropriately determined based on the reliability of the first candidate position.

[0010] In the present invention, it is preferable that the partial point extraction process includes a histogram acquisition process for acquiring a depth histogram for each part, the depth histogram being a histogram of the depth of the multiple points included in the point cloud whose three-dimensional positions are acquired by the object position acquisition means and whose two-dimensional positions correspond to each part recognized by the partial recognition process, and a central part extraction process for extracting, as the partial point cloud, points corresponding to the central part of the depth distribution range in the depth histogram acquired by the histogram acquisition process. According to this, by extracting a point cloud corresponding to the central part of the histogram, a point cloud that appropriately represents each part of the crop can be extracted.

[0011] In the present invention, it is preferable that the image data includes a plurality of frames arranged in the order of time when the depth camera photographs the crop, and the histogram acquisition process includes a frame-by-frame histogram acquisition process for acquiring a frame-by-frame histogram, which is a histogram of depth of the plurality of points whose two-dimensional positions correspond to each part of the crop acquired by the portion recognition process for each of the plurality of frames, and a histogram calculation process for acquiring the depth histogram based on the plurality of frame-by-frame histograms acquired by the frame-by-frame histogram acquisition process for the plurality of frames. This reduces adverse effects on the partial point cloud caused by obstacles that occur only in a specific frame, compared to the case where a partial point cloud is extracted from a histogram related to one frame.

[0012] In the present invention, it is preferable that the calculation unit further executes a frame acquisition process for acquiring a rectangular frame surrounding a crop based on the image data, and the partial recognition process is performed on an area in the image data that is set based on the rectangular frame acquired by the frame acquisition process. In this way, the area where the partial recognition process is performed is set to an appropriate range.

[0013] In the present invention, it is preferable that the part recognition process includes a process for recognizing a branch bearing fruit of a crop to be detected, and the target position estimation process estimates the position of the target point based on the positional relationship between the branch position recognized by the part recognition process and the non-target part. This makes it easier to accurately obtain the position of the target point based on the positional relationship between the non-target part and the fruiting branch.

[0014] A harvesting apparatus according to another aspect of the present invention comprises a crop position detection device, cutting means having a cutting tool and moving the cutting tool to the position of the target point detected by the crop position detection device and cutting the crop with the cutting tool, and traveling means supporting the position detection device and the cutting means and traveling to the vicinity of the crop.

[0015] According to the harvesting device of the present invention, appropriate harvesting of the crop is carried out based on the result of accurate detection of the position of the crop.

[0016] In the present invention, it is preferable that the partial recognition process includes a process of recognizing branches on which the crop whose position is detected by the position detection device bears, the partial point extraction process includes a process of extracting a partial point cloud, which is a point cloud corresponding to the branches recognized by the partial recognition process, from the point cloud whose three-dimensional positions are acquired by the object position acquisition means, and the cutting means adjusts the movement mode of the cutting tool based on the position of the partial point cloud extracted by the partial point extraction process for the branches. This reduces interference with the branches when cutting by the cutting means. [Brief description of the drawings]

[0017] [Figure 1] 1 is a schematic diagram of a grape harvesting apparatus according to one embodiment of the present invention; [Diagram 2] 2 is a block diagram showing the functional configuration of the grape harvesting apparatus of FIG. 1. [Diagram 3] 2 is a flowchart showing a series of processing steps performed by the grape harvesting apparatus of FIG. 1. [Figure 4]3 is a flowchart showing a series of processing steps performed by the base camera recognition processing unit in FIG. 2. [Diagram 5] 3 is a flowchart showing a series of processing steps performed by the hand camera recognition processing unit in FIG. 2. [Figure 6] 3 is an example of an image captured by the base camera or the hand camera in FIG. 2. [Figure 7] 7 is an example of a result of object detection performed on the image capture result of FIG. 6. [Figure 8] 8 is an example of an ROI image obtained based on the bounding box by the object detection in FIG. 7. [Figure 9] 9 is an example of a result of performing a segmentation process on the ROI image of FIG. 8. [Figure 10] 9 is another example of a result of performing segmentation processing on the ROI image of FIG. 8. [Figure 11] 3 is an example of a histogram created based on an ROI image generated by the base camera recognition processing unit or the hand camera recognition processing unit in FIG. 2. [Figure 12] 3 is a graph showing point clouds of each part of grapes and various cut point positions acquired by the hand camera recognition processing unit of Figure 2. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] A grape harvesting device 1 according to one embodiment of the present invention will be described below with reference to Figs. 1 to 12. In this embodiment, as an example, it is assumed that harvesting takes place in a farm field trained as a hedge where grape berries (corresponding to the "crop" of the present invention), branches and leaves, supports, wires, etc. are intricately intertwined. The grape harvesting device 1 has a harvesting section 100 and a traveling carriage 200 as shown in Figs. 1 and 2.

[0019] The harvesting unit 100 (corresponding to the "position detection device" of the present invention) has a robotic hand 110, a robotic arm 120, a base camera 130, a hand camera 140, a harvesting control unit 150, and a base 160 that supports these, as shown in Figures 1 and 2.

[0020] The robot hand 110 includes a drive mechanism including a plurality of movable claws 111, a cutter 112 (corresponding to the "cutting tool" of the present invention) fixed to the claws 111, and a motor for driving the claws 111. The drive mechanism can rotate and open / close the plurality of claws 111. This allows the robot hand 110 to cut the cob Gp1 of the grapes Gp while gripping it with the claws 111 and adjusting the angle of the claws 111 with respect to the cob Gp1. The robot arm 120 includes a drive mechanism including a plurality of links, joints connecting the links, a motor for driving the links, and the like. The robot hand 110 is supported at the tip of the robot arm 120. The rear end of the robot arm 120 is supported by a base 160. The robot arm 120 moves the robot hand 110 freely within its movable range. The robot hand 110 and the robot arm 120 correspond to the "cutting means" of the present invention.

[0021] The base camera 130 is directly installed on the base 160. The hand camera 140 is installed at the tip of the robot arm 120. Both the base camera 130 and the hand camera 140 are depth cameras. The depth camera has an imaging element using a CCD (Charge Coupled Device) system or a CMOS (Complementary Metal Oxide Semiconductor) system, and a depth sensor using a stereo system or a ToF (Time of Flight) system. The imaging element generates and outputs a two-dimensional digital image (hereinafter referred to as a "two-dimensional image") showing the result of imaging the subject. The depth sensor outputs the result of detecting the distance from the sensor to each point on the surface of the subject. Each of the base camera 130 and the hand camera 140 performs appropriate calculations on these outputs to output each of the XYZ coordinate values ​​representing the three-dimensional position of the point cloud corresponding to the subject within the shooting range and the two-dimensional image to the harvest control unit 150. Of these, the Z coordinate represents the depth. The function of acquiring the coordinates of the three-dimensional positions of the point clouds corresponding to the objects within the shooting range of each of the base camera 130 and the hand camera 140 corresponds to the function of the "object position acquisition means" of the present invention. One output from these cameras corresponds to one frame. Each camera outputs one frame at a predetermined time interval.

[0022] The harvesting control unit 150 (corresponding to the "calculation unit" of the present invention) controls the robotic hand 110 and the robotic arm 120 based on the output results from the base camera 130 and the hand camera 140. In this way, grape harvesting work is carried out by the robotic hand 110 and the robotic arm 120.

[0023] The harvest control unit 150 includes a processor that executes various types of arithmetic processing, and a memory that stores program data and other data to be executed by the processor. The processor executes arithmetic processing based on the program data, that is, software including a program causes hardware including a processor to execute arithmetic processing, thereby realizing various functions of the harvest control unit 150.

[0024] The harvest control unit 150 has a base camera recognition processing unit 151 and a hand camera recognition processing unit 152. The base camera recognition processing unit 151 stores in a memory the output results from the base camera 130 corresponding to the most recent frames arranged in chronological order of the images taken by the base camera 130. The base camera recognition processing unit 151 then controls the robot arm 120 based on the results of processing the output results for the multiple frames stored in the memory. The hand camera recognition processing unit 152 stores in a memory the output results from the hand camera 140 corresponding to the most recent frames arranged in chronological order of the images taken by the hand camera 140. The hand camera recognition processing unit 152 then controls the robot hand 110 and the robot arm 120 based on the results of processing the output results for the multiple frames stored in the memory. The detailed processing contents of the harvest control unit 150 including the base camera recognition processing unit 151 and the hand camera recognition processing unit 152 will be described later.

[0025] The harvest control unit 150 transmits the grape detection status based on the output from the base camera recognition processing unit 151 to a travel control unit 210 of the traveling carriage 200 described below. In addition, the harvest control unit 150 acquires the travel status of the traveling carriage 200 from the travel control unit 220, and manages the harvesting work based on the acquired results.

[0026] The traveling carriage 200 (corresponding to the "traveling means" of the present invention) is a device that travels by itself while supporting the harvesting section 100. As shown in FIG. 1 and FIG. 2, the traveling carriage 200 has a traveling control section 210, a carriage drive section 220, and a container 230. The carriage drive section 220 has a plurality of wheels and a wheel drive mechanism that drives the wheels. The traveling control section 210 controls the carriage drive section 220 based on the grape detection status from the harvesting control section 150 to make the traveling carriage 200 travel to an appropriate position relative to the grapes. In addition, the traveling control section 210 transmits the traveling status, for example, a status of traveling, a status of traveling to an appropriate position relative to the grapes, and a status of stopping, to the harvesting control section 150. The container 230 is a container that contains grapes harvested by the harvesting section 100.

[0027] The processing contents of the harvest control unit 150 and the travel control unit 210 will be described in detail below with reference to Figs.

[0028] First, the overall flow of the processing will be described with reference to FIG. 3. The travel control unit 210 controls the carriage drive unit 220 so that the travel carriage 200 travels in the field along the grape hedge (S101). The travel control unit 210 controls the carriage drive unit 220 so that the travel carriage 200 travels until grapes are detected in the vicinity of the grape harvesting device 1 by the harvest control unit 150 (S102, No → S101). This detection may be performed, for example, based on the base camera recognition processing unit 151 executing the processing of S201 to S204 described later and detecting grapes in S204. Alternatively, the base camera recognition processing unit 151 may execute the processing of S201 to S218 described later and determining in S218 that grapes are in a position where they can be harvested.

[0029] When the harvest control unit 150 detects grapes (S102, Yes), the traveling control unit 210 controls the carriage driving unit 220 to stop the traveling carriage 200 (S103). Next, the base camera recognition processing unit 151 executes the processes of S201 to S218 described later (S104). Next, based on the result of the process of S104, the base camera recognition processing unit 151 moves the robot hand 110 to the front of the grapes (S105). This process corresponds to the process of S219 described later. Next, the hand camera recognition processing unit 152 executes the processes of S301 to S331 described later (S106). Next, based on the result of the process of S106, the hand camera recognition processing unit 152 controls the robot arm 120 to move the robot hand 110 to the vicinity of the position of the grape cut point (corresponding to the "destination point" of the present invention) and harvests it (S107). Specifically, the robot hand 110 is moved to a position where the grape cob can be cut at the confirmed cut point position acquired in S330 described later. Next, the hand camera recognition processor 152 causes the robot hand 110 to grip the grape cob and cut it at the confirmed cut point position (S108). Next, the hand camera recognition processor 152 controls the robot hand 110 and the robot arm 120 to transfer the cut grapes to the container 230 (S109). Note that the processes in S105 to S109 correspond to the process in S332 described later.

[0030] Next, a series of processes performed by the base camera recognition processing unit 151 will be described with reference to FIG. 4. First, the base camera recognition processing unit 151 acquires an output result from the base camera 130 (S201). The base camera recognition processing unit 151 waits until acquisition of the output results for the necessary frames is completed (S202, No→S201). When acquisition of the output results from the base camera 130 is completed (S202, Yes), the base camera recognition processing unit 151 acquires, from the acquired output results, two-dimensional images showing crops and other objects within the range of shooting by the base camera 130 and three-dimensional positions of point clouds corresponding to these objects (S203). FIG. 6 shows an example of a two-dimensional image acquired in this process.

[0031] Next, the base camera recognition processing unit 151 executes an object detection process on the two-dimensional image acquired in S203 (S204). This object detection process is performed using a convolutional neural network (CNN) that has been trained to appropriately recognize grape berries under the hedge training conditions assumed in this embodiment. By the grape berry recognition process using the convolutional neural network, a bounding box (corresponding to the "rectangular frame" of the present invention) that individually surrounds the grape berries in the two-dimensional image is acquired as a detection result of the object detection process. FIG. 7 shows an example of a bounding box acquired by this process. Note that a unique model adjusted for this object detection process may be applied to the convolutional neural network, or an existing model such as Regional CNN, YOLO, or SSD (Single Shot MultiBox Detector) may be applied. Note that the process of S204 corresponds to the "frame acquisition process" of the present invention.

[0032] Next, the base camera recognition processing unit 151 judges whether or not the detection of grapes by the object detection processing of S204 has been successful, and if it judges that the detection has failed (S205, No), it returns to the processing of S201. On the other hand, if it judges that the detection has been successful (S205, Yes), the base camera recognition processing unit 151 acquires an ROI (Region of Interest) image from the two-dimensional image acquired in S203 (S206). In this processing, the base camera recognition processing unit 151 sets a detection frame for the segmentation processing of S207 based on the bounding box acquired in S204. Then, the base camera recognition processing unit 151 cuts out a portion corresponding to the detection frame from the two-dimensional image to set it as an ROI image. As an example, the detection frame specifies a square area (corresponding to the "area in image data" of the present invention) of a predetermined size having the same center as the center of the bounding box. The detection frame is set for each bounding box. FIG. 8 shows an example of an ROI image cut out in this processing.

[0033] Next, the base camera recognition processing unit 151 executes a segmentation process on the ROI image acquired in S206 (S207). The segmentation process is a process of classifying each pixel in the ROI image into a type of object. The segmentation process in S207 is performed using a convolutional neural network for semantic segmentation that has been trained to appropriately classify each pixel of the two-dimensional image into each part of the grape bunch, the stalk, the branch, etc. under the hedge training situation assumed in this embodiment. As a result, the ROI image is divided into each region (hereinafter, referred to as a divided region) representing the type of object. FIG. 9 shows an example of an ROI image divided into each divided region of the bunch region, the stalk region, and the branch region by this process. The process in S207 corresponds to the "partial recognition process" of the present invention.

[0034] Next, the base camera recognition processing unit 151 judges whether the segmentation processing in S207 has been successful, and if it judges that the detection has failed (S208, No), it returns to the processing in S201. On the other hand, if it judges that the detection has been successful (S209, Yes), the base camera recognition processing unit 151 performs processing to individualize the divided area acquired in S207 (S209). This processing is processing to individualize the bunch of grapes. Specifically, the area outside the bounding box acquired in S204 is excluded from the divided area showing the two overlapping bunches of grapes in the ROI image. As shown in an example in FIG. 10, this processing is effective when the divided areas related to a plurality of bunches of grapes are overlapped in the ROI image by the processing in S207. By this processing, in the divided area of ​​the bunch of grapes, the area inside the bounding box shown in FIG. 10 becomes the target of the subsequent processing, and the other areas are excluded.

[0035] Next, the base camera recognition processing unit 151 judges whether the individualization processing of S209 has been successful, and if it judges that the processing has failed (S210, No), it returns to the processing of S201. On the other hand, if it judges that the processing has been successful (S210, Yes), the base camera recognition processing unit 151 continues to perform reduction processing and enlargement processing of the ROI image that has been subjected to the individualization processing of S209 (S211). This removes noise contained in the ROI image.

[0036] Next, the base camera recognition processing unit 151 acquires the coordinates of the three-dimensional position (each of XYZ coordinates) in the division area related to the bunch of grapes in the ROI image (S212). Specifically, from the point cloud acquired in S203, a point cloud whose two-dimensional position coordinates (XY) correspond to each pixel in the division area related to the bunch of grapes is extracted. Next, the base camera recognition processing unit 151 generates a histogram of the Z coordinates for those points having valid Z coordinates from the point cloud extracted in S212 (S213). A valid Z coordinate corresponds to, for example, that the Z coordinate is within a range where grapes are considered to exist. The graph in FIG. 11 shows an example of a histogram generated by this process. In FIG. 11, the classes on the horizontal axis are set with a width of 0.002 m. The vertical axis indicates the frequency, that is, the number of points that fall into each class among the points having the above-mentioned valid Z coordinates. The process in S213 corresponds to the "histogram acquisition process" of the present invention.

[0037] In this embodiment, the base camera recognition processing unit 151 generates histograms for multiple frames (for example, 10 frames) by repeating the generation of one histogram (corresponding to the "frame-by-frame histogram" of the present invention) for each frame a predetermined number of times. Then, the base camera recognition processing unit 151 takes the average of the frequencies for each class for the histograms for the multiple frames, and sets the average as the histogram (corresponding to the "depth histogram" of the present invention) to be used for the processing from S214 onwards. In this way, by using the average of multiple frames, more accurate processing is possible than when only one frame is used. For example, if processing is performed based only on a frame in which an obstacle happens to be reflected, the result may be inaccurate. On the other hand, if the average of multiple frames including the frames before or after that frame is taken, the adverse effect on processing due to an obstacle reflected in one frame is reduced. Note that the processing of acquiring a histogram for each frame as described above corresponds to the "frame-by-frame histogram acquisition processing" of the present invention. Also, the processing of acquiring the average of histograms for multiple frames corresponds to the "histogram calculation processing" of the present invention.

[0038] Next, the base camera recognition processing unit 151 extracts a point group corresponding to the center part of the distribution range of the Z coordinate in the histogram acquired in S213 (S214). Specifically, the base camera recognition processing unit 151 extracts a point group (corresponding to the "partial point group" of the present invention) consisting of points within a predetermined range centered on a reference value in the histogram from among the points related to the histogram (i.e., points having valid Z coordinates among the points extracted in S212). An example of the reference value is the median of the Z coordinate in the distribution range of the histogram. Another example is the average value of the Z coordinate in the distribution range of the histogram. Yet another example is the median value of the Z coordinate of the class with the highest frequency in the histogram. An example of the predetermined range is a range including a number of points equivalent to P% of the total number of points related to the histogram on both sides of the reference value. Another example is a range including a range of a length equivalent to P% of the entire distribution range of the histogram on both sides of the reference value with respect to the length in the Z direction. P may be appropriately adjusted to a size that is just right for properly grasping the range of the grape bunch, and as an example, P=15. The five bars surrounded by the dashed line in FIG. 11 are an example of the central portion. The processes of S213 and S214 correspond to the "partial point extraction process" of the present invention. Also, the process of S214 corresponds to the "central portion extraction process" of the present invention.

[0039] Next, the base camera recognition processing unit 151 calculates the average value of the Z coordinate of the point group corresponding to the central part of the histogram extracted in S214 (S215). Next, the base camera recognition processing unit 151 calculates the average value of each of the X coordinate and Y coordinate of the point group corresponding to the central part of the histogram extracted in S214 (S216).

[0040] Next, the base camera recognition processing unit 151 performs low-pass filtering to remove noise on the average values ​​of the X, Y, and Z coordinates acquired in S215 and S216 (S217). Specifically, the average value of the Z coordinate is calculated based on the following formula 1. In formula 1, Z^ t is the average value of the Z coordinate after noise removal at time t (i.e., the average value processed in S217 in the past). Zt is the average value of the Z coordinate before noise removal at time t (i.e., the average value obtained by the process of S215). k is a constant related to the low-pass filter. In one embodiment, k=0.2. t-1 is the average value of the noise-removed Z coordinate at time (t-1). The noise-removed X, Y, and Z coordinates are then treated as the 3D positions of the grapes in subsequent processing.

[0041]

number

[0042] Similarly, the average values ​​of the X and Y coordinates are calculated based on the following formulas 2 and 3.

[0043]

number

[0044]

number

[0045] Next, the base camera recognition processing unit 151 determines whether the three-dimensional position of the grapes acquired in S217 is within a harvestable range (S218). This determination is made based on predefined conditions in terms of the structures of the robotic hand 110 and the robot arm 120. In other words, the determination is made based on whether the three-dimensional position of the grapes is within a range where the robotic arm 120 can appropriately move the robotic hand 110 to a position where the grapes can be grasped and cut. If it is determined that the position of the grapes is not within a harvestable range (S218, No), the base camera recognition processing unit 151 returns to the processing of S201.

[0046] In S218, when it is determined that the position of the grapes is within a harvestable range (S218, Yes), the base camera recognition processing unit 151 controls the robot arm 120 to move the robot hand 110 to the robot arm 120 so as to approach the grapes to a position near the front of the grapes. As a result, the hand camera 140 installed on the robot hand 110 is positioned at an appropriate shooting position close to the grapes.

[0047] Next, a series of processes performed by the hand camera recognition processing unit 152 will be described with reference to Fig. 5. Many steps including object detection, segmentation, and histogram creation used in the image processing performed by the hand camera recognition processing unit 152 are similar to the steps performed by the above-mentioned base camera recognition processing unit 151. Therefore, in the following, when the processing content is similar to that of the above-mentioned base camera recognition processing unit 151, detailed description of the specific processing content will be omitted as appropriate.

[0048] First, the hand camera recognition processing unit 152 acquires the output result from the hand camera 140 (S301). The hand camera recognition processing unit 152 waits until acquisition of the output results for the necessary frames is completed (S302, No → S301). When acquisition of the output results from the hand camera 140 is completed (S303, Yes), the hand camera recognition processing unit 152 acquires, from the acquired output results, two-dimensional images showing crops and other objects within the range of photography by the hand camera 140 and three-dimensional positions of point clouds corresponding to these objects (S303).

[0049] Next, the hand camera recognition processing unit 152 executes an object detection process on the two-dimensional image acquired in S303 (S304). This object detection process is performed using a convolutional neural network that has been trained to appropriately recognize grape berries under hedge training conditions, similar to the base camera recognition processing unit 151. The process of S304 corresponds to the "frame acquisition process" of the present invention.

[0050] Next, the hand camera recognition processing unit 152 judges whether or not the detection of grapes by the object detection processing of S304 has been successful, and if it judges that the detection has failed (S305, No), the process returns to S301. On the other hand, if it judges that the detection has been successful (S305, Yes), the hand camera recognition processing unit 152 acquires an ROI image from the bounding box of the two-dimensional image acquired in S303 (S306).

[0051] Next, the hand camera recognition processing unit 152 executes a segmentation process on the ROI image acquired in S303 (S307). The process of S307 corresponds to the "partial recognition process" of the present invention. Next, the hand camera recognition processing unit 152 determines whether the segmentation process of S307 was successful, and if it determines that the detection was unsuccessful (S308, No), the process returns to S301. On the other hand, if it determines that the detection was successful (S309, Yes), the hand camera recognition processing unit 152 executes a process of individualizing the divided regions acquired in S307 (S309).

[0052] Next, the hand camera recognition processing unit 152 judges whether the individualization process of S309 was successful, and if it judges that the process failed (S310, No), the process returns to S301. On the other hand, if it judges that the detection was successful (S310, Yes), the hand camera recognition processing unit 152 continues to perform reduction and enlargement processes of the ROI image that has been subjected to the individualization process of S309 (S311). This removes noise contained in the ROI image.

[0053] Next, the hand camera recognition processing unit 152 removes the division areas related to the branches and rachises recognized in S307 from the ROI image from which noise has been removed in S311 (S312). As a result, the ROI image mainly contains only the bunches of grapes. Next, the hand camera recognition processing unit 152 extracts the coordinates of the three-dimensional position of the division area related to the bunches of grapes in the ROI image from the coordinates of the three-dimensional position of the point cloud acquired in S303 (S313). Next, the hand camera recognition processing unit 152 generates a histogram of the Z coordinate for those points having valid Z coordinates among the points extracted in S313 (S314). Specifically, the average of the histograms for multiple frames is calculated, similar to the above-mentioned process by the base camera recognition processing unit 151.

[0054] Next, the hand camera recognition processing unit 152 extracts a point group corresponding to the center part of the distribution range of the Z coordinate in the histogram acquired in S314 (S315). FIG. 12 shows an example of a graph in which the extracted point group related to the grape bunch is plotted on an XY coordinate plane. Next, the hand camera recognition processing unit 152 calculates the average value of each of the X and Z coordinates of the point group corresponding to the center part of the histogram extracted in S315 (S316). Next, the hand camera recognition processing unit 152 acquires the Y coordinate of the point located at the top of the point group corresponding to the center part of the histogram extracted in S314 (S317). For example, when the +Y direction corresponds to the upward direction, the point with the maximum Y coordinate among the points in the point group in the center part corresponds to the point located at the top.

[0055] Next, the hand camera recognition processing unit 152 performs low-pass filtering to remove noise from the X, Y, and Z coordinate values ​​acquired in S316 and S317 (S318). This processing is performed using formulas 1 to 3, similar to the above-described processing by the base camera recognition processing unit 151. The X, Y, and Z coordinates acquired in this way are treated as the top position of the bunch of grapes in the subsequent processing.

[0056] Next, the hand camera recognition processor 152 acquires the coordinates of the three-dimensional position (each of XYZ coordinates) in the division area related to the cob in the ROI image from which noise has been removed in S311 (S319). Specifically, from the point cloud acquired in S303, a point cloud is extracted whose two-dimensional coordinates (XY) correspond to each pixel in the division area related to the cob in the ROI image. Next, the hand camera recognition processor 152 generates a histogram related to the Z coordinate for those points having a valid Z coordinate from the point cloud extracted in S319 (S320). Specifically, the average of the histograms for multiple frames is calculated, similar to the above-mentioned process by the base camera recognition processor 151.

[0057] Next, the hand camera recognition processor 152 extracts a point cloud corresponding to the center of the distribution range of the Z coordinate in the histogram acquired in S320 (S321). Specifically, the base camera recognition processor 151 extracts a point cloud (corresponding to the "partial point cloud" of the present invention) consisting of points within a predetermined range centered on a reference value in the histogram from among the points related to the histogram (i.e., points having valid Z coordinates among the points extracted in S319). The reference value and the predetermined range centered on the reference value are the same as those in the above-mentioned processing by the base camera recognition processor 151. FIG. 12 shows an example of a graph in which the extracted point cloud related to the cob is plotted on the XY coordinate plane.

[0058] Next, the hand camera recognition processing unit 152 calculates the average value of the Z coordinate of the point group corresponding to the central part of the histogram extracted in S321 (S322). Next, the hand camera recognition processing unit 152 calculates the average value of each of the X coordinate and Y coordinate of the point group corresponding to the central part of the histogram extracted in S321 (S323).

[0059] Next, the hand camera recognition processing unit 152 performs low-pass filter processing for noise removal on the average values ​​of the X, Y, and Z coordinates acquired in S322 and S323 (S324). This processing is performed using formulas 1 to 3, as in the above-mentioned processing by the base camera recognition processing unit 151. The X, Y, and Z coordinates after noise removal acquired in this way are treated as detection cut point positions (corresponding to the "first candidate position" of the present invention). FIG. 12 shows an example of the acquired detection cut point positions.

[0060] Next, the hand camera recognition processing unit 152 acquires the coordinates of the three-dimensional position in the division area related to the branch in the ROI image from which noise has been removed in S311 (S325). Next, the hand camera recognition processing unit 152 clusters the coordinates of the three-dimensional position in the division area related to the branch in the ROI image (S326). The clustering is a process of classifying the point cloud extracted in S325 for the division area related to the branch into one or more clusters. Specifically, if the distance between points included in the point cloud is within a predetermined threshold, the points are associated with a common cluster, and the point cloud is divided into clusters corresponding to each individual branch. If it is determined that the clustering in S326 has failed (S327, No), the hand camera recognition processing unit 152 returns to the process of S301. If it is determined that the clustering in S326 is successful (Yes in S327), the hand camera recognition processing unit 152 searches for an area relating to the branch that is the shortest distance from the top position of the grape bunch acquired in S318 among one or more divided areas relating to the clustered branches (S328). The branch that is the shortest distance from the top position of the grape bunch corresponds to the branch on which the grape berries have borne fruit, that is, the fruiting branch.

[0061] Next, the hand camera recognition processing unit 152 acquires the midpoint of the line segment that connects the area related to the resultant branch extracted in the search of S328 and the top position of the grape bunch at the shortest distance as an estimated cut point position (corresponding to the "second candidate position" of the present invention) (S329). An example of the acquired estimated cut point position is shown in Fig. 12. The process of S329 corresponds to the "destination position estimation process" of the present invention.

[0062] Next, the hand camera recognition processing unit 152 compares the detected cut point position acquired in S324 with the estimated cut point position acquired in S329, and determines the one with the higher reliability as the determined cut point position (S330).

[0063] The reliability is determined based on at least one of the following two criteria: (1) The number of points included in the point cloud corresponding to the central portion extracted in S321, i.e., the point cloud representing the cob. For example, if the number of points is smaller than a threshold, the reliability of the detected cut point position is determined to be lower than the reliability of the estimated cut point position. On the other hand, if the number of points is equal to or greater than the threshold, the reliability of the detected cut point position is determined to be higher than the reliability of the estimated cut point position.

[0064] (2) A criterion for the relationship between the three-dimensional position of the point cloud representing the rachis and the three-dimensional position of the point cloud corresponding to the central part extracted in S315, i.e., the point cloud representing the bunch of grapes. For example, based on the relationship between these three-dimensional positions, if the point cloud of the rachis is not within a predetermined range with respect to the point cloud of the bunch of grapes, it is determined that the reliability of the detected cut point position is lower than the reliability of the estimated cut point position. On the other hand, if the point cloud of the rachis is within a predetermined range with respect to the point cloud of the bunch of grapes, it is determined that the reliability of the detected cut point position is higher than the reliability of the estimated cut point position. The predetermined range is, for example, a range consisting of the inside of a predetermined cone whose apex and lowest point are the center of gravity of the point cloud of the bunch of grapes, and whose base is located above the apex.

[0065] Next, the hand camera recognition processing unit 152 judges whether or not the confirmed cut point position acquired in S330 is within a harvestable range (S331). This judgment is performed in the same manner as the judgment by the base camera recognition processing unit 151 described above. If it is judged that the confirmed cut point position is not within a harvestable range (S331, No), the hand camera recognition processing unit 152 returns to the processing of S301. If it is judged that the confirmed cut point position is within a harvestable range (S331, Yes), the hand camera recognition processing unit 152 executes the cob cutting processing by the robot hand 110 based on the confirmed cut point position as described above. At this time, it is preferable that the hand camera recognition processing unit 152 adjusts the movement mode of the robot hand 110 based on the division area related to the fruiting branch in the ROI image acquired in S307. For example, based on the position and inclination of the fruiting branch, the hand camera recognition processing unit 152 may rotate the claws 111 of the robotic hand 110 so that the angle is such that the claws 111 are less likely to interfere with the fruiting branch, while bringing the claws 111 closer to the cob, and then cause the robotic hand 110 to cut the cob.

[0066] According to the grape harvesting device 1 of the present embodiment described above, the cut point position on the grape cob (corresponding to the "target part" of the present invention) is estimated based on the point cloud extracted for the grape bunch (corresponding to the "non-target part" of the present invention). In this embodiment, a hedge-trained environment in which objects other than grape berries are arranged in a complex manner is assumed, and the cob including the cut point is thinner than the other parts. For this reason, there are frequent situations in which it is not easy to directly detect the cut point by image recognition. In this way, in a situation in which it is not easy to accurately detect the cut point position, by performing position estimation based on the grape bunch, which is more noticeable than the cob and easier to detect by image recognition, it becomes easier to obtain the accurate position of the cut point. This also makes it easier to properly perform grape harvesting by the harvesting unit 100.

[0067] In addition, in this embodiment, when determining the cut point position, it is possible to select either the detected cut point position directly detected by image recognition or the estimated cut point position based on the grape bunch. Therefore, since it is possible to select either the appropriate method, it is easy to obtain a more accurate position. Furthermore, the selection is made based on the number of points in the point cloud representing the rachis and the positional relationship between the point cloud representing the rachis and the point cloud representing the bunch. These are indices that indicate the reliability of the detected cut point position. Therefore, the cut point position is appropriately determined based on the reliability of the detected cut point position.

[0068] In this embodiment, when extracting a point cloud representing each part of the grapes, the center part of the point cloud histogram in the Z direction is extracted, thereby extracting a point cloud that appropriately represents each part of the grapes.

[0069] In this embodiment, the segmentation process is performed on the ROI image, whereby the region on which the segmentation process is performed is set to an appropriate range.

[0070] In the present embodiment, the estimated cut point positions are obtained based on the positional relationship between the grape clusters and the fruiting branches, which makes it easier to estimate the cut point positions accurately.

[0071] <Modification> The above is a description of a preferred embodiment of the present invention, but the present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the means for solving the problems.

[0072] In the above embodiment, semantic segmentation is used as the segmentation process. Alternatively, other segmentation-based processes may be used. For example, processes based on instance segmentation or panoptic segmentation may be used.

[0073] In the above embodiment, when determining the position from the detected cut point position and the estimated cut point position, one of them is selected based on the reliability. Alternatively, the result of calculation using both the detected cut point position and the estimated cut point position may be determined as the determined cut point position. For example, the result of taking a weighted average of both the X coordinate and the Y coordinate may be determined as the determined cut point position. In this case, the weight may be set according to the reliability.

[0074] In the above embodiment, the estimated cut point position is obtained based on the positional relationship between the top of the grape bunch and the fruiting branch (S329 in FIG. 5). Alternatively, the estimated cut point position may be a position a predetermined distance above the top of the grape bunch.

[0075] In the above embodiment, the present invention is applied to the harvesting of grapes, but the present invention may be applied to the harvesting of other crops. As other crops, various fruits and vegetables other than grapes may be targeted. In this case, for example, a point cloud representing the main body of the fruit may be used to estimate the cut point position of the stem connected to the fruit. [Explanation of symbols]

[0076] 1 Grape harvesting equipment 100 Harvesting Department 110 Robot Hand 111 Cutter 120 Robot Arm 130 Base Camera 140 Hand Camera 150 Harvest Control Unit 151 Base camera recognition processing unit 152 Hand camera recognition processing unit 200 Travelling cart

Claims

1. A depth camera and an object position acquisition means for acquiring a three-dimensional position of a point cloud corresponding to an object in a space including the crop based on the result of photographing the crop by the depth camera; A calculation unit, The calculation unit, a part recognition process for recognizing each part of the crop based on image data representing the result of photographing the crop by the depth camera; a partial point extraction process for extracting a partial point cloud, which is a point cloud corresponding to each of the parts recognized by the partial recognition process, from the point cloud whose three-dimensional positions have been acquired by the object position acquisition means; and a target position estimation process for estimating the position of the target point based on the partial point group extracted by the partial point extraction process for non-target parts, which are parts of the crop that do not contain the target point that is the subject of position detection.

2. The calculation unit, The crop position detection device of claim 1, further comprising a target position determination process for determining, as the position of the target point, either a first candidate position, which is the position of the target point obtained based on the partial point group extracted by the partial point extraction process for a target part, which is a part of the crop including the target point, or the second candidate position, which is the position of the target point estimated by the target position estimation process, or for determining the result of calculating both of them as the detection result of the position of the target point.

3. The calculation unit, in the destination position estimation process, The crop position detection device according to claim 2, characterized in that either the first candidate position or the second candidate position is determined to be the position of the target point based on at least one of the number of points in the partial point cloud extracted by the partial point extraction process for the target portion and the positional relationship between the partial point cloud extracted by the partial point extraction process for the target portion and the partial point cloud extracted by the partial point extraction process for the non-target portion.

4. The partial point extraction process includes: a histogram acquisition process for acquiring a depth histogram for each of the parts, the depth histogram being a histogram regarding depths of the multiple points included in the point cloud whose three-dimensional positions have been acquired by the object position acquisition means, the multiple points corresponding to two-dimensional positions of the respective parts recognized by the partial recognition process; The crop position detection device according to claim 1, further comprising a central portion extraction process for extracting, as the partial point cloud, a point corresponding to a central portion of a depth distribution range in the depth histogram acquired by the histogram acquisition process.

5. The image data includes a plurality of frames arranged in chronological order in which the depth camera photographs a crop; The histogram acquisition process includes: a frame-by-frame histogram acquisition process for acquiring a frame-by-frame histogram, which is a histogram regarding depth of the plurality of points whose two-dimensional positions correspond to each portion of the crop acquired by the portion recognition process for each of the plurality of frames; The crop position detection device according to claim 4, further comprising a histogram calculation process for acquiring the depth histogram based on the plurality of frame-based histograms acquired by the frame-based histogram acquisition process for the plurality of frames.

6. The calculation unit further executes a frame acquisition process for acquiring a rectangular frame surrounding a crop based on the image data, 2. The crop position detection device according to claim 1, wherein the portion recognition process is performed on an area in the image data that is set based on the rectangular frame acquired by the frame acquisition process.

7. the part recognition process includes a process of recognizing a branch on which a crop to be detected as a position is grown; The destination position estimation process includes:

2. The crop position detection device according to claim 1, further comprising a process of estimating a position of the target point based on a positional relationship between the branch position recognized by the part recognition process and the non-target part.

8. A crop position detection device according to any one of claims 1 to 7; a cutting means for moving the cutting tool to the position of the target point detected by the plant position detection device and cutting the plant with the cutting tool; A harvesting device comprising: a traveling means for supporting the position detection device and the cutting means and traveling to the vicinity of the crop.

9. the part recognition process includes a process of recognizing a branch bearing a crop that is a target of position detection by the position detection device, the partial point extraction process includes a process of extracting a partial point cloud, which is a point cloud corresponding to the branch recognized by the partial recognition process, from the point cloud whose three-dimensional positions have been acquired by the object position acquisition means, 9. The harvesting device according to claim 8, wherein the cutting means adjusts a movement mode of the cutting tool based on the positions of the group of partial points extracted by the partial point extraction process with respect to the branch.

Citation Information

Patent Citations

  • Harvesting method and harvesting device

    JP2023040516A

Cited By

  • Tree Management System

    JP7856828B1