Agricultural cutting system and method for generating cutting points
The agricultural cutting system automates pruning by generating precise cutting points using image processing and AI, improving reliability and reducing manual labor, thus enhancing vine health and grape quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-19
- Publication Date
- 2026-03-16
AI Technical Summary
Conventional agricultural cutting operations, such as pruning grapevines, are manual, costly, and lack reliability and consistency, affecting vine health and grape quality.
An agricultural cutting system and method that uses image capture, depth estimation, and artificial intelligence to generate precise two- and three-dimensional cutting points, employing a camera, robotic arm, and cutting tool for automated pruning.
Enables reliable and efficient cutting operations with improved consistency and reduced manual effort, enhancing vine health and grape quality.
Smart Images

Figure 2026508991000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Application No. 63 / 447,471, filed on February 22, 2023. The entire content of this application is incorporated herein by reference.
[0002] Background 1. Field of the Invention The present invention relates to an agricultural cutting system and a method for generating an agricultural cutting point.
Background Art
[0003] 2. Description of Related Art Conventionally, cutting operations in agriculture have been manual, which are expensive and time - consuming. For example, when the target crop is a grapevine, in the agricultural cutting operation of pruning the grapevine, a skilled person has to walk through the vineyard and manually prune the grapevine. Furthermore, since the technique of pruning grapevines can vary from person to person, it may reduce the reliability and consistency of grapevine pruning. Such low reliability and lack of consistency are undesirable because they may have an adverse effect on the health and growth of the grapevine and the quality of the grapes produced from the grapevine.
Summary of the Invention
[0004] For the reasons described above, there is a need for an agricultural cutting system and a method for generating an agricultural cutting point that can be used cheaply and reliably to perform a cutting operation on a target crop and generate a cutting point on the target crop.
Problems to be Solved by the Invention
[0005] Summary of the Invention Preferred embodiments of the present invention relate to an agricultural cutting system and a method for generating an agricultural cutting point.
Means for Solving the Problems
[0006] A preferred embodiment of the present invention provides a method for generating agricultural cutting points for a crop, comprising: capturing an image of the crop; generating a depth estimate of the crop; generating a segmented image that identifies different segments of the crop by segmenting the image of the crop; detecting agricultural features of the crop based on the image of the crop; generating two-dimensional cutting points based on the segmented image and the agricultural features; and generating three-dimensional cutting points based on the two-dimensional cutting points and the depth estimate of the crop.
[0007] In a method according to a preferred embodiment of the present invention, the generation of the depth estimation of the crop, the segmentation of the image of the crop, and the detection of the agricultural features are performed simultaneously.
[0008] In a method according to a preferred embodiment of the present invention, capturing the image of the crop includes capturing a plurality of images of the crop from a plurality of viewpoints, wherein the plurality of images are captured using a camera that is moved to the plurality of viewpoints.
[0009] In a method according to a preferred embodiment of the present invention, generating the depth estimation of the crops includes generating disparity estimations using an artificial intelligence disparity estimation model.
[0010] A method according to a preferred embodiment of the present invention further comprises generating a point cloud based on the depth estimation of the crop, and removing one or more points in the point cloud when one or more of those points have a depth greater than a depth-based threshold.
[0011] In a method according to a preferred embodiment of the present invention, the depth-based threshold is set based on the working range of the cutting system used to perform the cutting operation at the three-dimensional cutting point.
[0012] A method according to a preferred embodiment of the present invention further comprises generating a point cloud based on the depth estimation of the crop, and removing one or more points from the point cloud based on the density of points included in the point cloud.
[0013] In a method according to a preferred embodiment of the present invention, the segmented image is generated using an instance segmentation artificial intelligence architecture.
[0014] A method according to a preferred embodiment of the present invention further comprises training an instance segmentation artificial intelligence architecture using a segmentation dataset tailored to an instant segmentation task relating to the crop, wherein the segmentation dataset comprises a plurality of annotated images of the crop, the plurality of annotated images comprises a mask formed around a segment of the crop, and at least one of the plurality of annotated images comprises discrete portions of the same segment to which the same label has been assigned.
[0015] In a method according to a preferred embodiment of the present invention, the method further includes determining the agricultural feature locations of a plurality of agricultural features of the crop, including the agricultural features, using an object detection model that receives an image of the crop and detects the agricultural features within the image of the crop.
[0016] A method according to a preferred embodiment of the present invention further comprises determining the agricultural feature locations of a plurality of agricultural features of the crop, including the agricultural features, wherein generating the two-dimensional cut points comprises associating the plurality of agricultural features with a specific segment of the different segments of the crop based on the agricultural feature locations, assigning an identifier to each of the plurality of agricultural features relating to the specific segment to which the plurality of agricultural features is associated, and generating the two-dimensional cut points based on the identifiers assigned to the plurality of agricultural features.
[0017] In a method according to a preferred embodiment of the present invention, the segmented image includes a mask for identifying the different segments of the crop, the mask for identifying the different segments includes a specific mask for identifying the specific segment, and the plurality of agricultural features are associated with the specific segment when the agricultural feature locations of the plurality of agricultural features are within the specific mask or within a predetermined distance from the specific mask.
[0018] In a method according to a preferred embodiment of the present invention, the identifier is assigned to the plurality of agricultural features based on the distance between the agricultural feature location and a point on the particular mask.
[0019] In a method according to a preferred embodiment of the present invention, the two-dimensional cutting point is generated at a point between two of the plurality of agricultural features based on the identifiers assigned to the plurality of agricultural features.
[0020] In a method according to a preferred embodiment of the present invention, the method further includes moving the two-dimensional cut point generated at the points between the plurality of agricultural features so that the two-dimensional cut point is located within the particular mask when the two-dimensional cut point is not located within the particular mask.
[0021] A method according to a preferred embodiment of the present invention further includes determining the angle of the portion of the particular segment from which the two agricultural feature positions among the plurality of agricultural features are generated, and determining the cutting point angle of the two agricultural feature from which the two agricultural feature positions are generated.
[0022] In a method according to a preferred embodiment of the present invention, the agricultural characteristics of the crop are detected based on the segmented image.
[0023] In the method according to a preferred embodiment of the present invention, the method further includes combining a plurality of three-dimensional cutting points to generate a mega three-dimensional cutting point. Imaging the image of the crop includes imaging a plurality of images of the crop from a plurality of viewpoints. Generating the depth estimation of the crop includes generating a plurality of depth estimations of the crop corresponding to the plurality of images respectively. Segmenting the image of the crop includes generating a plurality of segmented images corresponding to the plurality of images respectively. Detecting the agricultural characteristics of the crop includes detecting the agricultural characteristics of the crop in each of the plurality of images. Generating the two-dimensional cutting points includes generating a plurality of two-dimensional cutting points based on the plurality of segmented images and the agricultural characteristics. Each of the plurality of two-dimensional cutting points corresponds to the plurality of viewpoints respectively. Generating the three-dimensional cutting points includes generating the plurality of three-dimensional cutting points based on the plurality of two-dimensional cutting points and the plurality of depth estimations of the crop. Each of the plurality of three-dimensional cutting points corresponds to the plurality of viewpoints respectively.
[0024] In the method according to a preferred embodiment of the present invention, combining the plurality of three-dimensional cutting points to generate the mega three-dimensional cutting point includes assigning a search radius to each of the plurality of three-dimensional cutting points, performing one or more spatial transformations to align the plurality of three-dimensional cutting points, and merging the plurality of three-dimensional cutting points into the mega three-dimensional cutting point when the plurality of three-dimensional cutting points are located within the search radius assigned to one of the plurality of three-dimensional cutting points.
[0025] In the method according to a preferred embodiment of the present invention, the method includes generating a plurality of point clouds each corresponding to the plurality of viewpoints based on the plurality of depth estimations of the crop, combining the plurality of point clouds to generate a mega point cloud, merging the mega 3D cut points and the mega point cloud, tracing a part of the mega point cloud corresponding to the segment of the crop where the mega 3D cut points are located to determine whether additional mega 3D cut points are located on the segment of the crop, and removing the additional mega 3D cut points when it is determined that the additional mega 3D cut points are on the segment of the crop.
[0026] In the method according to a preferred embodiment of the present invention, the method further includes determining whether additional mega 3D cut points are located on the segment of the crop where the mega 3D cut points are located, and deleting the additional mega 3D cut points when it is determined that the additional mega 3D cut points are on the segment of the crop.
[0027] In the method according to a preferred embodiment of the present invention, the method further includes determining a plurality of cut point angles corresponding to the plurality of 2D cut points, wherein the plurality of cut point angles are determined based on the angle of a part of a specific segment of the crop where the plurality of 2D cut points are generated, and determining a mega cut point angle of the mega 3D cut points based on the plurality of cut point angles.
[0028] In the method according to a preferred embodiment of the present invention, generating the depth estimation of the crop includes generating a parallax estimation based on the image of the crop.
[0029] In the method according to a preferred embodiment of the present invention, generating the depth estimation of the crop includes acquiring point cloud data from a LiDAR sensor.
[0030] A system for generating agricultural cutting points for crops according to a preferred embodiment of the present invention comprises a camera for capturing an image of the crop, and a processor configured or programmed to generate a segmented image that identifies different segments of the crop by segmenting the image of the crop, detect agricultural features of the crop based on the image of the crop, generate two-dimensional cutting points based on the segmented image and the agricultural features, and generate three-dimensional cutting points based on the two-dimensional cutting points and depth estimation of the crop.
[0031] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to simultaneously generate the depth estimate of the crop, segment the image of the crop, and detect the agricultural features.
[0032] In a system according to a preferred embodiment of the present invention, the image of the crop includes a plurality of images of the crop from a plurality of viewpoints, and the plurality of images are captured using the camera which is moved to the plurality of viewpoints.
[0033] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate the depth estimate of the crop by generating a disparity estimate using an artificial intelligence disparity estimation model.
[0034] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate a point cloud based on the depth estimation of the crop, and to remove one or more points in the point cloud if those points have a depth greater than a depth-based threshold.
[0035] In a system according to a preferred embodiment of the present invention, the depth-based threshold is set based on the working range in which a cutting operation can be performed at the three-dimensional cutting point.
[0036] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate a point cloud based on the depth estimation of the crop, and to remove one or more points from the point cloud based on the density of points included in the point cloud.
[0037] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate the segmented image using an instance segmentation artificial intelligence architecture.
[0038] In a system according to a preferred embodiment of the present invention, an instance segmentation artificial intelligence architecture is trained using a segmentation dataset tailored to an instant segmentation task relating to the crop, the segmentation dataset comprising a plurality of annotated images of the crop, the plurality of annotated images comprising a mask around the segments of the crop, and at least one of the plurality of annotated images comprising discrete portions of the same segment to which the same label has been assigned.
[0039] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to determine the agricultural feature positions of a plurality of agricultural features of the crop, including the agricultural features, based on the image of the crop.
[0040] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to determine the agricultural feature locations of a plurality of agricultural features of the crop, including the agricultural features, and the processor is configured or programmed to associate the plurality of agricultural features with a specific segment of the different segments of the crop based on the agricultural feature locations, assign an identifier to each of the plurality of agricultural features relating to the specific segment to which the plurality of agricultural features is associated, and generate the two-dimensional cut points based on the identifiers assigned to the plurality of agricultural features.
[0041] In a system according to a preferred embodiment of the present invention, the segmented image includes a mask for identifying the different segments of the crop, the mask for identifying the different segments includes a specific mask for identifying the particular segment, and the processor is configured or programmed to associate the plurality of agricultural features with the particular segment when the agricultural feature locations of the plurality of agricultural features are within the specific mask or within a predetermined distance from the specific mask.
[0042] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to assign the identifier to the plurality of agricultural features based on the distance between each agricultural feature location and a point on a particular mask.
[0043] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate the two-dimensional cut point at a point between two of the plurality of agricultural features, based on the identifiers assigned to the plurality of agricultural features.
[0044] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to move the two-dimensional cut point, generated at the points between the plurality of agricultural features, so that the two-dimensional cut point is located within the specific mask, if the two-dimensional cut point is not located within the specific mask.
[0045] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to determine the angle of the portion of the particular segment from which the two agricultural feature positions among the plurality of agricultural features are generated, and to determine the cutting point angle of the two agricultural feature from which the two agricultural feature positions are generated.
[0046] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to detect the agricultural characteristics of the crop based on the segmented image.
[0047] In a system according to a preferred embodiment of the present invention, the camera is operable to capture multiple images of the crop from multiple viewpoints, and the processor is configured or programmed to: combine multiple three-dimensional cutting points to generate a mega three-dimensional cutting point; generate multiple segmented images corresponding to each of the multiple images; detect the agricultural features of the crop in each of the multiple images; generate multiple two-dimensional cutting points based on the multiple segmented images and the agricultural features, where each of the multiple two-dimensional cutting points corresponds to one of the multiple viewpoints; and generate multiple three-dimensional cutting points based on the multiple two-dimensional cutting points and multiple depth estimations of the crop, where each of the multiple three-dimensional cutting points corresponds to one of the multiple viewpoints.
[0048] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to combine the plurality of three-dimensional cuts to generate the mega three-dimensional cut, assign a search radius to each of the plurality of three-dimensional cuts, perform one or more spatial transformations to align the plurality of three-dimensional cuts, and merge the plurality of three-dimensional cuts into the mega three-dimensional cut when the plurality of three-dimensional cuts are located within the search radius assigned to one of the plurality of three-dimensional cuts.
[0049] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate a plurality of point clouds, each corresponding to a plurality of viewpoints, based on the plurality of depth estimates of the crop; combine the plurality of point clouds to generate a mega point cloud; merge the mega 3D cut points with the mega point cloud; trace a portion of the mega point cloud corresponding to the segment of the crop where the mega 3D cut points are located to determine whether additional mega 3D cut points are located on the segment of the crop; and, if it is determined that additional mega 3D cut points are located on the segment of the crop, remove the additional mega 3D cut points.
[0050] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to determine whether an additional mega 3D cut point is located on the segment of the crop where the mega 3D cut point is located, and to delete the additional mega 3D cut point if it is determined that an additional mega 3D cut point is located on the segment of the crop.
[0051] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to: determine a plurality of cutting angles corresponding to the plurality of two-dimensional cutting points, where the plurality of cutting angles are determined based on the angles of the portion of a particular segment of the crop from which the plurality of two-dimensional cutting points are generated; and determine the mega cutting angle of the mega three-dimensional cutting point based on the plurality of cutting angles.
[0052] In a system according to a preferred embodiment of the present invention, the processor is configured or programmed to generate a disparity estimate based on the image of the crop in order to generate the depth estimate of the crop.
[0053] In a system according to a preferred embodiment of the present invention, the system comprises a LiDAR sensor, and the processor is configured or programmed to generate the depth estimation of the crop based on point cloud data obtained from the LiDAR sensor.
[0054] In a system according to a preferred embodiment of the present invention, the system comprises a mobile body, a robotic arm supported by the mobile body, and a cutting tool attached to the robotic arm, wherein the camera is supported by the mobile body, and the processor is configured or programmed to control the movement of the robotic arm to which the cutting tool is attached based on the three-dimensional cutting point.
[0055] The above and other features, elements, steps, configurations, characteristics, and advantages of the present invention will become more apparent from the following detailed description of preferred embodiments of the invention with reference to the accompanying drawings. [Brief explanation of the drawing]
[0056] A patent or application file includes at least one color drawing. The patent or application file also includes a black and white line drawing corresponding to each of the at least one color drawing.
[0057] Figure 1 shows a front perspective view of a cutting system according to a preferred embodiment of the present invention.
[0058] Figure 2 shows an enlarged view of a part of the cutting system according to a preferred embodiment of the present invention.
[0059] Figures 3A and 3B show an example block diagram of a cloud system including a cutting system, a cloud platform, and a user platform according to a preferred embodiment of the present invention.
[0060] Figure 4 is a flowchart showing the cutting point generation process according to a preferred embodiment of the present invention.
[0061] Figures 5A and 5B show examples of images captured in the imaging step according to a preferred embodiment of the present invention, where Figure 5A shows an example of an image captured in the imaging step in black and white, and Figure 5B shows an example of an image captured in the imaging step in color.
[0062] Figures 6A and 6B are flowcharts illustrating an example of the disparity estimation step according to a preferred embodiment of the present invention, where Figure 6A shows the flowchart in black and white and Figure 6B shows the flowchart in color.
[0063] Figures 7A and 7B are flowcharts illustrating other examples of the disparity estimation step according to a preferred embodiment of the present invention, where Figure 7A shows the flowchart in black and white and Figure 7B shows the flowchart in color.
[0064] Figures 8A and 8C show point cloud generation steps according to preferred embodiments of the present invention, with Figure 8A showing the point cloud generation steps in black and white and Figure 8C showing the point cloud generation steps in color.
[0065] Figures 8B and 8D show a point cloud registration step according to a preferred embodiment of the present invention, with Figure 8B showing the point cloud registration step in black and white and Figure 8D showing the point cloud registration step in color.
[0066] Figures 9A and 9C show an example of a depth-based thresholding process according to a preferred embodiment of the present invention, where Figure 9A shows the depth-based thresholding process in black and white, and Figure 9C shows the depth-based thresholding process in color.
[0067] Figures 9B and 9D show an example of an outlier removal process according to a preferred embodiment of the present invention, where Figure 9BA shows the outlier removal process in black and white, and Figure 9D shows the outlier removal process in color.
[0068] Figures 10A and 10B show an example of a component division step according to a preferred embodiment of the present invention, where Figure 10A shows the component division step in black and white and Figure 10B shows the component division step in color.
[0069] Figures 11A, 11B, and 11C show examples of segmented images generated in the component division step according to a preferred embodiment of the present invention, where Figures 11A and 11B show examples of segmented images in black and white, and Figure 11C shows an example of a segmented image in color.
[0070] Figures 12A and 12B show examples of images annotated using a computer-implemented labeling tool according to a preferred embodiment of the present invention, where Figure 12A shows the image in black and white and Figure 12B shows the image in color.
[0071] Figures 13A and 13B show an expansion process according to a preferred embodiment of the present invention, with Figure 13A showing the expansion process in black and white and Figure 13B showing the expansion process in color.
[0072] Figures 14A and 14B show examples of agricultural feature detection steps according to preferred embodiments of the present invention, with Figure 14A showing an example of the agricultural feature detection step in black and white and Figure 14B showing an example of the agricultural feature detection step in color.
[0073] Figures 15A and 15B show examples of feature images generated in the agricultural feature detection step according to a preferred embodiment of the present invention, where Figure 15A shows the feature image in black and white and Figure 15B shows the feature image in color.
[0074] Figures 16A and 16B show examples of images annotated using a computer-implemented labeling tool according to a preferred embodiment of the present invention, where Figure 16A shows the image in black and white and Figure 16B shows the image in color.
[0075] Figures 17A and 17B show examples of the cutting point generation step according to a preferred embodiment of the present invention, with Figure 17A showing an example of the cutting point generation step in black and white and Figure 17B showing an example of the cutting point generation step in color.
[0076] Figure 18 is a flowchart showing the cutting point generation step according to a preferred embodiment of the present invention.
[0077] Figure 19 is a flowchart showing the process for determining the cutting point angle according to a preferred embodiment of the present invention.
[0078] Figures 20A and 20B show feature images illustrating the process for determining the cutting point angle according to a preferred embodiment of the present invention, where Figure 20A shows the feature image in black and white and Figure 20B shows the feature image in color.
[0079] Figures 21A and 21B show examples of the cutting point projection step according to a preferred embodiment of the present invention, with Figure 21A showing an example of the cutting point projection step in black and white and Figure 21B showing an example of the cutting point projection step in color.
[0080] Figures 22A and 22B show examples of the cutting point registration step according to a preferred embodiment of the present invention, with Figure 22A showing an example of the cutting point registration step in black and white, and Figure 22B showing an example of the cutting point registration step in color.
[0081] Figures 23A and 23B show examples of trace modules according to preferred embodiments of the present invention, with Figure 23A showing an example of a trace module in black and white and Figure 23B showing an example of a trace module in color.
[0082] Figures 24A and 24B show examples of agricultural feature projection steps according to preferred embodiments of the present invention, with Figure 24A showing an example of the agricultural feature projection step in black and white and Figure 24B showing an example of the agricultural feature projection step in color.
[0083] Figures 25A and 25B show examples of agricultural feature registration steps according to preferred embodiments of the present invention, with Figure 25A showing an example of the agricultural feature registration step in black and white, and Figure 25B showing an example of the agricultural feature registration step in color. [Modes for carrying out the invention]
[0084] Detailed description of preferred embodiments Figure 1 shows a front perspective view of a cutting system 1 according to a preferred embodiment of the present invention. As shown in Figure 1, the cutting system 1 may have a moving body or the like. However, the cutting system 1 can be installed on a moving body or a cart that can be towed by a person, or on a self-propelled or self-propelled cart or moving body.
[0085] As shown in Figure 1, the cutting system 1 includes a base frame 10, side frames 12 and 14, a horizontal frame 16, and a vertical frame 18. The side frames 12 and 14 are mounted on the base frame 10, and the side frames 12 and 14 directly support the horizontal frame 16. The vertical frame 18 is mounted on the horizontal frame 16. One or more devices, such as a camera 20, a robotic arm 22, and / or a cutting tool 24, may be mounted and supported on, for example, the vertical frame 18, and / or other frames among frames 10, 12, 14, or 16.
[0086] The base frame 10 has a base frame motor 26 that can move the side frames 12 and 14 along the base frame 10 so that one or more devices can be moved in the depth direction (z-axis shown in Figure 1). The horizontal frame 16 has a horizontal frame motor 28 that can move the vertical frame 18 along the horizontal frame 16 so that one or more devices can be moved in the horizontal direction (x-axis shown in Figure 1). The vertical frame 18 has a vertical frame motor 30 that can move one or more devices along the vertical frame 18 in the vertical direction (y-axis shown in Figure 1). Each of the base frame motor 26, horizontal frame motor 28, and vertical frame motor 30 may be, for example, a screw motor. A screw motor can provide relatively high precision for precisely moving and positioning one or more devices. However, each of the base frame motor 26, horizontal frame motor 28, and vertical frame motor 30 may be any motor that provides a continuous torque of, for example, about 0.2 Nm or more, preferably about 0.3 Nm or more.
[0087] Each of the base frame motor 26, horizontal frame motor 28, and vertical frame motor 30 may be designed and / or sized according to the total weight of one or more devices. Furthermore, the couplers for each of the base frame motor 26, horizontal frame motor 28, and vertical frame motor 30 may be modified according to the diameter of the motor shaft and / or the corresponding mounting hole pattern.
[0088] The base frame 10 can be mounted on the base 32, and the base electronics system 34 can also be mounted on the base 32. Multiple wheels 36 can be mounted on the base 32. For example, as shown in Figures 3A and 3B, the multiple wheels 36 can be controlled by the base electronics system 34, which may have a power supply 35 for driving an electric motor 37, etc. As an example, the multiple wheels 36 can be driven by an electric motor 37 having a target capacity of about 65 kW to about 75 kW, and the power supply 35 for the electric motor 37 may be a battery with a capacity of about 100 kWh.
[0089] The base electronics system 34 also includes a processor and memory elements programmed or configured to perform autonomous navigation of the cutting system 1. Furthermore, as shown in Figure 1, a LiDAR (light detection and ranging) system 38 and a Global Navigation Satellite System (GNSS) 40 are installed or supported, for example, on the base frame 10 or base 32 and / or other frames among frames 10, 12, 14, or 16, so that the position data of the cutting system 1 can be determined. The LiDAR system 38 and GNSS 40 may be used for obstacle avoidance and navigation when the cutting system 1 is moving autonomously. Preferably, for example, the cutting system 1 is implemented using a remote control interface and can communicate via one or more of Ethernet, USB, wireless communication, and GPS RTK (real-time kinematics). The remote control interface and communication device may be included in either or both of the base electronics system 34 and the imaging electronics system 42 (described later). As shown in Figure 1, the cutting system 1 may include a display device 43 that displays data and / or images obtained by one or more devices and information provided by the base electronic system 34 (e.g., position, speed, battery life of the cutting system 100), or may be communicably connected to such a display device 43. Alternatively, the data and / or images obtained by one or more devices and provided by the base electronic system 34 may be displayed to the user through a user platform.
[0090] Figure 2 is an enlarged view of a part of the cutting system 1 having one or more devices. As shown in Figure 2, one or more devices include, for example, a camera 20, a robotic arm 22, and a cutting tool 24, which may be mounted on the vertical frame 18 and / or other frames among frames 10, 12, 14, or 16. Further devices among the one or more devices may also be mounted on, for example, the vertical frame 18 and / or other frames among frames 10, 12, 14, or 16.
[0091] Camera 20 may include a stereo camera, an RGB camera, etc. As shown in Figure 2, camera 20 may have a body 20a that includes a first camera / lens 20b (e.g., left camera / lens) and a second camera / lens 20c (e.g., right camera / lens). Alternatively, body 20a may include two or more cameras / lenses. The resolution of camera 20 may be, for example, 1536 × 2048 pixels or 2448 × 2048 pixels, but camera 20 may have a different resolution. Camera 20 may have, for example, a PointGrey CM3-U3-31S4C-CS or PointGrey CM3-U3-50S5C sensor and a 3.5mm f / 2.4 or 5mm f / 1.7 lens, and a field of view of 74.2535 × 90.5344 or 70.4870 × 80.3662. Camera 20 may have other sensors and lenses, and may also have different fields of view.
[0092] One or more light sources 21 may be mounted on one or more sides of the camera body 20a. The light sources 21 may have LED light sources oriented in the same direction as one or more devices, such as the camera 20, along the z-axis as shown in Figure 1. The light sources 21 may provide illumination to one or more objects imaged by the camera 20. For example, the light sources 21 may act as a flash to compensate for ambient light when imaged by the camera 20 during daytime operation. During nighttime operation, the light sources 21 may act as a flash for the camera 20, or the light sources may provide steady illumination to the camera 20. In a preferred embodiment, one or more light sources 21 may have, for example, 100-watt LED modules, but LED modules with different wattages (e.g., 40 watts or 60 watts) may be used.
[0093] The robot arm 22 may include robot arms known to those skilled in the art, such as the Universal Robot 3 e-series robot arms and the Universal Robot 5 e-series robot arms. For example, the robot arm 22, also known as an articulated robot arm, may have multiple joints that act as axes enabling degrees of freedom of motion. Here, the more rotary joints the robot arm 22 has, the higher the degrees of freedom of motion it has. For example, the robot arm 22 may have 4 to 6 joints, and these will provide the same number of axes of rotation for motion.
[0094] In a preferred embodiment of the present invention, the controller may be configured or programmed to control the movement of the robot arm 22. For example, the controller may be configured or programmed to control the movement of the robot arm 22 to which a cutting tool 24 is attached, and to position the cutting tool 24, according to steps described later. For example, the controller may be configured or programmed to control the movement of the robot arm 22 based on the position of the cutting point located on the target crop.
[0095] In a preferred embodiment of the present invention, the cutting tool 24 has a body 24a and a blade portion 24b, as shown, for example, in Figure 2. The blade portion 24b may have a driven blade that moves relative to a fixed blade and is actuated to perform a cutting operation together with the fixed blade. The cutting tool 24 may have, for example, a cutting device disclosed in U.S. Patent Application No. 17 / 961,666, which is titled “End Effector Including cutting blade and pulley assembly,” the title of which is incorporated herein by reference.
[0096] In a preferred embodiment of the present invention, the cutting tool 24 may be attached to the robot arm 22 using a robot arm mount assembly 23. The robot arm mount assembly 23 may be, for example, the robot arm mount assembly disclosed in U.S. Patent Application No. 17 / 961,668, which is titled “Robotic Arm Mount Assembly including Rack and Pinion” and is incorporated herein by reference.
[0097] The cutting system 1 may have an imaging electronic system 42, which can be installed on the side frame 12 or side frame 14, for example, as shown in Figure 1. The imaging electronic system 42 may supply power to and control each of the base frame motor 26, the horizontal frame motor 28, and the vertical frame motor 30. That is, the imaging electronic system 42 may have a power supply that provides power to each of the base frame motor 26, the horizontal frame motor 28, and the vertical frame motor 30. The imaging electronic system 42 may also have a processor and memory element programmed or configured to control each of the base frame motor 26, the horizontal frame motor 28, and the vertical frame motor 30. The processor and memory element of the imaging electronic system 42 may also be configured or programmed to control one or more devices having a camera 20, a robot arm 22, a robot arm mount assembly 23, and a cutting tool 24. The processor and memory element of the imaging electronic system 42 may also be configured or programmed to process image data obtained by the camera 20.
[0098] As described above, the imaging electronic system 42 and the base electronic system 34 may include processors and memory elements. The processors may be hardware processors, multipurpose processors, microprocessors, dedicated processors, digital signal processors (DPS), and / or other types of processing components configured or programmed to process data. The memory elements may include one or more volatile, non-volatile, and / or interchangeable data storage elements. For example, the memory elements may include magnetic, optical, and / or flash storage elements that can be integrated whole or partially with the processor. The memory elements may store instructions and / or instruction sets or programs that can be read and / or executed by the processor.
[0099] In another preferred embodiment of the present invention, the imaging electronics system 42 may be partially or completely implemented by the base electronics system 34. For example, each of the base frame motor 26, the horizontal frame motor 28, and the vertical frame motor 30 may receive power from and / or be controlled by the base electronics system 34 rather than the imaging electronics system 42.
[0100] In a further preferred embodiment of the present invention, the imaging electronic system 42 may be connected to a separate power supply (one or more) from the base electronic system 34. For example, the power supply may be included in one or both of the imaging electronic system 42 and the base electronic system 34. The base frame 10 may also be detachably attached to the base 32 so that the base frame 10, side frames 12 and 14, horizontal frame 16, vertical frame 18, and the components installed thereon can be installed on another mobile body or the like.
[0101] The base frame motor 26, horizontal frame motor 28, and vertical frame motor 30 can move one or more devices in three separate directions or along three separate axes. However, in another preferred embodiment of the present invention, only a portion of one or more devices, such as a camera 20, a robotic arm 22, and a cutting tool 24, may be moved by the base frame motor 26, horizontal frame motor 28, and vertical frame motor 30. For example, the base frame motor 26, horizontal frame motor 28, and vertical frame motor 30 may move only the camera 20. Furthermore, the cutting system 1 may be configured to move the camera 20 linearly along only a single axis while the camera captures multiple images, as will be described later. For example, the horizontal frame motor 28 may be configured to move the camera 20 linearly across a target crop such as a grapevine, and the camera 20 may capture multiple images of the grapevine.
[0102] The imaging electron system 42 and the base electron system 32 of the cutting system 1 are, for example, VIDIA(R) JETSON TMThe mobile platform may be provided by being partially or fully implemented by edge computing, such as by an AGX computer. In a preferred embodiment of the present invention, edge computing provides all the computational and communication needs of the disconnected system 1. Figures 3A and 3B show an example block diagram of a cloud system that includes a mobile platform and interacts with a cloud platform and a user platform. As shown in Figures 3A and 3B, the edge computing of the mobile platform includes a cloud agent, which is a service-based component that facilitates communication between the mobile platform and the cloud platform. For example, the cloud agent can receive command and instruction data from the cloud platform (e.g., a web application on the cloud platform) and forward the command and instruction data to the corresponding component of the mobile platform. As another example, the cloud agent can send operational and production data to the cloud platform. Preferably, the cloud platform may include software components and data storage to maintain the overall operation of the cloud system. The cloud platform preferably provides enterprise-level services with on-demand capability, fault tolerance, and high availability (e.g., Amazon Web Services). TMThe cloud platform includes one or more application programming interfaces (APIs) for communicating with the mobile platform and the user platform. Preferably, the APIs are protected with a high level of security, and the capacity of each API can be automatically adjusted according to the computing load. The user platform controls the cloud system and provides a dashboard for receiving data acquired by the mobile platform and the cloud platform. The dashboard can be implemented by a web-based application (e.g., an internet browser), a mobile application, a desktop application, etc.
[0103] As an example, the edge computing of the mobile platform shown in Figures 3A and 3B can obtain data from hardware GPS (Global Positioning System) (e.g., GNSS 40) and LiDAR data (e.g., from LiDAR system 38). The mobile platform can also obtain data from camera 20. The edge computing of the mobile platform may have, for example, temporary storage for storing the raw data obtained by camera 20. The edge computing of the mobile platform may also have, for example, persistent storage for storing processed data. Specifically, camera data stored in temporary storage may be processed by an artificial intelligence (AI) model, the camera data may then be stored in persistent storage, and a cloud agent may retrieve and transmit the camera data from persistent storage.
[0104] Figure 4 is a flowchart showing a cutpoint generation process according to a preferred embodiment of the present invention. The cutpoint generation process shown in Figure 4 has multiple steps, including an imaging step S1, a disparity estimation step S2, a component division step S3, an agricultural feature detection step S4, a point cloud generation step S5, a point cloud registration step S6, a cutpoint generation step S7, a cutpoint projection step S8, a cutpoint registration step S9, a mega registration step S10, an operation step S11, an agricultural feature projection step S12, and an agricultural feature registration step S13. These will be described in more detail below.
[0105] In a preferred embodiment of the present invention, the disparity estimation step S2, the component division step S3, and the agricultural feature detection step S4 can be performed simultaneously. Alternatively, one or more of the disparity estimation step S2, the component division step S3, and the agricultural feature detection step S4 can be performed individually or in series.
[0106] In a preferred embodiment of the present invention, imaging step S1 includes the cutting system 1 moving to a waypoint located in front of the target crop (e.g., a grapevine). The waypoint may be pre-configured or programmed in the onboard memory of the cutting system 1 or retrieved from remote storage, and may be determined based on the distance or time from the previous waypoint. Upon reaching the waypoint located in front of the target crop, the cutting system 1 stops and uses the camera 20 to capture multiple images of the target crop.
[0107] In a preferred embodiment of the present invention, imaging step S1 includes capturing multiple images of the target crop from multiple viewpoints (e.g., multiple positions of the camera 20) using the camera 20. For example, at each viewpoint, the camera 20 is controlled to capture a first image (e.g., a left image) using a first lens 20a and a second image (e.g., a right image) using a second lens 20b. By controlling the horizontal frame motor 28, the camera 20 can be moved to multiple positions in front of the target crop in the horizontal direction (x-axis direction in Figure 1), thereby reaching multiple viewpoints (positions of the camera 20). The number of viewpoints can be determined based on the field of view of the camera 20 and the number of viewpoints required to capture an image of the entire target crop. In a preferred embodiment, the multiple images captured by the camera 20 are stored in the local storage of the cutting system 1.
[0108] Figures 5A and 5B show examples of multiple images captured in imaging step S1. For example, a first image (image L0) is captured from viewpoint 0 using the first lens 20a, and a second image (image R0) is captured using the second lens 20b. After images L0 and R0 are captured from viewpoint 0, the camera 20 is moved to viewpoint 1. From viewpoint 1, a first image (image L1) is captured using the first lens 20a, and a second image (image R1) is captured using the second lens 20b. Similarly, the first image (image L2) and the second image (image R2) are captured from viewpoint 2, the first image (image L3) and the second image (image R3) from viewpoint 3, the first image (image L4) and the second image (image R4) from viewpoint 4, the first image (image L5) and the second image (image R5) from viewpoint 5, and the first image (image L6) and the second image (image R6) from viewpoint 6. In the examples shown in Figures 5A and 5B, the camera 20 is moved from left to right to reach viewpoints 0 to 6, but it is also possible to move the camera 20 from right to left to reach viewpoints 0 to 6.
[0109] In a preferred embodiment of the present invention, imaging step S1 may include downsampling the image captured by the camera 20, for example, by a factor of 2. Imaging step S1 may also include rectifying each pair of stereo images (e.g., image L0 and image R0) captured using the first lens 20a and the second lens 20b. This involves the process of reprojecting the image planes (left image plane and right image plane) onto a common plane parallel to the line between the camera lenses.
[0110] In the imaging step S1, once an image of the target crop is captured, the disparity estimation step S2 can be performed. In a preferred embodiment of the present invention, the disparity estimation step S2 is an example of a depth estimation step that generates a depth estimate of the crop. The disparity estimation step S2 includes estimating the depth of pixels contained in the image captured in the imaging step S1 using a disparity estimation model. In a preferred embodiment, the disparity estimation model generates a disparity map 46 corresponding to the images captured from viewpoints 0 to 6 in the imaging step S1. The disparity estimation step S2 can be performed using a number of approaches, including artificial intelligence (AI) deep learning approaches or classical computer vision approaches, as will be described in more detail below.
[0111] Figures 6A and 6B are flowcharts illustrating an example of a disparity estimation step S2 performed using an AI deep learning approach. Figures 6A and 6B show an AI disparity estimation model 44 used to generate a disparity map 46-0 corresponding to viewpoint 0 shown in Figures 5A and 5B. The input to the AI disparity estimation model 44 includes a parallelized stereo image pair containing a first image (image L0) and a second image (image R0), respectively, captured from viewpoint 0 in the imaging step S1. The output of the AI disparity estimation model 44 includes a disparity map 46-0 corresponding to viewpoint 0.
[0112] Using the AI disparity estimation model 44, disparity maps 46 corresponding to each of viewpoints 0 to 6 can be generated. For example, using a parallelized stereo image pair containing a first image (image L1) and a second image (image R1) captured at viewpoint 1 in imaging step S1, disparity map 46-1 corresponding to viewpoint 1 can be generated. Similarly, using a parallelized stereo image pair containing a first image (image L2) and a second image (image R2) captured at viewpoint 2 in imaging step S1, disparity map 46-2 corresponding to viewpoint 2 can be generated. Furthermore, using images captured at viewpoints 3 to 6, disparity maps 46-3 to 46-6 corresponding to viewpoints 3 to 6 can be generated.
[0113] In a preferred embodiment of the present invention, the AI disparity estimation model 44 matches each pixel in a first image (e.g., image L0) with the corresponding pixel in a second image (e.g., image R0) based on a correspondence. This aims to determine that pairs of pixels in the first and second images are projections of the same physical point in space. The AI disparity estimation model 44 then calculates, for example, the distance between each pair of matching pixels. In a preferred embodiment of the present invention, the AI disparity estimation model 44 generates a disparity map 46 based on the fact that depth is inversely proportional to disparity, for example, that the greater the disparity, the closer the object in the image is. In a preferred embodiment, the disparity map 46 includes pixel differences mapped to real-world depth (depth) based on the configuration of the camera 20, including, for example, intrinsic and extrinsic parameters of the camera 20.
[0114] Figures 6A and 6B show examples of the disparity estimation model 44 including AI (deep learning) frameworks, such as stereo matching AI frameworks like the RAFT-Stereo architecture. The RAFT-Stereo architecture is based on optical flow and uses a recurring neural network approach. The RAFT-Stereo architecture may include, for example, feature encoders 44a and 44b, a context encoder 44c, a correlation pyramid 44d, and a disparity estimator component / gated recurrent unit (GRU) 44e. Feature encoder 44a is applied to the first image (e.g., image L0), and feature encoder 44b is applied to the second image (e.g., image R0). Feature encoders 44a and 44b map the first and second images, respectively, to a dense feature map used to construct a correlation volume. Feature encoders 44a and 44b determine individual features of the first and second images, including density, texture, and pixel intensity.
[0115] In a preferred embodiment of the present invention, the context encoder 44c is applied only to the first image (e.g., image L0). The context features generated by the context encoder 44c are used to initialize the hidden state of the update operator and are also injected into the GRU 44e during each iteration of the update operator. The correlation pyramid 44d constructs a three-dimensional correlation volume using the feature maps generated by the feature encoders 44a and 44b. In a preferred embodiment, the disparity estimator component / gated regressive unit (GRU) 44e estimates the disparity of each pixel in the image. For example, the GRU 44e predicts a set of disparity fields from the initial starting point using the current estimate of disparity in each iteration. Between each iteration, the correlation volume is indexed using the current estimate of disparity to generate a set of correlation features. The correlations, disparities, and context features are then concatenated and injected into the GRU 44e. The GRU 44e updates the hidden state and uses the updated hidden state to predict the disparity update.
[0116] In a preferred embodiment of the present invention, AI deep learning frameworks / approaches other than the RAFT-Stereo architecture can be used to perform the disparity estimation step S2 to generate disparity maps 46 corresponding to viewpoints (e.g., viewpoints 0-6). For example, disparity maps 46 can be generated by performing the disparity estimation step S2 using AI deep learning approaches such as EdgeStereo, HSM-Net, LEAStereo, MC-CNN, LocalExp, CRLE, HITNet, NOSS-ROB, HD3, gwcnet, PSMNet, GANet, and DSMNet. The neural networks of the above-mentioned AI approaches, including the RAFT-Stereo architecture, can be trained using artificially generated / synthesized datasets.
[0117] Alternatively, the disparity estimation step S2 can be performed using a classical computer vision approach. A classical computer vision approach may include a stereo semi-global block matching (SGMB) function 48, which is an intensity-based approach that generates a dense disparity map 46 for 3D reconstruction. More specifically, the SGMB function 48 is a geometric approach algorithm that uses the eigenparameters and exextrinsic parameters of a camera (e.g., camera 20) used to capture images for generating the disparity map 46.
[0118] Figures 7A and 7B are flowcharts illustrating an example in which the disparity estimation step S2 is performed using a classical computer vision approach. In Figures 7A and 7B, for example, the SGMB function 48 is used to generate the disparity map 46-0 corresponding to viewpoint 0 shown in Figures 5A and 5B. The input to the SGMB function 48 includes a parallelized stereo image pair containing a first image (e.g., image L0) and a second image (e.g., image R0) captured from viewpoint 0 in the imaging step S1. The output of the SGMB function 48 includes the disparity map 46-0 corresponding to viewpoint 0. The SGMB function 48 can be used to generate disparity maps 46 corresponding to viewpoints 0 to 6.
[0119] In a preferred embodiment of the present invention, a first image (e.g., image L0) and a second image (e.g., image R0) are input to the camera parallelization and destraining module 47 before the SGMB function 48, as shown, for example, in Figures 7A and 7B. The camera parallelization and destraining module 47 performs a camera parallelization step and a destraining step before generating a disparity map 46 using the SGMB function 48. In the camera parallelization step, a transformation process is performed to project the first image and the second image onto a common plane. In the image destraining step, a mapping process is performed to map the coordinates of the output destrained image to the input camera image using distortion coefficients. After the camera parallelization and destraining module 47 has performed the camera parallelization step and the destraining step, the camera parallelization and destraining module 47 outputs the pair of parallelized and destrained images to the SGMB function 48.
[0120] In a preferred embodiment of the present invention, parameters of the SGMB function 48, such as window size, minimum parallax, or maximum parallax, can be fine-tuned according to factors including image size, lighting conditions, and camera mounting angle of camera 20 in order to optimize parallax. The uniqueness ratio parameter can also be fine-tuned to filter out noise. The parameters of the SGMB function 48 may also include post-processing parameters used to avoid speckle artifacts, namely speckle window size and speckle range, which can also be fine-tuned based on operating conditions including lighting conditions and camera mounting angle of camera 20.
[0121] In a preferred embodiment of the present invention, the point cloud generation step S5 includes generating a point cloud 49 corresponding to the viewpoint from which the image was captured in the imaging step S1. For example, the point cloud generation step S5 includes generating point clouds 49-0 to 49-6 corresponding to viewpoints 0 to 6 shown in Figures 5A and 5B. Each of the point clouds 49-0 to 49-6 is a set of data points in space, each point having a position represented as a set of real-world Cartesian coordinates (x, y, z points).
[0122] As shown in Figures 8A and 8C, for example, the disparity map 46 generated in the disparity estimation step S2 can be reprojected onto three-dimensional data points by the point cloud generation module 491 to form a point cloud 49 generated in the point cloud generation step S5. The point cloud generation step S5 includes the process of converting the two-dimensional disparity (depth) map 46 into a set of points (X, Y, Z coordinates) of the point cloud 49. In a preferred embodiment, the point cloud generation module 491 can convert the disparity (depth) map 46 into three-dimensional points of the point cloud 49 using the camera parameters of the camera 20. For example, the equation Z = fB / d can be used to convert the disparity (depth) map 46 into three-dimensional points, where f is the focal length (in pixels), B is the baseline (in meters), and d is the disparity (depth) map 46. After Z is determined, X and Y can be calculated using the camera projection equations (1) X = uZ / f and (2) Y = vZ / f. Here, u and v are pixel positions in a 2D image space, X is the x-axis in the real world, Y is the y-axis in the real world, and Z is the z-axis in the real world.
[0123] In a preferred embodiment, the disparity map 46 can be transformed into three-dimensional points used to generate a point cloud 49 by using built-in API functions and an inverse projection matrix obtained using the intrinsic and exextrinsic parameters of the camera 20. For example, the values of X (real-world x-position), Y (real-world y-position), and Z (real-world z-position) can be determined based on the following matrix.
number
[0124] In the matrix above, X is the x-coordinate in the real world (x-axis), Y is the y-coordinate in the real world (y-axis), and Z is the z-coordinate in the real world (z-axis). Variables x and y are values corresponding to the coordinates in the calibrated 2D left or right image (e.g., image L0 or image R0) captured in imaging step S1, and variable z = 1. Disparity (x,y) is the disparity value determined from the disparity map 46. For example, the disparity map 46 can be a 1-channel 8-bit unsigned, 16-bit signed, 32-bit signed, or 32-bit floating-point disparity image. Variable Q can be a 4x4 viewpoint transformation matrix, which is a disparity-to-depth mapping matrix, and can be obtained using a program such as stereoRectify based on the following variables: That is, the eigenmatrices of the first camera (e.g., the first camera 20b), the first camera distortion parameters, the eigenmatrices of the second camera (e.g., the second camera 20c), the second camera distortion parameters, the image size used for stereo calibration, the rotation matrix from the coordinate system of the first camera to the coordinate system of the second camera, and the translation vector from the coordinate system of the first camera to the second camera. For example, the variable Q can be represented by the following matrix: c x1 c is the distance (in pixels) from the left edge of the parallelized 2D image (e.g., parallelized image L0) to the point where the optical axis (e.g., the axis between the center of the first camera / lens 20b and the physical object) intersects the image plane of the parallelized 2D image, and c x2 c is the distance (in pixels) from the left edge of the parallelized 2D image (e.g., parallelized image R0) to the point where the optical axis (e.g., the axis between the center of the second camera / lens 20c and the physical object) intersects the image plane of the parallelized 2D image, and c y F is the distance (in pixels) from the left edge of the parallelized 2D image (e.g., parallelized image L0) to the point where the optical axis (e.g., the axis between the center of the first camera / lens 20b or the second camera / lens 20c and the physical object) intersects the image plane of the parallelized 2D image, where f is the focal length (in pixels), and T xis the distance between the first camera 20b and the second camera 20c.
number
[0125] Therefore, the variable W can be expressed by the following equation. The variable W can be used to convert the values of X (real-world x-position), Y (real-world y-position), and Z (real-world z-position) from pixels to units of distance (e.g., millimeters).
number
[0126] The point cloud generation step S5 may include generating point clouds 49-0 to 49-6 corresponding to viewpoints 0 to 6, respectively, as shown in Figures 5A and 5B. For example, point cloud 49-0 can be generated based on disparity map 46-0 corresponding to viewpoint 0, point cloud 49-1 can be generated based on disparity map 46-1 corresponding to viewpoint 1, point cloud 49-2 can be generated based on disparity map 46-2 corresponding to viewpoint 2, and similarly, point clouds 49-3 to 49-6 can be generated based on disparity maps 46-3 to 46-6 corresponding to viewpoints 3 to 6, respectively.
[0127] In a preferred embodiment of the present invention, the point cloud registration step S6 includes determining one or more spatial transformations (e.g., scaling, rotation, and / or translation) to combine / align the point clouds (e.g., point clouds 49-0 to 49-6) generated in the point cloud generation step S5. More specifically, the point cloud registration module 1161 is used to align the point cloud 49 generated in the point cloud generation step S5 to generate a mega point cloud 116, as shown in Figures 8B and 8D, for example.
[0128] In a preferred embodiment, the point cloud registration step S6 may be performed based on one or more assumptions, including that the horizontal frame 16 is precisely horizontal and oriented correctly, and that the physical distance between each viewpoint (e.g., viewpoints 0-6) is a predetermined value, such as approximately 15 cm or approximately 20 cm. Based on such one or more assumptions, it may be sufficient to perform translation along the X-axis (the axis of the horizontal frame 16) to obtain the mega point cloud 116. In a preferred embodiment, a 4x4 transformation matrix can be used to transform individual point clouds 49 from one viewpoint to individual point clouds 49 from another viewpoint, such that each element of the transformation matrix represents translation and rotation information. For example, in the point cloud registration step S6, a 4x4 transformation matrix can be used to sequentially transform each of the point clouds (e.g., point clouds 49-0 to 49-6) generated in the point cloud generation step S5 to generate the mega point cloud 116.
[0129] In a preferred embodiment of the present invention, a depth-based thresholding step can be performed after the point cloud generation step S5. The depth-based thresholding step includes removing points from the point cloud 49 that have a depth greater than a set depth-based threshold. The depth-based threshold is, for example, a user-configurable depth value (a value in the z-direction shown in Figure 1). For example, the depth-based threshold can be set based on the length of the robot arm 22 or the working space of the robot arm 22. For example, if the length of the robot arm is 1.5 meters, or if the working space of the robot arm extends 1.5 meters in the depth direction (z-direction in Figure 1), the depth-based threshold can be set to a value of 1.5 meters. In other words, the depth-based threshold can be set based on the working range in the depth direction of the cutting system 1.
[0130] Since a disparity map 46 is generated using the image captured in the imaging step S1, each point cloud 49 generated in the point cloud generation step S5 is generated using the disparity map 46, which includes both the foreground and background. For example, the target crop (e.g., grapevines) is included in the foreground of the disparity map 46, and the background of the disparity map 46 is not of interest. The depth-based thresholding step can remove points from the point cloud 49 that correspond to the background of the disparity map 46. Figures 9A and 9C show examples of the point cloud 49A before the depth-based thresholding step and the same point cloud 49B after the depth-based thresholding step. The depth-based thresholding step can reduce the number of points included in the point cloud 49. Therefore, the computation time and memory requirements, which are affected by the number of points included in the point cloud 49, can be reduced.
[0131] In a preferred embodiment of the present invention, a statistical outlier removal step can be performed after the point cloud generation step S5. The statistical outlier removal step can be performed after the depth-based thresholding step, or it can be performed before or simultaneously with the depth-based thresholding step. The statistical outlier removal step includes a process of removing trailing points and dense points generated in the disparity estimation step S2 from undesirable regions of the point cloud 49. For example, the statistical outlier removal step may include a process of removing trailing points and dense points from undesirable regions of the point cloud 49 that include portions of the point cloud 49 corresponding to the edges of an object, for example, the edges of a vine.
[0132] In a preferred embodiment, the statistical outlier removal step includes removing points that are farther from their neighbors compared to the mean of the point cloud 49. For example, the mean distance between a given point and its neighbors is calculated by calculating the distance between that point and a predetermined number of neighbors for each given point in the point cloud 49. Parameters of the statistical outlier removal step include a neighbor parameter and a ratio parameter. The neighbor parameter sets how many neighbors to consider when calculating the mean distance for a given point. The ratio parameter can set a threshold level based on the standard deviation of the mean distance of the entire point cloud 49, and determines the extent to which the statistical outlier removal step removes points from the point cloud 49. In a preferred embodiment, the lower the ratio parameter, the more aggressively the statistical outlier removal step filters / removes points from the point cloud 49. Figures 9B and 9D show examples of front views of point cloud 49C before and after the statistical outlier removal step, as well as side views of point cloud 49E before and after the statistical outlier removal step. A statistical outlier removal step can be used to reduce noise and the number of points from undesirable regions in the point cloud 49, thereby reducing computation time.
[0133] In the example described above, the depth-based thresholding step and statistical outlier removal step are performed on the individual point clouds (e.g., point clouds 49-0 to 49-6) generated in the point cloud generation step S5. However, in addition to, or alternatively to, performing the depth-based thresholding step and statistical outlier removal step on the individual point clouds, the depth-based thresholding step and statistical outlier removal step can also be performed on the megapoint cloud 116 generated by the point cloud registration step S6.
[0134] In a preferred embodiment of the present invention, component division step S3 includes identifying different segments (e.g., individual components) of the crop in question. For example, if the crop in question is a grapevine, component division step S3 may include identifying different segments of the grapevine, including the trunk, each individual cordon, each individual spur, and each individual cane.
[0135] In a preferred embodiment, the component segmentation step S3 is performed using an instance segmentation AI architecture 50. The instance segmentation AI architecture 50 may include a fully convolutional network (FCN) and may rely on an instance mask representation scheme that dynamically segments each instance in an image. Figures 10A and 10B show an example of the component segmentation step S3 using the instance segmentation AI architecture 50 to identify different segments of a target crop (e.g., a grapevine). The input to the instance segmentation AI architecture 50 includes an image of the target crop. For example, as shown in Figures 10A and 10B, the input to the instance segmentation AI architecture 50 includes an image (e.g., image L2) captured in the imaging step S1. The instance segmentation AI architecture 50 receives the image input and outputs a segmented image 51. The segmented image 51 includes one or more masks that identify different segments / individual components of the target crop contained in the image input to the instance segmentation AI architecture 50. For example, Figures 10A and 10B show that the instance segmentation AI architecture 50 outputs a segmented image 51 that includes masks that identify different segments of a grapevine, such as the main trunk, each individual main branch, each individual short shoot, and each individual branch. Figures 10A and 10B show that the segmented image 51 includes masks that include a main trunk mask 52, a main branch mask 54 that masks individual main branches, a short shoot mask 56 that masks individual short shoots, and a branch mask 58 that masks individual branches.
[0136] In a preferred embodiment of the present invention, the instance segmentation AI architecture 50 may include mask generation, which is separated into mask kernel prediction and mask feature learning, each generating a convolution kernel and a feature map to be convolved. The instance segmentation AI architecture 50 can significantly reduce or prevent inference overhead by a novel non-maximum suppression (NMS) technique for matrices. This technique takes an image (e.g., image L2 shown in Figures 10A and 10B) as input and directly outputs instance masks (e.g., trunk mask 52, main branch mask 54, short branch mask 56, and branch mask 58) and corresponding class probabilities in a fully convolutional, box-free, and grouping-free paradigm.
[0137] In a preferred embodiment of the present invention, the component division step S3 includes using the instance segmentation AI architecture 50 to identify different segments of the target crop (e.g., a grapevine) that are included in one or more of the multiple images captured in the imaging step S1. For example, Figures 11A to C show images L0 to L6 captured in the imaging step S1 and input to the instance segmentation AI architecture 50 in the component division step S3, and segmented images 51-0 to 51-6 output by the instance segmentation AI architecture 50 when images L0 to L6 are input to the instance segmentation AI architecture 50.
[0138] In a preferred embodiment, the instance segmentation AI architecture 50 employs adaptive learning and dynamic convolutional kernels for mask prediction, and a deformable convolutional network (DCN) is used. For example, the SoloV2 instance segmentation framework can be used to perform the component partitioning step S3. However, the instance segmentation AI architecture 50 may include instance segmentation frameworks other than the SoloV2 framework to perform the component partitioning step S3. For example, the instance segmentation AI architecture 50 may include the Mask-RCNN framework, which includes a deep neural network that can be used to perform the component partitioning step S3. Alternatively, the instance segmentation AI architecture 50 may also include instance segmentation frameworks such as SOLO, TrnsorMask, YOLACT, PolarMask, and BlendMask to perform the component partitioning step S3.
[0139] In a preferred embodiment of the present invention, the instance segmentation AI architecture 50 is trained using a segmentation dataset tailored to an instant segmentation task for a specific target crop. For example, if the target crop is a grapevine, the segmentation dataset is tailored to an instant segmentation task for a grapevine. The segmentation dataset includes multiple images selected based on factors such as whether the images were captured under appropriate operating conditions and whether the images contain an appropriate level of variety. Once multiple images to be included in the segmentation dataset have been selected, the multiple images are cleansed and annotated. For example, the multiple images in the segmentation dataset can be manually annotated using a computer-implemented labeling tool, as will be described in detail below.
[0140] Figures 12A and 12B show an example of image 60 of a segmentation dataset annotated using a computer-implemented labeling tool. The computer-implemented labeling tool includes a user interface that allows the formation of polygon masks around segments / individual components of the target crop. For example, if the target crop is a grapevine, the labeling tool's user interface can form polygon masks around different segments of the grapevine, including the main trunk, each individual main branch, each individual short shoot, and each individual branch. Polygon masks can also be formed around other objects included in image 60, such as poles or trellises used to support parts of the target crop. Each polygon mask formed around a segment of the target crop or other object is assigned a label that indicates the instance of the target crop or object segment around which the polygon mask is formed. For example, Figures 12A and 12B show a trunk polygon mask 62 formed around the main trunk, a main branch polygon mask 64 formed around individual main branches, a short branch polygon mask 66 formed around individual short branches, and a branch polygon mask 68 formed around individual branches.
[0141] In a preferred embodiment of the present invention, the labeling tool enables a specific type of annotation called group-identification-based labeling, which can be used to annotate discrete portions of the same segment / individual component using the same label. In other words, group-identification-based labeling can be used to annotate discrete portions of the same instance using the same label. Figures 12A and 12B show an example where the crop in question is a grapevine and group-identification-based labeling can be used to annotate discrete portions of the same branch using the same label. For example, in image 60 shown in Figures 12A and 12B, a first branch 70 overlaps / intersects with a second branch 72 in image 60, and image 60 contains a first discrete portion 72a and a second discrete portion 72b, which are spaced apart from each other in image 60 but are parts of the same second branch 72. Group identification-based labeling creates a first polygon mask 74 around the first discrete portion 72a, a second polygon mask 76 around the second discrete portion 72b, and the same label is assigned to the first polygon mask 74 and the second polygon mask 76, i.e., a common label, to indicate that the first discrete portion 72a and the second discrete portion 72b are parts of the same second branch 72.
[0142] In a preferred embodiment of the present invention, approximately 80% of the segmentation dataset is used as a training set for training and teaching the network of the instance segmentation AI architecture, and approximately 20% of the segmentation dataset is used as a validation / test set for the network included in the instance segmentation AI architecture 50. However, these percentages can be adjusted so that more or less of the segmentation dataset is used as the training set and the validation / test set.
[0143] In a preferred embodiment of the present invention, an augmentation process can be used to create additional images for the segmentation dataset from existing images included in the segmentation dataset. As shown in Figures 13A and 13B, the augmentation process may include editing and modifying the original captured images 78 to create new images that can be included in the segmentation dataset to create a good distribution of images for the network of the instance segmentation AI architecture 50 to learn / train. By including multiple different relative augments applied to the original captured images 78 in the augmentation process, the network of the instance segmentation AI architecture 50 can learn and generalize a wide range of lighting conditions, textures, and spatial augments.
[0144] Figures 13A and 13B show examples of augments that can be performed on the original captured image 78 during the augmentation process. For example, the augmentation process may include non-perpendicular augments such as color jitter augmentation 80, equalization augmentation 82, Gaussian blur augmentation 84, and sharpening augmentation 86, and / or spatial augments such as perpendicular augmentation 88 and affine augmentation 90. Non-perpendicular augments can be included in a custom data loader that operates on the fly and reduces memory constraints. Spatial augments can be manually added and saved before the network of the instance segmentation AI architecture 50 is trained with the updated segmentation dataset.
[0145] In a preferred embodiment of the present invention, agricultural feature detection step S4 includes detecting specific agricultural features of a target crop. For example, if the target crop is a grapevine, agricultural feature detection step S4 may include detecting one or more buds on the grapevine. Agricultural feature detection step S4 can be performed using an object detection model 92, such as an AI deep learning object detection model. Figures 14A and 14B show an example of agricultural feature detection step S4 in which an object detection model 92 is used to detect / identify specific agricultural features of a target crop (e.g., a grapevine). The input to the object detection model 92 includes an image of the target crop. For example, as shown in Figures 14A and 14B, the input to the object detection model 92 may include a first image (e.g., image L2) captured in imaging step S1. The object detection model 92 receives the input of the first image and outputs a feature image 94 which includes bounding boxes 96 surrounding the specific agricultural features shown in the first image. For example, Figures 14A and 14B show that the object detection model 92 outputs a feature image 94 that includes bounding boxes 96 surrounding the buds contained in the first image.
[0146] In a preferred embodiment of the present invention, the agricultural feature position 95 of an agricultural feature (e.g., a sprout) can be defined by the x and y coordinates of the center point of the bounding box 96 surrounding the agricultural feature. For example, the agricultural feature position 95 can be defined by the x and y coordinates of a pixel in the feature image 94 that includes the center point of the bounding box 96 surrounding the agricultural feature. Alternatively, the agricultural feature position 95 can be defined using the x and y coordinates of another point within or on the bounding box 96 (e.g., the lower left corner, lower right corner, upper left corner, or upper right corner of the bounding box 96). Thus, the agricultural feature position 95 can be determined for each agricultural feature (e.g., a sprout) detected in the agricultural feature detection step S4.
[0147] In a preferred embodiment of the present invention, the agricultural feature detection step S4 includes detecting / identifying agricultural features contained in each of a plurality of images captured in the imaging step S1 using an object detection model 92. For example, Figures 15A and 15B show images L0 to L6 captured in the imaging step S1 and input to the object detection model 92 in the agricultural feature detection step S4, and feature images 94-0 to 94-6 output by the object detection model 92 based on images L0 to L6. The object detection model 92 may include a model backbone, a model neck, and a model head. The model backbone is mainly used to extract important features from a given input image (e.g., image L2 in Figures 14A and 14B). In a preferred embodiment, a cross-stage partial (CSP) network can be used as the model backbone to extract useful features from the input image. The model neck is mainly used to generate a feature pyramid. The feature pyramid helps the object detection model 92 to generalize well when scaling objects of agricultural features (e.g., grape buds). The performance of the object detection model 92 is improved by identifying the same object (e.g., grape buds) at different scales and sizes. The model head is primarily used for the final detection of agricultural features. The model head applies anchor boxes to agricultural features contained in the image features and generates a final output vector with class probabilities, object scores, and bounding boxes 96 for the feature image 94.
[0148] In a preferred embodiment of the present invention, the agricultural feature detection step S4 is performed using an object detection model 92 such as YoloV5. However, the agricultural feature detection step S4 can also be performed using other models such as Yolov4. The trained object detection model 92 can be converted into a TensorRT optimization engine for faster inference.
[0149] The object detection model 92 can be trained using a detection dataset tailored to an object detection task related to a particular agricultural feature. For example, if the agricultural feature is grape buds, the detection dataset can be tailored to an object detection task related to grape buds. The detection dataset contains multiple images, selected based on factors such as whether the images were captured under appropriate operating conditions and whether the images contain an appropriate level of variety. Once the multiple images to be included in the detection dataset have been selected, the images are cleansed and annotated. For example, images in a detection dataset tailored to an object detection task related to grape buds can be manually annotated using a computer-implemented labeling tool.
[0150] Figures 16A and 16B show an example of image 98 included in a detection dataset annotated using a computer-implemented labeling tool. The computer-implemented labeling tool includes a user interface that allows polygon masks to be formed around specific agricultural features 100 of a target crop. For example, if the agricultural feature 100 is a grapevine bud, the labeling tool's user interface can form a polygon mask 102 around each grapevine bud. In a preferred embodiment, polygon masks 102 of different sizes can be formed around agricultural features 100 in image 98. For example, the size of the polygon mask 102 can be determined based on the size of the specific agricultural feature 100 around which the polygon mask 102 should be formed. For example, if the size of the specific agricultural feature 100 in image 98 is small due to a large distance between the specific agricultural feature 100 and the camera used to capture image 98, then the size of the polygon mask 102 formed around the specific agricultural feature 100 will be small. More specifically, in a preferred embodiment, the size of each polygon mask 102 formed around an agricultural feature 100 in the image 98 can be determined / adjusted based on a predetermined ratio of the pixel area of the agricultural feature 100 to the total pixel area of the polygon mask 102. For example, the size of the polygon mask 102 formed around an agricultural feature 100 in the image 98 can be determined / adjusted such that the ratio between the pixel area of the agricultural feature 100 and the total pixel area of the polygon mask 102 is a predetermined ratio of 50% (i.e., the area of the agricultural feature 100 is 50% of the total area of the polygon mask 102). Alternatively, each polygon mask 102 can be the same size regardless of the size of the particular agricultural feature 100 around which the polygon mask 102 is to be formed.
[0151] In a preferred embodiment of the present invention, approximately 80% of the detection dataset is used as a training set for training and teaching the network of the object detection model 92, and approximately 20% of the detection dataset is used as a validation / test set for the network of the object detection model 92. However, these percentages can be adjusted so that a large or small portion of the dataset is used as the training set and the validation / test set.
[0152] In a preferred embodiment of the present invention, an augmentation process can be used to create additional images for the detection dataset from existing images included in the detection dataset. As shown in Figures 13A and 13B, the augmentation process may include editing and modifying the original captured images 78 to create new images to be included in the detection dataset in order to create a good distribution of images in the detection dataset for the object detection model 92 network to train / learn. By including multiple different relative augments applied to the original captured images 78, the augmentation process can enable the object detection model 92 network to learn and generalize across a wide range of lighting conditions, textures, and spatial augments.
[0153] Figures 13A and 13B show examples of augments that can be performed on the original captured image 78 during the augmentation process. For example, Figures 13A and 13B show that the augmentation process can include non-perspective augments such as color jitter augmentation 80, equalization augmentation 82, Gaussian blur augmentation 84, and sharpening augmentation 86, and / or spatial augments such as perspective augmentation 88 and affine augmentation 90. In a preferred embodiment, non-perspective augments can be included in a custom data loader that operates on the fly and reduces memory constraints. Spatial augments can be manually added and saved before the object detection model 92 network is trained with the updated dataset.
[0154] In a preferred embodiment of the present invention, the cutting point generation step S7 includes generating two-dimensional cutting points 108 using a cutting point generation module 104. If the target crop is a grapevine, the cutting point generation module 104 generates two-dimensional cutting points 108 for a branch of the grapevine. Preferably, the cutting point generation module 104 generates two-dimensional cutting points 108 for each branch included in the grapevine. As an example, Figures 17A and 17B show an example of a two-dimensional cutting point 108 on a cutting point image 106. The location of the two-dimensional cutting point 108 can be represented by x and y coordinates. For example, the location of the two-dimensional cutting point 108 can be defined by the x and y coordinates of a pixel in the cutting point image 106 that contains the two-dimensional cutting point 108.
[0155] As shown in Figures 17A and 17B, the cut-point generation module 104 receives input including a mask from the segmented image 51 generated by the instance segmentation AI architecture 50 in the component division step S3, and the agricultural feature locations 95 of agricultural features (e.g., buds) detected in the agricultural feature detection step S4. For example, Figures 17A and 17B show that the input to the cut-point generation module 104 includes a mask from the segmented image 51 (e.g., segmented image 51-2) and the agricultural feature locations 95 of agricultural features (buds) contained in the corresponding feature image 94 (e.g., feature image 94-2). Both of these are generated using the image L2 captured from viewpoint 2 in the imaging step S1.
[0156] In a preferred embodiment of the present invention, the cut point generation module 104 generates a two-dimensional cut point 108 by performing an agricultural feature association step S18-1, an agricultural feature identification step S18-2, and a cut point generation step S18-3. Figure 18 shows a flowchart of the cut point generation step S7, which includes the agricultural feature association step S18-1, the agricultural feature identification step S18-2, and the cut point generation step S18-3.
[0157] In the agricultural feature association step S18-1, the agricultural features detected in the agricultural feature detection step S4 are associated with specific segments / individual components of the target crop identified in the component division step S3. For example, if the agricultural feature is a grapevine bud, each bud detected in the agricultural feature detection step S4 is associated with a specific branch of the grapevine identified in the component division step S3. In the example shown in Figures 17A and 17B, when the bud position 95 is compared with the branch mask 58 of the segmented image 51, if the agricultural feature position 95 (bud position 95) falls within / exists within the specific branch mask 58, then the bud associated with bud position 95 is considered to be located on / attached to the branch associated with the specific branch mask 58. For example, if the pixel of agricultural feature position 95 in feature image 94 corresponds to a pixel in the branch mask 58 of segmented image 51, then it can be determined that the agricultural feature position 95 (bud position 95) falls within / exists within the branch mask 58. In this way, the buds detected in the agricultural feature detection step S4 can be associated with the specific branch / branch mask 58 identified in the component division step S3.
[0158] When comparing the bud position 95 with the branch mask 58 of the segmented image 51, it is possible that the agricultural feature position 95 (bud position 95) does not fall within / exist within a particular branch mask 58. For example, because buds are attached to the outer surface of branches, the agricultural feature position 95 (bud position 95) may be adjacent to the branch mask 58 and therefore not fall within / exist within the branch mask 58. To address this problem, a search radius is assigned to the agricultural feature position 95. If it is determined that the agricultural feature position 95 is located within the area of the branch mask 58, the agricultural feature position 95 is maintained. However, if it is determined that the agricultural feature position 95 is not located within the area of the branch mask 58, the search radius is used to determine whether the agricultural feature position 95 is located within a predetermined distance from the branch mask 58. Using the search radius, if it is determined that a branch mask 58 is located within a predetermined distance from the agricultural feature position 95, the position of the agricultural feature position 95 is moved to a point within the area of the branch mask 58, for example, the nearest point within the area of the branch mask 58. On the other hand, if the search radius is used to determine that the branch mask 58 is not located within a predetermined distance from the agricultural feature location 95, then it is determined that the agricultural feature location 95 is either not located on the branch mask 58 or is not associated with the branch mask 58.
[0159] The agricultural feature identification step S18-2 includes assigning each agricultural feature an identifier relating to a specific segment / individual component of the target crop to which the agricultural feature was associated in the agricultural feature association step S18-1. For example, if the agricultural feature is a grape bud, each bud is assigned an identifier relating to a specific branch / branch mask 58 to which the bud was associated in the agricultural feature association step S18-1.
[0160] The agricultural feature identification step S18-2 may include identifying the starting point 57 of the branch mask 58 located at the connection point between the short-prune mask 56 and the branch mask 58. For example, the connection point between the short-prune mask 56 and the branch mask 58 can be identified by pixels that fall into both the short-prune mask 56 and the branch mask 58. This connection point indicates the overlap between the short-prune mask 56 and the branch mask 58. Once the starting point 57 of the branch mask 58 is identified, each bud detected in the agricultural feature detection step S4 can be assigned an identifier relating to the specific branch / branch mask 58 associated with the bud in the agricultural feature association step S18-1, based on the distance from the starting point 57 of the branch mask 58 to the respective bud. In the examples shown in Figures 17A and 17B, agricultural feature position 95-1 is closest to the starting point 57 of the branch mask 58 (the connection point between the short-prune mask 56 and the branch mask 58), agricultural feature position 95-2 is second closest to the starting point 57 of the branch mask 58, and agricultural feature position 95-3 is third closest to the starting point 57 of the branch mask 58. Agricultural feature positions 95-1, 95-2, and 95-3 are illustrated on the cutting point image 106 in Figures 17A and 17B.
[0161] Based on the respective distances from the starting point 57 of the branch mask 58 to the agricultural feature locations 95-1, 95-2, and 95-3, each agricultural feature can be assigned an identifier relating to a specific segment / individual component of the target crop to which the agricultural feature is associated. For example, a bud with agricultural feature location 95-1 can be assigned as the first bud of the branch associated with the branch mask 58, a bud with agricultural feature location 95-2 can be assigned as the second bud of the branch associated with the branch mask 58, and a bud with agricultural feature location 95-3 can be assigned as the third bud of the branch associated with the branch mask 58.
[0162] Step S18-3, the cut point generation step, includes executing a cut point generation algorithm to generate a two-dimensional cut point 108. The cut point generation algorithm uses one or more rules to generate a two-dimensional cut point 108 based on one or more identifiers assigned to agricultural features in the agricultural feature identification step S18-2. For example, if the agricultural feature is a grapevine bud and the specific segment / individual component of the target crop is a specific branch / branch mask 58 of the grapevine, the rule may include generating a two-dimensional cut point 108 between a first bud having agricultural feature position 95-1 and a second bud having agricultural feature position 95-2 when the branch contains multiple buds (when multiple agricultural feature positions 95 are located within the branch mask 58). More specifically, the rule may include generating a cut point 108 at the midpoint (approximately 50th percentile) between agricultural feature position 95-1 and agricultural feature position 95-2. Alternatively, the rule may include generating a breakpoint 108 at another point between agricultural feature location 95-1 and agricultural feature location 95-2 (e.g., approximately the 30th percentile or approximately the 70th percentile). Alternatively, the rule may include generating a breakpoint 108 at a predetermined distance from agricultural feature location 95-1. One or more of the above rules may also include not generating a breakpoint if the branch contains a single bud or does not contain a bud, for example, if a single agricultural feature location 95 is located within the branch mask 58 or if agricultural feature location 95 is not located there.
[0163] In preferred embodiments of the present invention, one or more of the above rules may differ from or be modified from the rules described above. For example, if a branch contains more than two buds (if more than two agricultural feature positions 95 are located within the branch mask 58), one or more of the above rules may include generating a cutting point 108 between a second bud having agricultural feature position 95-2 which is the second closest to the starting point 57 of the branch mask 58, and a third bud having agricultural feature position 95-3 which is the third closest to the starting point 57 of the branch mask 58.
[0164] In a preferred embodiment of the present invention, the two-dimensional cut point 108 generated in the cut point generation step S18-3 may not be located on a branch or within the branch mask 58. For example, if the cut point 108 is generated at the midpoint (approximately 50th percentile) between agricultural feature position 95-1 and agricultural feature position 95-2, and the branch between agricultural feature position 95-1 and agricultural feature position 95-2 is bent or curved, the generated cut point 108 may not be located on a branch or within the branch mask 58. To address this problem, a search radius is assigned to the cut point 108. If it is determined that the cut point 108 generated in the cut point generation step S18-3 is located within the area of the branch mask 58, the position of the cut point 108 is maintained. On the other hand, if it is determined that the cut point 108 generated in the cut point generation step S18-3 is not located within the area of the branch mask 58, the search radius is used to determine whether the cut point 108 generated in the cut point generation step S18-3 is located within a predetermined distance from the branch mask 58. If the search radius is used to determine that the cutting point 108 is located within a predetermined distance from the branch mask 58, the position of the cutting point 108 is moved to a point within the area of the branch mask 58, for example, to the point within the area of the branch mask 58 closest to the cutting point 108 generated in the cutting point generation step S18-3. On the other hand, if the search radius is used to determine that the cutting point 108 is not located within a predetermined distance from the branch mask 58, the cutting point 108 is deleted.
[0165] In a preferred embodiment of the present invention, the cutting point angle is determined for a two-dimensional cutting point 108. An example of the process used to determine the cutting point angle is shown in the flowchart of Figure 19. In step S19-1, it is identified which agricultural feature positions 95 the cutting point 108 was generated between. For example, as shown in Figures 20A and 20B, it is identified that the cutting point 108 was generated between agricultural feature positions 95-1 and 95-2. In step S19-2, the angle of the portion of a particular segment / individual component of a crop where the cutting point 108 is located is determined using the agricultural feature positions 95 identified in step S19-1. For example, the angle of the portion of the branch where the cutting point 108 is located is determined by forming a line 126 connecting agricultural feature positions 95-1 and 95-2. Once the angle of the portion of the crop that is the specific segment / individual component where the cutting point 108 is located is determined in step S19-2, the cutting point angle of the cutting point 108 can be determined in step S19-3 by forming a line 127 perpendicular to line 126 at a certain angle to line 126, for example. Line 127 may also be formed at another angle to line 126, for example, 30 degrees or 45 degrees to line 126. The angle of line 127 defines the cutting point angle of the cutting point 108, which is the angle with respect to the specific segment / individual component of the crop where the cutting point 108 is located.
[0166] In a preferred embodiment of the present invention, the cutting point generation step S7 includes generating a set of two-dimensional cutting points 108 using a cutting point generation module 104 with a plurality of images captured from a plurality of viewpoints (e.g., viewpoints 0 to 6) in the imaging step S1. For example, the cutting point generation step S7 may include generating a set of two-dimensional cutting points 108 for each viewpoint from which an image was captured in the imaging step S1 (e.g., one cutting point 108 for each branch) using the cutting point generation module 104. The cutting point generation module 104 generates a first set of cutting points 108 based on the mask of segmented image 51-0 (see Figures 11A-C) and agricultural feature positions 95 from feature image 94-0 (see Figures 15A and 15B), generates a second set of cutting points 108 based on the mask of segmented image 51-1 (see Figures 11A-C) and agricultural feature positions 95 from feature image 94-1 (see Figures 15A and 15B), generates a third set of cutting points 108 based on the mask of segmented image 51-2 (see Figures 11A-C) and agricultural feature positions 95 from feature image 94-2 (see Figures 15A and 15B), and generates a first set of cutting points 108 based on the mask of segmented image 51-3 (see Figures 11A-C) and feature image 94-3 (see Figure 15A and 15B). A fourth set of cutting points 108 can be generated based on agricultural feature locations 95 from (see A and 15B), a fifth set of cutting points 108 can be generated based on the mask of segmented image 51-4 (see Figures 11A-C) and agricultural feature locations 95 from feature image 94-4 (see Figures 15A and 15B), a sixth set of cutting points 108 can be generated based on the mask of segmented image 51-5 (see Figures 11A-C) and agricultural feature locations 95 from feature image 94-5 (see Figures 15A and 15B), and a seventh set of cutting points 108 can be generated based on the mask of segmented image 51-6 (see Figures 11A-C) and agricultural feature locations 95 from feature image 94-6 (see Figures 15A and 15B).
[0167] In a preferred embodiment of the present invention, the cutting point projection step S8 includes generating a three-dimensional cutting point 114 using a cutting point projection module 110. As shown in Figures 21A and 21B, the cutting point projection module 110 receives an input including a set of two-dimensional cutting points 108 generated in the cutting point generation step S7 and a corresponding disparity map 46 generated in the disparity estimation step S2. In Figures 21A and 21B, the set of two-dimensional cutting points 108 is shown on the cutting point image 106. For example, the input to the cutting point projection module 110 may include a third cutting point image 106 containing a set of two-dimensional cutting points 108 generated in the cutting point generation step S7 based on the mask of the segmented image 51-2 and the agricultural feature positions 95 of the feature image 94-2, and a corresponding disparity map 46 generated in the disparity estimation step S2 based on images L2 and R2. In other words, both the set of 2D cutting points 108 and the corresponding disparity map 46 are generated based on images taken from the same viewpoint, for example, viewpoint 2 shown in Figures 5A and 5B.
[0168] The cutting point projection module 110 outputs a three-dimensional cutting point 114, as shown, for example, in Figures 21A and 21B. For illustrative purposes, Figures 21A and 21B show the three-dimensional cutting point 114 on a three-dimensional cutting point cluster 112. The cutting point projection module 110 generates a three-dimensional cutting point 114 corresponding to the two-dimensional cutting point 108 by slicing the position of the two-dimensional cutting point 108 from the disparity map 46 and reprojecting the sliced disparity with a known camera configuration of a camera (e.g., camera 20). For example, pixels in the cutting point image 106 containing the two-dimensional cutting point 108 can be identified, and the corresponding pixels in the disparity map 46 can be identified. The depth value of the corresponding pixel from the disparity map 46 can be used as the depth value of the two-dimensional cutting point 108. In this way, the two-dimensional cutting point 108 can be projected onto a three-dimensional cutting point 114 including X, Y, and Z coordinates.
[0169] In an alternative preferred embodiment of the present invention, the cutting point projection module 110 receives input including a set of two-dimensional cutting points 108 generated in the cutting point generation step S7 and a crop depth estimate obtained from a LiDAR sensor (e.g., LiDAR system 38), a time-of-flight (TOF) sensor, or another depth sensor capable of generating crop depth estimates. For example, the crop depth estimate can be obtained from point cloud data generated by a LiDAR sensor calibrated to have a coordinate system aligned with the coordinate system of camera 20, and the set of two-dimensional cutting points 108 can be generated based on images captured using camera 20, including an RGB camera. The cutting point projection module 110 generates three-dimensional cutting points 114 by determining the depth values of the two-dimensional cutting points 108 based on the crop depth estimate, thereby generating three-dimensional cutting points 114 corresponding to the two-dimensional cutting points 108. For example, the coordinates (pixels) of the cutting point image 106, which includes the two-dimensional cutting point 108, can be identified, and the corresponding coordinates in the depth estimation of crops, such as the corresponding coordinates in the point cloud data generated by the LiDAR sensor, can be identified. The depth values of the corresponding coordinates from the depth estimation of crops can be used as the depth values of the two-dimensional cutting point 108. In this way, the two-dimensional cutting point 108 can be projected onto a three-dimensional cutting point 114 that includes X, Y, and Z coordinates.
[0170] In a preferred embodiment of the present invention, the cutting point projection step S8 includes generating a set of three-dimensional cutting points 114 for each of the multiple viewpoints (e.g., viewpoints 0 to 6) whose images were captured by the camera 20 in the imaging step S1. For example, using the cutting point projection module 110, a first set of three-dimensional cutting points 114 can be generated using a first set of two-dimensional cutting points 108 and a parallax map 46-0; a second set of three-dimensional cutting points 114 can be generated using a second set of two-dimensional cutting points 108 and a parallax map 46-1; a third set of three-dimensional cutting points 114 can be generated using a third set of two-dimensional cutting points 108 and a parallax map 46-2; a fourth set of three-dimensional cutting points 114 can be generated using a fourth set of two-dimensional cutting points 108 and a parallax map 46-3; a fifth set of three-dimensional cutting points 114 can be generated using a fifth set of two-dimensional cutting points 108 and a parallax map 46-4; a sixth set of three-dimensional cutting points 114 can be generated using a sixth set of two-dimensional cutting points 108 and a parallax map 46-5; and a seventh set of three-dimensional cutting points 114 can be generated using a seventh set of two-dimensional cutting points 108 and a parallax map 46-6.
[0171] In a preferred embodiment of the present invention, when a set of three-dimensional cutting points 114 (for example, the first to seventh sets of three-dimensional cutting points 114) is generated in the cutting point projection step S8, the set of three-dimensional cutting points 114 is joined / aligned with each other in the cutting point registration step S9 to form a set of mega cutting points 115. For illustrative purposes, Figures 22A and 22B show a set of three-dimensional cutting points 114 on a three-dimensional cutting point group 112 corresponding to multiple viewpoints, and a set of mega cutting points 115 on a mega cutting point group 117. The mega cutting point group 117 can be formed by merging a set of mega cutting points 115 with a mega point group 116 generated in the point group registration step S6.
[0172] In a preferred embodiment, a set of three-dimensional cutting points 114 are joined / aligned to one another by a cutting point registration module 1151 that determines one or more spatial transformations (e.g., scaling, rotation, and translation) to align the set of three-dimensional cutting points 114. For example, similar to the point cloud registration step S6, the cutting point registration step S9 may be performed based on one or more assumptions, including that the horizontal frame 16 is exactly horizontal and oriented correctly, and that the physical distance between each viewpoint (e.g., viewpoints 0-6) is a predetermined value. Based on one or more such assumptions, it may be sufficient to perform translation along the X-axis (the axis of the horizontal frame 16) to obtain a set of mega-cutting points 115. In a preferred embodiment, a 4x4 transformation matrix can be used to transform individual sets of three-dimensional cutting points 114 from one viewpoint to another, such that each element of the transformation matrix represents translation and rotation information. For example, each of the sets of three-dimensional cutting points 114 can be sequentially transformed using a 4x4 transformation matrix to complete the cutting point registration step S9 and generate a set of mega-cutting points 115.
[0173] The set of 3D cutting points 114 is generated based on images taken from different viewpoints (e.g., viewpoints 0-6 in Figures 5A and 5B). Therefore, even after a spatial transformation intended to align the set of 3D cutting points 114 is performed in the cutting point registration step S9, the set of 3D cutting points 114 may not be perfectly aligned with each other. Thus, to identify 3D cutting points 114 that represent the same cutting point but belong to different sets of cutting points 114, i.e., 3D cutting points 114 that represent the same cutting point but are still slightly misaligned with each other even after the transformation of the set of 3D cutting points 114, a search radius (e.g., 4 cm) is assigned to each of the 3D cutting points 114. When sets of 3D cutting points 114 are combined / aligned to generate a set of mega-cutting points 115, the search radius of the 3D cutting point 114 is used to determine whether one or more other 3D cutting points 114 from another set of 3D cutting points 114 are located within the search radius of that 3D cutting point 114. If one or more other 3D cut points 114 are located within the search radius of that 3D cut point 114, that 3D cut point 114 and the one or more other 3D cut points 114 are merged into a megacut point 115 that is included in the set of megacut points 115.
[0174] In a preferred embodiment of the present invention, two or more 3D cuts 114 from different sets of 3D cuts 114 must be merged to generate a single megacut 115. For example, if, when a set of 3D cuts 114 is merged / aligned, there are no other 3D cuts 114 from another set of 3D cuts 114 that are located within the search radius of a given 3D cut 114, then no megacut 115 is generated. As another example, three or more 3D cuts 114 from different sets of 3D cuts 114 must be merged to generate a single megacut 115. Alternatively, a megacut 115 may be generated based on a single 3D cut 114 even if, when a set of 3D cuts 114 is merged / aligned, there are no other 3D cuts 114 from another set of 3D cuts 114 that are located within the search radius of that 3D cut 114.
[0175] The mega-cutting points 115 are generated by combining / aligning sets of 3D cutting points 114 that are generated based on images taken from different viewpoints (e.g., viewpoints 0-6 in Figures 5A and 5B). However, in some cases, due to the different viewpoints from which the images were taken, the first 3D cutting point included in the first set of 3D cutting points 114 and the second 3D cutting point included in the second set of 3D cutting points 114 may be located in significantly different positions, even though the first and second 3D cutting points are located on the same specific segment / individual component of the crop (e.g., the same branch). For example, buds detected based on images taken from one viewpoint (e.g., viewpoint 6) on a particular branch may be different from buds detected based on images taken from another viewpoint (e.g., viewpoint 2). As a result, the agricultural features detected from one viewpoint (e.g., viewpoint 6) may differ from those detected from another viewpoint (e.g., viewpoint 2). Consequently, the position of the first 3D crosspoint generated based on the image captured from one viewpoint may differ significantly from the position of the second 3D crosspoint generated based on the image captured from the other viewpoint. For example, an agricultural feature (e.g., a sprout) detected based on an image captured from one viewpoint (e.g., viewpoint 6) may be hidden or otherwise invisible in an image captured from another viewpoint (e.g., viewpoint 2). In such cases, the same agricultural feature will not be detected in the agricultural feature detection step S4 for the image captured from the other viewpoint (viewpoint 2). Furthermore, an agricultural feature may be incorrectly detected in the agricultural feature detection step S4 for an image captured from one viewpoint (e.g., viewpoint 6), and that agricultural feature may not be detected in the agricultural feature detection step S4 for an image captured from the other viewpoint (viewpoint 2). In each of these cases, the first 3D cutting point included in the first set of 3D cutting points 114 and the second 3D cutting point included in the second set of 3D cutting points 114 will be in significantly different locations, even though the first and second 3D cutting points are located on the same specific segment / individual component (e.g., the same branch) of the crop.Therefore, when the set of 3D cutting points 114 are joined / aligned with each other in the cutting point registration step S9, the first 3D cutting point and the second 3D cutting point are not located within each other's search radius and are not merged with each other in the cutting point registration step S9. As a result, the first mega-cutting point 115-1 is generated based on the first 3D cutting point, and the second mega-cutting point 115-2 for the same branch is generated based on the second 3D cutting point. For example, Figures 23A and 23B show the first mega-cutting point 115-1 generated based on the first 3D cutting point generated based on an image taken from one viewpoint, and the second mega-cutting point 115-2 generated based on the second 3D cutting point generated based on an image taken from another viewpoint.
[0176] In a preferred embodiment of the present invention, it is desirable to have only one megacut point 115 for each specific segment / individual component of a crop. That is, it is desirable to have only one megacut point 115 for each branch of a grapevine. Accordingly, a preferred embodiment of the present invention includes a trace module 120 that can be used to identify and remove one or more megacut points 115 when multiple megacut points 115 are assigned to a specific segment / individual component of the crop in question. For example, the trace module 120 can be used to identify and remove one or more megacut points 115 when multiple megacut points 115 are assigned to a branch of a grapevine.
[0177] In a preferred embodiment of the present invention, the mega-cut points 115 generated in the cut point registration step S9 are merged with the mega-point cloud 116 generated in the point cloud registration step S6 to form the mega-cut point cloud 117 in the mega-registration step S10. The mega-cut point cloud 117 is used by the trace module 120. As shown in Figures 23A and 23B, for example, the trace module 120 fits a cylinder 122 around a specific segment of the target crop and traces the specific segment, starting from a first mega-cut point 115 (first mega-cut point 115-1) that is closest to the connection point between the short shoot and the branch. The trace module 120 can determine that the mega-cut point 115-1 is closest to the connection point between the short shoot and the branch by using the short shoot mask 56 and branch mask 58 included in one or more of the segmented images 51 generated in the component division step S3. Branch masks 58 included in one or more of the segmented images 51 are projected onto three-dimensional coordinates using one or more corresponding disparity maps 46, allowing the three-dimensional space of the branches that the cylinder 122 traces around to be determined. The trace module 120 uses the cylinder 122 to trace a specific segment of the target crop from a first megacut point 115-1 to the free end 124 of that particular segment. If multiple megacut points 115 exist in the area traced by the cylinder 122, one or more megacut points following the first megacut point can be identified as false megacut points and removed from the set of megacut points 115. In the example shown in Figures 23A and 23B, the second megacut point 115-2 is identified as a false megacut point and removed from the set of megacut points 115. As a result, the first megacut point 115-1 remains as the only remaining megacut point for the particular branch traced by the trace module 120. In a preferred embodiment, each branch represented in the megacut point group 117 can be traced simultaneously by different cylinders 122 of the trace module 120. Alternatively, the trace module 120 can be used to trace each branch in series (one after the other) until each branch is traced by the trace module 120.
[0178] In a preferred embodiment of the present invention, a mega-cutting angle can be determined for each of one or more mega-cutting points 115. The mega-cutting angle is the angle at which the blade portion 24b of the cutting tool 24 is directed when a cutting operation is performed at the mega-cutting point 115. In a preferred embodiment, the mega-cutting angle can be determined based on the cutting angles of the cutting points 108 corresponding to the mega-cutting point 115. For example, if the mega-cutting point 115 corresponds to cutting points 108 generated from each of a plurality of viewpoints, the cutting angles of these cutting points 108 are averaged to determine the mega-cutting angle. Alternatively, the mega-cutting angle can be determined by averaging the angles of the portion of the branch where the cutting point 108 is located.
[0179] In a preferred embodiment of the present invention, the operation step S11 shown in Figure 4 can be performed based on a set of mega-cutting points 115. Operation step S11 includes controlling one or more of the horizontal frame motor 28, vertical frame motor 30, robot arm 22, or robot arm mount assembly 23 to position the blade portion 24b of the cutting tool 24 and perform a cutting operation at the mega-cutting point 115. In a preferred embodiment, one or more of the horizontal frame motor 28, vertical frame motor 30, robot arm 22, or robot arm mount assembly 23 are controlled via a robot operating system (ROS) and a free-space motion planning framework such as "MoveIt!". This framework is used to plan the movement of the robot arm 22 and cutting tool 24 between two points in space without collision. For example, the free-space motion planning framework can use information from the mega-cutting points 115 and a set of mega-cutting points 117 that provide the real-world coordinates of the target crop to plan the movement of the robot arm 22 and cutting tool 24 between two points in space without colliding with any part of the target crop. More specifically, operation step S11 may include positioning the blade portion 24b of the cutting tool 24 based on a cutting point mark, where the mega cutting point 115 coincides with a position on the blade portion 24b of the cutting tool 24.
[0180] In the preferred embodiment of the present invention described above, the agricultural feature detection step S4, in which specific agricultural features of the target crop are detected, is different from the component segmentation step S3. However, in another preferred embodiment of the present invention, the component segmentation step S3 may include identifying specific agricultural features of the target crop. For example, if the target crop is a grapevine, the component segmentation step S3 may include identifying buds of the grapevine when identifying different segments of the grapevine. For example, the component segmentation step S3 can be performed using an instance segmentation AI architecture 50 that identifies different segments of the grapevine, including the main trunk, each individual main branch, each individual short shoot, each individual branch, and each individual bud. In this case, the agricultural feature location 95 can be determined based on the results of the component segmentation step S3, such as an agricultural feature mask (bud mask) output by the instance segmentation AI architecture 50. Therefore, it is not necessary to provide a separate agricultural feature detection step S4.
[0181] In a preferred embodiment of the present invention, the agricultural feature locations 95 of the agricultural features detected in the agricultural feature detection step S4 are defined in two dimensions. For example, the agricultural feature locations 95 are defined by the x and y coordinates of the points of the bounding box 96 surrounding the agricultural feature. The agricultural feature projection step S12 includes generating three-dimensional agricultural features 130 using the agricultural feature projection module 1301. As shown in Figures 24A and 24B, the agricultural feature projection module 1301 receives input including a set of two-dimensional agricultural feature locations 95 generated in the agricultural feature detection step S4 and a corresponding disparity map 46 generated in the disparity estimation step S2. In Figures 24A and 24B, the set of two-dimensional agricultural feature locations 95 is shown on the feature image 94. For example, the input to the agricultural feature projection module 1301 may include agricultural feature locations 95 detected in the agricultural feature detection step S4 based on image L0 and a corresponding disparity map 46 generated in the disparity estimation step S2 based on images L0 and R0. In other words, both the agricultural feature locations 95 and the corresponding disparity maps 46 are generated based on images taken from the same viewpoint, for example, viewpoint 0 shown in Figures 5A and 5B.
[0182] In a preferred embodiment, the agricultural feature projection module 1301 outputs a three-dimensional agricultural feature 130. For example, in Figures 24A and 24B, the three-dimensional agricultural feature 130 is shown on a group of three-dimensional agricultural features 132. The agricultural feature projection module 1301 generates the three-dimensional agricultural feature 130 by slicing the agricultural feature locations (agricultural feature locations 95) from the disparity map 46, reprojecting the sliced disparity using a known camera configuration of a camera (e.g., camera 20), and generating a three-dimensional agricultural feature 130 corresponding to an agricultural feature having a two-dimensional agricultural feature location 95. For example, pixels in a feature image 94 containing the two-dimensional agricultural feature location 95 can be identified, and the corresponding pixels in the disparity map 46 can be identified. The depth value of the corresponding pixel from the disparity map 46 can be used as the depth value of the two-dimensional agricultural feature having the agricultural feature location 95. In this way, a two-dimensional agricultural feature can be projected onto a three-dimensional agricultural feature 130 having X, Y, and Z coordinates.
[0183] In a preferred embodiment of the present invention, the agricultural feature projection step S12 includes generating a set of three-dimensional agricultural features 130 for each of a plurality of viewpoints (e.g., viewpoints 0 to 6) whose images were captured by the camera 20 in the imaging step S1. For example, using the agricultural feature projection module 1301, a first set of three-dimensional agricultural features 130 is generated using agricultural feature locations 95 from feature image 94-0 and a disparity map 46-0; a second set of three-dimensional agricultural features 130 is generated using agricultural feature locations 95 from feature image 94-1 and a disparity map 46-1; a third set of three-dimensional agricultural features 130 is generated using agricultural feature locations 95 from feature image 94-2 and a disparity map 46-2; and agricultural feature locations 9 from feature image 94-3 A fourth set of 3D agricultural features 130 can be generated using 5 and the disparity map 46-3; a fifth set of 3D agricultural features 130 can be generated using agricultural feature locations 95 from feature images 94-4 and the disparity map 46-4; a sixth set of 3D agricultural features 130 can be generated using agricultural feature locations 95 from feature images 94-5 and the disparity map 46-5; and a seventh set of 3D agricultural features 130 can be generated using agricultural feature locations 95 from feature images 94-6 and the disparity map 46-6.
[0184] When a set of 3D agricultural features 130 (for example, the first to seventh sets of 3D agricultural features 130) is generated in the agricultural feature projection step S12, the sets of 3D agricultural features 130 are combined / aligned with each other in the agricultural feature registration step S13 to form a set of mega agricultural features 134. As an example, Figures 25A and 25B show a set of 3D agricultural features 130 on a 3D agricultural feature group 132 corresponding to multiple viewpoints, and a set of mega agricultural features 134 on a mega agricultural feature group 136. The mega agricultural feature group 136 can be formed by merging the set of mega agricultural features 134 with the mega point cloud 116 generated in the point cloud registration step S6.
[0185] In a preferred embodiment, the agricultural feature registration module 1341 is used to combine / align a set of three-dimensional agricultural features 130 by determining one or more spatial transformations (e.g., scaling, rotation, and translation) that align the set of three-dimensional agricultural features 130. For example, similar to the point cloud registration step S6 and the cut point registration step S9, the agricultural feature registration step S13 may be performed based on one or more assumptions, including that the horizontal frame 16 is exactly horizontal and oriented correctly, and that the physical distance between each viewpoint (e.g., viewpoints 0-6) is a predetermined value. Based on one or more such assumptions, it may be sufficient to perform translation along the X-axis (the axis of the horizontal frame 16) to obtain a set of mega-agricultural features 134. In a preferred embodiment, a 4x4 transformation matrix can be used to transform individual sets of three-dimensional agricultural features 130 from one viewpoint to another, such that each element of the transformation matrix represents translation and rotation information. For example, to complete the agricultural feature registration step S13 and generate a set of mega-agricultural features 134, each of the three-dimensional agricultural features 130 can be sequentially transformed using a 4x4 transformation matrix.
[0186] The set of 3D agricultural features 130 is generated based on images taken from different viewpoints (e.g., viewpoints 0-6 in Figures 5A and 5B). Therefore, even after one or more spatial transformations are performed in the agricultural feature registration step S13 with the intention of aligning the set of 3D agricultural features 130, the set of 3D agricultural features 130 may not be perfectly aligned with one another. Accordingly, in order to identify 3D agricultural features 130 that belong to different sets of agricultural features but represent the same agricultural feature, i.e., 3D agricultural features 130 that represent the same agricultural feature but are still slightly misaligned with one another even after the set of 3D agricultural features 130 has been transformed, each of the 3D agricultural features 130 is assigned a search radius (e.g., approximately 4 cm). When a set of 3D agricultural features 130 is combined / aligned to generate a set of mega agricultural features 134, the search radius of the 3D agricultural features 130 is used to determine whether one or more other 3D agricultural features 130 from another set of 3D agricultural features 130 are located within the search radius of that 3D agricultural feature 130. If one or more other 3D agricultural features 130 are located within the search radius of that 3D agricultural feature 130, then that 3D agricultural feature 130 and the one or more other 3D agricultural features 130 are merged into a mega agricultural feature 134 included in the set of mega agricultural features 134.
[0187] In a preferred embodiment of the present invention, two or more 3D agricultural features 130 from different sets of 3D agricultural features must be merged to generate one mega-agricultural feature 134. For example, if, when a set of agricultural features 130 is merged / aligned, there are no other 3D agricultural features 130 from another set of 3D agricultural features 130 that are located within the search radius of a certain 3D agricultural feature 130, then the mega-agricultural feature 134 is not generated. As another example, three or more 3D agricultural features 130 from different sets of 3D agricultural features 130 are merged to generate one mega-agricultural feature 134. Alternatively, the mega-agricultural feature 134 may be generated based on a single 3D agricultural feature 130 even if, when a set of 3D agricultural features 130 is merged / aligned, there are no other 3D agricultural features 130 from another set of 3D agricultural features 130 that are located within the search radius of that 3D agricultural feature 130.
[0188] In a preferred embodiment of the present invention, the image captured in imaging step S1, the disparity map 46, the segmented image 51, the feature image 94, the point cloud 49, the mega point cloud 116, the cutting point image 106, the 3D cutting point cloud 112, the mega cutting point cloud 117, the 3D agricultural feature group 132, and the mega agricultural feature group 136, or a portion thereof, can be stored as a data structure for performing the various steps described above. However, one or more of the image captured in imaging step S1, the disparity map 46, the segmented image 51, the feature image 94, the point cloud 49, the mega point cloud 116, the cutting point image 106, the 3D cutting point cloud 112, the mega cutting point cloud 117, the 3D agricultural feature group 132, and the mega agricultural feature group 136, or a portion thereof, can also be displayed to the user, for example, on a display device 43 or via a user platform.
[0189] As described above, the processor and memory elements of the imaging electronic system 42 may be configured or programmed to control one or more devices, including the camera 20, the robot arm 22, the robot arm mount assembly 23, and the cutting tool 24, and may also be configured or programmed to process image data obtained by the camera 20. In a preferred embodiment of the present invention, the processor and memory elements of the imaging electronic system 42 may be configured or programmed to perform the functions described above, including a disparity estimation step S2, a component splitting step S3, an agricultural feature detection step S4, a point cloud generation step S5, a point cloud registration step S6, a cut point generation step S7, a cut point projection step S8, a cut point registration step S9, a mega registration step S10, an operation step S11, an agricultural feature projection step S12, and an agricultural feature registration step S13. In other words, the processor and memory elements of the imaging electronic system 42 can be defined, configured, or programmed to function as components including the AI disparity estimation model 44, instance segmentation AI architecture 50, object detection model 92, point cloud generation module 491, point cloud registration module 1161, cutpoint generation module 104, cutpoint projection module 110, cutpoint registration module 1151, trace module 120, agricultural feature projection module 1301, and agricultural feature registration module 1341.
[0190] In the preferred embodiments of the present invention described above, the target crop is a grapevine. However, preferred embodiments of the present invention are also applicable to other target crops such as fruit trees and flowering plants such as rose bushes.
[0191] It should be understood that the foregoing description is merely illustrative of the present invention. Various alternatives and modifications can be devised by those skilled in the art without departing from the present invention. Accordingly, the present invention is intended to encompass all such alternatives, modifications, and variations that fall within the scope of the appended claims.
Claims
1. A method for generating agricultural cutting points for crops, To capture images of the aforementioned agricultural products, To generate depth estimates for the aforementioned crops, By segmenting the image of the crop, a segmented image is generated that identifies different segments of the crop. Based on the aforementioned image of the crop, the agricultural characteristics of the crop are detected. To generate two-dimensional cutting points based on the segmented image and the agricultural features, Based on the aforementioned two-dimensional cutting points and the depth estimation of the crops, a three-dimensional cutting point is generated. Methods that include...
2. The method according to claim 1, wherein generating the depth estimation of the crop, segmenting the image of the crop, and detecting the agricultural features are performed simultaneously.
3. The act of capturing the image of the crop includes capturing multiple images of the crop from multiple viewpoints, The aforementioned multiple images are captured using cameras that are moved to the aforementioned multiple viewpoints. The method according to claim 1.
4. The method according to claim 1, wherein generating the depth estimation of the crops includes generating a disparity estimation using an artificial intelligence disparity estimation model.
5. To generate a point cloud based on the depth estimation of the aforementioned crops, When one or more points in the point cloud have a depth greater than a depth-based threshold, remove the one or more points. The method according to claim 1, further comprising:
6. The method according to claim 5, wherein the depth-based threshold is set based on the working range of the cutting system used to perform the cutting operation at the three-dimensional cutting point.
7. To generate a point cloud based on the depth estimation of the aforementioned crops, Based on the density of points included in the point cloud, one or more points are removed from the point cloud. The method according to claim 1, further comprising:
8. The method according to claim 1, wherein the segmented image is generated using an instance segmentation artificial intelligence architecture.
9. This further includes training an instance segmentation artificial intelligence architecture using a segmentation dataset tailored to the instant segmentation task for the aforementioned crops, The segmentation dataset includes multiple annotated images of the crops, The aforementioned multiple annotated images include a mask formed around the segment of the crop, At least one of the aforementioned annotated images includes discrete portions of the same segment to which the same label has been assigned, The method according to claim 8.
10. Using an object detection model that receives the image of the crop and detects the agricultural features within the image of the crop, the agricultural feature locations of multiple agricultural features of the crop, including the agricultural features, are determined. The method according to claim 1, further comprising:
11. The method further includes determining the agricultural feature locations of multiple agricultural features of the crop, including the aforementioned agricultural features. The generation of the two-dimensional cutting point is Associating the aforementioned multiple agricultural features with a specific segment among the different segments of the crop based on the location of the agricultural features, Assigning an identifier to each of the aforementioned multiple agricultural features relating to the specific segment to which the aforementioned multiple agricultural features are associated, Based on the identifiers assigned to the plurality of agricultural features, the two-dimensional cutting points are generated. including, The method according to claim 1.
12. The segmented image includes a mask that identifies the different segments of the crop, The mask that identifies the different segments includes a specific mask that identifies the specific segment, When the agricultural feature locations of the plurality of agricultural features are within the specific mask or within a predetermined distance from the specific mask, the plurality of agricultural features are associated with the specific segment. The method according to claim 11.
13. The method according to claim 12, wherein the identifier is assigned to the plurality of agricultural features based on the distance between the agricultural feature location and a point on the particular mask.
14. The method according to claim 13, wherein the two-dimensional cutting point is generated at a point between two of the plurality of agricultural features based on the identifiers assigned to the plurality of agricultural features.
15. The method according to claim 14, further comprising moving the two-dimensional cut point so that it is located within the specific mask when the two-dimensional cut point generated at the points between the plurality of agricultural features is not located within the specific mask.
16. Based on the positions of two of the aforementioned agricultural features, the angle of the portion of the specific segment from which the two-dimensional cutting point is generated is determined. The cutting point angle of the two-dimensional cutting point is determined based on the angle of the portion of the specific segment in which the two-dimensional cutting point is generated. The method according to claim 14, further comprising:
17. The method according to claim 1, wherein the agricultural characteristics of the crop are detected based on the segmented image.
18. This further includes generating mega 3D cuts by combining multiple 3D cuts, The act of capturing the image of the crop includes capturing multiple images of the crop from multiple viewpoints. The generation of the depth estimates of the crops includes generating multiple depth estimates of the crops corresponding to each of the multiple images, Segmenting the images of the crops includes generating a plurality of segmented images corresponding to each of the plurality of images, The detection of the agricultural characteristics of the crop includes detecting the agricultural characteristics of the crop in each of the plurality of images. The generation of the two-dimensional cutting points includes generating a plurality of two-dimensional cutting points based on the plurality of segmented images and the agricultural features, wherein each of the plurality of two-dimensional cutting points corresponds to the plurality of viewpoints. The generation of the three-dimensional cutting points includes generating the plurality of three-dimensional cutting points based on the plurality of two-dimensional cutting points and the plurality of depth estimates of the crop, wherein each of the plurality of three-dimensional cutting points corresponds to the plurality of viewpoints. The method according to claim 1.
19. The process of generating the mega 3D cutting point by combining the aforementioned plurality of 3D cutting points is as follows: Assigning a search radius to each of the aforementioned multiple three-dimensional cutting points, Aligning the plurality of three-dimensional cutting points by performing one or more spatial transformations, The process includes merging the plurality of three-dimensional cutting points into the mega three-dimensional cutting point when the plurality of three-dimensional cutting points are located within the search radius assigned to one of the plurality of three-dimensional cutting points. The method according to claim 18.
20. Based on the aforementioned multiple depth estimations of the crop, a plurality of point clouds are generated, each corresponding to one of the plurality of viewpoints. The process involves combining the aforementioned multiple point clouds to generate a megapoint cloud, Merging the aforementioned mega 3D cutting points and the aforementioned mega point cloud, The process involves tracing a portion of the mega point cloud corresponding to the segment of the crop where the mega 3D cutting point is located, and determining whether an additional mega 3D cutting point is located on the segment of the crop. When it is determined that the additional mega 3D cutting point is located on the segment of the crop, the additional mega 3D cutting point is removed. The method according to claim 18, further comprising:
21. To determine whether an additional mega 3D cutting point is located on the segment of the crop where the aforementioned mega 3D cutting point is located, When it is determined that the additional mega 3D cut point lies on the segment of the crop, the additional mega 3D cut point is deleted. The method according to claim 18, further comprising:
22. The process involves determining a plurality of cutting point angles corresponding to the plurality of two-dimensional cutting points, wherein the plurality of cutting point angles are determined based on the angles of the portion of a specific segment of the crop where the plurality of two-dimensional cutting points are generated. Based on the aforementioned multiple cutting point angles, the mega cutting point angle of the mega 3D cutting point is determined, The method according to claim 18, further comprising:
23. The method according to claim 1, wherein generating the depth estimation of the crop comprises generating a disparity estimation based on the image of the crop.
24. The method according to claim 1, wherein generating the depth estimation of the crop includes acquiring point cloud data from a LiDAR sensor.
25. A system for generating agricultural cutting points for crops, A camera for capturing images of the aforementioned crops, By segmenting the image of the crop, a segmented image is generated that identifies different segments of the crop. Based on the aforementioned image of the crop, the agricultural characteristics of the crop are detected. Based on the segmented image and the agricultural features, a two-dimensional cutting point is generated. Based on the two-dimensional cutting points and the depth estimation of the crops, a three-dimensional cutting point is generated. A processor configured or programmed in such a way, A system that includes these features.