Autonomous under-canopy vehicle for pest detection and treatment

WO2026177781A2PCT designated stage Publication Date: 2026-08-27KANSAS STATE UNIV RES FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/053781
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-06-04
Filing Date
2025-11-03
Publication Date
2026-08-27

Smart Images

  • Figure US2025053781_27082026_PF_FP_ABST
    Figure US2025053781_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A method of operating an autonomous robotic vehicle, the method includes a step of receiving a two-dimensional (2D) image via a camera of the autonomous robotic vehicle. The method further includes a step of detecting an object represented in the 2D image via a detection system of the autonomous robotic vehicle. The method further includes a step of determining three-dimensional (3D) world coordinates of the object based on the detecting step via the detection system of the autonomous robotic vehicle. The method further includes a step of activating a component of the autonomous robotic vehicle according to the 3D world coordinates.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. 61333-PCTAUTONOMOUS UNDER-CANOPY VEHICLE FOR PEST DETECTION AND TREATMENTGOVERNMENT INTERESTS

[0001] This invention was made with government support under Award No. 2019-67021-28995, awarded by the U.S. Department of Agriculture. The government has certain rights in the invention.BACKGROUND

[0002] The agricultural production system faces two major environmental issues: the overuse of pesticides and the accumulation of plastic waste. Traditional methods employ large self-propelled sprayers for applying herbicides, insecticides, and fungicides without adequate knowledge of the spatial variability of pests within the field, leading to overuse of pesticides. The need for precision agriculture solutions that can accurately detect pests and apply targeted treatments is pressing, particularly in the case of aphid infestations in crops like Sorghum.

[0003] Aphids are notorious pests that can cause significant damage to crops, and their timely detection is crucial for effective pest management. However, detecting aphids in real-time poses significant technical challenges. Traditional approaches to neural network deployment, which are commonly used for image-based pest detection, are limited by wireless communication constraints. These limitations hinder the development of autonomous systems that can detect pests and apply targeted treatments in real-time.

[0004] Several researchers have attempted to address these challenges. In particular, a method for pest detection and positioning based on binocular stereo vision to guide a robot for automatic pesticide spraying in greenhouses has been proposed. This approach utilized color feature extraction and image segmentation to identify pests, combined with binocular stereo vision techniques to obtain 3D positioning of the pests. While innovative, this method relied on a stereo camera setup, which can be more complex and expensive than monocular systems.

[0005] In addition, navigation of ground-based vehicles that traverse under the canopy of the growing crops presents challenges. Because the vehicles travel underneathDocket No. 61333-PCTthe canopy, reliance upon GPS for navigation is impractical as the canopy can block GPS reception. Therefore, a different navigation system is required that does not rely on external input for directing the vehicle through the field.SUMMARY OF THE INVENTION

[0006] According to one or more embodiments, a ground-based robotic vehicle is provided that can detect certain objects or conditions related to crop health and apply targeted treatment or perform other actions accordingly. The ground-based robotic vehicle can also navigate under the canopy formed between row crops. The vehicle includes an on-board vision navigation system that allows the vehicle to autonomously guide itself in between rows of crops and conduct phenotyping activities, such as to identifying one or more conditions associated with the health of individual plants making up the row crops. This may include insect infestations, fungal infections, viral infections, bacterial infestations, heat damage, cold or frost damage, and other conditions.

[0007] The vehicle is sized to traverse a field between rows of crops and under the canopy without causing untoward damage to the plants. As used herein, the term “canopy” means the aboveground portion of a plant cropping formed by the collection of individual plant crows. In certain embodiments, the vehicle has a track width that is substantially less than the width between adjacent crop rows. In particular embodiments, the track width of the vehicle is such that it maintains a clearance of at least 1 inch, at least 2 inches, at least 4 inches, or from about 4 to about 5 inches of clearance on either side of the vehicle as it traverses the space in between crops rows. Conventionally, corn, for example, is planted to give a 30 inch row spacing. Thus, in specific embodiments, the vehicle can be about 21 inches wide.

[0008] The vehicle comprises apparatus making up a neural network trained to identify observable characteristics associated with the row crop plants including pest infestation and plant nutrition. Should concerning characteristics be identified, the vehicle is configured to deploy a treatment to the affected plant(s). The treatment may comprise deploying a pesticide, fertilizer, fungicide, herbicide, or the like to the affected plant using one or more onboard sprayers or delivery mechanisms. The treatment may be a liquid, a powder, a granule, a gel, a gas, or any other suitable form. In the context of a pestDocket No. 61333-PCTinfestation, such as an aphid infestation, the vehicle can a computer vision system to identify the presence of the infestation and use the neural network to generate a three-dimensional map of the precise location of the infestation. The neural network can then cause the vehicle to deploy an appropriate amount of pesticide directly onto the identified location of the infestation without necessarily requiring that the entire plant receive a pesticide treatment. Thus, the vehicle can provide highly specific and targeted treatments to only those portions of the plant that require such thereby avoiding unnecessary use of the pesticide. The vehicle further carries wireless communication equipment permitting transmission of data gathered onboard, such as information related to plant phenotyping, treatments, geocoordinates, vehicle operating parameters, and the like, to a central control center where the operation of the vehicle can be monitored by an operator.

[0009] The onboard computer vision system detects crop phenotypes in real-time using monocular depth estimation to provide 3D localization of pests and hardware acceleration techniques with the aim of significantly reducing pesticide usage in crop management while maintaining effective pest control. The monocular depth estimation technique eliminates the need for complex stereo vision systems.

[0010] Generally, performing 3D localization of pests using monocular depth estimation is accomplished by taking advantage of a generally accepted or “fixed” size parameter for the target object. For example, if the target is a sorghum aphid, the size of the aphid can be assumed to be approximately 3 mm, which is an accepted average size for this pest. The depth estimation is then based on the principle of perspective projection where the relationship between an object’s actual size, its size in an image, and its distance from the camera can be used to determine the object’s depth in the 2D image. Once the depth is determined, the object’s full 3D position can be determined.

[0011] The process of determining the object’s full 3D position involves a pixel projection step in which the center of the detection bounding box is projected into 3D space. This returns a unit vector in the camera coordinate frame, pointing in the direction of the rectified pixel in the image plane. The unit vector obtained from the pixel projection step is then scaled using the previously calculated depth. The scaling transforms the unit vector into a 3D point representing the object’s position relative to the camera.Docket No. 61333-PCT

[0012] Once the 3D localization of the object has been determined, the vehicle can orchestrate the precise application of pesticides based on the detected pest locations. The pest’s coordinates are transformed to a frame for the respective sprayer so that the pest’s location relative to the sprayer boom can be determined thereby enabling precise targeting. Then, the vehicle determines whether each pest is located within predefined spraying bounds. These bounds are determined by two key parameters: the pre-target and post-target distances. The pre-target distance defines the point at which the sprayer should activate in anticipation of reaching the pest, while the post-target distance indicates when the sprayer should deactivate after passing the pest’s position. When the target pest is determined to be within the target range for the sprayer, one or more sprayer nozzles are activated causing pesticide to be applied to the precise location of the pests. Alternatively, pesticide may be applied to the entire plant deemed to be affected by the pest.

[0013] In one or more embodiments, the sprayer system comprises a tank, a pump, a flow control assembly, and a distribution system. In addition, the sprayer system can also include a pressure gauge and an agitator. The spraying system can be configured to deliver the liquid at a pressure of up to 300 psig at an application rate of less than 1 to more than 100 gallons per acre. The distribution system terminates in at least one spraying boom. In certain embodiments, the vehicle comprises a delivery boom for each of the crop-facing sides of the vehicle. The booms preferably are located aft on the vehicle at an opposite end from the cameras used to image the crops. The booms can be configured with multiple nozzles to permit spraying of a liquid treatment at various heights above the ground. For example, the boom can be configured with nozzles placed from 5 to 25 inches, 8 to 20 inches, 12 to 17 inches, or about 15 inches apart along the length of the boom.

[0014] When the robotic vehicle is operating under the row crop canopy, GPS signal reception can be intermittent. Therefore, in one or more embodiments, the present invention provides a system for autonomous navigation of the robotic vehicle without relying upon GPS coordinates. The navigation system is a visual-based system that utilizes forward and / or side-facing cameras in order to determine the spatial relationship between the robotic vehicle and the crops. Images acquired by the cameras areDocket No. 61333-PCTprocessed through a backbone network. After passing through the backbone, a global average pooling (GAP) layer is applied. The resulting features are then fed into three separate multi-layer perceptrons (MLPs), each comprising two layers. The MLPs are responsible for generating confidence scores and predicting polynomial coefficients for various polynomials representing navigational coordinates U and P.

[0015] Similar to the identification of and 3D localization of pests, the navigational system relies upon using a known value, here height, to project 2D coordinates and predict spatial coordinates representative of the left and right boundaries of the row crops. With these coordinates determined, the vehicle can guide itself through the space between adjacent crop rows without the need to rely upon external navigational information, such as that provided by GPS.

[0016] Camera calibration parameters tend to drift over time, impacting the accuracy of information derived from images. To maintain precision, these parameters must be continually estimated. For an under-canopy robot that detects crop rows as splines or polynomials, such as that described herein, the camera calibration parameters can be effectively estimated using a least squares approach under the assumption that all the rows detected are in the same plane and the distance between the rows is approximately the same.

[0017] Camera parameters are characteristics that define how a camera captures images. Examples of camera parameters include intrinsic parameters that are specific to the camera and define the camera’s lens and sensor, and extrinsic parameters that are external to the camera and may change with respect to the world frame (i.e., a fixed coordinate system for representing objects in the real world). Intrinsic parameters characterize the optical, geometric, and digital characteristics of the camera and include focal length, transformation between camera frame and pixel coordinates, and geometric distortion introduced by the lens. Extrinsic parameters recall the fundamental equations of perspective projection and are made up of a rotation and a translation.

[0018] In certain embodiments, the neural network carried onboard of the vehicle can re-estimate one or more camera parameters using images of the row crops gathered by one or more onboard cameras taken at different times. Since the distance of the vehicle (or camera) to the crops in the image is known, the changes in camera parametersDocket No. 61333-PCTbetween images can be estimated. The re-estimation of camera parameters can be carried out while the robotic vehicle is in operation within the field. The frequency at which re-estimation of camera parameters often depends on several factors, including factors specific to the camera(s) being used. However, in one or more embodiments, the reestimation of camera parameters can be programmed to occur on a schedule, such as once a day, every second day, or once a week. This re-estimation of camera parameters on a timely basis ensures accurate identification of plant phenotypes, application of any plant treatment, and vehicle navigation between the row crops.

[0019] An embodiment of the present invention is a method of operating an autonomous robotic vehicle. The method includes a step of receiving a two dimensional (2D) image via a camera of the autonomous robotic vehicle. The method further includes a step of detecting an object represented in the 2D image via a detection system of the autonomous robotic vehicle. The method further includes a step of determining three dimensional (3D) world coordinates of the object based on the detecting step via the detection system of the autonomous robotic vehicle. The method further includes a step of activating a component of the autonomous robotic vehicle according to the 3D world coordinates.

[0020] Another embodiment of the present invention is a method of navigating an autonomous robotic vehicle between rows of crops. The method includes a step of predicting polynomial coefficients for detection of the rows via a navigation system of the autonomous robotic vehicle. The method further includes a step of predicting a plurality of polynomial candidates via the navigation system of the autonomous robotic vehicle. The method further includes a step of predicting a classification probability for occurrence of one of the rows via the navigation system of the autonomous robotic vehicle. The method further includes a step of training a neural network based on a loss function via the navigation system of the autonomous robotic vehicle. The method further includes a step of projecting the polynomial coefficients, the polynomial candidates, and the classification probability onto a ground plane via the navigation system of the autonomous robotic vehicle. The method further includes a step of directing the autonomous robotic vehicle to traverse a ground surface according to the projections.Docket No. 61333-PCT

[0021] Yet another embodiment of the present invention is an autonomous robotic vehicle broadly comprising a frame, a camera mounted on the frame, and a drive system including a number of wheels configured to traverse a ground surface and a motor configured to drive at least one of the wheels. The vehicle further comprises a sprayer system including a tank mounted on the frame and configured to hold a treatment, a boom mounted on the frame, a nozzle mounted on the boom and configured to deliver the treatment, and a control system configured to direct delivery of the treatment. The vehicle further comprises a detection system configured to receive a two dimensional (2D) image via the camera, detect an object represented in the 2D image, determine three dimensional (3D) world coordinates of the object based on the detection, and instruct the control system of the sprayer system to direct delivery of the treatment according to the 3D world coordinates.

[0022] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Other aspects and advantages of the current invention will be apparent from the following detailed description of the embodiments and the accompanying drawing figures.BRIEF DESCRIPTION OF THE DRAWING FIGURES

[0023] Embodiments of the present invention are described in detail below with reference to the attached drawing figures, wherein:

[0024] FIG. 1 is an environmental view of an autonomous robotic vehicle constructed in accordance with an embodiment of the invention;

[0025] FIG. 2 is a perspective view of the autonomous robotic vehicle of FIG. 1 ;

[0026] FIG. 3 is a schematic diagram of certain components of the autonomous robotic vehicle;

[0027] FIG. 4 is a schematic diagram of certain components of the autonomous robotic vehicle;

[0028] FIG. 5 is a schematic diagram of certain components of the autonomous robotic vehicle;Docket No. 61333-PCT

[0029] FIG. 6 is a perspective view of a rending of the autonomous robotic vehicle according to a navigation and detection system thereof;

[0030] FIG. 7 is a schematic diagram of a kinematic organization of certain components of the autonomous robotic vehicle according to the navigation and detection system thereof;

[0031] FIG. 8A is a flow diagram depicting certain method steps in accordance with an embodiment of the invention;

[0032] FIG. 8B is a flow diagram depicting certain method steps in accordance with an embodiment of the invention;

[0033] FIG. 8C is a flow diagram depicting certain method steps in accordance with an embodiment of the invention;

[0034] FIG. 8D is a flow diagram depicting certain method steps in accordance with an embodiment of the invention; and

[0035] FIG. 9 is a flow diagram depicting certain method steps in accordance with an embodiment of the invention.

[0036] The drawing figures do not limit the current invention to the specific embodiments disclosed and described herein. The drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the invention.DETAILED DESCRIPTION

[0037] The present invention is directed to systems and methods for visual navigation and visual detection for performing agricultural tasks such as pest detection for crop treatment implemented by an autonomous robot vehicle. Additional information regarding various aspects of the present invention can be found in Application Serial No.63 / 716,020, filed November 4, 2024, titled “AUTONOMOUS UNDER-CANOPY VEHICLE FOR PEST DETECTION AND TREATMENT”, incorporated by reference herein in its entirety.

[0038] Turning to FIGS. 1-4, an exemplary autonomous robotic vehicle 100 will now be described. The autonomous robotic vehicle 100 broadly comprises a frame 102, a housing 104, a power supply 106, a drive system 108, a radio communication system 110, a sprayer system 112, and a navigation and detection system 114.Docket No. 61333-PCT

[0039] The frame 102 may be a lightweight body formed of extruded metal and aluminum sheeting, or any other suitable construction such as a molded components. An extruded metal construction provides modularity for addition, subtraction, modification, or rearrangement of various components as needed. The frame 102 supports the housing 104, the sprayer system 112, and other components and systems of the autonomous robotic vehicle 100.

[0040] The housing 104 may be mounted on the frame and may enclose the power supply 106 and certain (particularly electronic) components of the drive system 108, the radio communication system 110, the sprayer system 112, and the navigation and detection system 114. The housing 104 may include a door or panel for providing access to these components.

[0041] Turning to FIG. 5, the power supply 106 provides electrical power to the drive system 108, the radio communication system 110, the sprayer system 112, and the navigation and detection system 114. The power supply 106 may include one or more batteries. In one embodiment, the power supply 106 includes two 24V, 60Ah Lithium-Ion batteries connected in parallel, thus providing a total capacity of 120Ah.

[0042] The drive system 108 propels the autonomous robotic vehicle 100 and may include a plurality of wheels 116, a plurality of motors 118, and a plurality of motor controllers 120.

[0043] The wheels 116 may be drivably coupled to the motors 118 to traverse a ground surface. The wheels 116 may have skid-steer capabilities. In on embodiment, the wheels include four driven wheels. In other embodiments, more or fewer wheels, and unpowered wheels may be used. The wheels 116 may support the frame 102 and housing 104 via a suspension system as seen in FIG. 2.

[0044] The motors 118 may drive the wheels 116 and may be brushless DC motors or any other suitable type of motor. In one embodiment, the motors 118 include four brushless DC motors.

[0045] The motor controllers 120 may activate the motors 118 according to activation signals from the navigation and detection system 114. The motor controllers 120 may be independent of each other to effect differential control of the wheels 116.Docket No. 61333-PCT

[0046] The radio communication system 110 may be communicatively coupled with the drive system 108, the sprayer system 112, the navigation and detection system 114, or other systems for transmitting data to a remote computing system and receiving data and instructions therefrom. To that end, the radio communication system 110 may include a transceiver, antenna, or any other suitable wireless communication component.

[0047] Turning to FIGS. 3 and 4, the sprayer system 112 may be configured to apply a liquid treatment, such as a liquid insecticide or fertilizer and may be mounted on the frame 102 such as near a rearward portion of the frame 102. The sprayer system 112 may include a tank 122, one or more booms 124, a distribution system 126, and a control system 128. The sprayer system 112 may also include a pressure gauge 130 and an agitator (not shown).

[0048] The tank 122 may be a twenty gallon tank or any other suitable size and may be made of corrosion-resistant material. In one embodiment, the tank is a twenty gallon Ace Roto-Mold SF-approved rectangular application tank made of polyethylene plastic with dimensions of 28 inches by 14 inches by 12 inches. The tank 122 may allow easy filling and draining.

[0049] The booms 124 may be mounted on the frame 102 and may be vertically oriented, although other mounting configurations may be used. In one embodiment, the booms 124 may have a length of approximately 1.2 meters. The booms 124 may be equipped with a linear actuation capability of seven inches from their base. The booms 124 may be configured with multiple nozzles (described below) to permit spraying of a liquid treatment at various heights above the ground, as described below. The booms 124 may be located on an aft end (opposite the end on which the crop imaging cameras are mounted) of the autonomous robotic vehicle 100 and may include one boom on each crop-facing side thereof. In one embodiment, the booms 124 may be equipped with a linear actuation capability of 7 inches with a Linak (LA 25) actuator. The booms 124 may be made of stainless steel to prevent corrosion and provide better stability. The mounting was made of aluminum sheet to make it lightweight. A 3D-printed sleeve may be used to hold each boom 124 in the correct position. The distance between left and right booms may be 6 inches. The booms may include three nozzles, with the distance between each nozzle being 40 inches.Docket No. 61333-PCT

[0050] The distribution system 126 may include a pump 134, one or more valves 136, and one or more nozzles 138. The distribution system 126 may terminate in at least one boom 124 and may be configured to deliver liquid at a pressure of up to 300 psi at an application rate of less than 1 to more than 100 gallons per acre.

[0051] The pump 134 draws fluid liquid from the tank 122 and should have a capacity to deliver liquid at a specified pressure to the valves 136 with consistent distribution and agitation. The capacity should be twenty percent greater than the most significant volume required by the nozzles 138 to account for agitation requirements and losses due to pump wear. The pump 134 may be a diagram pump, which maintains constant pressure throughout spraying and does not depend on flow variation. In one embodiment, the pump 134 may be a Delavan 5850-111 E Power FLO electric diaphragm pump with a flow rate of 3.2 GPM at 40 PSI and flow limit of 5 GPM at a maximum rated pressure of 60 PSI. The pump 134 may have Viton valves and a Santoprene diaphragm, making it corrosion-resistant. The pump 134 may have inlet and outlet ports of % inch and 1 inch sizes, respectively. The pump 134 may be preceded upstream by a strainer such as a T-line strainer AA122-3 / 4-PP50 with a mesh size of 50 from Teejet Technologies and a % inch ball valve to keep foreign debris out of the pump 134 and to control flow from the tank 122 to the pump 134.

[0052] The valves 136 ensure liquid is distributed according to signals from the control system 128. The valves 136 may be solenoid valves to regulate nozzle flow based on the selected duty cycle. The valves 136 may be 12V e-chemsaver (115880) solenoid valves from Teejet Technologies, which provide individual nozzle control (ON / OFF capabilities). The valves 136 may be diaphragm check valves, which are normally closed and opens when the solenoid is energized.

[0053] The nozzles 138 may be mounted on the booms 124 and may have hollow cone bodies. The nozzles 138 help to determine the flow rate, atomizes the mixture, and distributes these droplets in a particular pattern. In one embodiment, the nozzles 138 are TXVS-1 stainless steel cone jet visible hollow cone spray tip nozzles from Teejet Technologies with a spray angle of 80 degrees. The nozzles 138 may have a flow rate of 0.033 GPM at 40 PSI pressure. The nozzles 138 may be placed from 5 to 25 inches, 8 to 20 inches, 12 to 17 inches, or about 15 inches apart along a length of the booms 124.Docket No. 61333-PCT

[0054] On the delivery side, the pump 134 may be connected to a three-way connector, with one side being connected to the tank 122 via a throttling valve, and the other side being connected to a delivery header. The delivery header may comprise an1 / 2 inch AA122-1 / 2-PP16 T-strainer of 16 mesh from Teejet Technologies, for example, to protect the nozzles 138 from clogging. The pressure gauge 130 may be a digital pressure gauge with a plastic case and a 304 stainless steel bottom connection with % NPT male connection to measure the pressure. The pressure gauge 130 may handle pressure from 0 to 100 PSI. The 14 inch gravity liquid flow sensor measures the amount of insecticide applied as well as the flow rate of the sprayer system 112. This helps measure the amount of insecticides applied during spraying. The non-return valve in the delivery line may also help maintain pressure at the booms 124 after spraying. The main delivery header may be attached to the booms 124 via a 90-degree connector.

[0055] With reference to FIG. 4, the control system 128 directs distribution of liquid according to instructions from the navigation and detection system 114 and may include an electronic control unit (ECU) 140, a microcontroller 142, a plurality of sensors 144, and a CAN transceiver 146.

[0056] The sprayer control system 128 may draw 12V from the power supply 106, via a DC-to-DC buck converter (24V to 12V). The 12V input may be supplied to the pump 134 via a relay. The relay was used to control pump operation based on the signal received from the microcontroller 142. A solid-state relay from Schneider Electric (861SSR115-DD) may be used to handle the current up to 20 A.

[0057] The ECU 140 may be a Teejet Technologies Dynajet Interface module (78-05136) Electronic Control Unit (ECU), which serves as the primary hub for the sprayer system 112, connecting and exchanging signals with different components of the sprayer system 112. The ECU 140 may be powered by a 12 V supply from the buck converter. The CAN transceiver 146 may be used between the microcontroller 142 and the ECU 140 to support CAN communication. The ECU 140 may support CAN 2.0 B communication protocol with 29-bit(extended) identifiers. The duty cycle and individual nozzle control command can be set in the ECU 140 to activate the solenoids. A Dynajet HF Driver (78-05124) module from Teejet Technologies may provide power and transmit PWM commands to the individual solenoid valves 136 from the ECU 140. The ECU 140Docket No. 61333-PCTmay send commands to the driver and activate the solenoids based on the input duty cycle and nozzle number. A 120 Ohm terminating resistor may be provided in the driver to establish the CAN communication.

[0058] The microcontroller 142 may send signals to the pump 134, the valves 136, and other components for dictating delivery of liquid from the nozzles 138. The microcontroller 142 may be an STM-32 microcontroller and may have a Micro-Robotic Operating System (Micro-ROS) as an agent between the navigation and detection system 114 and the microcontroller 142. The microcontroller 142 may be an STM32G4 with Arm® Cortex®-M4 at 170MHz in an LQFP64 package, with 512kb of Flash memory and 32kb of SRAM from ST Microelectronics, or may be any other suitable controller. The microcontroller 142 may utilize micro-ROS with CAN communication, and specifically CAN Flexible Data Rate (CAN FD) as a transport layer between the microcontroller and the navigation and detection system 114 described below. The microcontroller 142 may interface with the sensors 144 and the solenoid drivers for activating valves 136.

[0059] The sensors 144 may include a pressure sensor, a flowmeter, a liquid level sensor, and any other suitable sensor. In one embodiment, the pressure sensor may be an M32JM-000105-100PG model from TE Connectivity to ensure the consistency of the pressure while the liquid is being delivered. A 1 / 2-inch pressure regulating valve (23120) from Teejet Technologies may be connected from the delivery head to the tank to maintain and adjust line pressure at the different operating ranges. The liquid level sensor may be an ultrasonic sensor having a detection range from 1.18 to 177 inches to help measure the tank level with a total height of 12 inches. In one embodiment, the ultrasonic sensor may be an SEN 0311 from DF Robot.

[0060] The above pressure sensors, flow meters, and liquid level sensors may be required to ensure the proper functioning and evaluation of the sprayer system 112. The pressure sensor may be configured with STM with the help of an analog-to-digital converter ADS 1115 (16-bit precision). The ADS 1115 allows four sensors to connect at the same time. Inter-Integrated Circuit(l2C) may be used as a connection protocol between STM and ADC. The continuous monitoring of spraying pressure may be important to evaluate the performance of the spraying. The pressure sensor data may be sampled at a frequency of 100 Hz. The set pressure for the sprayer operation may be 40 PSI. TheDocket No. 61333-PCTwater flow and liquid level sensors (ultrasonic sensors) may be configured with STM32 with a GPIO (General Purpose Input / Output) pin and universal asynchronous receiver / transmitter(UART) protocol, respectively. The flow sensor may measure the NPN pulse signal to determine the flow of the liquid. A liter of liquid flowing through the sensor may be equivalent to 150 pulses. The liquid level sensor may determine the tank's liquid level by measuring the time lapses between sending and receiving the electronic pulse. The number of pulse signals received may be converted to the amount of liquid flow. The sampling frequency for both flow and level sensors may be 100 Hz. The sensors 144 may be calibrated before use.

[0061] The sprayer system 112 may utilize the following software architecture. The The sprayer micro-ROS node described below may be the main governing node for the sprayer operation. Micro-ROS humble may be installed in the STM with Zephyr RTOS as the operating system. This node may be used to send the duty cycle and PWM commands to activate the solenoids. Teejet documentation for Dynajet ECU may be used as a reference to send the commands. The message to Dynajet ECU may be sent using CAN 2.0 B communication protocol with 29-bit extended identifier at a bit rate of 250,000 bits / sec. The first byte in this message’s 8-byte data packet may be the message identifier. The message identifier for ‘Set Target PWM’ and ‘Send Extended Boom Status’ may be used to command the target PWM and activate the individual nozzle respectively. The set nozzle status service may send these commands and activate the sprayer system 112. In the set nozzle status service, the entire left boom may be treated as ‘0’ and the right boom as ‘1 ’. The node may also publish the data from pressure, flow, and liquid level sensors with different individual topics. The sprayer nozzle status may give information about the activation of the booms 124 during the spraying.

[0062] Turning again to FIG. 5, the navigation and detection system 114 broadly comprises a plurality of sensors (e.g., encoders 148, an inertial measurement unit (IMU) 150, and a GPS unit 152), a plurality of cameras 154, and a primary computer 156.

[0063] The encoders 148 may be configured to sense position, rotation, or status of the wheels and generate wheel encoder data therefrom. The encoders 148 may be Hall encoders, or any other suitable encoders.Docket No. 61333-PCT

[0064] The IMU 150 may be configured to sense inertial movement of the autonomous robotic vehicle 100 and generate orientation or acceleration data (IMU data) therefrom.

[0065] The GPS unit 152 may be configured to receive GPS data from remote satellites to generate GPS data. At least two of the wheel encoder data, the IMU data, and the GPS data may be fused to provide a local odometry source (wheel encoder data and IMU data) and a global odometry source (fusing the local odometry with GPS data).

[0066] The cameras 154 may be Arducam AR0234 global shutter camera modules, connected to a Luxonis OAK-FFC 4P PoE development kit via a MIPI interface. The OAK- FFC 4P PoE is a camera interface development kit that allows for the connection of up to four camera modules via MIPI. It features two 2-lane MIPI connections and two 4-lane MIPI connections, and it is equipped with an Intel Movidius Myriad X VPU. The OAK-FFC 4P PoE development kit provides flexibility in testing different camera sensors by simply swapping different Luxonis FFC camera modules. The 4 MIPI ports allow for potential upgrades, such as the addition of two cameras on each side of the autonomous robotic vehicle 100 for stereo vision depth perception. In such a scenario, four cameras could be connected to one development kit, with a single PoE connection to the Jetson. The Boxer 8640-ai, with its four PoE ports, can directly power the OAK-FFC 4P PoE.

[0067] The primary computer 156 may be an NVIDIA Orin (8-core ARM v8.264-bit CPU processor with NVIDIA Ampere architecture), serving as the computational backbone for the robotic platform to suffice the real time processing needs, while Robotic Operating System (ROS) may serve as the middle ware to integrate and ensure seamless communication among subsystems. An Nvidia Isaac ROS GPU accelerated package may be used to compute an efficient image processing pipeline on the Nvidia AGX Orin, as well to deploy neural network models. The Orin is a compact, fan-less embedded Al system that can perform real-time object detection and expedited data transmittance, and has rugged elemental features that bring the whole Al application to the edge.

[0068] The Orin combines an ARM processor with a powerful GPU, which ensures exceptional model inference performance compared to CPU-based solutions, enabling real-time processing and feedback for the autonomous robotic vehicle 100. In addition toDocket No. 61333-PCTdeploying the neural network on the autonomous robotic vehicle 100, the present invention also focuses on efficiently handling image frames captured by the cameras 154. The GPU in the Jetson platform may be utilized to accelerate the image pipeline and leverage the hardware encoder to compress image frames to H.264 format, resulting in streamlined data recording. This comprehensive approach enhances the vehicle’s capabilities while ensuring energy efficiency, a critical factor for field operations. The entire computing system is seamlessly integrated into ROS 2, and benefits from Nvidia’s Isaac ROS packages, which provide GPU-accelerated ROS nodes. Furthermore, the use of hardware acceleration enables substantial computational power while maintaining efficiency.

[0069] Moreover, Nvidia Jetson provides comprehensive documentation and examples, making it an accessible and user-friendly platform for applications that require hardware acceleration and neural network deployment. While platforms with an x86 CPU and a discrete Nvidia GPU may offer superior software support and documentation, they are not feasible for a battery-powered robot due to their power consumption and physical size. As previously mentioned, the robot’s software system may use ROS 2, and Nvidia provides ROS 2 packages for hardware-accelerated image processing, neural network deployment, H264 image encoding, and more with the Jetson Orin. This is a significant advantage, as ROS 2 nodes eliminate the need to develop these hardware-accelerated functionalities independently. In one embodiment, the AGX Orin development kit may be used, which includes 275 Tops of Al Performance, a 2048-core NVIDIA Ampere architecture GPU with 64 Tensor Cores, and a 12-core Arm® Cortex®-A78AE v8.2 64-bit CPU. The development kit may include 32GB to 64GB of RAM. In one embodiment, a third-party carrier board and fanless enclosure, the AAEON Boxer 8640AI, powered by the Nvidia Jetson AGX Orin with 32GB of RAM may be used. Although this version has slightly less performance than the development kit, it may be suitable with 200 Tops of Al Performance, an 8-core CPU, a 1792-core NVIDIA Ampere architecture GPU with 56 Tensor Cores, and 32GB of RAM. The industrial computer form factor and fanless design prevent dust from entering the computer’s fan, which is beneficial for field operations.Docket No. 61333-PCTAutonomous Robotic Vehicle Computational Description

[0070] To accurately represent the autonomous robotic vehicle’s physical structure and properties in the ROS 2 environment, Unified Robot Description Format (URDF) may be utilized. URDF is an XML-based format that describes a robot’s physical characteristics, including its kinematic and dynamic properties, visual representations, and collision models.

[0071] For the autonomous robotic vehicle 100, a URDF file may be exported from a CAD model using an ROS SolidWorks URDF Exporter plugin. This approach ensures that the ROS 2 representation of the autonomous robotic vehicle 100 accurately reflects its real-world counterpart, which is crucial fortasks such as motion planning, visualization, and simulation.

[0072] The URDF file defines the autonomous robotic vehicle’s link structure, joint connections, and geometric properties. This information is essential for the proper functioning of various ROS 2 components, including a powerful tool called the Transform Library (TF2), and robot localization systems.

[0073] FIG. 6 provides a visual representation of the robot’s URDF model. This visualization demonstrates how the URDF accurately captures the autonomous robotic vehicle’s structure and geometry, allowing for precise spatial reasoning and transformation computations within the ROS 2 ecosystem.

[0074] This graphical output not only serves as a visual confirmation of the URDF’s accuracy but also aids in debugging and verifying the autonomous robotic vehicle’s configuration within the ROS 2 environment. It allows for visual inspection of the autonomous robotic vehicle’s structure, joint placements, and overall geometry, ensuring that all components are correctly positioned and scaled.

[0075] To accurately make visual detections in the field, a consistent coordinate frame for the autonomous robotic vehicle 100 and its environment may be established. The TF2 manages coordinate frames and transformations between them. With reference to FIG. 7, below is a description of each frame in the tree:• base link 160: The root frame of the autonomous robotic vehicle 100, typically located at the center of the autonomous robotic vehicle’s frame 102.Docket No. 61333-PCT• chassis link (not shown): Represents the frame 102 of the autonomous robotic vehicle 100.• wheel[0-3] link 162: Four frames representing each of the autonomous robotic vehicle’s wheels 116. These are crucial for odometry calculations.• imu link 164: Frame for the IMU 150, used for orientation and acceleration data.• gps link 166: Frame for the GPS unit 152, used for global positioning.• [left-right] sprayer link 168, 170: Represents the sprayer booms 124 on the autonomous robotic vehicle 100.• camera mount 172: Frame for the navigation camera 154.• left camera mount link 174: Represents the mounting point for the detection camera 154, connected to the autonomous robotic vehicle 100 via a revolute joint for angle adjustment.• left camera link 176: Represents the physical body of the detection camera 154.• left camera optical link 178: Represents the optical center of the detection camera 154, with its coordinate system aligned to the camera’s view.• Odom 180: A frame representing the autonomous robotic vehicle’s position relative to its starting point, based on local odometry.• Map 182: A global reference frame, typically used with GPS data for global localization.

[0076] These frames and their relationships, as defined in the URDF, allow for precise tracking of the autonomous robotic vehicle’s components and sensors in 3D space. It is important to note the relationships and differences between the camera-related frames:• left camera mount link 174 is connected to the autonomous robotic vehicle’s frame 102 via a revolute joint, allowing for adjustment of the camera angle. This flexibility is crucial for optim izing the field of view for detection in various crop configurations.• left camera link 176 is a child of the left camera mount link and represents the physical body of the camera 154. Its orientation typically aligns with the mount’s coordinate system.• left camera optical link 178 is a child of the left camera link and represents the optical center of the camera 154. In this frame, the Z-axis points along the camera’sDocket No. 61333-PCTviewing direction, the X-axis points to the right in the image, and the Y-axis points down. This frame is crucial for computer vision operations as it aligns with the camera’s perspective.

[0077] The transformation chain from left camera mount link 174 to left camera optical link 178 accounts for both the adjustable mounting angle and any rotational offset between the camera’s body orientation and its optical axis. This setup allows for precise 3D perception and accurate projection of points from the camera’s image plane into the robot’s 3D world space, which is essential for detection actions.Autonomous Robotic Vehicle Localization

[0078] To obtain an accurate odometry source for the autonomous robotic vehicle 100, it is necessary to fuse data from different sensors. Individual sensors may not provide accurate odometry data, or the error in the odometry data may accumulate over time. The robot localization package provides an Extended Kalman Filter as well as an Unscented Kalman filter for fusing odometry data from multiple sources, such as wheel encoders, IMU, and GPS, to estimate the autonomous robotic vehicle’s pose in the world frame. Autonomous Robotic Vehicle Localization

[0079] A pose, in robotics, refers to the combination of position and orientation of an object in 3D space. It typically consists of six degrees of freedom: three for position (x, y, z coordinates) and three for orientation (roll, pitch, yaw angles).

[0080] The robot localization node subscribes to the wheel encoder data and IMU data and fuses them using a Kalman filter to estimate the autonomous robotic vehicle’s pose in the world frame. Two instances of the Extended Kalman Filter are used for the navigation system of the autonomous robotic vehicle 100: one fuses wheel encoder data with IMU, which provides a local odometry source and another fuses the aforementioned local odometry with GPS odometry, providing a global odometry source.

[0081] Although in one embodiment, only the local odometry may be used with no GPS signal. The two Extended Kalman Filters also may create coordinate frames called “odom” for the local odometry and ’’map” for the global odometry.

[0082] In ROS, the odom frame and map frame serve different purposes in autonomous robotic vehicle navigation. The odom frame, derived from the autonomous robotic vehicle’s internal sensors, provides a continuous, high-frequency estimate of theDocket No. 61333-PCTautonomous robotic vehicle’s position relative to its starting point. It is ideal for short-term localization and local navigation tasks but is prone to drift over time. On the other hand, the map frame represents a fixed, global coordinate system, typically aligned with a known map of the environment. It offers a stable, long-term reference frame, useful for global localization and navigation across multiple runs or autonomous robotic vehicles. While the odom frame excels in providing smooth updates for immediate motion planning and obstacle avoidance, the map frame is crucial for consistent long-term navigation and positioning within a known environment. In practice, ROS often utilizes both frames in tandem, with the odom frame handling local, short-term movements and the map frame providing global context and drift correction.

[0083] In one embodiment, the odometry (odom) frame may be selected as the world reference due to its superior suitability for precise, real-time robotic interventions in agricultural settings. While high-precision GPS solutions like RTK (Real-Time Kinematic) can significantly improve global localization accuracy, the odom frame still offers distinct advantages. The odom frame provides high-frequency, locally consistent position updates based on the autonomous robotic vehicle’s internal sensors, which are crucial for detecting and responding to small targets like aphids with minimal latency. Even with RTK corrections, GPS updates typically occur at lower frequencies (often 1-10 Hz) compared to odometry updates (which can exceed 100 Hz), making the odom frame more suitable for real-time control and rapid decision-making. Additionally, the odom frame is immune to GPS signal interruptions that can occur even with RTK systems, such as when operating under dense foliage or near structures. The map frame, while globally more accurate overtime, often incorporates sensor fusion algorithms (e.g., Kalman filters) that can introduce smoothing effects and small, sudden adjustments as new sensor data is integrated. These adjustments, while beneficial for long-term accuracy, can create discontinuities that are undesirable for precise, short-term operations like targeted spraying. In contrast, the odom frame provides a continuous and smooth trajectory, ensuring consistent tracking of the autonomous robotic vehicle’s position relative to recently detected infestations. This local consistency is paramount for the immediate task of identifying and treating infestations in the autonomous robotic vehicle’s vicinity, for example, where relative positioning is more critical than absolute global positioning.Docket No. 61333-PCTFurthermore, the odom frame supports faster, more deterministic real-time transformations, which is essential for responsive spraying actions. While the odom frame may accumulate drift over extended periods, this drift is typically small over short durations or within confined areas, localized operations involved in this precision agriculture application. Therefore, even with the availability of high-precision RTK GPS, the odom frame remains the optimal choice for maintaining the necessary accuracy, responsiveness, and consistency required for targeted actions.Visual Detection Method

[0084] Turning to FIGS. 8A-C, a method 200 of visual detection via the autonomous robotic vehicle 100 will now be described in detail. The method 200 utilizes image processing that leverages hardware acceleration and neural network inferencing to achieve real-time performance in object detection and localization.

[0085] First, a camera driver node 202 publishes raw RGB images and camera information topics, including camera intrinsics data. The camera driver node 202 may include initiating with the DepthAI ROS driver, which is the driver compatible with the OAK-FFC 4P PoE camera, for example.

[0086] The following nodes may be hardware accelerated nodes, leveraging ROS 2 packages such as NVIDIA ISAAC Transport for ROS (NITROS) framework and NVIDIA ISAAC ROS packages. The NITROS framework provides efficient high-speed, low-latency data transfer between these GPU-accelerated nodes.

[0087] A rectify node 204 may then perform image rectification, correcting for lens distortion and preparing the image for subsequent processing steps. The rectify node 204 may subscribe to both raw RGB images and camera information and may publish the rectified image and updated camera information.

[0088] An H264 encoder node 206 may run parallel to the main processing sequence to compress raw image data to H264 format. This compression may enable efficient recording in ROS 2 bags for later replay or dataset creation.

[0089] A resize node 208 may then adjust image dimensions to match input requirements of the neural network model. The resize node 208 may publish the resized image with corresponding camera information.Docket No. 61333-PCT

[0090] A DNN encoder node 210 may then prepare the resized image for neural network inference by converting the resized image to a tensor format and normalizing it based on provided mean and standard deviation parameters.

[0091] One or more inference nodes 212 may then perform inference on the resized image to make detections. This may involve pre-processing steps such as normalizing the image, converting the image to a tensor, and decoding model output (e.g, segmentation mask or detection array. The inference node 212 may execute a detection model using NVIDIA’s tensorRT, for example, for optimized inference on the GPU. The inference node 212 will be described in more detail below.

[0092] A decoder node 214 (e.g., RT-DETR node) may then receive tensors via the transport (e.g., NITROS transport), for example. The decoder node 214 may decode model output into a detection array transferring data back to the CPU for further processing. That is, the decoder node 214 may output detection data to the 2D to 3D point estimation node described below.

[0093] The following nodes operate outside the hardware accelerated transport graph. They do not utilize GPU or Vision Programming Interface (VPA) acceleration.

[0094] A 2D to 3D point estimation node 216 (also referred to as a detection to pose node) may then convert 2D detections to 3D world coordinates. The 2D to 3D point estimation node 216 may estimate localization with respect to the world frame, using the output from the inference node(s) 212. The 2D to 3D point estimation node may utilize transform ( / tf) and static transform ( / tf static) information to convert 2D detections into 3D poses. The 2D to 3D point estimation node 216 will be described in more detail below.

[0095] A sprayer manager node 218 may then process 3D locations and make sprayer decisions (i.e., determinations of when to spray liquid to achieve the desired result). The sprayer manager node 218 may receive estimated locations then determine when to spray based on the autonomous robotic vehicle’s sprayer system 112 approaching the locations to be sprayed. More specifically, the sprayer manager node 218 may receive 3D poses along with transform data and thereby make decisions that result in service requests to the micro-ROS agent described below.

[0096] The method may conclude with a micro-ROS agent 220 interfacing with a the sprayer manager node 218 and running on the microcontroller 142, which may receiveDocket No. 61333-PCTcommands from the sprayer manager node 218 to control physical spraying by the sprayer system 112. The micro-ROS agent 220 may interface directly with the hardware of the sprayer system 112 to activate or deactivate the nozzles 138. In other words, this final stage translates high-level spraying decisions into physical actuation of the nozzles 138.

[0097] The inference node 212 will now be described in more detail in terms of aphid detection on infested crops. To detect aphids in sorghum plants, a neural network architecture may be used to process the images captured by the cameras 154. An objective is to achieve high accuracy in detecting aphids, while also having low inference time to run in real-time on the hardware platform. These two goals are usually conflicting, as more complex models tend to have higher accuracy but also higher inference time.

[0098] Models should be optimized to the specific hardware on which they are deployed. Inference libraries like ONNX Runtime and NVIDIA’s TensorRT may be used in optimizing deep learning models for efficient deployment on various hardware platforms. TensorRT, specifically designed for NVIDIA GPUs, is an SDK that optimizes neural network models for faster and more efficient execution. It employs several techniques to enhance performance, including precision calibration (reducing from FP32 to FP16 or INT8), layer and tensor fusion, kernel auto tuning, and dynamic tensor memory management. These optimizations target key aspects such as inference latency, throughput, memory footprint, and energy efficiency. TensorRT analyzes the computational graph of a model, fusing multiple layers into single operations where possible, selects the most efficient GPU kernels, and eliminates unnecessary operations. In the context of aphid detection on the autonomous robotic vehicle 100, these optimizations enable real-time processing of high-resolution images from multiple cameras 154, allow for the deployment of more complex and accurate models within the computational constraints of an embedded system, and reduce power consumption, ultimately extending the operational time of the autonomous robotic vehicle 100 and increasing the area that can be covered efficiently.

[0099] In one embodiment, the U-Net network architecture may be used. U-Net is a popular convolutional neural network (CNN) architecture for image segmentation tasks. Its main application is for medical image segmentation, but it has been used in otherDocket No. 61333-PCTapplications as well. A PyTorch implementation of U-Net may be trained with an image resolution of 1024x736 and normalized according to the mean and standard deviation of the dataset as indicated by a previous study that utilized the same dataset to perform segmentic segmentation on aphids. The model may be exported to an ONNX file to later be deployed on the Jetson AGX Orin using the inference node 212, with ONNX runtime as its backend. Nvidia Triton is an open-source inference serving software that simplifies the deployment of Al models at scale in production environments. It is designed to help developers deploy trained Al models from any framework (such as TensorFlow, PyTorch, ONNX Runtime, or TensorRT) on any GPU or CPU-based infrastructure (cloud, data center, or edge). The Nvidia Isaac Tensor-RT node may be used, however, since the model may have layers not compatible with Tensor-RT, it may be necessary to use the Nvidia Isaac Triton ROS 2 node with ONNX runtime as a backend, and partial Tensor-RT optimization and Float16 quantization may be applied.

[0100] U-Net Semantic Segmentation may achieve approximately 12 FPS on the Jetson AGX Orin embedded device. While this performance is commendable, higher frame rates may be desired to accommodate increased vehicle speeds and ensure comprehensive aphid detection across all captured frames. Furthermore, a shift from segmentation to object detection may be considered advantageous, as it would enable the estimation of 3D aphid locations from 2D images using methods to be discussed later. Although instance segmentation could have addressed this need, it may be unnecessarily computationally expensive for the task at hand. Consequently, RT-DETR (Real-Time Detection Transformer) may be used for its exceptional balance of inference speed and accuracy. It has impressive performance, with the ResNet101 backbone achieving around 40 FPS and the ResNet50 backbone reaching approximately 60 FPS. These models may be deployed using the NVIDIA Isaac TensorRT ROS 2 node with FP16 quantization, further optimizing inference speed. The substantial improvement in frame rate not only allows for higher rover speeds but also provides a comfortable margin to ensure consistent aphid detection.

[0101] Turning to FIG. 8B, the 2D to 3D point estimation node 216 will now be described in more detail also in terms of aphid detection on infested crops. Depth from 2D images may be estimated assuming a fixed aphid size, as shown in block 216A. ThisDocket No. 61333-PCTmay be driven by practical considerations and certain constraints. While stereo cameras and time-of-flight sensors are common choices for depth perception in robotics, they present certain limitations and complexities that were deemed unsuitable for this application. Stereo cameras have a minimum distance threshold below which depth estimation becomes unreliable.

[0102] Time-of-flight cameras, while capable of providing accurate depth information, introduce additional complexities. Aligning the depth data from a time-of-flight camera with RGB data from a separate camera is a non-trivial task, requiring precise calibration and registration processes. This alignment procedure not only adds to the overall system complexity but also introduces potential points of failure, such as misalignment due to vibrations or environmental factors. Furthermore, specialized hardware that seamlessly integrates RGB and depth data can be prohibitively expensive. Considering these factors and other constraints, a monocular depth estimation approach may be used. This offers a simpler, more robust solution that aligns well with this example’s goal of identifying and treating aphid-infested plants. While this approach assumes a fixed aphid size, which introduces some inherent inaccuracies, it may provide sufficient accuracy for the task of localized spraying if the given fixed aphid size was in the average of a sorghum aphid. The depth estimation is based on the principle of perspective projection, where the relationship between an object’s actual size, its size in the image, and its distance from the camera can be expressed as:

[0103] Z(depth) = (Object Size * Focal Length) / Object Pixel Size.

[0104] This forms the basis of the monocular depth estimation approach for aphid detection. After estimating the depth of the aphid from the camera 154, the aphid’s full 3D position may be determined, as shown in block 216B. To implement this, a custom ROS 2 node may be used. With reference to FIG. 8C, this node performs the following operations:

[0105] Subscription: This node subscribes to a detection 2D array topic, as shown in block 216B1. This contains the output from the decoder node 214, thus providing the 2D bounding boxes of detected aphids.Docket No. 61333-PCT

[0106] Camera Information: This node obtains the camera’s intrinsic parameters from the camera info message, as shown in block 216B2. This is crucial for accurate 3D projections.

[0107] Pixel Projection: The center of the detection bounding box may be projected into 3D space using a projectPixelTo3dRay function from the ROS 2 image geometry package, as shown in block 216B3. This function returns a unit vector in the camera coordinate frame, pointing in the direction of the rectified pixel (u,v) in the image plane. The resulting vector has z=1.

[0108] Unit Vector Computation: Using the image geometry ROS 2 package, this node computes the unit vector in the direction of each detected aphid, as shown in block 216B4.

[0109] Depth Calculation: This node calculates the depth of each aphid with respect to the camera 154 using the formula described earlier, assuming a fixed aphid size, as shown in block 216B5.

[0110] 3D Point Calculation: The unit vector may be scaled with the calculated depth, thereby obtaining the 3D point coordinates of each aphid relative to the camera 154, as shown in block 216B6.

[0111] Coordinate Frame Transformation: The 3D point coordinates are loaded into a pose message. The node then transforms this pose from the camera’s optical frame to the world frame, as shown in block 216B7. This ensures consistency with the overall robotic system’s coordinate system.

[0112] Pose Array Publication: For each aphid in the detection 2D array, its pose with respect to the world frame is appended to a pose array message. This array, containing the 3D positions of all detected aphids, may then be published for use by other nodes in the system, as shown in block 216B8.

[0113] The above sub-steps effectively bridge the gap between 2D image detections and 3D world coordinates, enabling precise localization of aphids in the autonomous robotic vehicle’s operational space. The use of ROS 2’s image geometry package and coordinate frame transformations ensures that the resulting 3D positions are accurate and consistent with the autonomous robotic vehicle’s understanding of its environment.Docket No. 61333-PCT

[0114] The implementation of the 2D to 3D point estimation node 216 demonstrates the integration of computer vision techniques with robotics principles, showcasing how 2D image processing results can be meaningfully translated into actionable 3D information for tasks such as targeted pesticide application.

[0115] While effective, some discrepancies may be observed in the depth estimation. When running the 2D detection to 3D point node in normal mode (without averaging, publishing all poses for each frame), the z-coordinates in the camera frame may show noticeable variations. The estimated positions of aphids may appear spread out in depth, although not significantly. This may be due to the natural variation in aphid sizes, with individual aphids being either smaller or larger than the given fixed size parameter. This size variation directly affects the precision of the depth estimation.

[0116] However, these variations may not substantially impact the application’s effectiveness, as the autonomous robotic vehicle 100 primarily relies on the x-coordinate for determining plant infestation and timing the sprayer activation. The y and z coordinates, while less critical for the spraying decision, still play a role in the overall localization. The y-coordinate is less crucial since the entire plant is treated if infested, regardless of the specific infested area. The z-coordinate’s primary importance lies in scaling a unit vector. To address potential inaccuracies, pre-target and post-target spraying distances in the sprayer manager node 218 can be adjusted to compensate for depth estimation variations.

[0117] It is worth noting that the object size parameters used for depth estimation may be implemented as ROS 2 parameters for easy modification. In the case of aphid detection, typically, a size of 3mm may be used, reflecting the average size of sorghum aphids (2-4mm), with occasional use of 2mm for comparison.

[0118] The autonomous robotic vehicle 100 thus effectively detects and localizes aphids as the plant moves across the camera’s field of view.

[0119] Turning to FIG. 8D, the sprayer manager node 218 will now be described in more detail, also continuing the example of aphid detection on infested crops. The sprayer manager node 218 may orchestrate the precise application of pesticides based on the detected aphid locations. The sprayer manager node 218 may work in conjunction withDocket No. 61333-PCTthe micro-ROS agent 220, running on the microcontroller 142, to translate the 3D localization data into targeted spraying actions.

[0120] The sprayer manager node 218 may subscribe to two pose array topics, one for each camera (left and right), which contain the 3D positions of detected aphids in the world coordinate frame, as shown in block 218A. Upon receiving a pose array message, the sprayer manager node 218 populates a buffer with the individual poses, maintaining separate buffers for the left and right camera detections, as shown in block 218B. This buffering mechanism allows for efficient processing of multiple aphid detections and helps manage the temporal aspects of the spraying operation.

[0121] A dedicated thread within the sprayer manager node 218 continuously processes the buffered poses, as shown in block 218C. It begins by examining the first pose in each buffer, transforming the aphid coordinates from the world frame to the respective sprayer frame (left or right, depending on the source camera). This transformation is critical for determining the aphid’s position relative to the sprayer boom, enabling precise targeting. The node then evaluates whether each aphid is within the predefined spraying bounds. These bounds are determined by two key parameters: the pre-target and post-target distances. The pre-target distance defines the point at which the sprayer system 112 should activate in anticipation of reaching the aphid, while the post-target distance indicates when the sprayer should deactivate after passing the aphid’s position.

[0122] When an aphid is detected within the pre-target and post-target distance, the sprayer manager node 218 interacts with the micro-ROS agent 220. This interaction occurs through a service called / set nozzle status, which accepts a nozzle status message containing nozzle id and duty cycle fields. A simplified approach may be used: the entire left boom may be treated as nozzle id 0, and the right boom may be treated as nozzle id 1. As such, the whole plant may be sprayed if aphids are detected, regardless of their specific location on the plant.

[0123] The micro-ROS agent 220 may be hardcoded to control the three nozzles on each side according to the respective nozzle id. Furthermore, for simplicity, the node interprets any non-zero duty cycle as 40%, effectively creating an on / off control for each boom. When an aphid is within the target range, the sprayer manager node 218 mayDocket No. 61333-PCTactivate all nozzles on the corresponding side (left or right) by setting the appropriate nozzle id with a non-zero duty cycle.

[0124] If an aphid’s position exceeds the post-target distance, indicating that the rover has moved past it, the corresponding pose may be removed from the front of the buffer, and the sprayer manager node 218 may proceed to process the next pose in line. Simultaneously, the sprayer manager node 218 may send a command to turn off the solenoid for that side if no other aphids in the buffer require spraying. Conversely, if aphids remain within the target range, their poses may be retained in the buffer for reassessment in subsequent processing cycles.

[0125] This implementation strikes a balance between precision and practicality. While it doesn’t provide individual nozzle control or variable spray intensity, it ensures that entire plants with detected aphids receive treatment. The separation of left and right boom control allows for some degree of targeted application, reducing unnecessary pesticide use compared to blanket spraying approaches.

[0126] By implementing this decision-making process, the sprayer manager node 218, in conjunction with the micro-ROS agent 220, ensures that pesticide application is both responsive and efficient. It minimizes wastage by spraying only when aphids are detected within the effective range of the sprayer booms 124. This exemplifies the integration of robotic control, real-time decision making, and agricultural precision.Navigable Path Detection

[0127] Turning to FIG. 9, a method of navigable path detection 300 will now be described in more detail. For general purpose navigation of robots using a camera, a costmap is generated using machine learning method called segmentation. This method generates the boundary between the navigable path using polynomials and splines. Instead of using segmentation masks, detecting keypoints or bounding boxes the present invention detects polynomials in image space. This enables better accuracy and lower compute requirement.

[0128] Under the canopy robot navigation has seen a significant demand due to increased interest in agricultural community for crop phenotyping, crop water stress assessment, liquid application, and the like. Traditional GPS fails under canopies, posing challenges. The present invention introduces a novel approach using smooth polynomialsDocket No. 61333-PCTfor row and canopy identification, leveraging image space prediction and a customized loss function to enhance generalization. This method delineates navigation boundaries with polynomial functions in image space and extrapolates seamlessly into real-world navigation. The present invention achieves 1 ms latency on edge devices and is evaluated using the Intersection over Union (IOU) metric, ensuring high accuracy for under-canopy navigation tasks.

[0129] For nearly two decades, GPS-based autonomy has been extensively researched and implemented. RTK and GPS-based systems have been used to guide tractors and harvesters with high accuracy. However, these technologies face significant challenges under dense canopy cover due to GPS multi-pathing errors, which render them less effective in such environments. A prime example is the RoboBotanist project, which encountered navigation difficulties beneath the canopy when leaf coverage interfered with GPS signals.

[0130] To address these challenges, vision-based strategies have been developed. Although these strategies have thus far exhibited deficiencies making them unsuitable for several applications, especially where curved rows or dense canopy situations are involved.

[0131] On the other hand, lane detection has been studied vastly for autonomous driving. There are various methods for detecting the lane marking, notably the ones which achieve fast real-time performance are methods that use polynomial regression-based loss to detect lanes. These techniques effectively identify lanes by modeling them as polynomials precisely as a function of U = F(V) or by piecewise prediction of the polynomials they also attain very low latency compared to segmentation methods.

[0132] The navigable path detection method 300 uses polynomials for row detection. Unlike the lane detection networks, which predict multiple piecewise or U as a function of Y for polynomials for each lane, the method includes predicting two polynomial coefficients (block 302), namely:

[0133] [U (“A), P ("A)].

[0134] If (u, v) are the normalized pixels of the said row, then:

[0135] U ("A) = u and P (“A) = 1.Docket No. 61333-PCT

[0136] These polynomials are third-degree polynomials parameterized by A G [0, 1], This allows the prediction of smooth polynomials in the image space and effectively decouples the prediction of U from V. To overcome the issue of multiple rows, the method includes predicting a plurality (e.g., 10) of polynomial candidates (block 304) and predicting a classification probability for the occurrence of a row (block 306). Moreover, to eliminate the need for Non-Maximum Suppression (NMS), the method includes training the neural network based on Hungarian matching loss (block 308). Assuming that the robot is always on a ground plane, given the height of the camera from the ground plane (after correcting for roll from the IMU), the method includes projecting back the predictions onto the ground plane from image space (block 310). This approach achieves a latency of 1.5ms on Jetson Orin’s deep learning accelerator. Evaluation of accuracy may be performed via the Intersection over Union (loU) metric with trapezoidal integration between the predicted polynomial and the original polynomial. Exemplary results indicate a 0.04 pixel error in U and V. However, the polynomials fit very well for predicting the navigation boundaries.Dataset and Augmentations

[0137] An exemplary exercise for accumulating a robust dataset for training the neural network will now be described. Multiple ROS2 Bags may be recorded while the autonomous robotic vehicle 100 may be manually controlled around and underneath the canopy, capturing images. The labeling process may be conducted using the platform. Each row may be labeled as a polyline, distinguishing between the navigable area and the canopy area. Third-degree polynomials (U (A), V (A)) may be fitted to these points, parameterized by A, which represents the normalized curvature of the (u, v). From the collected dataset, every 25th (approximately 0.8sec) frame may be selected to avoid redundant data points at the same location, resulting in approximately 2000 images. This dataset may then be split into training and testing sets in an 80-20 ratio. The training set may further be divided into training and validation subsets, also in an 80-20 ratio. This means that the final training process may use 64% of the total dataset, with the remaining 16% serving as the validation set.

[0138] Data augmentation techniques using Kornia may then be applied to enhance the robustness of the model. These techniques may include random rotations,Docket No. 61333-PCTflips, random cut out, rain, snow and color jittering, ensuring that the neural network could generalize well to various real-world scenarios. The augmented dataset provides diverse training examples, which is crucial for the model’s performance in different environmental conditions.

[0139] The training of the models may be conducted using the initial 80% of the dataset, ensuring that the neural network has ample data to learn from diverse scenarios. The split and augmentation strategies may maximize the model’s ability to accurately detect and navigate crop rows under varying conditions.Model Architecture

[0140] Input images may be resized to 512x512 pixels and then processed through a backbone network. Among the various backbones tested, the EfficientNet backbone may yield the best results, likely due to its use of Bi-FPNs (Bidirectional Feature Pyramid Networks). Other backbones such as ResNet-18 and RegNet may also be used.

[0141] After passing through the backbone, a global average pooling (GAP) layer may be applied. The resulting features may then be fed into three separate Multi-Layer Perceptrons (MLPs), each comprising two layers. These MLPs may be responsible for predicting the following:• Confidence scores: This MLP may use a sigmoid activation function to output the confidence levels for the detected polynomials.• Polynomial coefficients for U: This MLP, responsible for predicting the coefficients of 10 polynomials representing U, may use an Exponential Linear Unit (ELU) activation function in the first layer and a linear activation function in the final layer.• Polynomial coefficients for P: Similar to the U MLP, this MLP may predict the coefficients for 10 polynomials representing P, also utilizing an ELU activation function in the first layer and a linear activation function in the final layer followed by an activation which satisfies constraints on P which will be discussed below. Predictions

[0142] It may be best to predict polynomials (x, y, z) directly in the camera frame, where z represents depth and y represents height from the image. However, if it is assumed that the camera is positioned on a plane and the height is known, the prediction can be simplified by focusing on (u, v) pixel coordinates and then using calibrationDocket No. 61333-PCTparameters to project and predict x, y. Given the assumptions, the coordinates may be expressed as (U(A), 1 / P(A)), where A parameterizes the polynomial. Since y is a known constant height h, the coordinates in the camera frame can be derived as follows:

[0143] (x, y, z) = (U(A) P(A)) / h, h, (P(A)) / h.

[0144] This approach leverages the known height to project the (u, v) coordinates and predict the spatial (x, y, z) coordinates, simplifying the calculation while maintaining accuracy.Neural Network Learning

[0145] The neural network may be trained on the following loss function:

[0146] L = LU + LV + LC (1)

[0147] where LC is the Binary Cross-Entropy (BCE) Loss between the predictions and the targets. The targets may be selected by a Hungarian matcher, which assigns polynomials in a way that minimizes the total loss. This approach removes the need for Non-Maximum Suppression (NMS).

[0148] The loss function for the polynomial predictions U defined as PolyLoss can be defined as:- / (U(A) -- (2)

[0149]

[0150] where 0(A) is the predicted polynomial and 17(A) is the target.

[0151] The loss function defined as CircleLoss for the polynomial predictions P can be defined as:= / (V(A)P(A) - 1V dX (3)

[00152] °

[0153] Where V (A) is the target and P (A) is the polynomial. This approach may be necessary since there is no closed-form solution for polynomials of degree greater than five. Hence, it is not possible to find the solution analytically.Camera Calibration

[0154] Camera calibration parameters tend to drift over time, impacting the accuracy of information derived from images. To maintain precision, these parameters must be continually estimated. For an under-canopy autonomous robotic vehicle that detects crop rows as splines or polynomials, the camera calibration parameters can be effectivelyDocket No. 61333-PCTestimated using a least squares approach under the assumption that all the rows detected are in the same plane.

[0155] If an autonomous robotic vehicle is traveling beneath the crop canopy, equipped with a camera and sufficient compute power to run a small neural network (such as the case with autonomous robotic vehicle 100), camera calibration parameters can be obtained by the following method.

[0156] A neural network predicts at least two crop rows in pixel coordinates for each incoming image frame. These predictions are then mapped onto the ground plane, utilizing the known distance between the crop rows. By aligning the predictions from the current frame (t1 ) with those from the previous frame (tO), the camera parameters can be estimated for a pinhole camera model.

[0157] The present invention may utilize prior knowledge of the crop row widths to estimate the camera parameters. Once these parameters are estimated rows detected by the neural networks can be projected on to the ground plane with better accuracy hence reducing the need for manually resetting the navigation system by aligning it with the crop row. This can greatly enhance under canopy rover navigation using pure vision-based sensors.Additional Considerations

[0158] Throughout this specification, references to “one embodiment”, “an embodiment”, or “embodiments” mean that the feature or features being referred to are included in at least one embodiment of the technology. Separate references to “one embodiment”, “an embodiment”, or “embodiments” in this description do not necessarily refer to the same embodiment and are also not mutually exclusive unless so stated and / or except as will be readily apparent to those skilled in the art from the description. For example, a feature, structure, act, etc. described in one embodiment may also be included in other embodiments, but is not necessarily included. Thus, the current invention can include a variety of combinations and / or integrations of the embodiments described herein.

[0159] Although the present application sets forth a detailed description of numerous different embodiments, it should be understood that the legal scope of the description is defined by the words of the claims set forth at the end of this patent andDocket No. 61333-PCTequivalents. The detailed description is to be construed as exemplary only and does not describe every possible embodiment since describing every possible embodiment would be impractical. Numerous alternative embodiments may be implemented, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.

[0160] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.

[0161] Certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as computer hardware that operates to perform certain operations as described herein.

[0162] In various embodiments, computer hardware, such as a processor, may be implemented as special purpose or as general purpose. For example, the processor may comprise dedicated circuitry or logic that is permanently configured, such as an application-specific integrated circuit (ASIC), or indefinitely configured, such as an FPGA, to perform certain operations. The processor may also comprise programmable logic or circuitry (e.g., as encompassed within a general-purpose processor or otherDocket No. 61333-PCTprogrammable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement the processor as special purpose, in dedicated and permanently configured circuitry, or as general purpose (e.g., configured by software) may be driven by cost and time considerations.

[0163] Accordingly, the term “processor” or equivalents should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which the processor is temporarily configured (e.g., programmed), each of the processors need not be configured or instantiated at any one instance in time. For example, where the processor comprises a general-purpose processor configured using software, the general-purpose processor may be configured as respective different processors at different times. Software may accordingly configure the processor to constitute a particular hardware configuration at one instance of time and to constitute a different hardware configuration at a different instance of time.

[0164] Computer hardware components, such as communication elements, memory elements, processors, and the like, may provide information to, and receive information from, other computer hardware components. Accordingly, the described computer hardware components may be regarded as being communicatively coupled. Where multiple of such computer hardware components exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the computer hardware components. In embodiments in which multiple computer hardware components are configured or instantiated at different times, communications between such computer hardware components may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple computer hardware components have access. For example, one computer hardware component may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further computer hardware component may then, at a later time, access the memory device to retrieve and process the stored output. Computer hardware components may also initiateDocket No. 61333-PCTcommunications with input or output devices, and may operate on a resource (e.g., a collection of information).

[0165] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, comprise processor-implemented modules.

[0166] Similarly, the methods or routines described herein may be at least partially processor-implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processors may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processors may be distributed across a number of locations.

[0167] Unless specifically stated otherwise, discussions herein using words such as “processing”, “computing”, “calculating”, “determining”, “presenting”, “displaying”, or the like may refer to actions or processes of a machine (e.g., a computer with a processor and other computer hardware components) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0168] As used herein, the terms “comprises”, “comprising”, “includes”, “including”, “has”, “having”, or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus.Docket No. 61333-PCT

[0169] Patent claims stemming from this application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s).

[0170] Although the technology has been described with reference to the embodiments illustrated in the attached drawing figures, it is noted that equivalents may be employed and substitutions made herein without departing from the scope of the invention as recited in any claims ensuing from this provisional patent application.

[0171] Having thus described various embodiments of the invention, what is claimed as new and desired to be protected by Letters Patent includes the following:

Claims

Docket No. 61333-PCTWe claim:

1. A method of operating an autonomous robotic vehicle, the method comprising steps of:receiving a two dimensional (2D) image via a camera of the autonomous robotic vehicle;detecting an object represented in the 2D image via a detection system of the autonomous robotic vehicle;determining three dimensional (3D) world coordinates of the object based on the detecting step via the detection system of the autonomous robotic vehicle; andactivating a component of the autonomous robotic vehicle according to the 3D world coordinates.

2. The method of claim 1 , further comprising a step of rectifying the image.

3. The method of claim 1 , further comprising a step of compressing the image.

4. The method of claim 1 , further comprising a step of resizing the image.

5. The method of claim 4, further comprising a step of converting the resized image to a tensor format and normalizing the resized image.

6. The method of claim 4, wherein the detecting step includes performing inference on the resized image via a neural network.

7. The method of claim 6, further comprising a step of decoding tensor output of the neural network thereby generating detection data, wherein the step of determining 3D world coordinates includes interpreting the detection data.Docket No. 61333-PCT8. The method of claim 1, wherein the step of determining 3D world coordinates includes:receiving data representing a 2D detection bounding box;projecting a center of the detection bounding box into 3D space, thereby receiving a unit vector in a coordinate frame of the camera, the unit vector pointing in a direction of a rectified pixel in an image plane of the detection bounding box;determining the unit vector in a direction of the object;determine a depth of the object with respect to the camera;scaling the unit vector according to the depth of the object, thereby transforming the unit vector into a 3D point including 3D point coordinates representing the object’s position relative to the camera; andtransforming a pose including the 3D point coordinates to the 3D world coordinates.

9. The method of claim 1, further comprising a step of determining a time to activate the component of the autonomous robotic vehicle based on a position of the camera and a position of the component of the autonomous robotic vehicle.

10. The method of claim 9, wherein the step of determining a time to activate the component includes determining whether the object is within bounds associated with the component of the autonomous robotic vehicle.

11. The method of claim 1 , wherein the object is a pest on a plant and the component is a sprayer system.Docket No. 61333-PCT12. A method of navigating an autonomous robotic vehicle between rows of crops, the method comprising steps of:predicting polynomial coefficients for detection of the rows via a navigation system of the autonomous robotic vehicle;predicting a plurality of polynomial candidates via the navigation system of the autonomous robotic vehicle;predicting a classification probability for occurrence of one of the rows via the navigation system of the autonomous robotic vehicle;training a neural network based on a loss function via the navigation system of the autonomous robotic vehicle;projecting the polynomial coefficients, the polynomial candidates, and the classification probability onto a ground plane via the navigation system of the autonomous robotic vehicle; anddirecting the autonomous robotic vehicle to traverse a ground surface according to the projections.

13. The method of claim 12, wherein the step of predicting a plurality of polynomial candidates includes predicting ten polynomial candidates.

14. The method of claim 12, further comprising step of predicting confidence levels for the polynomial candidates.

15. The method of claim 12, wherein the step of predicting polynomial coefficients includes utilizing at least one of an exponential linear unit (ELU) activation function and a linear activation function.

16. The method of claim 12, wherein the step of predicting polynomial candidates includes predicting polynomials directly in a camera frame.Docket No. 61333-PCT17. The method of claim 16, wherein the step of predicting polynomial candidates includes assuming a height of a position of a camera of the autonomous robotic vehicle and utilizing calibration parameters to predict the polynomial candidates.Docket No. 61333-PCT18. An autonomous robotic vehicle comprising:a frame;a camera mounted on the frame;a drive system including:a plurality of wheels configured to traverse a ground surface; and a motor configured to drive at least one of the plurality of wheels;a sprayer system including:a tank mounted on the frame and configured to hold a treatment;a boom mounted on the frame;a nozzle mounted on the boom and configured to deliver the treatment; and a control system configured to direct delivery of the treatment; and a detection system configured to:receive a two dimensional (2D) image via the camera;detect an object represented in the 2D image;determine three dimensional (3D) world coordinates of the object based on the detection; andinstruct the control system of the sprayer system to direct delivery of the treatment according to the 3D world coordinates.

19. The autonomous robotic vehicle of claim 18, the detection system being a navigation and detection system further configured to:predict polynomial coefficients for detection of rows of crops;predict a plurality of polynomial candidates;predict a classification probability for occurrence of one of the rows;train a neural network based on a loss function;project the polynomial coefficients, the polynomial candidates, and the classification probability onto a ground plane; anddirect the autonomous robotic vehicle between the rows according to the projections.Docket No. 61333-PCT20. The autonomous robotic vehicle of claim 19, further comprising a plurality of sensors, the navigation and detection system being further configured to fuse odometry data from the sensors to estimate a pose of the autonomous robotic vehicle in a 3D world coordinate frame.