Device, apparatus and method for agricultural data collection and agricultural operations
Patent Information
- Application Number
- BR112020003713
- Authority / Receiving Office
- BR · BR
- Patent Type
- Patents
- Current Assignee / Owner
- Publication Date
- 2026-08-11
Smart Images

Figure 00000151_0000 
Figure 00000152_0000 
Figure 00000153_0000
Abstract
Description
[001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 688,885, filed June 22, 2018, the disclosure of which is incorporated by reference herein in its entirety (including each appendix attached hereto).
[002] This application also claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 596,506, filed December 8, 2017, disclosure of which is incorporated by reference herein in its entirety (including each appendix attached hereto).
[003] This application also claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 550,271, filed August 25, 2017, disclosure of which is incorporated by reference herein in its entirety (including each appendix attached hereto). STATEMENT OF RESEARCH SPONSORED BY THE FEDERATION
[004] This invention was made with government support under the TERRA-MEPP: DOE-DE-AR0000598 grant by Department of Energy. This invention was also made with government support under DE-AR0000598 granted by the Energy Advanced Research Projects Agency. This invention was also made with government support under DE-AR0000598 granted by the United States Department of Energy; and under 1720695 granted by the National Science Foundation. The government has certain rights to the invention. FIELD OF DISSEMINATION
[005] The disclosure in question generally refers to a Petition 870260066391, dated 06 / 07 / 2026, page 15 / 315 2 / 139 device and a method for collecting agricultural data and agricultural operations. More specifically, several modalities of the disclosure in question relate to robot-based crop stem width estimation (e.g., in a highly disordered field environment). In addition, several modalities of the disclosure in question relate to robot-based phenotyping using deep learning. BACKGROUND
[006] Plant phenotyping is the quantification of the effects of genotypic differences (i.e., differences in genetic makeup) and the environment on the exhibited phenotype (i.e., the appearance and behavior of the plant) [29A] (various references identified here are sometimes referred to by number followed by a letter, e.g., 1A, 2A, 1B, 2B, etc.). According to the According to the Food and Agriculture Organization of the United Nations, large-scale plant phenotyping experiments are a key factor in breeding better crops that are needed to feed a growing population and provide biomass for energy, using less water, land and fertilizers. The need for larger-scale, more comprehensive, and efficient phenotyping has become increasingly pressing recently due to an ever-evolving climate [22A] and demographic changes in rural areas. However, conventional phenotyping methods are mainly limited to manual field measurements, which is laborious, time-consuming, and lacks sufficiency and accuracy. This has created a so-called phenotyping bottleneck in increasing agricultural yield [13A], [2A].
[007] In recent years, several attempts have been made to automate the plant phenotyping process using a wide range of sensors involving multispectral remote sensing. Petition 870260066391, dated 06 / 07 / 2026, page 16 / 315 3 / 139 and hyperspectral, thermal infrared imaging, fluorescence imaging, 3D imaging, tomographic imaging and visible light imaging [19A]. Visible imaging is a practical, energy-efficient and economical way to measure various plant phenotypes. Certain recent approaches [10A], [17A], [23A], [15A], [28A], [6A], [32A], [8A] attempt to model plants using 3D imaging and reconstruction techniques. However, these approaches have typically been tested in simulated environments or in extensively monitored environments, such as a greenhouse. The aforementioned methods and algorithms have not typically been implemented in real agricultural fields, where the level of uncertainty is very high (e.g., due to changes in lighting conditions during different times of day and during different seasons, variation in plant color and size at different growth stages, background clutter, and various other factors). HoyosVillegas et al. (see [16A] V. Hoyos-Villegas, JH Houx, SK Singh and FB Fritschi. Ground-based digital imaging as a tool to assess soybean growth and yield. Crop Science, 54: 1756-1768, 2014. doi: 10.2135 / cropci2013.08.0540) and Chen et al. (see [7A] Yuhao Chen, Javier Ribera, Christopher Boomsma, and Edward Delp. Locating crop plant centers from UAV-based RGB imagery. In the IEEE International Conference on Computer Vision (ICCV), October 2017) attempt to conduct experiments under field conditions and use digital images to evaluate soybean and sorghum, respectively: Hoyos-Villegas et al. [16A] developed a portable digital imaging tool to evaluate soybean yield, while Chen et al. [7A] uses RGB images based on Unmanned Aerial Systems (UAS) to locate sorghum plant centers in a field, but achieves an accuracy of only 64% to 66%.
[008] Although there is enormous potential for deep learning Petition 870260066391, dated 06 / 07 / 2026, page 17 / 315 4 / 139 and computer vision in plant phenotyping, it became clear that the challenges that arise in plant phenotyping differ significantly from the usual tasks addressed by the computer vision community [29A]. In the context of robotic phenotyping, one challenge is the lack of labeled and curated datasets available to train deep networks under realistic field conditions.
[009] Minervini et al. (see [24A] Massimo Minervini, Andreas Fischbach, Hanno Scharr and Sotirios A. Tsaftaris. Annotated fine grain datasets for image-based plant phenotyping. Pattern Recognition Letters, 81:80-89, 2016. ISSN 0167-8655. doi: https: / / doi.org / 10.1016 / j.patrec.2015.10.013. URL http: / / www.sciencedirect.com / science / article / pii / S0167865515003645) provides a dataset of potted rosette plants throughout various growth stages, where each leaf segment of the plants is labeled with a different color (see [21A] M. Minervini, A. Fischbach, H. Scharr and SA Tsaftaris. Plant phenotyping datasets, 2015. URL http: / / www.plantphenotyping.org / datasets). Many recent works have used this dataset to achieve high accuracy in leaf counting and segmentation tasks [1A], [33A], [30A], [9A]. Giuffrida et al. (see [14A] Mario Valerio Giuffrida, Hanno Scharr and Sotirios A. Tsaftaris. ARIGAN: synthetic arabidopsis plants using generative adversarial network. CoRR, abs / 1709.00938, 2017. URL http: / / arxiv.org / abs / 1709.00938) augment this dataset by generating realistic images of rosette plants using generative adversarial networks. Pound et al. (see [27A] Michael Pound, Jonathan A. Atkinson, Darren M. Wells, Tony P. Pridmore, and Andrew P. French. Deep learning for multi-task plant phenotyping. bioRxiv, 2017. doi: 10.1101 / 204552. URL Petition 870260066391, dated 06 / 07 / 2026, page 18 / 315 5 / 139 https: / / www.biorxiv.org / content / early / 2017 / 10 / 17 / 204552) provides a new dataset for wheat ears, analyze the ears, and count them. However, the conditions in this dataset differ significantly from those in the field, with data obtained in the field from a moving robot having a high level of clutter, object similarity, variable sunlight, motion blur, and occlusions.
[010] With regard now in particular to the estimation of stem width, it is noted that the stem width of fuel plants is an important phenotype that determines the biomass content of the plant. Stem width is also important for identifying whether there are any stunted plants in a given field. Despite its importance, it is believed that conventionally there are no efficient field practices for measuring stem width. Conventional practice typically involves trained agronomists manually going into the fields and measuring stems using Vernier calipers (see, for example, FIG. 2). This technique is slow, inaccurate, risky and labor-intensive. Jin and Zakhor (see [18A] Jihui Jin and Avideh Zakhor. Point cloud based approach to stem width extraction of sorghum.(2017) propose an algorithm to estimate stem width from 3D point cloud data collected by a robot equipped with a depth sensor (ToF sensor). Baharav et al. (see [3A] Tavor Baharav, Mohini Bariya and Avideh Zakhor. In situ height and width estimation of sorghum plants from 2.5d infrared images) use 2 infrared cameras mounted on a robot and apply image processing techniques to estimate the height and width of the plants. However, it is believed that none of these existing width estimation algorithms have yet been validated in field settings for accuracy and. Petition 870260066391, dated 06 / 07 / 2026, p. 19 / 315 6 / 139 validity under high disorder and changes in field conditions.
[011] Furthermore, plant population plays a vital role in agricultural systems due to its strong influence on grain yield [1B] - [6B]. Several studies have concluded that grain yield per unit area follows a parabolic function of plant population [1B], [3B], [7B]. In other words, there is an ideal plant population that maximizes grain yield, where the ideal value depends on a number of environmental factors such as nitrogen, soil and precipitation [8B] [10B]. Therefore, accurate measurement of plant population can provide valuable information for estimating grain yield.
[012] The predominant technique in the current industry for counting corn stands in early growth stages is manual counting. This process requires a lot of intensive labor and is subject to errors. In the literature, several approaches to counting corn plants have been proposed. Croppers often employ a mechanical sensor consisting of a spring-loaded rod coupled to a rotating potentiometer [11B]. However, the method is not applicable to early growth stages as it is destructive. Thorp et al. [12B] developed an algorithm to estimate corn plant stand density using hyperspectral aerial images, but this method cannot be used in later growth stages, i.e., when the canopy closes. In contrast to aerial methods, ground-based methods can be used throughout the growing season.In ground-based techniques, the sensors investigated include Lidar [13B], [14B], Time-of-Flight (ToF) cameras [15B] and laser pointers [16B], [17B]. However, these sensors typically do not provide sufficient discrimination between corn and surrounding material and are therefore subject to. Petition 870260066391, dated 06 / 07 / 2026, page 20 / 315 7 / 139 large measurement errors. In particular, they cannot easily differentiate between corn leaves, debris, and weeds, which can trigger a signal similar to that of corn stalks in the measurement. Therefore, studies involving these sensors have been limited to weed-free fields (which is impractical in production fields) or have suffered from severe weed interference errors. Monocular RGB cameras, on the other hand, have the potential to identify corn stalks against a complex background. However, differentiating corn plants in cluttered agricultural environments in the presence of weeds, overlapping leaves, and varying lighting conditions is a highly challenging machine vision problem. Shrestha and Steward calculated an ellipsoidal decision surface in the RGB color space to segment background vegetation in their attempt to count corn stands [18B].Vegetation located beyond a certain distance threshold from the average location was considered weed. This heuristic is highly simplified and not robust to real-world conditions, therefore (it is believed) it has only been demonstrated under low-weed conditions.
[013] Conventional machine learning techniques require considerable domain-specific knowledge to carefully design a feature extractor to transform raw data (e.g., pixel values from an image) into an appropriate feature space where classifiers can detect patterns in the input. In contrast, deep learning is a set of methods that allows for end-to-end training and prediction. Deep learning models receive raw data input and automatically learn representations of the data's internal structure. Typically, a Petition 870260066391, dated 06 / 07 / 2026, page 21 / 315 8 / 139 The deep learning model consists of several modules, each of which slightly increases the level of abstraction of the previous representation. When these modules are used sufficiently, very intricate structures can be learned so that relevant patterns are recognized, while irrelevant variations are suppressed. Deep learning has made tremendous progress in areas that have puzzled traditional machine learning for many years, including image recognition [20B], [22B].
[014] Deep learning methods have shown themselves capable of recognizing complex structures and features in the presence of heavy noise. Today, deep neural networks are approaching human-level image recognition in Internet data [19B] - [22B]. However, it is believed that there is currently no machine vision-based corn stand counting algorithm (using deep learning or not) that is robust to real-world noise, variable lighting conditions, and deployable in real time on an ultra-compact moving robot.
[015] Convolutional neural networks (CNNs or ConvNets) are a class of deep learning methods that process multidimensional array data. The structure of a typical ConvNet consists of several layers of neurons connected serially, and the neurons in each layer are organized into feature maps. However, instead of dense connectivity in fully connected neural networks, the neurons in ConvNets are connected only to a local bundle of their predecessors. These neurons perform a weighted sum (convolution) on the feature maps of the previous layer. The sum is then activated by a non-linear function before moving on to the next layer. All Petition 870260066391, dated 06 / 07 / 2026, page 22 / 315 9 / 139 The neurons in a feature map share the same set of weights (filter), and different feature maps have different filters. The arrangement of local connectivity and weight sharing exploits image features where local groups of values are often highly correlated and local patterns are invariant to location. This architecture gives ConvNets distinct advantages in image recognition. For example, AlexNet [20B], a 5-layer ConvNet, won the 2012 ImageNet Large Scale Visual Recognition Competition (ILSVRC) [25B] with significantly superior performance compared to other competing approaches. Since then, ConvNets have become ubiquitous in various computer vision-related applications [26B] - [30B]. In recent years, several ConvNet architectures have been proposed [19B], [21B], [22B], [31B], [32B]. The best model registers a 3.57% error rate among the top 5 in image classification [19B], with an average human around 5%.
[016] In addition, agricultural systems experience a variety of stressors that reduce yield throughout their life cycle. These stressors include external agents such as diseases, insects, animals, etc., nutrient deficiency stress such as nitrogen, phosphorus deficiency, etc., or stress from external factors such as damage due to heavy field equipment or weather. It is important to detect these stressors quickly so that management tasks to counteract their effects can be informed.
[017] If a vehicle / robot / organism drives / walks / operates on a crop, it is said to damage the crop if its activities leave a permanent and lasting effect on the plant. Petition 870260066391, dated 06 / 07 / 2026, page 23 / 315 10 / 139 that affects its health, reduces its yield, or kills it completely. For example, FIG. 35 shows tire tracks in a field that led to permanent and direct damage due to the passage of heavy equipment. Equipment can also indirectly damage the crop; for example, FIG. 36 shows soil compaction due to heavy equipment, which in this case resulted in zero yield, but in others can significantly reduce yield.
[018] Furthermore, the lack of low-cost, high-throughput technologies for field phenotyping is one of the most significant obstacles affecting crop breeding. As described in this document, phenotyping is the measurement of plants, including simple metrics such as emergence, stem width, and plant height; more sophisticated features such as total biomass, leaf angles, and leaf area index; and complex properties such as hyperspectral reflectance and fluorescence emission spectra. Phenotyping allows seed breeders and agricultural scientists to identify and select genotypes that lead to desirable traits. The inability to collect phenotypic data reliably and inexpensively can create a bottleneck in the progress of seed breeding.This critical gap can significantly harm agricultural yields by limiting the reproduction of crops with higher yield potential—crops that can effectively withstand abiotic stresses such as heat, drought, floods, etc.—and preventing accurate yield forecasting.
[019] In addition, for producers, measuring simple phenotypes, such as crop stand count and biomass, provides an accurate estimate of yield when combined with crop yield models. Simple phenotypes, such as Petition 870260066391, dated 06 / 07 / 2026, page 24 / 315 11 / 139 Plant stand and biomass counts are early and accurate indicators of crop yield. The ability to reliably predict crop yield early in the growing season can have a significant impact on the financial profit a grower can expect. The most significant factor affecting the predictive capacity of statistical and biophysical crop yield models has been the lack of high-resolution agronomic datasets. Collecting this data manually is laborious and expensive. Skilled and willing agricultural labor is declining.
[020] Laboratory phenotyping for cells and other small organisms has been done, as has phenotyping using conveyor belts or other mechanisms in a greenhouse. However, these methods have proven difficult to move into fields for high-throughput phenotyping. The lightbox method of phenotyping involves moving a box with active radiation over a plant to obtain phenotypic information with sensors. However, this method is difficult to move to field settings. Furthermore, it requires its own light source, unlike several modalities disclosed here, which work with ambient light sources. Large-scale beam-mounted phenotyping platforms are available; however, these systems are very expensive and require quite heavy and elaborate infrastructure. This makes it impractical for a wide variety of breeders and producers to use these systems.Large tractor-mounted equipment or equipment mounted on large robots can permanently damage the plant, even killing it, if the tractor / robot / equipment runs over the plant. Furthermore, phenotyping typically requires repeated trips across the field multiple times. Petition 870260066391, dated 06 / 07 / 2026, p. 25 / 315 12 / 139 season; heavy equipment (e.g., tractor-mounted equipment) can compact the soil, which is undesirable for yield. Several software programs and algorithms are available to analyze phenotypic images from different sensors or through remote sensing. However, remote sensing data typically cannot penetrate the canopy with high enough resolution.
[021] Crop system management includes activities such as pruning, mulching, weeding, sampling, cultivation, spraying, seeding, thinning, and tilling. Traditionally, these activities are performed manually, which is labor-intensive, or by means of tractor-pulled devices. Intelligent tractor-pulled devices that can identify and selectively spray chemicals on undesirable plants in crop systems are available. However, tractor-pulled systems are impractical when the crop canopy grows. Furthermore, these heavy systems cause soil compaction, which is undesirable for yield. Additionally, these systems can damage plants when they drive over or brush against them. The ROWBOT system provides a robotic mechanism for nitrogen application. However, systems like the ROWBOT system can damage plants if they drive over them. Koselka et al.It teaches a robotic system for observation and management in vineyards. However, the robot does not guarantee that the cultivated plants will not be damaged during operation or if it passes over them.
[022] Agricultural products are highly commoditized. In the absence of any significant differentiator, farmers are pressured to compete on the prices of their products. The Petition 870260066391, dated 06 / 07 / 2026, page 26 / 315 13 / 139 Increased agricultural yields have created an abundant supply of agricultural products in certain geographies. Grains can be stored for long periods in grain silos. Vegetables and fruits can be transported quickly from their production site to a consumption site at a reasonable cost in many geographies. As a result, producers are under increasing price pressure. The farmer's profit margin is the net profit he can expect to obtain after paying for seeds, inputs (such as fertilizers, pesticides, insecticides, and other chemicals), and management costs, which includes labor costs and equipment financing. The increasing cost pressure makes it important for producers to have access to low-cost observation mechanisms that can inform decisions about whether or not to apply inputs. The cost of a ground-based observation mechanism includes its manufacturing cost, distribution cost, and operating cost.The manufacturing cost depends, among other things, on the material it is constructed from, the propulsion mechanism if it is self-propelled, the complexity of the actuation mechanism for its rotation and traversal of the field, and the electronic components it may use. Heavier robots are typically more expensive to maneuver, as are complex robots that require rack and pinion and other mechanisms to rotate. Heavier robots are also typically difficult to transport, requiring higher transportation and distribution costs. Heavier robots are also typically expensive to operate, as they can easily damage plants if they run over them, they can also damage property or people if they collide with them, and they require more force to pull their weight, leading to increased fuel / electricity costs. Petition 870260066391, dated 06 / 07 / 2026, page 27 / 315 14 / 139 also typically compact the soil and damage plants, which will result in costs due to reduced yield.
[023] A different class of observation mechanism from ground-based robots are aerial robots. There has been aerial observation work of agricultural fields for phenotyping and other agronomic functions to inform field management. Aerial observation can be carried out using unmanned aerial vehicles (e.g., drones), manned aircraft, or satellites. Each platform has unique advantages and disadvantages compared to the others. However, aerial observation only reveals stress symptoms that are visible from the top of the canopy. This may include changes in leaf color, overtly damaged plants, or areas of poor growth. Aerial observation can be conducted using RGB (visual) spectrum images or images with multiple spectra that are not visible to the naked eye.Aerial observation typically fails to reveal canopy characteristics, minor defects, or early indicators of plant stress, especially when plants are just emerging and not easily visible in aerial images. Furthermore, state-of-the-art aerial drones often have very limited endurance, particularly multi-rotor drones (which typically only have an endurance of 10-30 minutes), making it more difficult for them to cover larger areas in detail. Additionally, operating these drones can be expensive due to regulations, limited durability, higher manufacturing costs, and safety concerns.
[024] Emerging airborne detection methods include sensors based on active radiation, including radars, lidars, and sonar sensors. However, in addition to the typically prohibitive costs, these airborne sensors usually lack sufficient resolution. Petition 870260066391, dated 06 / 07 / 2026, page 28 / 315 15 / 139 to reveal the first indicators of stress. These airborne sensors typically have limited range and, as such, can only be used on manned or unmanned aircraft at low altitudes, where the weight, power, and sensor update rate usually make the use of these sensors impractical.
[025] Many plant diseases, insect infestations, and nutrient stresses manifest at the soil-plant interface and are not normally visible through aerial observation-based techniques. Indicators of some stressors are also visible closer to the stem. These early indicators are also not normally detectable through aerial observation. Therefore, exploration close to the ground or under the canopy is essential for the early detection of agricultural stressors.
[026] As a result of the shortcomings of aerial observation and the lack of satisfactory soil exploration robots that are low-cost and do not damage plants, the predominant method for agricultural observation is manual or on foot, where a trained agronomist walks the field. The agronomist may visually inspect plants or take plant or soil samples. Agronomists typically have at least a bachelor's degree and many hold higher degrees. This method is colloquially referred to as "boots-on-the-ground" observation. However, this method of observation is very labor-intensive and can be expensive due to the high costs associated with hiring and supporting agronomists. Furthermore, wet and muddy fields are impractical for an agronomist to traverse, and large fields take a long time to cross. Additionally, there is no data from field observation activity beyond the qualitative reports provided by the agronomist.There has been a growing interest in collecting data under the canopy and near-ground data that can be used. Petition 870260066391, dated 06 / 07 / 2026, page 29 / 315 16 / 139 in data-driven pipelines to inform better agricultural decisions.
[027] This has led to a growing interest in the use of robotic platforms for agricultural exploration. However, there are significant unresolved challenges in long-term robotic agricultural scouting without risking damage to plants in a plurality of crops, geographies, and environments. The common practice for agriculture is row cropping. But row size is variable. Crops also have various growth stages. When they are very young, they are small and in the seedling phase, as they begin to grow the plants grow differently. Some crops, such as grass-based cereals, have been cultivated to grow vertically; but others, such as soybeans and other vine-based crops, tend to grow more horizontally. Row spacing in modern agriculture is also highly variable.For corn in the United States, row spacing can range from 28 inches (71.12 centimeters) to 32 inches (81.28 centimeters), while for soybeans in the United States it ranges from less than 15 inches (35.56 centimeters) to 30 inches (76.2 centimeters). Row spacing also varies in different countries and geographies; for example, in Southeast Asia, soybean row spacing of 6 inches (15.24 centimeters) or less can be found.
[028] Furthermore, emerging agricultural practices, such as polycultures or cover crops, lead to agricultural fields with poorly defined or non-existent rows. Some farms may not have well-defined row spacing if they do not use GPS-guided tractors. The Robotanist is an agricultural phenotyping robot developed by CMU researchers in 2016. The robot Petition 870260066391, dated 06 / 07 / 2026, page 30 / 315 The 17 / 139 is built on a metal frame and carries a number of sensors necessary for highly accurate phenotyping. However, this robot weighs approximately 300 pounds (136.08 kilograms) and could significantly damage crops as it walks over them.
[029] U.S. Patent No. 8,381,501 issued February 26, 2013 describes an agricultural reconnaissance robot for operation in vineyards. It appears that this robot does not guarantee that the cultivated plants will not be damaged during operation or if it passes over them. Furthermore, it appears that the robot does not necessarily have the ability to rotate 180° in a row. In fact, it appears that there are no upper limits on size or weight specifications.
[030] There are commercially available walkers that can fit within some crop rows (28–32 inches (71.12–81.28 centimeters) wide), however, these walkers are not typically designed specifically for agricultural data collection and phenotyping. Clear Path robotics has developed several walkers for outdoor applications. The Clear Path Jackal series starts at around $10,000 for the empty platform with some computing capacity. Additional sensors can be added at an extra cost. The walker typically does not come with the ability to process the sensor data to provide the necessary phenotypic measurement. However, the Jackal is only about 2.5 inches (6.35 centimeters) off the ground, which combined with its relatively high platform cost does not make it an acceptable candidate for agricultural exploration.Husky is a larger walker on the Clear Path, but with a width of 26.4 inches (67.05 centimeters), it is difficult for him to move comfortably in a typical cornfield. Petition 870260066391, dated 06 / 07 / 2026, page 31 / 315 18 / 139 (corn) and cannot move in typical soybean rows. Furthermore, Clear Path walkers typically do not come integrated with the sensors and systems necessary for agricultural measurement tasks such as automated stand counting, stalk angle determination, stalk width estimation, or biomass estimation.
[031] Robotic robots are available from Robotnik, Omron Adept technologies, and ROWBOT. More traditional low-backlash track-based robots are also available from QinetiQ, Naio, Fendth, and iRobot. While many off-the-shelf options are available, these devices typically lack an integrated suite of sensors, autonomous technologies, software, data recording capabilities, and sensor integration modularity for plant phenotyping, specifically plant phenotyping under the canopy.
[032] The ROWBOT system is designed to be inserted into agricultural crop rows for fertilizer application. However, the ROWBOT is heavy, powered by diesel engines, and can permanently damage crops if it runs over them. It appears the system is poorly suited to corn rows. Corn row spacing can range from 28-32 inches (71.12 - 81.28 centimeters), and as mentioned earlier, crop row spacing may be much smaller or nonexistent in prevailing agricultural practices. The system also carries heavy fertilizer tanks. As such, it is not feasible to explore agricultural fields with this system without risking damage to the plants. Furthermore, it is believed that the ROWBOT cannot turn 180° in a row. The system uses contact sensors to locate the edges of crop rows and remain in the middle of the rows. However, these types of sensors do not Petition 870260066391, dated 06 / 07 / 2026, page 32 / 315 19 / 139 works when the cultivated plants are very small. BRIEF DESCRIPTION OF THE DRAWINGS
[033] Reference will now be made to the attached drawings, which are not necessarily drawn to scale and where:
[034] FIG. 1A represents an image 101 of a rosette plant from the dataset provided by Minervini et al. (see [24A] Massimo Minervini, Andreas Fischbach, Hanno Scharr and Sotirios A. Tsaftaris. Annotated fine grain datasets for image-based plant phenotyping. Pattern Recognition Letters, 81:80-89, 2016. ISSN 0167-8655. doi: https: / / doi.org / 10.1016 / j.patrec.2015.10013. URL http: / / www.sciencedirect.com / science / article / pii / S0167865515003645).
[035] FIG. 1B represents an image 111 acquired from a robot (according to a modality) under real field conditions with a high level of clutter, variable sunlight, motion blur and occlusion (as seen in comparison with FIG. 1A, the images obtained by robots are very different from the available datasets).
[036] FIG. 2 represents an image 201 of an example of a conventional practice of manually measuring stem width using vernier calipers (this manual practice is normally complicated, inefficient and inaccurate).
[037] FIG. 3A represents a CAD drawing of an embodiment of a phenotyping robot 300 (sometimes referred to in this document as TerraSentia). This type of robot was used for the acquisition of certain data, as described in this document. As seen in this Figure, the robot includes the following: GPS antenna (reference number 301); Bayspec hyperspectral sensor (reference number 302); Bayspec hyperspectral sensor (facing towards Petition 870260066391, dated 06 / 07 / 2026, page 33 / 315 20 / 139 per side) (reference number 303); Radiator for liquid cooling system (reference number 304); Mounting for Intel RealSense 3D Sensor (reference number 305); Integrated visual sensor (reference number 306); LIDAR sensor (reference number 307); Integrated visual sensor (reference number 308); GPS support for RedEdge multispectral sensor (reference number 309); RedEdge multispectral sensor (reference number 310).
[038] FIG. 3B depicts the TerraSentia 320 robot of a modality moving autonomously through a heavily disordered 30-inch (76.2 centimetre) wide sorghum row.
[039] FIG. 4 represents a front view of the TerraSentia 402 robot in a modality on a 30-inch (76.2 centimeter) track between crop row 1 and crop row 2. This Fig. shows the placement of camera 404, light 406 and LIDAR 408 (in this example, the camera's field of view is 600 and the lateral spacing between the wheels is 14 inches (35.56 centimeters)).
[040] FIG. 5 shows a 501 image of an aerial view of an 80-acre (32.375-hectare) sorghum field (Maxwell Field, Savoy, IL, USA), consisting of 960 sorghum plots of different sorghum genotypes.
[041] FIG. 6 shows a 601 algorithmic structure according to a modality. This algorithmic structure according to a modality includes: (1) Foreground extraction (see Algorithm 1A (shown as pseudocode) provided below); (2) SFM camera motion estimation (see Algorithm 2A (shown as pseudocode) provided below); (3) Lateral distance estimation using LIDAR or other sensors. Petition 870260066391, dated 06 / 07 / 2026, page 34 / 315 21 / 139 range; (4) Width estimation using LIDAR and SFM (see Algorithm 3A (shown as pseudo-code) provided below).
[042] FIG. 7 represents a video frame 701 according to one modality. As can be seen in this figure, window 702 (rectangular marker) is placed on the left side of the video frame. The center of the frame appears blurred, but no blur is visible within the window on the left.
[043] FIGS. 8A-8F represent sequential steps of a modality for foreground extraction from a cropped window of a current video frame (see the pseudocode for Algorithm 1A provided below). As seen in these FIGS. (showing images 801, 803, 805, 807, 809 and 811), these sequential steps include: FIG. 8A - Unprocessed fn window with high clutter; FIG. 8B - Edges of the image from FIG. 8A after Canny edge detection; FIG. 8C - Dilation of the image from FIG. 8B; FIG. 8D - Erosion of the image from FIG. 8C; FIG. 8E - Inversion of the image from FIG. 8D; FIG. 8F - Foreground extraction and smoothing removing unwanted components from the image from FIG. 8E after CCL (connected component labeling) and convex hull approach, respectively.
[044] FIGS. 9A-9D (showing images 901, 903, 905 and 907) represent the foreground smoothing of a modality: FIGS. 9A and 9C represent (for a first image and a second image, respectively) a rough foreground after CCL; FIG. 9B (corresponding to the first image of FIG. 9A) represents a smooth foreground after approximating the convex hull around the object of interest; FIG. 9D (corresponding to the second image of FIG. 9C) represents a smooth foreground after approximating the convex hull around the object of interest.
[045] FIGS. 10A-10D (showing images 1001, 1003, 1005 and Petition 870260066391, dated 06 / 07 / 2026, page 35 / 315 22 / 139 1007) represent the grayscale visualization of dense optical flow using Farneback ([11A] Gunnar Farneback. Two-frame motion estimation based on polynomial expansion. In Proceedings of the 13th Scandinavian Conference on Image Analysis, SCIA'03, pages 363-370, Berlin, Heidelberg, 2003. Springer-Verlag. ISBN 3-540-40601-8. URL http: / / dl.acm.org / citation.cfm?Id=1763974.1764031) algorithm in the direction of the robot's movement (according to a modality) by 5 consecutive windows. Lighter color means higher speed (see Algorithm 2A (shown as pseudocode) provided below).
[046] FIG. 11 represents an image 1101 showing the pixel width ( ) according to a modality in N different locations (N = 8) of a sorghum stem 1102. In this example, the pixel widths (from top to bottom) are 33, 36, 41, 47, 43, 39, 39 and 38 (only two of these pixel widths are identified separately in Fig.).
[047] FIG. 12 represents a 1201 graph showing the distribution of manual measurements, LIDAR Width according to a modality and SFM Width according to a modality in 18 sorghum plots in the Maxwell field.
[048] FIG. 13 represents a graph 1301 showing the variation in manual measurement per plant in lot 17MW0159. 3 measurements were taken per plant for all plants in the lot. Standard deviation: 0.256 inches (0.650 centimeters).
[049] FIGS. 14A-14D represent, according to one embodiment, an estimate of width (in inches) for sorghum: FIG 14A - A typical video frame 1401 for batch 17MW0159 with window 1402 marked. Wsi (dark) and Wli (light) are observed within window 1402. Ws (dark) and Wl (light) are placed in the lower left corner. V: Instantaneous robot speed, D: Petition 870260066391, dated 06 / 07 / 2026, page 36 / 315 23 / 139 FIG. 14B - Distribution of the SfM width in all frames of the video; FIG. 14C - Distribution of the LIDAR width in all frames of the video; FIG. 14D - Distribution of the manual measurement width for all plants in the plot.
[050] FIGS. 15A-15C represent, according to one embodiment, an estimate of width (in inches) for corn: FIG. 15A - A typical video frame 1501 with window 1502 marked. Wsi (dark) and Wli (light) are observed within window 1502. Ws (dark) and Wl (light) are placed in the lower left corner. V: Instantaneous robot speed, D: Distance between the robot and the plant measured by LIDAR); FIG. 15B - Distribution of the width of SfM across all video frames; FIG. 15C - Distribution of LIDAR width across all video frames. x-axis: width in inches, y-axis: normalized frequency.
[051] FIGS. 16A-16C represent, according to one embodiment, an estimate of width (in inches) for hemp: FIG. 16A A typical video frame 1601 with window 1602 marked. Wsi (dark) and Wli (light) are observed within window 1602. Ws (dark) and Wl (light) are placed in the lower left corner. V: Instantaneous robot speed, D: Distance between the robot and the plant measured by LIDAR); FIG. 16B - SfM width distribution across all video frames; FIG. 16C - LIDAR width distribution across all video frames. x-axis: width in inches, y-axis: normalized frequency.
[052] FIGS. 17A-17C represent images 1701, 1703 and 1705 showing several problematic scenarios. FIG. 17A represents a leaf that looks like a stem (similar size, shape, eccentricity and orientation), causing a false detection. FIG. 17B represents a stem almost completely occluded by a leaf, Petition 870260066391, dated 06 / 07 / 2026, p. 37 / 315 24 / 139 a situation in which detection fails. FIG. 17C represents irregular sunlight, which will likely produce inconsistent widths.
[053] FIG. 18 represents an illustrative method 1801 according to an embodiment (this method can operate, for example, in the system of FIG. 21).
[054] FIG. 19 represents an illustrative method 1901 according to an embodiment (this method can operate, for example, in the system of FIG. 21).
[055] FIG. 20 represents an illustrative method 2001 according to an embodiment (this method can operate, for example, in the system of FIG. 21).
[056] FIG. 21 represents an illustrative diagrammatic representation 2100 of a machine in the form of a computer system within which a set of instructions, when executed, can cause the machine to perform any one or more of the methodologies disclosed in this document.
[057] FIGS. 22A-22D represent various illustrations of a ground robot according to an embodiment. An RGB camera mounted on one side of the robot records video while the robot traverses two rows of corn. The camera (in this embodiment) has a 60° field of view and points downwards at 35°. FIG. 22A shows a CAD drawing of the robot (of this embodiment) with the camera facing sideways. FIG. 22B shows the robot (of this embodiment) in a cornfield between two rows. FIG. 22C shows a top view of the robot (of this embodiment). FIG. 22D shows a rear view (of the robot of this embodiment).
[058] FIGS. 23A and 23B represent a schematic illustration of a corn recognition signal (see FIG. 23B) for a particular environment (see FIG. 23A). The ROI (region of interest) Petition 870260066391, dated 06 / 07 / 2026, p. 38 / 315 25 / 139 moves with the camera (see the ROI shown as the vertical rectangle in the middle section of FIG. 23A). As the ROI sweeps across the row of corn plants, the model returns a positive sign when corn is present in the ROI and a negative sign when corn is absent from the ROI.
[059] FIG. 24 represents corn plants whose spacing is less than the ROI width.
[060] FIG. 25 represents a MobileNet architecture. Each convolution layer is followed by batch normalization and ReLU activation. [33B]
[061] FIGS. 26A and 26B represent (according to various modalities) examples of: (a) normal cases (see FIG. 26A) where a single corn appears in the ROI (here, the relative motion between the camera and the corn T « w; and (b) exceptions (see FIG. 26B) where neighboring corn plants are too close to be separated (in this case, T « d + w).
[062] FIG. 27 represents a graph related to a certain Validation (counting corn in the field).
[063] FIG. 28 represents a graph related to a certain Validation (counting corn in the field).
[064] FIGS. 29A and 29B represent examples of near-total obscuration by leaves. The images are not recoverable by the deep learning algorithm, which leads to an underestimation of the population.
[065] FIG. 30 represents a flowchart of a method (according to a modality) for determining the plant population.
[066] FIG. 31 represents a block diagram of an embodiment of an apparatus for determining a plant population for a plant field.
[067] FIG. 32 represents a block diagram of a Petition 870260066391, dated 06 / 07 / 2026, page 39 / 315 26 / 139 type of device for calculating a plant population.
[068] FIG. 33 represents an overview of an adaptive camera angle and robot path control approach according to a modality.
[069] FIGS. 34 and 356 represent crop damage caused by agricultural equipment (FIG. 34 shows that the tractor leaves permanent damage to the stands. FIG. 35 shows soil compaction due to heavy agricultural equipment; compacted soil yields less).
[070] FIG. 36 represents an illustrative embodiment of a device that can operate as an agricultural robot to collect data without damaging crops (this FIG. shows a drawing of the robot with a set of sensors: 3601 - GPS antenna; 3602 - Bayspec hyperspectral sensor; 3603 - Bayspec hyperspectral sensor (facing sideways); 3604 - Radiator for liquid cooling system; 3605 - Support for Intel 3D sensor RealSense; 3606 - Integrated visual sensor; 3607 - LIDAR sensor; 3608 - Integrated visual sensor; 3609 - GPS support for RedEdge multispectral sensor; 3610 - RedEdge multispectral sensor.
[071] FIG. 37 represents illustrative wheel designs that are not used in the robot of FIG. 36 (examples of faulty or undesirable wheel designs - from the left, very low ground clearance and too much pressure, middle resulted in soil excavation and damage to plants, on the right too much slippage causing damage).
[072] FIG. 38 represents a schematic diagram of a cross-section of the agricultural robot of FIG. 36 (this FIG. shows an internal electronic diagram showing how the design houses a set of electronics in a very compact space).
[073] FIG. 39 represents a schematic diagram of several Petition 870260066391, dated 06 / 07 / 2026, page 40 / 315 27 / 139 components of the agricultural robot in FIG. 36 including processors, sensors, motors and a base station.
[074] FIG. 40 represents a schematic top view of the agricultural robot in FIG. 36 illustrating a center of gravity.
[075] FIGS. 41A-41E represent various views of an illustrative embodiment of a device that can operate as an agricultural robot to collect data without damaging crops.
[076] FIGS. 42A-42D represent various views of an illustrative embodiment of a device that can operate as an agricultural robot to collect data without damaging crops (as seen in these FIGS.In one example: the robot body height could be 5.5 inches (13.97 centimeters) (see dimension A); the total width of the robot could be 18 inches (45.72 centimeters) (see dimension B); the width of the motor / shaft could be 3.25 inches (8.25 centimeters) (see dimension C); the height of the robot (from the ground to the top of the body) could be 11.25 inches (28.57 centimeters) (see dimension D); the diameter of the wheel could be 7.5 inches (19.05 centimeters) (see dimension E); the total height of the robot (including GPS mast and antenna) could be 18.25 inches (46.35 centimeters) (see dimension F); the height of the GPS mast and antenna could be 7 inches (17.78 centimeters) (see dimension G); and the ground clearance can be 6 inches (15.24 centimeters) (see dimension “H”).
[077] FIG. 43 represents the tracking performance for a device (according to a modality) that can operate as an agricultural robot to collect data without damaging crops.
[078] FIGS. 44A and 44B represent the tracking performance for a device (according to a modality) that Petition 870260066391, dated 06 / 07 / 2026, page 41 / 315 28 / 139 can operate as an agricultural robot to collect data without damaging crops.
[079] FIG. 45 represents an illustrative embodiment of a device that can operate as an agricultural robot for the application of chemicals. DETAILED DESCRIPTION
[080] As described in this document, these are new algorithms for estimating crop stem width (e.g., under high disorder in an agricultural field using a small mobile robot). The sensors used according to one example are low-cost, consisting of a side-facing monocular RGB camera (ELPUSBFHD01M, USA), a 2D LIDAR (2-D Hokuyo UTM-30LX), and wheel encoders to estimate robot speed. The sensors are mounted on a small (e.g., <15 lb) phenotyping robot (sometimes referred to here as TerraSentia) that can traverse crop rows in an agricultural field using the LIDAR.
[081] This document provides, according to one embodiment, an image processing algorithm designed to extract the foreground in the presence of significant leaf and stem clutter, views of other lines, and variable illumination. The extraction can be based on data from a side-facing USB camera on a moving robot. Using the extracted foreground, an algorithm provided in this document uses the ratio between the estimated speed of the robot from the wheel encoders and the pixel speed of the dense optical flow to estimate depth using a frame-of-motion (SfM) approach (with respect to the wheel encoders, it is noted that GPS does not work well under the crop canopy due to multipath errors and attenuation, therefore, the Petition 870260066391, dated 06 / 07 / 2026, page 42 / 315 29 / 139 encoder speed is an acceptable estimate of robot speed, especially at low speeds when the wheels do not slip excessively. The SfM is adapted as described in this document for phenotyping in cluttered and unstructured field conditions.
[082] As described in this document, this is a validation of several techniques in biomass sorghum fields. Algorithms according to various modalities were compared with manually measured plots available in an 80-acre (32.375-hectare) best-practice field test at a leading university. A trained agronomist used industry-standard practices to measure the average stem width of 18 plots. Several algorithms provided in this document match the agronomist's measurements within the allowable error range defined by a supervisory agency (8%). The width estimate match found was 92.5% (using vision only) and 98.2% (using vision and LIDAR), where the total processing time for running both algorithms is 0.04 seconds per video frame.
[083] It is believed that measurements using various modalities described herein may be potentially more accurate than those of the human agronomist, especially given that the agronomist can typically measure only limited plants and only in a few places, whereas the algorithms provided in this document may be more exhaustive.
[084] Several algorithms presented here are quite general in nature due to their use of fundamental machine vision principles. As such, the algorithms presented here can be used with little modification in other plants, as demonstrated by the experiments described. Petition 870260066391, dated 06 / 07 / 2026, page 43 / 315 30 / 139 below in corn (Zea mays) and hemp (Cannibis) fields. Therefore, the results presented here establish the feasibility of using small autonomous robots for stem width estimation in realistic field environments.
[085] In other examples, the foreground extraction and robot depth-to-line techniques described here can be used to automate other phenotypic measurements.
[086] Reference will now be made to various aspects of an experimental setup according to certain embodiments. More particularly, reference will first be made to a description of the robot according to one embodiment. In this embodiment (see, for example, FIGS. 3A and 3B) a robot that is used for data acquisition is a lightweight (e.g., <15 lb), ultra-compact, 3D-printed autonomous field phenotyping robot (sometimes referred to herein as TerraSentia). The lightness and careful construction of the wheels help prevent permanent damage to any part of the plant, even if the robot accidentally runs over them. The compactness allows the robot to easily traverse between narrow crop rows, especially in maize, hemp, and sorghum, where row spacing of 30” or greater is common. Each wheel of the robot is powered by a separate motor with encoders, and the encoder average values provide a reasonable estimate of the robot speed.The robot traveled in several instances at a speed of about 1.3 feet per second (0.4 m / s). At this speed, the robot covers a single 3m by 3m row of a variety of crops in less than 10 seconds and can cover multiple plots in a reasonable amount of time. Furthermore, at this speed, motion blur was not considered significant in the camera (ELPUSBFHD01M, USA) used on the robot. Cameras with higher frame rates might be more noticeable. Petition 870260066391, dated 06 / 07 / 2026, page 44 / 315 31 / 139 allow for increased robot speed.
[087] Still discussing aspects of the experimental setup, reference will now be made to a camera and light source according to an embodiment. In this mode, as the robot traverses the crop lines, video data is acquired at 90 frames per second with a frame resolution of 640 x 480 pixels by the robot's camera. The acquisition is performed, in this example, with a low-cost RGB monocular digital camera (ELPUSBFHD01M, USA) mounted on the side of the robot's chassis. The camera, in this example, has a field of view of 600. A common and inexpensive LED light source, having a color temperature of 3000K and providing 60 lumens of light, is installed on the robot's chassis near the camera position to ensure ample brightness under the dark sorghum canopy. This is done to prevent the camera firmware from increasing the exposure time, which can cause excessive blurring. In this robot configuration, there was no control over the camera firmware. The camera and light positions are shown in FIG. 4.
[088] Still discussing aspects of the experimental setup, reference will now be made to Data Acquisition according to a modality. In this modality, video data, encoder readings (for instantaneous robot speed estimation) and LIDAR point cloud data (for lateral distance estimation) are acquired at 90 fps, 5 fps and 20 fps, respectively. GPS data is also recorded; however, GPS accuracy varies widely under the canopy, therefore GPS data is not used. There was no inertial measurement unit on the robot in this embodiment. All results discussed in this paper are from an offline setup with data retrieved using WiFi; however, high speed and the Petition 870260066391, dated 06 / 07 / 2026, page 45 / 315 32 / 139 The low computational requirements of the presented algorithms indicate that in another embodiment they could be used on board the robot. In either case, there is little value lost by making off-board estimates, since the data usually needs to be retrieved for other purposes anyway. In this embodiment, Python 2.7.12 with OpenCV 2.4.9.1 was used for the development of all the code.
[089] Continuing the discussion of aspects of the experimental setup, reference will now be made to a description of the sorghum fields that were used. FIG. 5 shows an aerial view of an 80-acre (32.375-hectare) sorghum field (Maxwell Field, Savoy, IL, USA), consisting of 960 sorghum plots of different sorghum genotypes. All sorghum experiments discussed here were conducted at the 80-acre (32.375-hectare) Maxwells Field, Savoy, Illinois, during August to November 2017. The field consists of 960 plots of sorghum measuring 3m x 3m with a row spacing of 30 inches (76.2 centimeters). The aerial image in FIG. 5 shows the marked height difference due to the different genotypes in each plot.
[090] Reference will now be made to an Algorithmic Framework according to various modalities. Two algorithmic approaches for robust estimation of crop stem width under highly uncertain field conditions are presented. The algorithms have been divided into phases. Phase 1 (see those elements above the line marked Phase 1 in FIG. 6) is the same for both algorithms and involves a common image processing algorithm for foreground extraction (see the pseudocode of Algorithm 1A provided below). Phase 2 (see the elements below the line marked Phase 1 and above the line marked Phase 2 in FIG. 6) consists of estimating Petition 870260066391, dated 06 / 07 / 2026, page 46 / 315 33 / 139 depth using motion structure and LIDAR point clouds for each approach, respectively (the final results of “Phase 2” are not separately identified as a Phase). The structure was summarized in FIG. 6 and described in detail below.
[091] Referring now in particular to Phase 1: Foreground Extraction, it is noted that a fixed-size window is defined next to (in this example, on the left side) each video frame and only that region is processed to extract a stem boundary, if present. The window size used in this example was as follows: Width = frame_width / 4; Height = 10 * frame_height / 11 (since the algorithm used in this mode is robust to slight variations in window size, other dimensions can be used). The choice of this size is based on two assumptions valid for most crop fields: (1) The stem width does not exceed 4 inches (10.16 centimeters); (2) The stems never approach closer than 3 inches (7.62 centimeters) from the camera lens without blurring.This window size worked for validation testing, and therefore it is possible to avoid using a variable-size window – this avoids unnecessary and redundant calculations (in another example, a variable-size window could be used; in another example, the window could extend the full height of the frame). The window is positioned laterally (in this example) instead of centrally because the video frames sometimes blur towards the center (for example, due to the small line spacing), despite the high frame rate (e.g., 90 fps) and lighting (see, for example, the light on the robot in FIG. 4). As mentioned, in this example, the window was placed on the left side, but choosing the right side is also acceptable. The placement of window 702 in this embodiment is shown in FIG. 7. Petition 870260066391, dated 06 / 07 / 2026, page 47 / 315 34 / 139 placed in each frame of the raw video is processed to determine if an unoccluded portion of the stem is visible. The remainder of the video frame does not need to be used, and this is done (in this example) to reduce computational overhead. In this mode, discarding the rest of the frame does not cause any loss of valuable information, since the robot traverses the entire row visiting each plant one by one, so that all parts of the frame pass through the chosen window at some point in time. In this mode, a multi-window approach is avoided to reduce the chance of double-counting the same plant multiple times (which would distort the estimated width distribution across the entire plot).
[092] With reference now to FIGS. 8A-8F, these represent the sequential steps of a embodiment for foreground extraction from a clipped window of a current video frame (see the pseudocode of Algorithm 1A provided below). As seen, in this embodiment, these sequential steps include: FIG. 8A Unprocessed fn window with high clutter; FIG. 8B - Edges of the FIG. 8A image after Canny edge detection; FIG. 8C - Dilation of the FIG. 8B image; FIG. 8D - Erosion of the FIG. 8C image; FIG. 8E - Inversion of the FIG. 8D image; FIG. 8F - Foreground extraction and smoothing removing unwanted components of the FIG. 8E image after CCL (connected component labeling) and convex hull approximation, respectively.
[093] In general, the pseudocode for Algorithm 1A can provide foreground extraction of the clipped video frame from the cluttered background as follows: 1: procedure EDGES(fn). Function to return edges (fe) of the cut video frame fn Petition 870260066391, dated 06 / 07 / 2026, page 48 / 315 35 / 139 2: procedure MORPH(fe).fe2: Function to perform dilation, erosion, edge inversion (fe), and return of the first disordered plane (fe2) 3: procedure FOREGROUND (fe2). Function to perform connected component labeling in fe2 to remove clutter and smooth the foreground mask. Returns a clean and smooth full foreground mask.
[094] The pseudo-code details of the AI Algorithm (extraction of foreground from cropped video frame from disordered background) are as follows: Algorithm 1A: Extracting Foreground from Cropped Video Frame with Cluttered Background Petition 870260066391, dated 06 / 07 / 2026, page 49 / 315 36 / 139 1: pHKX'dlirv El)GES( / n) 2: t> fn: cropped window of the nth video frame 3: fh 4- equal izeHist(fn) 4: fg4— Gaussianfílur(fi,) 5: fe4-Canny(fg) 6: return fft> fe: the edges of f„ 1; procedure Morph( / J 8: ÁítI 4- kernel (25 x 3) 9: ker2 4- kernel(15 x 15) I ft fej 4— dilate(fe.ket\) II: / „4- erode(ftj.ker2) 12: fti 255-fee 13: ivturn f<2t> Jf2:After dilation, erosion, inversion 14: procedure Foreground! / r2) 15: CCL(ff2) labels 16: Props 4— Region Properties (Labels) 17: IbMask 4— zeros(size{labels)) 18: D> To contain individual components 19: f ullMask 4— zeros(size(labels)) 20: c> If IbMask passes the test, it is added to fullMask afterwards. 21: while lb = unique(labeis) do 22: > Repeat for each labeled component 23: cond\ 4- props(lb}.size > 500 24: C> Condition 1: size > 500 pixels 25: cond2 4— props(lb}.eccentricity > 0.9 26: cond3 4— (props(lb).orientation > —0.3) and (props(lb).orientation < 0.3) 27: if cWI and cond2 and cond3 then 28: / bMask 4—convexHull(IbMask) 29: full mask 4— full mask + IbMask 30: IbM ask 4—zeros (size (labels)) 31: return fid IM ask t> Returns the foreground
[095] Still with reference to FIGS. 8A-8F (and AI Algorithm): A. Canny Edge Detection: In this embodiment, a Canny edge detection technique (see, for example, [5A] J. Canny. A computational approach to edge detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-8 (6): 679698, November 1986. ISSN 0162-8828. doi: 10.1109 / Petition 870260066391, dated 06 / 07 / 2026, p. 50 / 315 37 / 139 TPAMI.1986.4767851) is used to discover edges as a first step (see, for example, FIG. 8B). This technique was adopted here because of its easy availability and robustness over other available edge detection techniques (such as Sobel, Laplacian (see, for example, [4A] Gary Bradski and Adrian Kaehler. Learning OpenCV: Computer Vision in C++ with the OpenCV Library. O'Reilly Media, Inc., 2nd edition, 2013. ISBN 1449314651, 9781449314651.). B. Morphological Operations: Due to noise and varying lighting conditions, the edges obtained in the previous step are often broken. The edges are closed using morphological dilation followed by erosion (see, for example, [4A] Gary Bradski and Adrian Kaehler. Learning OpenCV: Computer Vision in C++ with the OpenCV Library. O'Reilly Media, Inc., 2nd edition, 2013. ISBN 1449314651, 9781449314651) and inverting the image (see, for example, FIGS. 8C, 8D, and 8E). In this example, a rectangular kernel of size 25 x 3 is used for dilation followed by a kernel size of 15 x 15 for erosion. Using a long rectangular kernel helps restore horizontal width information; vertical gradients become distorted, which is not significant in this mode. C. Component Removal: Connected component tagging (CCL) is performed for stem extraction from a disordered image, as shown in FIG. 8E. This can be performed, for example, using the skimage library in Python (see, for example, [12A] Christophe Fiorio and Jens Gustedt. Two linear time union-find strategies for image processing. Theoretical Computer Science, 154 (2): 165-181, 1996. ISSN 03043975. doi: https: / / doi.org / 10.1016 / 0304-3975(94) 00262-2. URL http: / / www.sciencedirect.com / science / article / pii / Petition 870260066391, dated 06 / 07 / 2026, page 51 / 315 38 / 139 0304397594002622 and Kesheng Wu, Ekow Otoo and [34A] Kenji Suzuki. Optimizing two-pass connected-component labeling algorithms. Standard Anal. Appl., 12 (2): 117-135, February 2009. ISSN 1433-7541. doi: 10.1007 / s10044-008-0109-y. URL http: / / dx.doi.org / 10.1007 / s10044-008-0109-y). A labeled image is obtained in which each white component of FIG. 8E is indexed differently. Using skimage the essential features of each labeled component are measured and objects that do not have the desired stem features are removed. For example, a stem should be larger than the background clutter, have greater eccentricity than the unwanted leaf, and be more upright than the diverging branches. In one example, only the stem components that are cylindrical and projected as long rectangles in the video frames are needed. If an ellipse is approximated around them, the eccentricity of such ellipses is high (> 0.8).In one example, an orientation constraint of 400 is applied to the left and right. These three constraints (size, eccentricity, and orientation) remove most of the background clutter and non-stem components, such as leaves or other objects like road signs or a person's walking shoes. FIG. 8F shows the cleaned image after this step. D. Polygon Approximation: A rough mask of the plant stem is obtained after the above process, which is clean and free of unnecessary components. To smooth the mask, a convex hull is found from these sets of 2D points using the Sklanskys algorithm (see, for example, [31A] Jack Sklansky. Finding the convex hull of a simple polygon. Pattern Recogn. Lett., 1 (2): 79-83, December 1982. ISSN 0167-8655. doi: 10.1016 / 0167-8655(82)90016-2. URL Petition 870260066391, dated 06 / 07 / 2026, page 52 / 315 39 / 139 http: / / dx.doi.org / 10.1016 / 0167-8655(82)90016-2). This is shown in FIGS. 9A-9D: FIGS. 9A and 9C represent (for a first image and a second image, respectively) a rough foreground after CCL; FIG. 9B (corresponding to the first image of FIG. 9A) represents a smooth foreground after approaching the convex hull around the object of interest; FIG. 9D (corresponding to the second image of FIG. 9C) represents a smooth foreground after approaching the convex hull around the object of interest.
[096] Referring now in particular to Phase 2 (see, for example, FIG. 6), a discussion will be directed to camera motion estimation as follows: Structure of motion (SfM) is used in one embodiment to estimate the actual width of stems from the width in pixels. The SfM problem in computer vision is the problem of recovering the three-dimensional structure of a stationary scene from a set of projective measurements, represented as a collection of two-dimensional images, via camera motion estimation.In essence, SfM typically involves three main steps: (1) feature extraction in images and matching those features between images; (2) camera motion estimation (e.g., using relative paired camera positions estimated from the extracted features); and (3) 3D structure recovery using the motion and estimated features (e.g., minimizing the so-called re-projection error) (see, for example, [26A] Onur Ozyesil, Vladislav Voroninski, Ronen Basri and Amit Cantor. A survey on structure from motion. CoRR, abs / 1701.08493,. 2017. URL http: / / arxiv.org / abs / 1701.08493). Steps (1) and (2) mentioned above are used in a modality to determine the average speed of the clean foreground pixel in each Petition 870260066391, dated 06 / 07 / 2026, page 53 / 315 40 / 139 frame. Step (1) above is performed, for example, using Gunnar Farnebäck's algorithm (see, for example, [11A] Gunnar Farnebäck. Two-frame motion estimation based on polynomial expansion. In Proceedings of the 13th Scandinavian Conference on Image Analysis, SCIA'03, pages 363-370, Berlin, Heidelberg, 2003. Springer-Verlag. ISBN 3-540-40601-8. URL http: / / dl.acm.org / citation.cfm?id=1763974.1764031.). This algorithm (see the pseudocode for Algorithm 2A below) calculates the dense optical flux for all points in the window and provides a 2-channel matrix with optical flux vectors (Vx, Vy). FIGS. Figures 10A-10D show the grayscale visualization of the dense optical flow using the Farneback algorithm in the direction of the robot's movement (according to a mode) across 5 consecutive windows. Lighter color indicates higher speed (see Algorithm 2A (shown as pseudocode) provided below).A dense optical approach is adopted instead of sparse optical flow, which involves tracking characteristic points in video frames (such as the Lucas Kanade algorithm (see, for example, [20A] Bruce D. Lucas and Takeo Kanade. An iterative image registration technique with an application to stereo vision. In Proceedings of the 7th International Joint Conference on Artificial Intelligence - Volume 2, IJCAI '81, pages 674-679, San Francisco, CA, USA, 1981. Morgan Kaufmann Publishers Inc. URL http: / / dl.acm.org / citation.cfm?id=l623264.1623280)), due to the lack of distinctly different points on the stems and leaves in a sorghum field. Algorithm 2A Camera motion estimation using dense optical flow (motion structure) Petition 870260066391, dated 06 / 07 / 2026, page 54 / 315 41 / 139 1: procedure o ptflow 2: D> fn'· 3anela for nth video frame, instantaneous robot speed 3: (V'Y, Vv) <— denseOplica!Flow{f ramc\ 10) 4: [> For 10 consecutive frames 5: R <- VR / VS> Vx: movement in the horizontal direction ,. r / .tnrn J? κ Proportion to be used 6· return t\ t>for width calculation Referring now again specifically to Phase 2 (see, for example, FIG. 6), a discussion will be directed to Lateral distance estimation using LIDAR or other range sensors. FIG. 4 shows the LIDAR sensor mounted on the robot in a possible location. LIDAR is a type of active range sensor that uses a scanning mechanism to provide the distance of objects within its range. LIDAR returns are known as a point cloud. Although the range of the LIDAR sensor can be greater than 30m, meaningful readings do not exceed a few meters because the line spacing is (in this example) less than 1m. Frequent occlusion by leaves and weeds interferes with measurements ahead of the track, therefore points very close to the robot are discarded as clutter by the leaves. The LIDAR point cloud can be used to estimate the distance to the line using the points by a weighted average of the distances from the points to the robot's traverse line.
[097] Referring now again in particular to Phase 2 (see, for example, FIG. 6), a discussion will be directed to the Width Estimation of SFM and LIDAR data as follows: A clean white foreground mask on a black background, which has the same length and width as the fixed window, is obtained after Phase 1, which is the mask for the desired stem boundary. N imaginary horizontal lines are drawn on the mask image and for each line, the number of white pixels is Petition 870260066391, dated 06 / 07 / 2026, page 55 / 315 42 / 139 recorded. This gives the width in pixels (HAj at N different locations of the stem (as shown in FIG. 11). The process of estimating the actual width from the pixel width is described in more detail as follows (see also the pseudocode of Algorithm 3A below). Algorithm 3A - Width Estimation I: procedure PiXELWjDTH( / a / / Míwjt, N) 2: > fullMask: Output of Algorithm 1 3; I N 4i while i>0 do 5i lVp.[í] C <wiíWftüePí.x£ / s(co / [j]) 6; > eo / [í]: column numbers for width calculation 7: i (- I - 1 8: return Wí> 9: procedure SFMW[DTH(W7>, / f) ] Algorithm Output 2 12: «-WjtxP 13: 1¥s<-5MWj(Ws(.) 14: return Wj t> Wj: average width of the SFM 15: 16: procedure LIDARW1DTH(Wjl,D,F) 17: > D'. Output of Algorithm 3, ft focal length of the camera 18: W^^Wy.xD / F 19: WL<— return > W / J average width of the LIDAR
[098] Still referring to Phase 2, the width estimation of the SFM approach can operate as follows: To obtain the actual width of the stem Wsi, each WPi must be multiplied by a certain ratio R. The ratio of the average horizontal velocity of the foreground pixel (Vx) calculated as described above (see Camera motion estimation), to the actual instantaneous velocity (Vr) of the robot gives R, which is the desired ratio (Vr is obtained, for example, from encoder readings). Equations IA, 2A and 3A below show the steps for the calculation. Petition 870260066391, dated 06 / 07 / 2026, page 56 / 315 43 / 139 of the SfM width. The width (Ws) obtained for a given window is the average of all Wsi. (IA) (2I) (3A)
[099] This approach of using SfM eliminates the need to use complex sensors such as REALSENSE or stereo cameras to estimate depth information.
[100] Still referring to Phase 2, the width estimation of the LIDAR data approach can operate as follows: The instantaneous distance from the camera to the crop line under consideration (D) is obtained as described above (see Lateral distance estimation using LIDAR or other range sensors). The LIDAR width Wl is calculated according to the Equations 4A and 5A. WLi= WPx D + F (4A)Λ' 2=1 (5A)
[101] As described above, Ws and Wl are the balances of two Petition 870260066391, dated 06 / 07 / 2026, page 57 / 315 44 / 139 algorithmic approaches were presented. The results obtained after validating these estimates are discussed in detail below.
[102] Reference will now be made to certain results associated with various modalities described herein. A significant contribution of the work associated with various modalities described herein lies in the validation of the algorithms presented on sorghum biomass (Sorghum bicolor (L.) Moench) in real fields (all experiments described herein were carried out near the last growth stage of sorghum, when leaf disorder and occlusion are highest among all growth phases).
[103] Reference will now be made specifically to Experiment 1: Comparison with Agronomists. Some of the algorithms presented here were compared with manually measured plots available in the 80-acre (32.375-hectare) experimental sorghum field (Maxwell Fields in Savoy, Illinois). Each 3 x 3 meter plot consists of approximately 50 plants. A trained independent agronomist used industry standard practices to measure the average stem width of 20 plots dispersed throughout the field. The agronomist selected 3 representative plants from each plot and made a manual measurement with calipers of each of these 3 plants. This approach was designed to allow the agronomist to accumulate a reasonable amount of data from large fields within cost and time constraints. Industry practice is to use the average of these 3 readings to represent the average stem width for that plot.On the other hand, a type of robot, as described in this document, traversed those 18 plots and attempted to measure the stem width at various locations on each plant in the plot. The comparison is shown in FIG. 12 (in this Fig., the x-axis represents Plots, the y-axis... Petition 870260066391, dated 06 / 07 / 2026, p. 58 / 315 45 / 139 represents Width (inches), the line labeled A corresponds to Manual Measurement, the line labeled B corresponds to LIDAR, and the line identified as “C” corresponds to “SFM”). The percentage correspondence of the algorithms presented here with manual hand measurements by agronomists, considering all plots, is 78% using LIDAR and 76% using SFM. This discrepancy is not surprising, since the sparse manual measurements by agronomists typically do not reflect the truth, nor do they represent the true nature of plant width distribution. Manual measurements are usually limited by cost and time considerations. They typically do not take into account the fact that plant width varies significantly along its length; therefore, a single measurement does not reflect the actual plant width.Furthermore, the cross-section of the stem is typically elliptical, not circular, therefore the placement of vernier calipers affects the measurement. FIG. 13 (discussed in more detail below) represents the amount of width variation in a single batch (this conventional lack of precision and accuracy due to the high cost of trained manual labor is a proportion of which the industry is pursuing high-throughput robotic phenotyping, as described here with respect to various modalities). Therefore, to evaluate the algorithms described here in relation to a true representation of stem width, Experiment 2 was performed.
[104] Reference will now be made more specifically to this Experiment 2: Comparison with extensive hand measurements. To address the limited measurements obtained by the agronomist mentioned above, extensive hand measurements were carried out on a representative plot (17MW0159) in the sorghum field. 3 measurements were taken Petition 870260066391, dated 06 / 07 / 2026, p. 59 / 315 46 / 139 measurements of different lengths of each plant in a row of the plot consisting of 32 sorghum plants. FIG. 13 shows the variation in measurements of 32 x 3 (in this Fig., the x-axis represents Plant count, the y-axis represents Width (inches), bars A (one of which is labeled) correspond to “Width 1”, bars “B” (one of which is labeled) correspond to “Width 2”, and bars “C” (one of which is labeled) correspond to “Width 3”). The standard deviation of such a distribution is 0.256 inches (0.65 centimeters), clearly showing that averaging only 3 measurements per batch (as discussed above, which is conventional practice) is not very accurate. Even these latter measurements do not represent the “fundamental truth”; however, they are the best available dataset against which a comparison could be made. Table II (below) shows that robot-based algorithms of various modalities match exhaustive hand measurements by 91% when motion framework is used and 98% when vision and LIDAR are used (which is below the 8% tolerance set by a federal oversight agency). FIG.Figure 14A shows a processed video frame for batch 17MW0159 as the robot, according to a traversed modality, moves through the batch. The processing window 1502 after cropping the video frame is marked, indicated as the vertical rectangle on the left side. The dark and light colors of various numbers indicate the width using the motion structure and LIDAR data, respectively. The values within the window are the instantaneous width values (dark: Wsi and light: Wli) corresponding to different 'i's as described above, while the values in the lower left corner are the average values for each method (dark: Ws and light: Wl). All values in the figure are in inches. FIGS. Petition 870260066391, dated 06 / 07 / 2026, page 60 / 315 47 / 139 Figures 14B and 14C show the width distribution for the algorithms presented here (using SFM and LIDAR, respectively), and FIG. 14D shows the fundamental truth of manual measurement. Table 1(A) (below) shows the mean and variance of such measurements. TABLE I (A): Comparison of Mean and Variance of Width Estimates from Multimodal Algorithms and Manual Measurements Values for Lot 17MW0159 LIDAR SFM Manual Average (inches) 0.84528 0.93360 0.86014 Variance (inches) 0.13247 0.20301 0.06623
[105] The results obtained after comparing the width estimate results with the average measurements of batch 17MW0159 are tabulated in Table II (A) (below). The processing time is 0.04 seconds per video frame. TABLE II (A): Comparison of manual measurement with results obtained from Multimodal Algorithms for Lot 17MW0159 Manual LIDAR SFM Width (inches) 0.86014 0.84528 0.93360 % Match 98.27 91.4587
[106] Reference will now be made to a demonstration of generalization to other crops. The presented algorithms of various modalities have been extensively validated on sorghum in several plots consisting of different sorghum varieties, at different times over a period of days and under different climatic conditions (under strong sun and cloud cover). The results in all sorghum cases remain consistent. To further demonstrate robustness and generality, the algorithms have also been applied to other stem crops. Petition 870260066391, dated 06 / 07 / 2026, page 61 / 315 48 / 139 cylindrical: corn and hemp, without algorithmic modifications and some changes in the parameters for edge detection. The results on corn and hemp data obtained from fields in parts of Illinois and Colorado are shown in FIGS. 15A-15C (corn) and FIGS. 16A-16C (hemp). Several details of these FIGS. 15A-15C and 16A-16C are discussed below. The width estimates are within the expected ranges; however, a rigorous establishment of the fundamental truth about these crops was not performed due to a lack of manual resources. Regardless, the results demonstrate that the algorithms of the modalities, based on machine vision principles, generalize well across different crops with little or no change in parameters.
[107] With particular reference now to FIG. 15A, this represents a typical video frame with window 1502 (rectangular marker). This FIG. 15A shows Wsi (the 1st, 3rd, 5th, 7th, 9th and 11th numbers from top to bottom) and Wli (the 2nd, 4th, 6th, 8th, 10th and 12th numbers from top to bottom) annotated within window 1502. This FIG. 15A also shows that Ws (second number from bottom left) and Wl (number from bottom left) are placed in the lower left corner. This FIG. 15A also shows V as the instantaneous speed of the robot. This FIG. 15A also shows D as the distance between the robot and the plant measured by LIDAR. FIG. 15A shows the SfM width distribution across all video frames (x-axis: width in inches; y-axis: normalized frequency). FIG. 15C shows the LIDAR width distribution across all video frames (x-axis: width in inches; y-axis: normalized frequency).
[108] With reference now to FIG. 16A, this represents a typical video frame with window 1602 (rectangular marker). Petition 870260066391, dated 06 / 07 / 2026, p. 62 / 315 49 / 139 This FIG. 16A shows Wsi (the 1st, 3rd, 5th, 7th, 9th, and 11th numbers from top to bottom) and Wli (the 2nd, 4th, 6th, 8th, 10th, and 12th numbers from top to bottom) annotated within window 1602. This FIG. 16A also shows that Ws (second number from the bottom left) and Wl (number from the bottom left) are placed in the lower left corner. This FIG. 16A also shows V as the instantaneous speed of the robot. This FIG. 16A also shows D as the distance between the robot and the plant measured by LIDAR. Figure 16B shows the SfM width distribution across all frames in the video (x-axis: width in inches; y-axis: normalized frequency). Figure 16C shows the LIDAR width distribution across all frames in the video (x-axis: width in inches; y-axis: normalized frequency).
[109] As described in this document, certain results are more than 90% accurate. In addition, certain experiments were designed and evaluated on ground robots. These ground robots typically need to deal with more adverse conditions than those faced by a UAS (however, ground robots may be more desirable for high-throughput phenotyping (as opposed to a UAS), since these ground robots typically have a much closer and more detailed bottom view of the plant canopy) [25A].
[110] Reference will now be made to some challenging situations that can benefit from the use of certain enhancements, as described in this document. FIG. 17A shows a situation where even manual classification of the image as a stem or leaf is difficult. The image shows a leaf, possessing color, size, eccentricity, orientation, and shape identical to that of the stem. Therefore, the algorithm falsely detects this as a stem. FIG. 17B shows an almost completely stem Petition 870260066391, dated 06 / 07 / 2026, page 63 / 315 50 / 139 occluded by a leaf in front; there is no way to detect the stem in this case with the sensors that were used in the modalities described. FIG. 17B shows bright sunlight entering through the dense canopy of sorghum, causing stems and leaves to be only partially illuminated. In this case, only the partial outline is taken into account, leading to an incorrect width estimate. Training a machine learning algorithm could be a way to handle these situations, but it may require thousands of labeled frames of videos taken under extreme field conditions. Another disadvantage of using deep learning is that the algorithms become crop-specific, thus losing generality, unlike the approaches described here. Regardless, within a reasonable climate and crop cover, the results of the robot-based methods modalities should produce much more accurate and richer data than just manual measurements.
[111] As described in this document, these are algorithms for estimating crop stem width in small mobile robots. Stem width is an important phenotype needed by breeders and plant biologists to measure plant growth; however, its manual measurement is typically complicated, inaccurate, and inefficient. Several algorithms presented here use a common image processing core designed to extract the foreground in the presence of significant leaf and stem clutter, views of other rows, and variable illumination, from a side-facing USB camera on a small mobile robot. Using the extracted foreground, one algorithmic approach described here uses robot speed estimates from wheel encoders and structure Petition 870260066391, dated 06 / 07 / 2026, page 64 / 315 51 / 139 movement to estimate depth, while another approach described here employs the use of 2-D LIDAR point clouds to estimate depth. These algorithms were validated against available manual measurements on sorghum (Sorghum bicolor) biomass in real experimental fields. Experiments indicate that both methods are also applicable to other crops with cylindrical stems without significant modifications. As described in this document, the width estimation match in sorghum is 92.5% (using vision) and 98.2% (using vision and LIDAR) when compared to manual measurements by trained agronomists. Thus, the results described here clearly establish the feasibility of using small robots for stem width estimation in realistic field settings. Furthermore, the techniques presented in this document can be used to automate other phenotypic measurements.
[112] As described in this document, these are algorithms that are applicable to real field conditions under high disorder.
[113] As described in this document, these are algorithms that operate without the use of a Hough transform.
[114] As described in this document, several modalities utilize a motion algorithm structure that does not require a depth sensor.
[115] As described in this document, several modalities provide a general and computationally lightweight algorithm for autonomous width estimation of crops with cylindrical stems under highly uncertain field conditions.
[116] As described here, several results were validated with rigorous true soil measurements for sorghum.
[117] As described in this document, several modalities provide enhancements in the field of robotics for applications of Petition 870260066391, dated 06 / 07 / 2026, page 65 / 315 52 / 139 phenotyping.
[118] As described in this document, several modalities provide algorithms that can work directly with data obtained by robots under field conditions.
[119] As described in this document, several modalities provide algorithms for filtering useful content from noisy robot-obtained field images.
[120] As described in this document, several modalities can facilitate the generation of datasets that could feed future machine learning pipelines.
[121] In other modalities, several algorithmic approaches described here can be used to develop a dataset specifically for field conditions (e.g., with masked labels for stems and leaves). This would allow the use of machine learning in the context of phenotype estimation tasks targeting, for example, leaf area, leaf angle, leaf count, and enhanced stem width.
[122] With reference now to FIG. 18, several steps of a method 1801 according to an embodiment are shown. As can be seen in this FIG. 18, step 1803 comprises obtaining video data from a single monocular camera, wherein the video data comprise a plurality of frames, wherein the single monocular camera is attached to a mobile ground robot that is traveling along a track defined by a line of crops, wherein the line of crops comprises a first plant stem and wherein the plurality of frames includes a representation of the first plant stem. Next, step 1805 comprises obtaining robot speed data from one or more encoders, wherein one or more encoders are attached to the mobile ground robot that is traveling along the track. In Petition 870260066391, dated 06 / 07 / 2026, p. 66 / 315 53 / 139 then, step 1807 involves performing foreground extraction on each of the plurality of frames from the video data, where foreground extraction results in a plurality of foreground images. Next, step 1809 involves determining, based on the plurality of foreground images and based on the robot speed data, an estimated width of the first plant stem.
[123] With reference now to FIG. 19, several steps of a method 1901 according to an embodiment are shown. As can be seen in this FIG. 19, step 1903 comprises obtaining video data from a camera, wherein the video data comprise a plurality of frames, wherein the camera is attached to a mobile ground robot that is traveling along a track defined by a line of crops, wherein the line of crops comprises a first plant stem and wherein the plurality of frames includes a representation of the first plant stem. Next, step 1905 comprises performing foreground extraction on each of the plurality of frames from the video data, wherein the foreground extraction results in a plurality of foreground images.Next, step 1907 comprises obtaining sensor data from a light-sensing and ranging (LiDAR) sensor, wherein the LiDAR sensor is attached to the ground-based mobile robot that is traveling along the track defined by the crop row, and wherein the sensor data includes at least a portion of the crop row. Then, step 1909 comprises determining, based on the plurality of foreground images and based on the sensor data, an estimated width of the first plant stem.
[124] With reference now to FIG. 20, several steps of a 2001 method according to an embodiment are shown. As Petition 870260066391, dated 06 / 07 / 2026, p. 67 / 315 Figure 20, 54 / 139, shows that step 2003 involves obtaining video data from a single monocular camera, wherein the video data comprises a plurality of frames, and wherein the plurality of frames includes a representation of a first plant stem in a crop row. Next, step 2005 involves obtaining vehicle speed data from at least one of a plurality of wheels of a mobile vehicle. Next, step 2007 involves performing foreground extraction on each of the plurality of frames of the video data, wherein the foreground extraction results in a plurality of foreground images. Next, step 2009 involves determining, based on the plurality of foreground images and based on the vehicle speed data, an estimated width of the first plant stem.
[125] In another embodiment, a device is provided comprising: a processing system including a processor; and a memory that stores executable instructions which, when executed by the processing system, perform operations, the operations comprising: obtaining video data from a single monocular camera, wherein the video data comprise a plurality of frames, wherein the single monocular camera is attached to a mobile ground robot traveling along a track defined by a line of crops, wherein the line of crops comprises a first plant stem and wherein the plurality of frames includes a representation of the first plant stem; obtaining robot speed data from one or more encoders, wherein one or more encoders are attached to the mobile ground robot traveling along the track; performing foreground extraction on each of the plurality of frames of the video data, wherein the foreground extraction. Petition 870260066391, dated 06 / 07 / 2026, p. 68 / 315 55 / 139 results in a plurality of foreground images; and determine, based on the plurality of foreground images and based on the robot's speed data, an estimated width of the first stem of the plant.
[126] In one example, foreground extraction involves processing, for each of the plurality of frames in the video data, only a fixed-size window that is smaller than each of the plurality of frames.
[127] In another example, the fixed-size window associated with each of the plurality of frames is located off-center in each of the plurality of frames.
[128] In another example, foreground extraction comprises, for each of the plurality of frames of the video data: a first function to perform edge detection; a second function to perform morphological processing; and a third function to perform labeling of connected components.
[129] In another example, determining the estimated width of the first plant stem comprises: determining, based on the plurality of frames, an estimated camera movement using a motion process framework; determining a ratio R, where R = Vr / Vx, where Vr is an instantaneous speed of the robot obtained through the robot speed data and Vx is an average foreground horizontal pixel speed obtained through the motion process framework; determining a first width, in pixels, at a first location of the first plant stem, as represented in a first of the plurality of frames of the video data; and multiplying R times the first width, resulting in a first value.
[130] In another example, the first value is the estimated width. Petition 870260066391, dated 06 / 07 / 2026, page 69 / 315 56 / 139 of the first plant stem.
[131] In another example, the first width is determined as a horizontal width.
[132] In another example, determining the estimated width of the first plant stem further involves: determining a second width, in pixels, at a second location of the first plant stem as represented in the first of the plurality of frames of the video data; multiplying R times the second width, resulting in a second value; and averaging the first value and the second value, resulting in the estimated width of the first plant stem.
[133] In another example, the mobile ground robot comprises at least one wheel and one or more encoders determine the robot speed data by detecting a rotation of at least one wheel.
[134] In another example, the operations further comprise: obtaining additional video data from the single monocular camera, wherein the additional video data comprise an additional plurality of frames, wherein the crop row comprises a second plant stem and wherein the additional plurality of frames includes another representation of the second plant stem; obtaining additional robot speed data from one or more encoders; performing additional foreground extraction on each of the additional plurality of frames from the additional video data, wherein the additional foreground extraction results in an additional plurality of foreground images; and determining, based on the additional plurality of foreground images and based on the additional robot speed data, an estimated additional width of the second plant stem.
[135] In another form, a readable storage medium Petition 870260066391, dated 06 / 07 / 2026, p. 70 / 315 57 / 139 by non-transient computer is provided comprising executable instructions which, when executed by a processing system including a processor, perform operations, the operations comprising: obtaining video data from a camera, wherein the video data comprises a plurality of frames, wherein the camera is attached to a mobile ground robot that is traveling along a track defined by a line of crops, wherein the line of crops comprises a first plant stem and wherein the plurality of frames includes a representation of the first plant stem; performing foreground extraction on each of the plurality of frames of the video data, wherein the foreground extraction results in a plurality of foreground images;Obtaining sensor data from a light-sensing and ranging sensor (LiDAR), wherein the LiDAR sensor is attached to a mobile ground robot that is traveling along the track defined by the crop row and wherein the sensor data includes at least a portion of the crop row; and determining, based on the plurality of foreground images and based on the sensor data, an estimated width of the first plant stem.
[136] In one example, foreground extraction comprises, for each of the plurality of frames of the video data: a first function to perform edge detection; a second function to perform morphological processing; and a third function to perform connected component labeling.
[137] In another example, determining the estimated width of the first plant stem comprises: determining, based on sensor data, an estimated distance D from the camera to the first plant stem; determining a first width, in pixels, at a first location of the first plant stem, as Petition 870260066391, dated 06 / 07 / 2026, page 71 / 315 58 / 139 represented in a first of the plurality of frames of the video data; and multiply the first width times D divided by a focal length of a camera lens, resulting in a first value.
[138] In another example, the first value is the estimated width of the first plant stem.
[139] In another example, the camera is a single monocular camera.
[140] In another example, the operations further comprise: obtaining additional video data from the camera, wherein the additional video data comprise an additional plurality of frames, wherein the crop row comprises a second plant stem and wherein the additional plurality of frames includes another representation of the second plant stem; performing additional foreground extraction on each of the additional plurality of frames from the additional video data, wherein the additional foreground extraction results in an additional plurality of foreground images; obtaining additional sensor data from the LiDAR sensor, wherein the additional data include at least one additional portion of the crop row; and determining, based on the additional plurality of foreground images and based on the additional sensor data, an estimated additional width of the second plant stem.
[141] In another example, the LiDAR sensor comprises a 2-D LiDAR sensor.
[142] In another embodiment, a mobile vehicle is provided comprising: a body; a single monocular camera attached to the body; a plurality of wheels fixed to the body; a processing system including a processor; and a memory that stores executable instructions which, when executed by the processing system, perform operations, the operations Petition 870260066391, dated 06 / 07 / 2026, page 72 / 315 59 / 139 comprising: obtaining video data from a single monocular camera, wherein the video data comprise a plurality of frames, and wherein the plurality of frames includes a representation of a first plant stem in a crop row; obtaining vehicle speed data from at least one of the wheels; performing foreground extraction on each of the plurality of frames from the video data, wherein the foreground extraction results in a plurality of foreground images; and determining, based on the plurality of foreground images and based on the vehicle speed data, an estimated width of the first plant stem.
[143] In another example, the mobile vehicle is an autonomous ground mobile robot and operations are performed without the use of global positioning system (GPS) data.
[144] In another example, the mobile vehicle further comprises at least one encoder, wherein at least one encoder obtains vehicle speed data and wherein the single monocular camera is located on the body, on the body or on any combination thereof.
[145] From the descriptions in this document, it would be evident to a person skilled in the art that the various embodiments can be modified, reduced, or improved without departing from the scope and spirit of the claims described below. For example, the mobile vehicle may comprise an airborne vehicle (e.g., drone, airplane, helicopter, or the like). This airborne vehicle may travel, for example, below the crop canopy. Other suitable modifications may be applied to the disclosure in question. Therefore, the reader is directed to the claims for a fuller understanding of the breadth and scope of the disclosure. Petition 870260066391, dated 06 / 07 / 2026, page 73 / 315 60 / 139 of the subject.
[146] FIG. 21 describes an exemplary diagrammatic representation of a machine in the form of a 2100 computer system within which a set of instructions, when executed, can cause the machine to execute any one or more of the methods discussed herein. In some embodiments, the machine can be connected (e.g., using a network) to other machines. In a network deployment, the machine can operate in the capacity of a server or a client user machine in a client-server user network environment, or as a peer-to-peer machine in a peer-to-peer (or distributed) network environment.
[147] The machine may comprise a server computer, a client user computer, a personal computer (PC), a tablet PC, a smartphone, a laptop computer, a desktop computer, a control system, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or not) that specify actions to be taken by that machine. It will be understood that a communication device of the disclosure in question broadly includes any electronic device that provides voice, video or data communication. Furthermore, although a single machine is illustrated, the term machine should also be understood as including any collection of machines that individually or collectively execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.
[148] The 2100 computer system may include a 2102 processor (for example, a central processing unit (CPU), a graphics processing unit (GPU or both), a 2104 main memory and a 2106 static memory, Petition 870260066391, dated 06 / 07 / 2026, page 74 / 315 61 / 139 that communicate with each other by means of a bus 2108. The computer system 2100 may also include a video display unit 2110 (for example, a liquid crystal display (LCD), a flat panel or a solid-state display). The computer system 2100 may include an input device 2112 (for example, a keyboard), a cursor control device 2114 (for example, a mouse), a disk drive unit 2116, a signal generation device 2118 (for example, a loudspeaker or remote control) and a network interface device 2120.
[149] The disk drive unit 2116 may include a tangible computer-readable storage medium 2122 in which one or more instruction sets (e.g., software 2124) incorporating any one or more of the methods or functions described herein, including the methods illustrated above, are stored. The instructions 2124 may also reside, wholly or at least partially, within main memory 2104, static memory 2106, and / or within the processor 2102 during their execution by the computer system 2100. Main memory 2104 and the processor 2102 may also constitute tangible computer-readable storage media.
[150] Dedicated hardware implementations, including but not limited to application-specific integrated circuits, programmable logic arrays, and other hardware devices, may likewise be constructed to implement the methods described in this document. Applications that may include appliances and systems of various embodiments broadly encompass a variety of electronic and computer systems. Some embodiments implement functions in two or more specific interconnected hardware modules or signaling devices. Petition 870260066391, dated 06 / 07 / 2026, page 75 / 315 62 / 139 control and related data communicated between and through the modules, or as portions of an application-specific integrated circuit. Thus, the example system is applicable to software, firmware, and hardware implementations.
[151] According to various embodiments of the disclosure in question, the methods described in this document are intended for operation as software programs running on a computer processor. In addition, software implementations may include, but are not limited to, distributed processing or distributed component / object processing, parallel processing or virtual machine processing may also be constructed to implement the methods described in this document.
[152] Although tangible computer-readable storage medium 2122 is shown in an example embodiment as a single medium, the term tangible computer-readable storage medium should be considered as including a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instruction sets. The term tangible computer-readable storage medium should also be understood as including any non-transient medium that is capable of storing or encoding an instruction set for execution by the machine and that causes the machine to execute any one or more of the methods of disclosure of the subject.
[153] The term tangible computer-readable storage medium should therefore be considered as including, but not limited to: solid-state memories, such as a memory card or other package housing one or more read-only (non-volatile) memories, random-access memories or other Petition 870260066391, dated 06 / 07 / 2026, page 76 / 315 63 / 139 rewritable (volatile) memories, a magneto-optical or optical medium, such as a disk or tape, or other tangible medium that can be used to store information. Therefore, the disclosure is considered to include any one or more tangible computer-readable storage media, as listed in this document and including recognized equivalents in the art and successor media, in which the software implementations in this document are stored.
[154] Although the present specification describes components and functions implemented in embodiments with reference to particular standards and protocols, disclosure is not limited to such standards and protocols. Each of the standards for the Internet and other packet-switched network transmission (e.g., TCP / IP, UDP / IP, HTML, HTTP) represents examples of the state of the art. Such standards are superseded from time to time by faster or more efficient equivalents having essentially the same functions. Wireless standards for device detection (e.g., RFID), short-range communications (e.g., Bluetooth, WiFi, Zigbee), and long-range communications (e.g., WiMAX, GSM, CDMA) are contemplated for use by the 2200 computer system.
[155] The illustrations of modalities described in this document are intended to provide a general understanding of the structure of various modalities, and are not intended to serve as a complete description of all the elements and features of devices and systems that can make use of the structures described herein. Many other modalities will be evident to those skilled in the art after reviewing the description herein. Other modalities may be used and derived from these, so that structural and logical substitutions and changes can be made without departing from them. Petition 870260066391, dated 06 / 07 / 2026, page 77 / 315 64 / 139 of the scope of this disclosure. The figures are also merely representative and cannot be drawn to scale. Certain proportions may be exaggerated, while others may be minimized. Consequently, the descriptive report and drawings should be considered in an illustrative and not restrictive sense.
[156] Although specific embodiments have been illustrated and described herein, it should be appreciated that any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments. Combinations of the various embodiments and other embodiments not specifically described in this document will be evident to those skilled in the art upon review of the above description. References 1A-34A:
[157] [1A] Shubhra Aich and Ian Stavness. Leaf counting with deep convolutional and deconvolutional networks. CoRR, abs / 1708.07570, 2017. URL http: / / arxiv.org / abs / 1708.07570.
[158] [2A] Jose Luis Araus and Jill E Cairns. High-yielding field phenotype: the new frontier of crop breeding. Trends in Plant Science, 19 (1): 52-61, 2014.
[159] [3A] Tavor Baharav, Mohini Bariya and Avideh Zakhor. In situ height and width estimation of sorghum plants from 2.5d infrared images.
[160] [4A] Gary Bradski and Adrian Kaehler. Learning OpenCV: Computer Vision in C++ with the OpenCV Library. O'Reilly Media, Inc., 2nd edition, 2013. ISBN 1449314651, 9781449314651.
[161] [5A] J. Canny. A computational approach to edge detection. IEEE Transactions in Pattern Analysis and Intelligence Petition 870260066391, dated 06 / 07 / 2026, p. 78 / 315 65 / 139 of Machine, PAMI-8 (6): 679-698, November 1986. ISSN 01628828. doi: 10.1109 / TPAMI.1986.4767851.
[162] [6A] Ayan Chaudhury, Christopher Ward, Ali Talasaz, Alexander G. Ivanov, Mark Brophy, Bernard Grodzinski, Norman PA Huner, Rajni V. Patel, and John L. Barron. Machine vision system for 3d plant phenotyping. CoRR, abs / 1705.00540, 2017. URL http: / / arxiv.org / abs / 1705. 00540.
[163] [7A] Yuhao Chen, Javier Ribera, Christopher Boomsma, e Edward Delp. Locating crop plant centers from uavbased rgb imagery. Na Conferência Internacional IEEE sobre Visão Computacional (ICCV), outubro de 2017.
[164] [8A] Sruti Das Choudhury, Saptarsi Goswami, Srinidhi Bashyam, A. Samal e Tala N. Awada. Automated stem angle determination for temporal plant phenotyping analysis.
[165] [9A] Andrei Dobrescu, Mario Valerio Giuffrida e Sotirios A. Tsaftaris. Leveraging multiple datasets for deep leaf counting. CoRR, abs / 1709.01472, 2017. URL http: / / arxiv.org / abs / 1709.01472.
[166] [10A] Lingfeng Duan, Wanneng Yang, Chenglong Huang e Qian Liu. A novel machine-vision-based facility for the automatic evaluation of yield-related traits in rice. Plant Methods, 7 (1): 44, dezembro de 2011. ISSN 1746-4811. doi: 10.1186 / 1746-4811-7-44. URL https: / / doi.org / 10.1186 / 17464811-7-44.
[167] [11A] Gunnar Farnebeack. Two-frame motion estimation based on polynomial expansion. Em Proceedings of the 13th Scandinavian Conference on Image Analysis, SCIA'03, páginas 363370, Berlin, Heidelberg, 2003. Springer-Verlag. ISBN 3-54040601-8. URL http: / / dl.acm.org / citation.cfm? Id = 1763974.1764031. Petição 870260066391, de 06 / 07 / 2026, pág. 79 / 315 66 / 139
[168] [12A] Christophe Fiorio e Jens Gustedt. Two linear time union-find strategies for image processing. Theoretical Computer Science, 154 (2): 165-181, 1996. ISSN 0304-3975. doi: https: / / doi.org / 10.1016 / 0304-3975(94) 00262-2. URL http: / / www.sciencedirect.com / science / article / pii / 0304397594002622.
[169] [13A] Robert T Furbank e Mark Tester. Phenomics- technologies to relieve the phenotyping bottleneck. Trends in plant science, 16 (12): 635-644, 2011.
[170] [14A] Mario Valerio Giuffrida, Hanno Scharr e Sotirios A. Tsaftaris. ARIGAN: synthetic arabidopsis plants using generative adversarial network. CoRR, abs / 1709.00938, 2017. URL http: / / arxiv.org / abs / 1709.00938.
[171] [15A] Mahmood R. Golzarian, Ross A. Frick, Karthika Rajendran, Bettina Berger, Stuart Roy, Mark Tester e Desmond S. Lun. Accurate inference of shoot biomass from high-throughput images of cereal plants. Plant Methods, 7 (1): 2, fevereiro de 2011. ISSN 1746-4811. doi: 10.1186 / 1746-4811-7-2. URL https: / / doi.org / 10.1186 / 1746-4811-7-2.
[172] [16A] V. Hoyos-Villegas, JH Houx, SK Singh e FB Fritschi. Ground-based digital imaging as a tool to assess soybean growth and yield. Crop Science, 54: 1756-1768, 2014. doi: 10.2135 / cropci2013.08.0540.
[173] [17A] Mayuko Ikeda, Yoshitsugu Hirose, Tomonori Takashi, Yosuke Shibata, Takuya Yamamura, Toshiro Komura, Kazuyuki Doi, Motoyuki Ashikari, Makoto Matsuoka e Hidemi Kitano. Analysis of rice panicle traits and detection of qtls using an image analyzing method. Breeding Science, 60 (1): 55-64, 2010. doi: 10.1270 / jsbbs.60.55.
[174] [18A] Jihui Jin e Avideh Zakhor. Point cloud based Petição 870260066391, de 06 / 07 / 2026, pág. 80 / 315 67 / 139 approach to stem width extraction of sorghum. 2017.
[175] [19A] Lei Li, Qin Zhang e Danfeng Huang. A review of imaging techniques for plant phenotyping. Sensors, 14 (11): 20078-20111, 2014. ISSN 1424-8220. doi: 10.3390 / s141120078. URL http: / / www.mdpi.com / 1424-8220 / 14 / 11 / 20078.
[176] [20A] Bruce D. Lucas e Takeo Kanade. An iterative image registration technique with an application to stereo vision. Em Proceedings of the 7th International Joint Conference on Artificial Intelligence - Volume 2, IJCAI '81, páginas 674-679, San Francisco, CA, EUA, 1981. Morgan Kaufmann Publishers Inc. URL http: / / dl.acm.org / citation.cfm? id = 1623264.1623280.
[177] [21A] M. Minervini, A. Fischbach, H. Scharr e SA Tsaftaris. Plant phenotyping datasets, 2015. URL http: / / www.plant-phenotyping.org / datasets.
[178] [22A] M. Minervini, H. Scharr e SA Tsaftaris. Image analysis: The new bottleneck in plant phenotyping [applications corner]. IEEE Signal Processing Magazine, 32 (4): 126-131, julho de 2015. ISSN 1053-5888. doi: 10.1109 / MSP.2015.2405111.
[179] [23A] Massimo Minervini, Mohammed Abdelsamea e Sotirios A. Tsaftaris. Image-based plant phenotyping with incremental learning and active contours. 23: 35-48, 09 2014.
[180] [24A] Massimo Minervini, Andreas Fischbach, Hanno Scharr and Sotirios A. Tsaftaris. Finely-grained annotated datasets for image-based plant phenotyping. Pattern Recognition Letters, 81:80 - 89, 2016. ISSN 0167-8655. doi: https: / / doi.org / 10.1016 / j.patrec.2015.10. 013. URL http: / / www.sciencedirect.com / science / article / pii / S016786551500 3645.
[181] [25A] T. Mueller-Sim, M. Jenkins, J. Abel, and G. Kantor. The robotanist: A ground-based agricultural robot for Petition 870260066391, dated 06 / 07 / 2026, page 81 / 315 68 / 139 High-throughput crop phenotyping. In 2017 IEEE International Conference on Robotics and Automation (ICRA), pages 3634-3639, May 2017. doi: 10.1109 / ICRA.2017.7989418.
[182] [26A] Onur Ozyesil, Vladislav Voroninski, Ronen Basri and Amit Singer. A survey on structure from motion. CoRR, abs / 1701.08493, 2017. URL http: / / arxiv.org / abs / 1701. 08493.
[183] [27A] Michael P Pound, Jonathan A. Atkinson, Darren M. Wells, Tony P. Pridmore e Andrew P. French. Deep learning for multi-task plant phenotyping. bioRxiv, 2017. doi: 10.1101 / 204552. URL https: / / www.biorxiv.org / content / cedo / 2017 / 10 / 17 / 204552.
[184] [28A] David Rousseau e Henricus J. Van de Zedde. Counting leaves without “finger-counting” by supervised multiscale frequency analysis of depth images from top view, 01 2015.
[185] [29A] Hanno Scharr, Tony P. Pridmore e Sotirios A. Tsaftaris. Editorial: Computer vision problems in plant phenotyping, cvppp 2017 - introduction to the cvppp 2017 workshop papers. Na Conferência Internacional IEEE sobre Visão Computacional (ICCV), outubro de 2017.
[186] [30A] Kyle Simek e Kobus Barnard. Gaussian process shape models for bayesian segmentation of plant leaves. Em H. Scharr SA Tsaftaris e T. Pridmore, editores, Proceedings of the Computer Vision Problems in Plant Phenotyping (CVPPP), páginas 4.1-4.11. BMVA Press, setembro de 2015. ISBN 1-901725-55-3. doi: 10.5244 / C.29.CVPPP.4. URL https: / / dx.doi.org / 10.5244 / C.29.CVPPP.4.
[187] [31A] Jack Sklansky. Finding the convex hull of a simple polygon. Pattern Recogn. Lett., 1 (2): 79-83, dezembro de 1982. ISSN 0167-8655. doi: 10.1016 / 0167-8655 (82) 90016-2. URL Petição 870260066391, de 06 / 07 / 2026, pág. 82 / 315 69 / 139 http: / / dx.doi.org / 10.1016 / 0167-8655(82)90016-2.
[188] [32A] Siddharth Srivastava, Swati Bhugra, Brejesh Lall e Santanu Chaudhury. Drought stress classification using 3d plant models. CoRR, abs / 1709.09496, 2017. URL http: / / arxiv.org / abs / 1709.09496.
[189] [33A] Jordan Ubbens, Mikolaj Cieslak, Przemyslaw Prusinkiewicz e Ian Stavness. The use of plant models in deep learning: an application to leaf counting in rosette plants. Plant Methods, 14 (1): 6, 2018. ISSN 1746-4811. doi: 10.1186 / s13007-018-0273-z. URL https: / / doi.org / 10.1186 / s13007-018-0273z.
[190] [34A] Kesheng Wu, Ekow Otoo e Kenji Suzuki. Optimizing two-pass connected-component labeling algorithms. Padrão Anal. Appl., 12 (2): 117-135, fevereiro de 2009. ISSN 1433-7541. doi: 10.1007 / s10044-008-0109-y. URL http: / / dx.doi.org / 10.1007 / s10044-008-0109-y.
[191] Reference will now be made to robot-based phenotyping using deep learning according to one or more modalities. As described in this document, it is an algorithm that can estimate in real time the count of plant stands from image sequences obtained from a side-facing camera on an ultracompact ground robot. This algorithm is demonstrated in the challenging problem of counting maize (Zea mays or corn) plants under field conditions. In other modalities, the algorithm can be used to count other plants, including, but not limited to, sorghum, wheat, soybeans, or vegetables. The algorithm leverages a state-of-the-art convolutional neural network architecture that runs efficiently on mobile platforms. Furthermore, the algorithm is data-efficient, meaning it does not require a large amount of data. Petition 870260066391, dated 06 / 07 / 2026, page 83 / 315 70 / 139 corn stands to be used in practice. This is achieved through a novel transfer learning technique. A support vector machine (SVM) classifier is trained to classify features extracted from a convolutional neural network with pre-trained weights. The result is a robust stand counting algorithm (according to a modality), where the only sensor required is a low-cost RGB camera (<$30). Extensive field tests show that detection is robust against noise such as corn leaves, weeds, varying lighting conditions, and residues from the previous year. The disclosed modality system achieves high accuracy and reliability throughout the growing season. In essence, the disclosed modality system enables real-time robotic phenotyping in breeding batches and production fields.Robotic phenotyping is a welcome contribution, as it can help overcome the phenotyping bottleneck [23B], [24B] that has slowed progress in breeding better crops.
[192] Reference will now be made to specific Data Acquisition methods according to one or more modalities. The platform used in this study is a four-wheeled unmanned vehicle. The robot is (in this example) 12 inches (30.48 centimeters) high x 20 inches (50.8 centimeters) long χ 14 inches (35.56 centimeters) wide, with 6 inches (15.24 centimeters) of ground clearance and weighs 14.5 pounds (6.57 kilograms). However, the system can use a robot with any dimensions that allow the robot to traverse between crop rows. This compact and lightweight design allows the robot to easily traverse between crop rows with typical row spacing for corn. The robot is powered Petition 870260066391, dated 06 / 07 / 2026, page 84 / 315 71 / 139 (in this example) by four lithium-ion batteries that offer up to 4 hours of runtime. The robot can alternatively use power sources, including but not limited to different lithium-ion battery types and solar power. Images are recoded (in this example) with an RGB digital camera (ELP USBFHD01M, USA) mounted on the side of the robot chassis, but other cameras can be used, either in visual or other spectrum modes. The camera's field of view (in this example) is 60°. In other modes, the camera's field of view can be between 10° and 360°. The number of corn plants captured in the image depends on the distance between the camera and the row of plants, as well as the spacing between adjacent plants. In a row of 30, for example, two to three corn plants typically appear in the image. The camera points downwards at an angle (in this example) of 35° to avoid observing distant corn rows.In other embodiments, the camera points downwards at an angle between 0° and 85°. FIGS. 23A-23D show illustrations of the data acquisition system configuration for this embodiment. The camera resolution is 640 x 480 and records at 30 frames per second. In other embodiments, the camera resolution is at least 240P and records at least 5 frames per second. The camera has a USB2 interface (although other interfaces also work) that connects to a Jetson TX2 (NVIDIA, USA) or other computer(s), an integrated module for fast and efficient deep neural network inference. In another embodiment, any module capable of efficient deep neural network inference can be used. The module (in this example) hosts 8 GB of memory that is shared between a CPU and GPU and is capable of processing image frames captured by the camera in real time. In other embodiments, the module houses between 1 GB. Petition 870260066391, dated 06 / 07 / 2026, page 85 / 315 72 / 139 and 1024 GB of memory.
[193] Reference will now be made to certain methods according to one or more modalities. Images captured in an outdoor environment are subject to a wide range of variations, such as sunlight, occlusion, and camera viewing angle, etc. In addition, corn plants undergo significant changes during the growth period. These variations make classification difficult using conventional approaches. Therefore, in the current disclosure, a deep learning model is trained by combining a convolutional neural network and a support vector machine to classify the presence or absence of corn.
[194] Deep learning can identify corn plants in an image. However, this is not sufficient for counting plants from a moving robot, as it is difficult to distinguish different corn plants to avoid double counting. One way to solve this is to draw a fixed region of interest (ROI) in the image whose width is, on average, smaller than the gap between neighboring corn plants. The deep learning algorithm is applied only to the pixels within the ROI to detect whether or not a plant is present. This leads to a binary signal that takes (for example) the value 1 for all frames where a corn stalk is detected in the ROI and a value -1 otherwise (including weeds, leaves, or other green matter). For convenience, the ROI is placed (in one modality) in the center of the image, but this is not necessary. It is also possible to move the ROI around the image to scan multiple corn plants.However, even with a fixed ROI, it is possible that several corn plants may appear within the ROI (see FIG. 25 showing the two plants within the ROI (where the ROI is indicated in the FIG. as the vertical rectangle)), especially now. Petition 870260066391, dated 06 / 07 / 2026, page 86 / 315 73 / 139 that plant spacing can vary significantly between varieties, fields, and planting equipment. To solve this problem, motion estimation techniques are used (in one modality) to determine the number of plants that passed through the ROI until the ROI stops detecting corn plants. The details of this algorithm are provided in Algorithm 1B. Algorithm 1B: Counting Algorithm Input; ln;=image frames Parameters; k := average window size w := ROI width in pixels image translation in pixels Output; C := corn stand count Initialize; C = Oj d = 0 foreach / . do Extract features using ConvNet —> G K.1034; Sort by SVM —> € { — 1, 1} if y, > 0 then Estimate translation of 4-1 to If-^d, ( Algorithm 2); d^d + de, end else if íj-i > 0 then Multiplicity Λί <— —' d ^0; end end end
[195] Reference will now be made to a convolutional neural network according to one or more modalities. Considering the fact that the algorithm of this modality must be implemented on mobile platforms, it is important to maintain a good balance between performance and efficiency. Energy consumption and memory footprint must be taken into account, which may limit the size of the architecture and its number. Petition 870260066391, dated 06 / 07 / 2026, page 87 / 315 74 / 139 arithmetic operations. While most networks aim for greater precision by increasing depth and width [22B], [31B], [32B], MobileNets [33B] are a family of networks specifically developed to meet the requirements of mobile and embedded applications. Due to the use of depth-separable convolution which uses between 8 and 9 times less computation than standard convolution, the model runs significantly faster than its more complicated counterparts with only a small compromise in accuracy. For example, MobileNet achieves 89.5% accuracy among the top 5 in ImageNet with 569 million floating-point operations [33B], while ResNet-152 uses 11.3 billion flops to achieve 93.2% [19B]. The MobileNet architecture contains a standard convolution layer followed by 13 depth-separable convolution layers. The model uses 224 x 224 RGB images as input and generates 1024-dimensional feature vectors. The fully connected and softmax layers in the original model are replaced by SVM for classification.Details of the architecture (which can be used by various modalities) are shown in FIG. 25. Note that, even for such a compact architecture, a huge amount of data is still needed to train the model, as it contains 4.2 million parameters. However, no such dataset exists specifically for agriculture. To alleviate the high training data requirement, the present disclosure uses (in various modalities) the principles of transfer learning. The model is adjusted based on the pre-trained weights in the ILSVRC dataset [34B].
[196] Reference will now be made to a Support Vector Machine according to one or more embodiments. A Support Vector Machine (SVM) is a widely used model for Petition 870260066391, dated 06 / 07 / 2026, p. 88 / 315 75 / 139 classification [35B] - [40B]. It follows a linear model that constructs a hyperplane as shown in Equation (1B). The ideal hyperplane maximizes the margin, which is defined as the shortest distance between the decision boundary and any of the samples. Given a data set, ' , where xi ∈ Rn and ti ∈ {-1,1}, the distance from a point xi to the decision boundary is given by Equation (2B). Thus, the maximum margin solution can be found by solving Equation (3B). If webs are subsequently scaled so that the nearest point of any class to the hyperplane is ±1, the optimization is then equivalent to a quadratic programming problem (Equation (4B)).
[197] Since the target function is quadratic, subject to a linear constraint, a unique global minimum is guaranteed. In practice, however, cases arise where the conditional class distributions may overlap and the exact separation of the training data may lead to poor generalization. For this reason, the soft margin SVM was subsequently introduced [41B]. Instead of strictly classifying each data point correctly, each data point is assigned a slack variable ξί = 0 when it is within or outside the correct decision boundary and ξί = | ti - yi | otherwise. The sum of all ξί is then weighted by a parameter C that controls the trade-off between the slack variable penalty and the margin. The optimization problem now takes the form given in Equation (5B). Note that convexity gives the SVM a desirable advantage over other methods, such as a feedforward neural network, which suffers from the existence of multiple local minima [42B], [43B].Furthermore, despite being a linear classifier, the kernel functions transform the input data into a space of... Petition 870260066391, dated 06 / 07 / 2026, p. 89 / 315 76 / 139 higher dimensional feature allows the SVM to also classify separable features non-linearly. Due to the advantages, several modalities use a soft-margin SVM to classify features extracted from MobileNet. (1B) (2B) arg max (3B) argmin ||w|| subject to subject to j * 0. N Aς N (4B) (5B)
[198] Reference will now be made to training according to one or more modalities. Images are captured (in these modalities) by the camera facing the side of the robot during the growing season. The camera points slightly downwards so that only the nearest line is visible. Spots are cropped from the images and labeled according to the presence or Petition 870260066391, dated 06 / 07 / 2026, page 90 / 315 77 / 139 absence of corn as positive and negative samples, respectively. The number of training and test samples (from a specific example) is listed in Table IB. Data augmentation was also employed to further augment the training data. The images (in this example) are rotated (± 10°), enlarged (88% to 112%), vertically shifted (± 15%), and horizontally rotated. Each image is augmented (in this example) 16 times by means of random list drawing transformations.
[199] The MobileNet weights are kept constant during fine-tuning. The hyperparameters for SVM are determined by grid search and cross-validation implemented in a Python package: scikit-learn [44B]. The linear kernel and the radial basis function (RBF) kernel are investigated. Table IIB lists the hyperparameter values that are considered in this example. All combinations are exhausted to identify the ideal parameters. TABLE I (B): Number of training and test samples at each growth stage VI VT R2 Training positive 348 328 519 negative 416 304 310 Testing positive 89 83 130 negative 105 77 78 TABLE II(B): SVM hyperparameters tested in grid search Linear RBF (7 1,10,100,1000 1,10,100,1000 r N / A ΙΟ”2,10--110-4,10“s Petition 870260066391, dated 06 / 07 / 2026, page 91 / 315 78 / 139
[200] Reference will now be made to an estimate of movement according to one or more modalities. Although (in these modalities) the width of the ROI is chosen so that only one corn plant appears in it, exceptions to this assumption arise due to variations in plant spacing. If the gap between two neighboring plants is less than the width of the ROI, the detection signal of these two corn plants will merge into a single plateau (see FIGS. 26A and 26B). Therefore, it is likely to underestimate the population by taking the number of plateaus as the count. These errors can increase, degrading the estimation performance. To determine the number of plants that each signal plateau represents, the hard transformation between two consecutive frames Ii and I2 is calculated when a positive binary classification signal is present. A hard transformation is a combination of rotation, translation, and reflection. Certain modes do not yet assume any reflection; a hard transformation on K2 is given in Equation (6B). Two pairs of points are needed to solve for the transformation matrix M. In practice, the system of these modes extracts and combines about 100 SURF feature points [45B] in both images and solves for the least squares solution for M2.Since the camera movement is predominantly in the X direction, the system in these modes sums the tx values from the first frame on the plateau until the last one is T. If the ROI finishes sweeping through a corn before encountering the next one, as is the case in most instances, T is approximately equal to the width of the ROI w. In the case of adjacent corn plants being close together, T = (n1) d + w, where n is the number of plants, d is the distance between neighboring plants (see FIGS. 26A and 26B). If d is not much smaller than w, then n ≠ —. In practical implementation, obtain the ratio w Petition 870260066391, dated 06 / 07 / 2026, p. 92 / 315 79 / 139 results in a floating-point number that is rounded to the nearest integer. In other words, inequality (7B) must be maintained. For example, n = 2, the algorithm can correctly count two adjacent corn plants as long as their distance is not less than half the width of the ROI. The lower bound on d increases with the number of adjacent corn plants n. However, in most cases, n is less than 5. — sin0 cos Θ 0 (6B) n — 0.5 < — < η +0.5 n — 1.5 d < n — 0.5 η — 1—Η'—n — 1 (7B) Algorithm 2B: Motion estimation algorithm Input : 7;i= images N ;= maximum feature points to be extracted in each image Output: MGR := hard transformation between images (breach ή do | Extract feature points JV SURF {+>'Mί=1end Find points of correspondence between {+A,[ θ ; Solve Equation (6} for Af
[201] Reference will now be made to SVM Training according to one or more modalities. The best hyperparameters for SVM found in the grid search (of an example) are Petition 870260066391, dated 06 / 07 / 2026, page 93 / 315 80 / 139 listed in Table III (B) with their performance metrics on the test data. Precision represents the probability of a sample being correctly labeled. Precision represents the probability of a predicted positive sample (tp + fp) being a true positive (tp), while recall is the probability of a true positive sample (tp) being identified among all positive samples (tp + fn), and F1 score is the harmonic mean of precision and recall. Their formulas are given in Equation (8B).
[202] The hyperparameters are almost the same for the three growth stages, except that the C value is different for V4, largely due to the difference in the appearance of the corn plants in V4 from VT and R2. The radial basis function kernel works better than the linear kernel, and the optimal γ value for the kernel is identical in all three stages. The results indicate that the characteristics of the corn images at different stages are comparable and it is possible to classify all images using a single unified classifier.
[203] Performance metrics are also consistently high for three growth stages. It is observed that the training samples exclude leaves and focus only on stems and stalks. The metrics demonstrate that the deep learning model is able to distinguish subtle differences in images and therefore reduces the inference of leaves and weeds. The average accuracy for all stages is 92.97%. Furthermore, the predictions are calculated by averaging its k neighbors (k is an odd integer). Assuming a corn stalk is present in the image frames Ii, ..., Ii + n (nhk), for the frames Ii + ki, ..., Ii + n -k-1, the accuracy 2 Petition 870260066391, dated 06 / 07 / 2026, page 94 / 315 81 / 139 becomes effective where P is the accuracy for a single frame. On average, a corn stalk remains in the ROI for about 10 frames, so a value of 5 is chosen for k. Then, the recognition accuracy is increased to up to 99.69%. TABLE III(B): Best hyperparameters found in grid search and their corresponding performance metrics in the test data V4 VT R2 kernel rbf rbf rbf ç 10 100 100 r 0.001 0.001 0.001 accuracy 91.75% 94.38% 92.79% precision 0.91 0.91 0.91 recall 0.95 0.94 0.95 Score-Fl 0.95 0.93 0.94 tp · + tfl accuracy =------------tp 4“ jp H-tn + Jri ~lP precision —---—— tp recall —----tp+fn * precision x recall = 2 -------------precision + recall
[204] With reference now to FIG. 30, this is a flowchart representing one embodiment of a method for determining a plant population. First, step 3005 comprises capturing, by a camera coupled to a processing system that includes a processor, a group of images associated with a Petition 870260066391, dated 06 / 07 / 2026, page 95 / 315 82 / 139 plant field, where the group of images is captured while the camera is moving in the field. Next, step 3010 involves the processing system applying a convolutional neural network to the group of images to extract feature vectors. Next, step 3015 involves the processing system applying a support vector machine to the feature vectors to classify the feature vectors, resulting in classified features. Next, step 3020 involves the processing system applying motion estimation based on a rigid transformation to the classified features, resulting in corrected count data. Finally, step 3025 involves the processing system determining a plant population for the plant field based on the classified features and the corrected count data.
[205] In one embodiment, the support vector machine comprises a smooth margin support vector machine that applies a slack variable to data points on or within a correct decision boundary.
[206] In one embodiment, the group of images is not an aerial view of the plant field and the camera view is at a downward angle.
[207] In one mode, the camera movement while the group of images is being captured is at a constant speed.
[208] In one embodiment, the convolutional neural network is applied only to a fixed region of interest in each image of the group of images.
[209] In one embodiment, the convolutional neural network comprises a standard convolutional layer followed by multiple Petition 870260066391, dated 06 / 07 / 2026, page 96 / 315 83 / 139 depth-separable convolutional layers
[210] In one mode, data augmentation is applied to the training data for the convolutional neural network. Data augmentation in this mode is based on image rotation, image magnification, vertical image displacement, horizontal image inversion, or a combination thereof.
[211] With reference now to FIG. 31, this shows a block diagram representing an embodiment of an apparatus for determining a plant population for a plant field. The apparatus comprises an unmanned vehicle 3110, a battery 3120, a processing system 3130 and a camera 3140. The processing system 3130 includes a processor. The battery 3120, the processing system 3130 and the camera 3140 are supported by the unmanned vehicle 3110. The unmanned vehicle 3110, the processing system 3130 and the camera 3140 are electronically coupled to the battery 3120. The camera 3140 is in electronic communication with the processing system 3130.
[212] In operation, battery 3120 provides electricity for unmanned vehicle 3110, processing system 3130, and camera 3140. Camera 3140 captures a group of images associated with a field of plants while unmanned vehicle 3110 travels along the field lines. The group of plant images captured by camera 3140 is electronically communicated to processing system 3130. Processing system 3130 implements a convolutional neural network, a support vector machine, and motion estimation relative to the group of images to determine a plant population for the field of plants.
[213] In one mode, the 3140 camera is pointed at a downward angle. Petition 870260066391, dated 06 / 07 / 2026, page 97 / 315 84 / 139
[214] In one embodiment, the support vector machine comprises a smooth margin support vector machine that applies a slack variable to data points on or within a correct decision boundary.
[215] In one embodiment, the unmanned vehicle 3110 travels along the field lines at a constant speed.
[216] In one embodiment, the convolutional neural network is applied only to a fixed region of interest in each image of the group of images.
[217] In one embodiment, the convolutional neural network comprises a standard convolution layer followed by multiple depth-separable convolution layers.
[218] In one embodiment, data augmentation is applied to the training data for the convolutional neural network and data augmentation is based on image rotation, image zoom, vertical image displacement, horizontal image inversion, or a combination thereof.
[219] In one embodiment, the 3140 camera mounted on the 3110 unmanned vehicle is on a gimbal or servo that automatically adjusts the camera's view based on the characteristics of the images collected by the 3140 camera.
[220] In one embodiment, an air jet (i.e., a mechanism for producing an airflow) is supported by the unmanned vehicle 3110 and supplied with electricity by the battery 3120. The air jet clears obstructions, such as leaves, from the camera frame 3140, so that unobstructed images of the plants can be captured by the camera 3140.
[221] With reference now to FIG. 32, this shows a block diagram representing an embodiment of an apparatus for calculating a plant population. The apparatus comprises a means Petition 870260066391, dated 06 / 07 / 2026, page 98 / 315 85 / 139 of non-transient machine-readable storage 3210, executable instructions 3215, a processing system 3220, and a processor 3225. The non-transient machine-readable storage medium 3210 is in electronic communication with the processing system 3220. The executable instructions 3215 are stored electronically in the non-transient machine-readable storage medium 3210. The processor 3225 is coupled to the processing system 3220.
[222] In operation, the non-transient machine-readable storage medium 3210 electronically communicates executable instructions 3215 to the processing system 3220. The processing system 3220 uses the processor 3225 to execute the executable instruction 3215 and facilitate the performance of operations. First, the processing system 3220 applies a convolutional neural network to a group of images to extract feature vectors. The group of images is associated with a plant field and is captured while a camera is moving in the field. Next, the processing system 3220 applies a support vector machine to the feature vectors to classify the feature vectors, resulting in classified features. Then, the processing system 3220 provides motion estimation based on a hard transformation for the classified features, resulting in corrected count data.Finally, the 3220 processing system determines a plant population for the plant field based on the classified resources and corrected count data.
[223] Reference will now be made to the robot path modulation and camera angle for optimal visual data collection. According to one modality, the objective in this task is to close Petition 870260066391, dated 06 / 07 / 2026, page 99 / 315 86 / 139 firmly establishes the loop between the robot's path, the camera angles (controlled by servos), and the suitability of the raw visual information being collected for phenotyping.
[224] Different phenotypes require different parts of the plant to be visible to allow machine vision algorithms to function correctly. For example, plant counting is simplified by capturing images of the plant-soil interface with a downward perspective, while stem width estimation requires an unobstructed view of the stem with a horizontal perspective. Pre-specified GPS paths cannot guarantee the collection of the best quality data. For example, some farms use furrow irrigation (i.e., U-shaped depressions, rather than flat ground, between plant rows), which leads to the camera orientation unpredictably oscillating around the ideal during execution.
[225] To meet these challenges, an adaptive system capable of adjusting the robot path, camera angle, and robot speed is implemented (according to various modalities) to ensure consistent, high-quality images. FIG. 33 represents an overview of the adaptive camera angle and robot path control approach according to a modality. The main feedback signal is the “image feature critical” signal, which takes as input the features extracted from an image feature extractor. A design of these subsystems (according to various modalities) is presented in detail below.
[226] Reference will first be made to an image feature extractor. The purpose of the image feature extractor is to identify features in the image that are needed or Petition 870260066391, dated 06 / 07 / 2026, page 100 / 315 87 / 139 desirable to quantify the desired phenotype. The image feature extractor can be designed in two main ways, the first being a heuristic approach where certain low-level features and specific image-level features are searched for in the image, such as root-soil vision, plant stem through edge detectors, or corn cob in the image; alternatively (or in combination), a neural network can be used to extract features from the images in a manner similar to the corn counting approach.
[227] Reference will now be made to an image feature critique. The goal of image feature critique algorithms and software (of various modalities) is to provide a feedback signal to the control system that can be used to control the robot's path and / or camera angle. The challenge here is to determine in the most efficient way whether or not the image has the desired features that are needed. The first step is to determine whether or not the image has a feature of interest. This can be done by extracting the feature and comparing it in pixel space with a library of desirable features using various metrics such as mean difference sum, norms in feature spaces, information distances, and / or binary classifiers. The most general approach is the classifier-based approach, where the potential challenge is to determine whether the classifier itself is sufficiently confident.To mitigate this, similar to the corn counting algorithm, a final layer support vector machine (and / or Gaussian Process with non-Gaussian likelihood algorithms) can be invoked [37B].
[228] Reference will now be made to a Camera Angle - Petition 870260066391, dated 06 / 07 / 2026, page 101 / 315 88 / 139 Robot Path Control and Control Mixer. The purpose of these modules is to control the robot's camera angle and / or path to ensure that desired features are in the proper view. Using feature-critical feedback, two clear cases can be visualized. The first is when the feature is visible, but an adjustment to the angle and / or path is needed to place it in the center of the image. This is accomplished with traditional feedback control mechanisms using image feature pixel distance as a continuous feedback signal. The second case is that the image feature is not in view. This requires the robot to search for the feature by adjusting its path and / or camera angle. The camera angle scanning method is used (in this mode) as it is faster and less risky. The camera angle will scan all possible angles to find the feature of interest.If this does not lead to the correct feature, the robot can use its line distance estimate and adjust the distance to values that previously led to successful image features. If both mechanisms fail to bring the feature into view, the remote operator can be alerted. The data is logged so that, over time, a learning system can be trained to ensure that features remain visible. Future reinforcement learning algorithms will learn from examples the best camera angle and the best robot path in a field, given features and field locations.
[229] As described in this document, it is a color image-based algorithm for estimating corn plant population. The latest deep learning architecture is used along with classic SVM with radial basis function kernel. Petition 870260066391, dated 06 / 07 / 2026, page 102 / 315 89 / 139 The algorithm robustly recognized corn stalks in the presence of leaf and weed interference, but requires only a relatively small amount of data for training. After recognition, motion estimation techniques were used to calculate the relative movement between the camera and the corn stalks. Finally, the corn population was derived from the relative movement. The method disclosed according to a modality was tested at three different growth stages (V4, VT, R2) and achieved average errors of 0.34% ± 6.27%, -5.14% ± 5.95%, and 0.00% ± 3.71%, respectively. One of the main sources of error was hanging leaves covering the camera lenses, which made the algorithm in this modality more prone to underestimating the population. Adding an air jet to the unmanned vehicle to mechanically move the leaves out of the camera's field of view reduces this error.
[230] Reference will now be made to some examples. Referring now to a field experiment, data were collected at the University of Illinois at Urbana-Champaign Energy Farm, Urbana, Illinois, from June to October. Rows of various lengths in the growth stages V4, V6, VT, R2, and R6 were randomly selected. The robot was driven at a constant speed by remote control for an arbitrary number of batches. For data collections prior to October, a small portion of the images was annotated to train the recognition model. Two additional datasets were collected in October to test the model's ability to generalize on unseen data from different field conditions. Table IV(B) summarizes the batch and experiment conditions. TABLE IV (B): Dates and conditions of data collection Petition 870260066391, dated 06 / 07 / 2026, p. 103 / 315 90 / 139 Date Location m . , Training Stage . Growth Batches 6 June, 2017 Assumption, 1L Yes V4 16 6 July, 2017 Champaign, 1L Yes VT 10 2 August, 2017 Champaign, 1L Yes R2 28 21 September, 2017 Ivesdale, 1L Yes R6 10 25 October, 2017 Lebanon, ID No V6 15 26 October, 2017 Martinsville, IL No R6 8
[231] Referring now to a certain validation (field corn counting), FIG. 27 shows the corn plant population per pot, by robot vs. human for each dataset. The robot's predictions agree well with the fundamental truth. The least squares fitted line across all data has a correlation coefficient R = 0.96. The box and whisker plot accuracy for each dataset is shown in FIG. 28. The algorithm achieves consistently high accuracy for all locations and growth stages. The average accuracy across all datasets is 89.74%. Note that the Martinsville and Lebanon data were processed without using the new data to train SVM. The results demonstrate that the recognition model generalizes well to unseen data and handles real-world variations effectively.
[232] The main sources of error include heavy occlusion, uneven lighting, and nearby corn plants. FIGS. 29A and 29B show two examples of leaves hanging between the rows almost completely covering the camera lens. Although the deep learning algorithm can correctly recognize the leaf stalks, it lacks the power to make reliable predictions when the camera lens suffers from heavy occlusion. In one embodiment, a mechanical mechanism such as an air jet [46B] is used to keep the leaves away from the vicinity of the Petition 870260066391, dated 06 / 07 / 2026, page 104 / 315 91 / 139 visual sensor.
[233] As described in this document, the disclosure aspects in question may include, for example, a method for determining a plant population in a field. The method in this embodiment comprises capturing, by a camera coupled to a processing system, a group of images associated with a plant field, wherein the group of images is captured while the camera is moving in the field. Then, the processing system applies a convolutional neural network to the group of images to extract feature vectors. Then, the processing system applies a support vector machine to the feature vectors to classify the feature vectors, resulting in classified features. Then, the processing system applies motion estimation based on a stiff transformation to the classified features, resulting in corrected count data.Finally, the processing system determines a plant population for the plant field based on the classified resources and corrected count data.
[234] In one embodiment, a method is provided comprising: capturing, by a camera coupled to a processing system that includes a processor, a group of images associated with a plant field, wherein the group of images is captured while the camera is moving through the field; applying, by the processing system, a convolutional neural network to the group of images to extract feature vectors; applying, by the processing system, a support vector machine to the feature vectors to classify the feature vectors resulting in classified features; applying, by the processing system, motion estimation based on a rigid transformation to Petition 870260066391, dated 06 / 07 / 2026, page 105 / 315 92 / 139 classified resources, resulting in corrected count data; and determine, by the processing system, a plant population for the plant field based on the classified resources and the corrected count data.
[235] In one example, the support vector machine comprises a smooth margin support vector machine that applies a slack variable to data points on or within a correct decision boundary.
[236] In another example, the group of images is not an aerial view of the plant field and one camera view is at a downward angle.
[237] In another example, the camera movement while the group of images is being captured is at a constant speed.
[238] In another example, camera images are processed to adjust the camera angle to improve counting accuracy.
[239] In another example, camera images are processed to adjust the path of camera movement in the field to improve counting accuracy.
[240] In another example, the convolutional neural network is applied only to a fixed region of interest in each image of the group of images.
[241] In another example, the fixed region of interest is automatically calculated from the camera images.
[242] In another example, the convolutional neural network comprises a standard convolution layer followed by several depth-separable convolution layers.
[243] In another example, data augmentation is applied to the training data for the convolutional neural network and the augmentation Petition 870260066391, dated 06 / 07 / 2026, page 106 / 315 93 / 139 of data is based on image rotation, image zoom, vertical image shift, horizontal image inversion, or a combination thereof.
[244] In another embodiment, an apparatus is provided comprising: an unmanned vehicle; a battery supported by the unmanned vehicle; a processing system including a processor and supported by the unmanned vehicle; and a camera supported by the unmanned vehicle, wherein the camera captures a group of images associated with a field of plants, wherein the group of images is captured while the unmanned vehicle travels along the lines of the field, and wherein the processing system implements a convolutional neural network system, a support vector machine and motion estimation with respect to the group of images to determine a plant population for the field of plants.
[245] In one example, the camera view is at a downward angle.
[246] In another example, the support vector machine comprises a smooth margin support vector machine that applies a slack variable to data points on or within a correct decision boundary.
[247] In another example, the unmanned vehicle travels along the field lines at a constant speed.
[248] In another example, the camera mounted on the unmanned vehicle is on a gimbal that automatically adjusts the camera view based on the camera images.
[249] In another example, the convolutional neural network is applied only to a fixed region of interest in each image in the group of images.
[250] In another example, the convolutional neural network Petition 870260066391, dated 06 / 07 / 2026, page 107 / 315 94 / 139 comprises a standard convolution layer followed by several depth-separable convolution layers.
[251] In another example, data augmentation is applied to the training data for the convolutional neural network and data augmentation is based on image rotation, image zoom, vertical image displacement, horizontal image inversion, or a combination thereof.
[252] In another example, an air jet is supported by the unmanned vehicle, where the air jet clears obstructions, such as leaves from the camera frame.
[253] In another embodiment, a non-transient machine-readable storage medium is provided, comprising executable instructions which, when executed by a processing system including a processor, facilitate the performance of operations, comprising: applying a convolutional neural network to a group of images to extract feature vectors, wherein the group of images is associated with a plant field and is captured while a camera is moving in the field; applying a support vector machine to the feature vectors to classify the feature vectors resulting in classified features; applying motion estimation based on a rigid transformation to the classified features resulting in corrected count data; and determining a plant population for the plant field based on the classified features and the corrected count data.
[254] In one example, the support vector machine comprises a smooth margin support vector machine that applies a slack variable to data points on or within a correct decision boundary.
[255] In another example, the group of images is not a view Petition 870260066391, dated 06 / 07 / 2026, page 108 / 315 95 / 139 aerial view of the plant field and a camera view at a downward angle.
[256] In another example, the convolutional neural network is applied only to a fixed region of interest in each image of the group of images.
[257] In another example, the convolutional neural network comprises a standard convolution layer followed by several depth-separable convolution layers.
[258] In another example, data augmentation is applied to the training data for the convolutional neural network and data augmentation is based on image rotation, image zoom, vertical image displacement, horizontal image inversion, or a combination thereof. References 1B-46B
[259] [1B] R. HOLLIDAY et al., Plant population and crop yield. Nature, vol. 186, no. 4718, pp. 22-4, 1960.
[260] [2B] W. Duncan, “The relationship between population and maize yield,” Agronomy Journal, vol. 50, no. 2, pp. 82-84, 1958.
[261] [3B] R. Willey and S. Heath, “The quantitative relationships between plant population and crop yield,” Advances in Agronomy, vol. 21, pp. 281-321, 1969.
[262] [4B] J. Lutz, H. Camper, and G. Jones, “Row spacing and population effects on corn yields,” Agronomy journal, vol. 63, no. 1, pp. 12-14, 1971.
[263] [5B] A. Lang, J. Pendleton, e G. Dungan, “Influence of population and nitrogen levels on yield and protein and oil contents of nine corn hybrids,” Agronomy Journal, vol. 48, no. 7, pp. 284-289,1956.
[264] [6B] P. Thomison, J. Johnson, e D. Eckert, “Nitrogen Petição 870260066391, de 06 / 07 / 2026, pág. 109 / 315 96 / 139 fertility interactions with plant population and hybrid plant type in corn,” Soil Fertility Research, pp. 28-34, 1992.
[265] [7B] W. Duncan, A theory to explain the relationship between corn population and grain yield,” Crop Science, vol. 24, no. 6, pp. 1141-1145, 1984.
[266] [8B] R. Holt e D. Timmons, Influence of precipitation, soil water, and plant population interactions on corn grain yields,” Agronomy Journal, vol. 60, no. 4, pp. 379-381, 1968.
[267] [9B] D. Karlen e C. Camp, Row spacing, plant population, and water management effects on corn in the atlantic coastal plain,” Agronomy Journal, vol. 77, no. 3, pp. 393-398, 1985.
[268] [10B] D. J. Eckert e V. L. Martin, Yield and nitrogen requirement of no-tillage corn as influenced by cultural practices,” Agronomy journal, vol. 86, no. 6, pp. 1119-1123, 1994.
[269] [11B] S. Birrell e KA Sudduth, sensor de população de milho para agricultura de precisão. ASAE, 1995.
[270] [12B] K. Thorp, B. Steward, A. Kaleita e W. Batchelor, Using aerial hyperspectral remote sensing imagery to estimate corn plant stand density Trans. ASABE, vol. 51, n° 1, pp. 311320, 2008, citado por 0.
[271] [13B] Y. Shi, N. Wang, R. Taylor, W. Raun, e J. Hardin, Automatic corn plant location and spacing measurement using laser line-scan technique,” Precis. Agric., Vol. 14, no. 5, pp. 478-494, 2013, citado por 0.
[272] [14B] Y. Shi, N. Wang, R. Taylor, and W. Raun, Improvement of a ground-lidar-based corn plant population and spacing measurement system,” Comput. Elétron. Agric., Vol. 112, pp. 92 - 101, 2015, agricultura de precisão. [Conectado]. Petição 870260066391, de 06 / 07 / 2026, pág. 110 / 315 97 / 139 Disponível: http: / / www.sciencedirect.com / science / article / pii / S016816991400 3093
[273] [15B] A. Nakarmi e L. Tang, “Automatic inter-plant spacing sensing at early growth stages using a 3d vision sensor,” Comput. Elétron. Agric., Vol. 82, pp. 23-31, 2012, citado por 25.
[274] [16B] J. D. Luck, S. K. Pitla, e S. A. Shearer, “Sensor ranging technique for determining corn plant population,” in 2008 Providence, Rhode Island, June 29-July 2, 2008. American Society of Agricultural and Biological Engineers, 2008, p. 1.
[275] [17B] JA Rascon Acuna, “Corn sensor development for by-plant management,” Ph.D Dissertation, Oklahoma State University, 2012.
[276] [18B] DS Shrestha and BL Steward, “Automatic corn plant population measurement using machine vision,” Transactions of the ASAE, vol. 46, no. 2, p. 559, 2003.
[277] [19B] K. He, X. Zhang, S. Ren and J. Sun, Deep residual learning for image recognition, in Conference Proceedings IEEE in Computer Vision and Pattern Recognition, 2016, pp. 770-778.
[278] [20B] A. Krizhevsky, I. Sutskever, and GE Hinton, “Image classification with deep convolutional neural networks,” in Advances in neural information processing systems, 2012, pp. 1097-1105.
[279] [21B] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke and A. Rabinovich, Going deep with convolutions, in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1-9.
[280] [22B] K. Simonyan and A. Zisserman, “Very deep Petition 870260066391, dated 06 / 07 / 2026, p. 111 / 315 98 / 139 Convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
[281] [23B] JL Araus and JE Cairns, “High-throughput field phenotyping: the new frontier of crop breeding”, Trends in plant science, vol. 19, no. 1, pp. 52-61, 2014.
[282] [24B] RT Furbank and M. Tester, “Phenomics - technologies to alleviate the phenotyping bottleneck,” Trends in plant science, vol. 16, no. 12, pp. 635-644, 2011.
[283] [25B] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision (IJCV), vol. 115, no. 3, pp. 211-252, 2015.
[284] [26B] R. Girshick, J. Donahue, T. Darrell and J. Malik, “Region-based convolutional networks for precise object detection and segmentation.” IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 38, No. 1, pp. 142-158, 2016.
[285] [27B] R. Girshick, “Fast r-cnn,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1440-1448.
[286] [28B] S. Ren, K. He, R. Girshick and J. Sun, Faster rcnn: Towards real-time object discovery with proposed region networks, in Advances in neural information processing systems, 2015, pp. 91-99.
[287] [29B] J. Long, E. Shelhamer and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3431-3440.
[288] [30B] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, Petition 870260066391, dated 06 / 07 / 2026, page 112 / 315 99 / 139 D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, Generative adversarial nets,” in Advances in neural information processing systems, 2014, pp. 2672-2680.
[289] [31B] C. Szegedy, S. Ioffe, V. Vanhoucke and AA Alemi, Inception-v4, inception-resnet and the impact of residual connections on learning. in AAAI, 2017, pp. 4278-4284.
[290] [32B] K. He, X. Zhang, S. Ren and J. Sun, Identity mappings in deep residual networks, at the European Conference on Computer Vision. Springer, 2016, pp. 630-645.
[291] [33B] AG Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto e H. Adam, Mobilenets: Efficient convolutional neural networks for mobile vision applications, arXiv preprint arXiv: 1704.04861, 2017.
[292] [34B] J. Pruegsanusak e A. Howard, https: / / github.com / tensorflow / models / blob / master / slim / nets / mobilenet v1.md, acessado: 12 / 08 / 2017.
[293] [35B] N. Cristianini e J. Shawe-Taylor, An introduction to support vector machines,” 2000.
[294] [36B] B. Scholkopf, K.-K. Sung, C. J. Burges, F. Girosi, P. Niyogi, T. Poggio, e V. Vapnik, Comparing support vector machines with Gaussian kernels to radial basis function classifiers,” Transações IEEE sobre Processamento de Sinal, vol. 45, no. 11, pp. 2758-2765, 1997.
[295] [37B] A. J. Smola e B. Sch''olkopf, On a kernel-based method for pattern recognition, regression, approximation, and operator inversion,” Algorithmica, vol. 22, no. 1-2, pp. 211231, 1998.
[296] [38B] AJ Smola, B. Scholkopf e K.-R. Muller, The connection between regularization operator and support vector kernels”, Neural networks, vol. 11, no. 4, pp. 637-649, 1998. Petição 870260066391, de 06 / 07 / 2026, pág. 113 / 315 100 / 139
[297] [39B] B. Scholkopf, S. Mika, CJ Burges, P. Knirsch, K.R. Muller, G. Ratsch e AJ Smola, “Input space versus feature space in kernel-based methods,” Transações IEEE sobre redes neurais, vol. 10, no. 5, pp. 1000-1017, 1999.
[298] [40B] B. Schoolkopf, CJ Burges e AJ Smola, Advances in kernel methods: support vector learning. Imprensa do MIT, 1999.
[299] [41B] C. Cortes e V. Vapnik, “Support-vector networks,” Machine learning, vol. 20, no. 3, pp. 273-297, 1995.
[300] [42B] J. A. Suykens, J. De Brabanter, L. Lukas, and J. Vandewalle, “Weighted least squares support vector machines: robustness and sparse approximation,” Neurocomputing, vol. 48, no. 1, pp. 85-105, 2002.
[301] [43B] CM Bishop, Neural networks for pattern recognition. Oxford University Press, 1995.
[302] [44B] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011.
[303] [45B] H. Bay, T. Tuytelaars and L. Van Gool, “Surf: Speeded up robust features,” Computer vision - ECCV 2006, pp. 404-417, 2006.
[304] [46B] B. Lobdell and J. Hummel, “Optical plant spacing and stalk diameter sensor with air-jet assist,” in Proc. 3rd Intl. Conferência sobre informação Geoespacial em Agricultura e Silvicultura, 2001.
[305] Reference will now be made to an agricultural robot that (in various forms) enables field-based phenotyping that overcomes previous limitations, such as high operating costs. Petition 870260066391, dated 06 / 07 / 2026, page 114 / 315 101 / 139 and maintenance, low footprint, safety, disruptive internal combustion engines, excessive manual or off-board processing of sensor data, and / or the need for experienced operators. As an example, an ultralight, low-cost, autonomous field phenotyping robot is provided that can navigate a variety of field conditions to overcome the aforementioned limitations with existing field-based phenotyping systems.
[306] In one or more embodiments, variability in crop rows, topology, environmental conditions, and / or size can be overcome by an agricultural robot that is capable of traversing a plurality of crops in different geographies without damaging the crop plants for phenotyping, observation, and / or during crop management operations. When plants are very small or when rows are unstructured or widely spaced, the agricultural robot can traverse the field without damaging the cultivated plants, even if it has to pass over them, and has a low cost of ownership. In one or more embodiments, an agricultural robot is manufactured and / or operates at a low cost for robotic exploration, phenotyping, and / or crop management tasks, which can be significant given the cost pressure faced by farmers.
[307] In one or more embodiments, a manufactured additive manufacturing process apparatus can traverse or operate in a plurality of crops, agricultural environments and / or geographies without damaging the crops and at a sufficiently low cost. In one or more embodiments, the apparatus can be manufactured and operated economically, while at the same time being capable of fully automated stand counting and phenotyping.
[308] In one or more forms, a plastic device Petition 870260066391, dated 06 / 07 / 2026, page 115 / 315 102 / 139 manufactured from composite material can traverse or operate in a plurality of crops, agricultural environments and / or geographies without damaging the crops and at a sufficiently low cost. In one or more embodiments, the apparatus can be manufactured and operated economically, while at the same time being capable of fully automated stand counting and phenotyping.
[309] In one or more embodiments, a device made of metal can traverse or operate in a plurality of crops, agricultural environments and / or geographies without damaging the crops and at a sufficiently low cost. In one or more embodiments, the device can be manufactured and operated economically, while at the same time being capable of fully automated stand counting and phenotyping.
[310] In one or more embodiments, plastic, metal, composites or other material is used in combination to construct the robot.
[311] In one or more forms, the robot can be constructed of compatible or soft materials.
[312] In one or more embodiments, a device is provided that will not damage crops, even if it passes over them, and does not use contact sensors to determine row edges or paths.
[313] In one or more embodiments, an autonomous ground robot is provided that is capable of tracking a pre-specified path with a high level of accuracy without human intervention. In one embodiment, an autonomous ground robot may also incorporate other higher-level behaviors, such as autonomously deciding when to recharge and / or collaborating with other autonomous ground robots to complete a task.
[314] GPS-guided tractors and cultivators that can follow prescribed paths are being adopted commercially. Petition 870260066391, dated 06 / 07 / 2026, page 116 / 315 103 / 139 by producers, however, the automation of large equipment only partially addresses the challenges of robotics in agriculture, especially since large equipment generally cannot be used when the crop canopy closes. In contrast, in one or more modalities, small ground robots are mechanized walkers that can autonomously track pre-specified paths in harsh outdoor environments. Their small size allows them to maneuver in rows and avoid problems such as soil compaction, through the use of compact, lightweight, and energy-efficient steering, sensing, and autonomy technologies.
[315] In one or more embodiments, an agricultural robot is provided that performs observational measurements and / or agricultural phenotyping and that does not permanently damage plants, even if it passes over them.This is enabled through various features, including one or more of the following: an ultralight robotic platform (e.g., a robot weighing less than 20 pounds (9.07 kilograms)); ultracompact construction (e.g., a compact-sized robot); a low-complexity, battery-operated robot with hours of endurance leading to low manufacturing and operating costs; a robot that does not damage plants, even when walking over them; a robot that includes a compact sensor, computing, and integrated analytical suite to perform agronomically relevant tasks and measurements onboard, including stand counting, stem angle determination, disease detection, and others on an ultralight robotic platform; a robot that can utilize a suite of web-based analytics to perform these measurements and / or enhance onboard measurement capabilities. an autonomous path tracking system that ensures highly accurate position tracking, which guarantees that the Petition 870260066391, dated 06 / 07 / 2026, page 117 / 315 104 / 139 robot stay on a prescribed path to minimize trampling on plants; an autonomous path-following system that does not require a high-precision RTK GPS signal; a human-interaction GUI (graphical user interface) implemented as an application on a smart device that allows a single human operator to control or direct one or more autonomous agricultural robots; an ultra-compact robot that can rotate 180° in a crop row; and / or a rotation mechanism that is not dependent on a rack and pinion mechanism, which can instead utilize a drag to rotate implemented by independent electric motor drive on all four wheels.
[316] One embodiment provides an apparatus with a frame, wheels rotatably connected to the frame, and a battery connected to the frame. The apparatus may have electric motors supported by the frame and connected to the battery, wherein each of the electric motors is connected to one of the wheels, wherein the electric motors enable the apparatus to rotate at least 180 degrees (e.g., to rotate 180 degrees within the space between adjacent rows of crops). The apparatus may have a processor supported by the frame and speed sensors supported by the frame and coupled to the processor, wherein each of the speed sensors is connected to one of the wheels and wherein the speed sensors transmit speed feedback signals associated with each of the wheels to the processor.The device may have a global navigation satellite system (GNSS) supported by the frame and comprising a GNSS antenna and GNSS computer coupled to the processor, where the global navigation satellite system receives, processes, and transmits position information to the processor. The device may have a gyroscope supported by the frame and coupled to the processor. Petition 870260066391, dated 06 / 07 / 2026, page 118 / 315 105 / 139 where the gyroscope transmits measurement signals including yaw rate measurements to the processor. The device may have a group of frame-supported sensors, where the sensor group provides non-contact measurements of surrounding crops and their characteristics for navigation purposes. The device may have a group of frame-supported sensors coupled to the processor, where the sensor group collects phenotypic data associated with the crops. The processor determines the control signals based on the desired angular and linear velocities according to the velocity feedback signals, positional information, and measurement signals, applying real-time horizon estimation and control. The processor transmits the control signals to each of the electric motors, causing the device to track a reference path.
[317] Another embodiment provides a device that has a frame, wheels rotatably connected to the frame, and a battery connected to the frame. The device may have electric motors supported by the frame and connected to the battery, where each of the electric motors is connected to one of the wheels, wherein the electric motors allow the device to rotate at least 180 degrees (for example, to rotate 180 degrees within the space between adjacent rows of crops). The device may have a processor supported by the frame. The device may have speed sensors supported by the frame and coupled to the processor, where each of the speed sensors is connected to one of the wheels and where the speed sensors transmit speed feedback signals associated with each of the wheels to the processor. The device may have a global satellite navigation system supported by the frame and Petition 870260066391, dated 06 / 07 / 2026, page 119 / 315 106 / 139 comprising a GNSS antenna and GNSS computer coupled to the processor, where the global navigation satellite system receives, processes, and transmits position information to the processor. The device may have a gyroscope supported by the frame and coupled to the processor, where the gyroscope transmits measurement signals including yaw rate measurements to the processor. The device may have a group of sensors supported by the frame and coupled to the processor, where the sensor group collects phenotypic data associated with the crops. The processor determines the control signals based on the desired angular and linear velocities according to the velocity feedback signals, position information, and measurement signals. The processor transmits the control signals to each of the electric motors, causing the device to track a reference path.The processor generates data collection signals to select or adjust one or more data collection algorithms, one or more sensor groups, or a combination thereof, thereby adjusting the collection of phenotypic data associated with crops. The data collection signals can be generated by the processor based on an analysis of at least some of the phenotypic data, other information obtained by the processor, or a combination thereof. The data collection signals can be generated by the processor without receiving command signals from a user on a remote source.
[318] Yet another embodiment provides a method involving receiving, by a processing system of a wheeled agricultural robot, speed feedback signals associated with each of the wheels, where the speed feedback signals are received from agricultural robot speed sensors and where each of the speed sensors is connected to one of the wheels. The method Petition 870260066391, dated 06 / 07 / 2026, p. 120 / 315 107 / 139 may include the processing system receiving position information from a global navigation satellite system (GNSS) of the agricultural robot. The method may include the processing system receiving measurement signals from a gyroscope of the agricultural robot, where the measurement signals include yaw rate measurements. The method may include the processing system determining control signals based on desired angular and linear velocities according to velocity feedback signals, positional information, and measurement signals by applying real-time horizon estimation and control.The method may involve the transmission, by the processing system, of control signals to a controller coupled to electric motors, causing the agricultural robot to track a reference path, where each of the electric motors is connected to one of the wheels and enables the agricultural robot to rotate at least 180 degrees (for example, to rotate 180 degrees within the space between adjacent rows of crops).
[319] With reference now to FIG. 36, an illustration of an embodiment of the robot with an attached set of sensors is shown. This robot is ultralight (less than 15 pounds (6.80 kilograms)) and compact. The ultralight nature of the robot ensures that it does not push plants into the ground due to weight and minimizes the robot's momentum while it drives. Existing agricultural robots on the market are generally significantly heavier. The ultralight weight is based on material selection, material composition, and structural design. Conventional wisdom for agricultural equipment is to manufacture them from metal. However, metal is heavy and expensive. In contrast, this embodiment of the robot is constructed by means of Petition 870260066391, dated 06 / 07 / 2026, page 121 / 315 108 / 139 Additive manufacturing (like 3D printing). Additive manufacturing creates complex designs by layering material, unlike traditional metalworking or traditional plastic injection molding. Leveraging additive manufacturing capabilities, this method minimizes weight while maintaining sufficient structural rigidity. The robot incorporates parts in a way that optimally distributes stress. The layering mechanism in additive manufacturing is used to control density, so that a denser and stronger construction is used for high-stress and high-wear parts. Additive manufacturing techniques can be applied to different parts based on their expected wear and stress levels. In one method, the robot may incorporate selected lightweight metal components, which are chosen to maximize the robot's field strength.
[320] In one embodiment, the robot's ground clearance is high enough to allow traversal of complex terrain. The motors can be placed close enough to the robot's wheels to minimize transmission loss. In another embodiment, wheel-close mechanisms can be used to ensure that little stress is placed on the motors and power is transmitted directly to the wheels with minimal (if any) gearing and transmission system to minimize losses.
[321] In one embodiment, wheels are used that achieve a performance compromise between traction, ground pressure, propensity to displace the ground, sufficient ground clearance, ease of manufacture, and mechanical strength. FIG. 37 shows some of the wheels that were considered undesirable in operation. The leftmost wheel failed due to low ground clearance. Petition 870260066391, dated 06 / 07 / 2026, page 122 / 315 109 / 139 to the ground and excessive pressure on the ground. The middle wheel failed due to too much soil disturbance and damage to plants due to sharp edges (leading to excessive pressure). The right wheel failed due to excessive slippage due to lack of traction, causing damage. Other considerations include the use of correct manufacturing filler patterns to create sufficient smoothness in the wheels, even when manufactured with additives.
[322] FIG. 38 shows an internal electronic layout according to one embodiment (this FIG. represents how a complex set of electronics can be integrated into a very small robot).
[323] The robot's electronics suite maximizes or increases the robot's endurance between loads, optimizes or increases space utilization, minimizes or reduces interference between devices, and ensures or promotes that high-performance electronics remain sufficiently cool.
[324] In one embodiment, a global navigation satellite system (GNSS) antenna is mounted directly in the center of the robot, and the real-time kinematic differential GNSS module with dual-frequency GPS capability (Piksi Multi, Swift Navigation, USA) is used to acquire centimeter-level precise positional information at a rate of 5 Hz. Another antenna and module are used as a portable base station and transmit differential corrections. A 3-axis gyroscope (STMicroelectronics L3G4200D) is used to obtain yaw rate measurements with an accuracy of 1 degree per second at a rate of 5 Hz. In one embodiment, the system can be modularized with any RTK GPS-enabled antenna.
[325] In one example, there may be four 12 V DC brushed motors with a 131.25:1 metal gearbox, which are capable of driving a coupled wheel at 80 revolutions per minute. Petition 870260066391, dated 06 / 07 / 2026, page 123 / 315 110 / 139 These motors provide the necessary torque and RPM without draining the battery too much, so the robot has at least 8 hours of endurance in field conditions.
[326] In one example, a two-channel Hall effect encoder (e.g., Pololu Corporation, USA) for each DC motor can be connected to measure the wheel speeds. A Sabertooth motor controller (e.g., Dimension engineering, USA) can be used, which is a two-channel motor driver that uses digital control signals to drive two motors per channel (left and right channel) and has a rated supply current of 12 A per channel. A Kangaroo x2 motion controller (e.g., Dimension engineering, USA) can be used, which is a two-channel self-tuning PID controller that uses feedback from the encoders to maintain the desired linear and angular speed commands of the robot.
[327] In an embodiment that refers further to FIG.39, a 1.2 GHz, 64-bit quad-core Raspberry Pi 3 Model B CPU acquires measurements from all available sensors and sends the desired command signals (e.g., desired linear and angular velocities) to the Kangaroo x2 motion controller in the form of two pulse-width modulation signals. The robot estimates its global position by feeding all available measurements from all its onboard sensors (GPS, gyroscope, and encoders) to an online state estimator. In one embodiment, each time the estimated states are updated, they can be fed to a path-following controller that uses these estimated states to calculate the desired linear and angular velocities, within the body structure, required for the robot to follow the path reference given by a trajectory generator. Petition 870260066391, dated 06 / 07 / 2026, page 124 / 315 111 / 139 calculated command speeds can then be sent to the Kangaroo Motion Controller (KMC) as reference command signals, in the form of Pulse-Width Modulation (PWM) signals. The KMC acts as the robot's low-level controller, using feedback from encoders attached to the motors to determine the control signals needed to track the provided reference command signals, ensuring that the robot's desired speeds are maintained. The KMC outputs the modified command signals to the Sabertooth Motor Controller (SMC), which correlates the provided control signals to the required output voltages for the motors. In one embodiment, a cooling system can provide liquid cooling for the NVIDIA Tegra GP-GPU. This cooling system may have an externally mounted radiator that is mounted in such a way as to avoid obstructing the wheels during movement.
[328] With further reference to FIG. 40, a schematic top view of the robot of this embodiment is illustrated showing the center of gravity (CG) and the DC motors / encoders (only one of which is labeled) that allow a desired level of control for traversing crop rows. In one embodiment, the wheels may be critical to ensure that the robot can navigate over crops without damaging them and can navigate over wet and muddy terrain. As an example, the wheels may have paddle-shaped extensions, structures, or arrangements that are designed to provide good traction on loose soils while minimizing the contact area. These paddle-shaped extensions may extend from the wheel and encircle the wheel to make contact with the ground. Unlike tracked robots, this wheel has significant advantages: (1) it does not lead to a large area that is placed under pressure and force as the robot moves, instead Petition 870260066391, dated 06 / 07 / 2026, page 125 / 315 112 / 139 In addition, it limits it to only a small contact area; (2) it is much simpler to manufacture and operate in the field; and (3) it is modular, in the sense that each wheel can be replaced if it breaks, instead of having to replace the entire track.
[329] In one embodiment, the drive mechanism may include motors mounted near each wheel to allow independent drive on all four wheels, without the need to distribute power from a central power unit. This is very different from typical existing equipment and vehicles, which generally use a single power plant and then transmit the power to different wheels. The four-wheel drive mechanism may allow the robot to turn by varying the speeds of the independent wheels. This is a much simpler mechanism, as it does not require complex rack and pinion or other similar drive mechanisms.
[330] Another feature of the wheel and support is the incorporation (in one example) of the suspension without having to increase the size of the robot. The suspensions can be embedded between the wheel mount and the chassis. This has the advantage that the chassis can be enlarged, while also allowing for a simple mechanism to handle uneven agricultural fields.
[331] Reference will now be made to certain experiments that have been carried out to ensure that the robot of a given type does not damage the plants: I. Driving over individual plants: one version of the robot described in this document was driven over individual plants, such as common weeds, soybeans at different growth stages, corn plants at different growth stages, legumes, and shrubs. Damage to the plants was measured by visible damage and the plant's ability to grow. Petition 870260066391, dated 06 / 07 / 2026, page 126 / 315 113 / 139 days after the robot passed and whether the plant thrived days after the driving event. The results were compared with driving over the factories with existing robots / vehicles, including a larger robot, small tractors, and vehicles weighing over 600 lb. All experiments demonstrated that the robot modality disclosed here does not damage the plants, while the other vehicles / robots do. II. Soil Compaction Tests: One embodiment of the robot disclosed in this document was conducted on different soil types and conditions, including clay-based soils in Southeast Asia that were saturated with water and silty clay soils in Illinois with varying moisture contents ranging from dry to supersaturated with water (the day after heavy rain). In all conditions, the robot in this embodiment left tread marks less than 0.3 inches (0.762 centimeters) deep; in drier conditions, tread marks were hardly visible in hard, dry soils. Soil compaction was tested by checking soil density using a probe. The soils driven by the robot embodiment disclosed here were actually found to be looser due to the way the wheels were designed. These results were compared with other robots, small tractors, large equipment, and vehicles.These other robots / vehicles resulted in much deeper soil compaction and tread marks than the robot model disclosed in this document, since none of these other robots / vehicles are designed to drive over plants without damaging them (unlike the robot model disclosed here).
[332] FIGS. 41A-41E represent an embodiment of a robotic platform constructed primarily of printed parts Petition 870260066391, dated 06 / 07 / 2026, page 127 / 315 114 / 139 in 3D (instead of primarily metal parts), which is battery-operated (lithium polymer batteries) and features a highly integrated set of sensors, computing devices, memory, and onboard control and analysis software for recognition. The total cost of this robot is an order of magnitude lower than existing agricultural robots, including commercially available options and many academic prototypes. Furthermore, the robot is also much more compact than many existing robots. This type of robot is capable of turning 180° in a crop row.
[333] FIGS. 42A-42D represent dimensions for one embodiment of the robot. The robot has a mechanism for including suspension on the robot's wheels. This robot is equipped with two side-facing visual spectrum cameras, one front-facing camera, and an Intel RealSense 3D sensor. The robot may also include LED lighting to assist in operation in shaded areas. A Paralyne coating can be provided for water resistance, where a benefit of this approach is that it does not significantly increase the weight or form factor of the robot while still providing water resistance. In this embodiment, the robot can provide autonomous path following with high precision. The robot is fully autonomous, capable of following predefined trajectories with a high level of accuracy. Precise steering through the planting rows prevents crop damage and is one of the most important features for an agricultural robot.In practice, variations in ground conditions can result in off-track navigation due to unknown soil traction coefficients. To solve this problem, the robot in this modality implements RHEC-based navigation applied to a fully autonomous mobile platform that has been rigorously evaluated under these conditions. Petition 870260066391, dated 06 / 07 / 2026, page 128 / 315 Realistic field measurements (115 / 139) yield significantly less than 5 inches (12.7 centimeters) of path tracking error. The guidance and control system merges measurements from an inertial sensor (gyroscope) with RTK-GPS and uses predictive control of the nonlinear model to perform path tracking control. All four wheels are controlled independently. A real-time retreat horizon estimation and control (RHEC) framework is developed and implemented on an embedded computer. In this example, RHE (retreat horizon estimation) is used to estimate constrained states and parameters, and RHC (retreat horizon control) is based on an adaptive system model containing time-varying parameters. The capabilities of the real-time RHEC framework are experimentally verified, and the results show accurate tracking performance on a rough, wet field, as illustrated in FIG. 43.The average values for the Euclidean error and the calculation time required for the RHEC structure are 1.66 inches (4.21 centimeters) and 0.88 milliseconds, respectively.
[334] In one or more embodiments, an agricultural robot is provided with shape, form, weight and material such that it does not damage plants, even if it passes over them or brushes them during operation throughout the season, from plant emergence to post-crop. A locomotion mechanism for the robot enables it to traverse agricultural fields without damaging the plants. An autonomous navigation system comprising sensors, computers and actuators enables the robot to plan and steer in a path in the agricultural field that does not damage the agricultural plants and to plan and conduct crop management actions in a manner that does not damage the plants. Petition 870260066391, dated 06 / 07 / 2026, page 129 / 315 116 / 139
[335] In one embodiment, the robot may be provided with an embedded sensor system, algorithms, and software for onboard agronomic functions without the need for cloud connectivity or connectivity with other computers / robots. The robot may perform various data collection functions, including stand counting, stem angle estimation, ear height and plant height estimation, disease identification (e.g., rust, mosaic, fungus), insect damage identification (e.g., leaf damage, stem damage, insect nest identification), insect detection, plant and soil sampling, and / or sample retrieval that can be returned to a ground station. In one embodiment, the robot may include a communication and coordination system that allows a remote user to control one or more autonomous robots in a manner that does not damage agricultural plants.The communication and coordination system can allow a team of robots to communicate and collaborate to obtain agronomic measurements without damaging plants, determining the least damaging paths and allocating portions of the field to specific robots to reduce the areas traveled and increase the speed of obtaining agronomic information.
[336] In one embodiment, the communication and coordination system enables a team of robots to communicate and collaborate to conduct field management, including, but not limited to, one or more of the following: weed removal, disposal, sampling, spraying, pruning, cultivation without damaging the agricultural plants, by determining the least damaging paths and allocating portions of the field to specific robots to reduce the areas traveled and increase the speed of performing management tasks. Petition 870260066391, dated 06 / 07 / 2026, page 130 / 315 117 / 139
[337] In one embodiment, a field-deployed automatic recharging system is provided that enables the robot to recharge its batteries including, but not limited to, one or more connected recharges, inductive charging, battery swapping, or wireless charging. In one or more embodiments, a field-deployable robot is capable of assisting other robots through activities such as repairing, rescuing, or recharging other robots in a manner that does not damage cultivated plants.
[338] Precise guidance through crop rows that avoids crop damage can be an important task for agricultural robots used in various field operations. In one or more modalities, variable soil conditions that normally result in off-track navigation due to unknown traction coefficients and that can typically cause crop damage can be accounted for by various modalities. For example, a real-time receding horizon estimation and control (RHEC) framework can be applied to a fully autonomous field robot to increase its steering accuracy.
[339] Faster and lower-cost microprocessors, as well as advances in solution methods for nonlinear optimization problems, enable nonlinear retreat horizon control (RHC) and retreat horizon estimation (RHE) methods to be used for field robots requiring high-frequency (millisecond) updates. In one embodiment, a real-time RHEC framework is developed and applied to a fully autonomous mobile robotic platform for field phenotyping applications, such as in sorghum fields (although embodiments can be used in various fields with multiple crops). RHE can be used to estimate states and parameters. Petition 870260066391, dated 06 / 07 / 2026, p. 131 / 315 118 / 139 restricted, and is formulated below: £ +Συ-«+1Ιύ™(4) P - P HN st z(t) = / min — *>(0 — / max Pmin — p(0 — Pmax (1C) where ξ, u, pe z are respectively the state, the input parameters and the output vectors. RHC is designed based on the adaptive system model with time-varying parameters and formulated below: | ΚΣ^+i1HfrCt.) - + IMt) - + HUt / c+iv) - ξ (tfc+«) II LI í =s I ζ-inin — ζ (0 — tmax I Umin — u(t) A Nuix (2C)
[340] In one embodiment, a 3D-printed field robot, as shown in FIGS. 43A-43D, can be formulated with the following equations: X = uvcos Θ (3C) y = μν sen θ (40) θ = κω (50) where χ and y denote the position of the field robot, denotes the yaw angle, v denotes the speed, ω denotes the yaw rate and μ and κ denote the traction parameters. Petition 870260066391, dated 06 / 07 / 2026, p. 132 / 315 119 / 139
[341] The results show accurate tracking performance on a field of uneven and wet soil (see FIGS. 44A and 44B). The average Euclidean error values for the RHEC-based and EKF-based RHCs are 0.0423 m and 0.0514 m, respectively. The RHEC framework benefits from traction parameter estimates and results in fewer errors compared to the EKF-based RHC. The available space on each side of the robot (in this embodiment) is limited to only 0.12 m; therefore, the error must be less than this limit to avoid crop damage and keep the robot centered on the row. The results of several experiments indicated that the RHEC framework does not violate this error constraint, while the EKF-based RHC violates it 17 times during straight-line path tracking. This demonstrates the capability of the RHEC framework. Furthermore, the required calculation time of the RHEC framework is respectively equal to 0.88 ms.
[342] In another embodiment, the nonlinear system and measurement models can be represented by the following equations: ξ(>) = ifçtl('! / >(')) x / (6C) (7C) Where ξ is the state vector, li is the control input, j) is the system parameter vector, Z is the measured output, R is the continuously differentiable state update function and / (O.Oj)) =OVr, e is the measurement function. The derivative of ξ with respect to t is denoted by í and KC
[343] A schematic diagram of a 3D-printed field robot of one embodiment is illustrated in FIG. 40. An adaptive nonlinear kinematic model is derived for the 3D-printed field robot as an extension of the traditional kinematic model. Petition 870260066391, dated 06 / 07 / 2026, page 133 / 315 120 / 139 as follows: gvcosG μ vjsin Θ KÚ) (8C) where xey denote the position of the field robot, Θ denotes the yaw angle, ω denotes the velocity, v denotes the yaw rate, and μ ek denote the traction parameters. The difference between the traditional model and the one developed above is two traction parameters that change slowly. These parameters provide for minimizing deviations between the real-time system and the 3D-printed field robot with an online parameter estimator. It is noteworthy that they must be (in this embodiment) between zero and one.
[344] The state, parameter, input and output vectors are respectively denoted as follows: (90) (10C) ω (11C)ΊT xyv ω (12C)
[345] Although model-based controllers need complete state and parameter information to generate a control signal, the number of sensors is less than the number of incommensurable states and parameters in practice. Therefore, state estimators are needed to estimate states and Petition 870260066391, dated 06 / 07 / 2026, page 134 / 315 121 / 139 incommensurable parameters. The extended Kalman filter is the best-known method for nonlinear systems. However, it is not capable of handling constrained nonlinear systems. Estimates of traction parameters can play a vital role for the 3D printed robot, and there are constraints on these parameters, which makes the extended Kalman filter inconvenient for this system. The RHE approach has the ability to handle state and parameter constraints and is formulated for the 3D printed robot (of this embodiment) as follows: £ > + Σ i=k-N+1 J st ξ(ζ)= / (ξ(ί),ι / (ζ),ρ) z(r) = ι(ξ(ι),ι<(ι),ρ) < μ and κ < 1 Ví E (13C) where the deviations of the state and parameter estimates before the estimation horizon are minimized by a symmetric positive semidefinite weighting matrix HN and the deviations of the measured results and the system at the estimation horizon are minimized by a symmetric positive semidefinite weighting matrix Hk. The estimation horizon is represented by N, and the lower and upper bounds on the traction parameters pk parameters are respectively defined as 0 and 1. The objective function in the RHE formulation consists of two parts: the arrival and the quadratic costs. The arrival cost represents the first measurements: Petition 870260066391, dated 06 / 07 / 2026, page 135 / 315 122 / 139 ' = KOΛ-ΛA 1 and the quadratic cost represents the recent measurements:
[0001] The measurements were perturbed by Gaussian noise with standard deviations of õx = õy = 0.03 m, õw = 0.0175 rad / s, õv = 0.05 m / s based on experimental analysis. Therefore, the following weighting matrices Hk and HN are used in nonlinear RHE: TT, · / 2 2 2 2\-l ulUg\(5x> C7y , Gj, , ) = ^(0.032,0.O32,O.52,O.O1752)-1(14C) Ηχ = diug(x2, y2, Θ2, v2, μ2fk2)-1= i / / ^(]0.02,10.02,0J2?L02,0.252,0.252)-](15C)
[346] The inputs to the RHE algorithm are the position values from the GNNS, the speed values from the encoders mounted on the CG engines, and the yaw rate values from the gyroscope. The outputs to the RHE are the position in the xey coordinate system, the yaw angle, the speed, and the thrust coefficients. The estimated values are then provided to the RHC.
[347] The RHC approach is usable for a system with fast dynamics. The advantage is that RHC has the ability to handle rigid state and input constraints, and online optimization allows updating costs, model parameters, and constraints. The following finite horizon optimal control formulation for the 3D printed robot (of this embodiment) Petition 870260066391, dated 06 / 07 / 2026, page 136 / 315 123 / 139 is resolved to obtain the current control action using the current state of the system as the initial state: I ( <·*+«-! 2({ F ΙΙξ<·(ίί)-ξ(6·)|Ιβ, + ||«,(6·)-«(Γ;)ΙΙβ} + II'^r!.!k+jv ) — SÍ'* hv ) ||)Λ. J S.t. ξ(ί*) = ξ(α) ξ(ί) = / (ξ(ί),»(Ο,ρ) — 0.1 rad / s < ω(ζ) < 0.1 rad / s ΐ e [fy+i ,4+iy_i] (16C) Where Qke R'^, R € R™ and are symmetric and positive semidefinite weighted matrices, ζ,Γe it, , are the reference state and input, ς and κ , are the states and inputs,(a is the current time, / V is the forecast horizon, çjf / j.) is the current estimate. The first term in the cost function is the stage cost and is the cost along the forecast horizon - the second term in the cost function is the terminal penalty and is the cost at the end of the forecast horizon. The terminal penalty is stated for stability reasons
[11] .
[348] The first element of the ideal control sequence is applied to the system: ''(A+h^t+i) =if*(A+i) (17C) and then the procedure is repeated for future sampling times, changing the prediction horizon to the subsequent time instant. It is important to note that the control input x' is exactly the same as it would be if all immeasurable states and parameters acquired values equal to their estimates based on the estimate up to the current time tk due to the certainty equivalence principle.
[349] The state reference for the 3D printed robot (this Petition 870260066391, dated 06 / 07 / 2026, p. 137 / 315 124 / 139 modality) is changed online and defined as follows: ξΓ= [à>, er]rand ur= (18C) ωΓwhere xr and yr are the position references, is the yaw rate reference calculated from the position references as follows: Θ}· = atan2 (yr, à> ) + λ π (19C) where λ describes the desired direction of the 3D printed field robot (λ = 0 forward and λ = 1 backward). The yaw rate reference can be calculated from the reference trajectory as the yaw angle reference. However, steady-state error may occur in the case of a mismatch between the system model and the 3D printed robot. Therefore, the recently measured yaw rate is used as the input reference to penalize the input rate in the objective function discussed in this document.
[350] In one or more embodiments, an agricultural robot is provided that does not damage cultivated plants, even when passing over them, through the use of a significantly small robot weight, although the robot can withstand adverse field conditions; the material from which the robot is constructed provides the correct ratio of strength to weight and does not damage the plants; the robot's structure has adequate ground clearance, absence of sharp edges, and a wheel design appropriate to be able to pass over plants without damaging them; control and analytics algorithms take into account the physical constraints of the robot and its environment; and / or the robot's manufacturability, operational paradigm, and / or use case reduces or minimizes the cost of manufacturing, operation, and / or Petition 870260066391, dated 06 / 07 / 2026, page 138 / 315 125 / 139 property.
[351] From the descriptions in this document, it would be evident to a technician with ordinary skill in the art that the various embodiments can be modified, reduced, or enhanced without departing from the scope and spirit of the claims described below. For example, the sensing functions can be implemented dynamically or controlled in another way. In one example, different groups of sensors can be triggered based on various factors such as weather, crop type, location, crop size, robot speed, and so on. In one embodiment, the dynamic control of the sensing functions can be autonomous without user action. In another embodiment, the dynamic control of the sensing functions can be based on user action or can be initiated after user authorization, such as a user responding to a message from the robot suggesting that a particular sensor or algorithm be triggered or executed.
[352] In one embodiment, the data collected by the robot can trigger the collection of different data. For example, a robot can detect a condition associated with crops and can trigger one or more sensors that attempt to detect the causes of that condition (e.g., diseases, insects, and so on). In one embodiment, the triggering of a new data collection technique by the robot can be autonomous without user action; based on user action; or it can be initiated after user authorization, such as a user responding to a message from the robot suggesting that a particular sensor or algorithm be triggered or executed.
[353] In one embodiment, a robot can provide notice to another robot(s) and / or to a central processing system indicating specific data that must be retrieved. For example, a Petition 870260066391, dated 06 / 07 / 2026, page 139 / 315 126 / 139 The first robot can collect data indicating the existence of a particular condition based on a threshold comparison associated with the collected data. The first robot can then wirelessly transmit a message to one or more second robots (with or without user notice or authorization) to cause the second robots to begin collecting the same data associated with the condition.
[354] In one embodiment, the robot may be part of a group of robots that are collecting data at a location. In another embodiment, the group of robots may be in a master / slave relationship such that one or more master robots may control one or more slave robots, such as control over the types of data collection, navigation paths, operational parameters (e.g., speed), and so forth. In one embodiment, the group of robots may be similar robots with similar operational capabilities (such as the ability to collect the same types of data). In another embodiment, the group of robots may be different robots with one or more different operational capabilities (such as the ability to collect different types of data).
[355] In one embodiment, the robot can implement different algorithms for data collection based on certain factors, such as time, location, weather, robot speed, data received from other sources (e.g., other robots, smart agricultural equipment, and so on).
[356] In one embodiment, a group of robots can collaborate to provide an optimized and / or efficient data collection technique for a specific location. For example, the location can be divided into parts, where each available robot performs all data collection for that part. In this example, the Petition 870260066391, dated 06 / 07 / 2026, page 140 / 315 127 / 139 portions can be of different sizes, depending on the different speeds of the robots. As another example, data collection can be divided into tasks where each available robot performs some of the tasks for all or part of the site. In this example, tasks can be shared among several robots, and the robots can be selected based on capabilities such as speed, processing resources, sensing components, and so on.
[357] In one embodiment, historical information can be used by the robot to adjust its operations. For example, the robot can select a particular set of data to be collected based on historical information indicating that the crop was previously susceptible to a particular insect or disease. Historical information can be obtained from various sources, including other robots in other locations. In one embodiment, the robot can travel at a speed of 1 mph, 2-3 mph, and / or 4-5 mph.
[358] In one embodiment, the robot and / or team of robots and / or a web-based computer uses machine learning to: detect and identify phenomena relevant to agriculture using data from a single or plurality of sensors, such as weeds, stressors, insects, diseases, etc.; and / or to measure agronomic quantities, such as stem width, stand count, plant height, ear height, etc.
[359] As described in this document, the aspects of the disclosure in question may include, for example, a device having a processing system that receives speed feedback signals associated with each of the device's wheels, which receives position information from a global navigation satellite system of the device, which receives measurement signals from a Petition 870260066391, dated 06 / 07 / 2026, page 141 / 315 The device's gyroscope, number 128 / 139, determines control signals based on desired angular and linear velocities according to velocity feedback signals, positional information, and measurement signals, applying real-time horizon estimation and control. It transmits these control signals to a controller coupled to electric motors, causing the device, operating as an agricultural robot, to track a reference path. Each electric motor is connected to one of the wheels, allowing the agricultural robot to rotate 180 degrees. Additional embodiments are disclosed.
[360] In one embodiment, an apparatus is provided comprising: a frame; wheels rotatably connected to the frame; a battery connected to the frame; electric motors supported by the frame and connected to the battery, wherein each of the electric motors is connected to one of the wheels, wherein the electric motors enable the apparatus to rotate 180 degrees; a processor supported by the frame; speed sensors supported by the frame and coupled to the processor, wherein each of the speed sensors is connected to one of the wheels, wherein the speed sensors transmit speed feedback signals associated with each of the wheels to the processor; a global navigation satellite system (GNSS) supported by the frame and comprising a GNSS antenna and GNSS computer coupled to the processor, wherein the global navigation satellite system receives, processes and transmits position information to the processor;a gyroscope supported by the structure and coupled to the processor, wherein the gyroscope transmits measurement signals including yaw rate measurements to the processor; and a group of sensors; Petition 870260066391, dated 06 / 07 / 2026, page 142 / 315 129 / 139 supported by the structure and coupled to the processor, where the sensor group collects phenotypic data associated with the crops, where the processor determines the control signals based on the desired angular and linear velocities according to the velocity feedback signals, the positional information and the measurement signals applying a real-time estimation and control of the receding horizon, and where the processor transmits the control signals to each of the electric motors, causing the device to track a reference path.
[361] In one example, the device weighs less than 20 pounds (9.07 kilograms).
[362] In another example, each of the wheels includes paddle structures that encircle the wheel and selectively contact the ground.
[363] In another example, the processor transmits the control signals without receiving command signals from a user on a remote source.
[364] In another example, the battery is chargeable by an implantable charging device via inductive charging, wireless charging, a physical connection, or a combination thereof.
[365] In another example, the processor selects an algorithm from among a group of algorithms associated with one of the sensor groups without receiving command signals from a user on a remote source.
[366] In another example, the processor selects an algorithm from among a group of algorithms associated with one of the sensor groups based on an analysis of at least some of the phenotypic data and without receiving command signals from a user on a remote source.
[367] In another example, the processor acts as one of the group of Petition 870260066391, dated 06 / 07 / 2026, page 143 / 315 130 / 139 sensors not receiving command signals from a user at a remote source.
[368] In another example, the processor acts on one of the sensor groups based on an analysis of at least some of the phenotypic data and without receiving command signals from a user on a remote source.
[369] In another example, the processor transmits a command signal to another device, wherein the command signal causes the other device to select an algorithm from among a group of algorithms, activate a sensor or a combination thereof, wherein the other device is an autonomous agricultural robot that collects data associated with crops.
[370] In another example, the processor generates the command signal based on an analysis of at least some of the phenotypic data and without receiving other command signals from a user on a remote source.
[371] In another example, the device is part of a group of devices in which each operates as autonomous agricultural robots that collect phenotypic data associated with crops, where the group of devices communicates with each other to assign data collection paths, data collection tasks, or a combination thereof.
[372] In another example, the device group is in a master / slave device relationship that allows one or more master devices in the device group to control one or more slave devices in the device group.
[373] In another example, the processor receives a remote control signal from a remote source that causes the processor to adjust the operation of the device.
[374] In another example, the components of the device, including Petition 870260066391, dated 06 / 07 / 2026, page 144 / 315 131 / 139 the frame, are made of 3D printed material.
[375] In another embodiment, an apparatus is provided comprising: a frame; wheels rotatably connected to the frame; a battery connected to the frame; electric motors supported by the frame and connected to the battery, wherein each of the electric motors is connected to one of the wheels, wherein the electric motors enable the apparatus to rotate 180 degrees; a processor supported by the frame; speed sensors supported by the frame and coupled to the processor, wherein each of the speed sensors is connected to one of the wheels, wherein the speed sensors transmit speed feedback signals associated with each of the wheels to the processor; a global navigation satellite system (GNSS) supported by the frame and comprising a GNSS antenna and GNSS computer coupled to the processor, wherein the global navigation satellite system receives, processes and transmits position information to the processor;a gyroscope supported by the structure and coupled to the processor, wherein the gyroscope transmits measurement signals including yaw rate measurements to the processor; and a group of sensors supported by the structure and coupled to the processor, wherein the sensor group collects phenotypic data associated with the crops, wherein the processor determines the control signals based on the desired angular and linear velocities according to the velocity feedback signals, positional information and measurement signals, wherein the processor transmits the control signals to each of the electric motors, causing the device to track a reference path, wherein the processor generates data collection signals to select or adjust one or more data collection algorithms, one or more of them; Petition 870260066391, dated 06 / 07 / 2026, page 145 / 315 132 / 139 the group of sensors or a combination thereof, thus adjusting the collection of phenotypic data associated with crops, wherein the data collection signals are generated by the processor according to an analysis of at least some of the phenotypic data, other information obtained by the processor, or a combination thereof, and wherein the data collection signals are genes evaluated by the processor without receiving command signals from a user on a remote source.
[376] In one example, the other information obtained by the processor comprises historical phenotypic data associated with a crop location, meteorological information, weather information, or a combination thereof.
[377] In another example, the device weighs less than 20 pounds (9.07 kilograms) and each of the wheels includes paddle structures that encircle the wheel and selectively contact the ground.
[378] In another embodiment, a method is provided comprising: receiving, by a processing system of a wheeled agricultural robot, velocity feedback signals associated with each of the wheels, wherein the velocity feedback signals are received from agricultural robot velocity sensors, wherein each of the velocity sensors is connected to one of the wheels; receiving, by the processing system, position information from a global navigation satellite system (GNSS) of the agricultural robot; receiving, by the processing system, measurement signals from a gyroscope of the agricultural robot, wherein the measurement signals include yaw rate measurements; determining, by the processing system, the control signals based on the desired angular and linear velocities according to the velocity feedback signals, the position information and the measurement signals by applying a Petition 870260066391, dated 06 / 07 / 2026, page 146 / 315 133 / 139 real-time estimation and control of the receding horizon; and transmit, through the processing system, the control signals to a controller coupled to electric motors, causing the agricultural robot to track a reference path, in which each of the electric motors is connected to one of the wheels and allows the agricultural robot to rotate 180 degrees.
[379] In one example, the method further comprises: receiving, by the processing system of a group of sensors of the agricultural robot, phenotypic data associated with the crops; and generating, by the processing system, data collection signals to select or adjust one or more data collection algorithms, one or more of the sensor group or a combination thereof, thus adjusting the collection of phenotypic data associated with the crops, wherein the data collection signals are generated by the processing system according to an analysis of at least some of the phenotypic data, other information obtained by the processing system or a combination thereof, and wherein the data collection signals are generated by the processing system without receiving command signals from a user on a remote source.
[380] With reference now to FIG. 45, an embodiment of a robotic chemical treatment system is illustrated that is capable of spraying chemicals under the canopy in various crops. The system may include a small autonomous robot 4502 (e.g., less than 50 lb and / or less than 20 inches (50.8 centimeters)) with various sensors and actuators. The system may include a replaceable container 4504 containing liquid, which may be pressurized and / or which may have multiple chambers containing different liquids. In one example, the container may be configured for mounting on, and / or on, the robot 4502. In another example, the canister may be configured to tow behind the robot 4502 (e.g., Petition 870260066391, dated 06 / 07 / 2026, page 147 / 315 134 / 139 in a trailer, not shown). The system may include mechanisms for accepting one or more containers (e.g., two). The system may include mechanisms for spraying the liquid, including one or more replaceable nozzles. The mechanism may include a pressurization system. The system may include an actuator system with one or more actuators to operate the spraying mechanism (e.g., four), controlling the rate, duration, direction, and force of spraying. The system may include a liquid pressurization system that may include motors, pumps, pressurized gas containers (e.g., CO2, N, or others). The system may include a remote operator interface connected via a network, allowing a user to control the robot and the spraying, including the ability to upload a prescription from another device on an interconnected network.The robot may have the capability to autonomously follow the prescribed path and spray according to the prescription. The robot may have the capability to automatically decide when, what, and where to spray, analyzing agronomic quantities through onboard sensors while autonomously following recognized paths. The system may generate an ideal path based on field conditions and / or data from other connected devices (e.g., the device in FIG. 36) to achieve prescription targets and / or mission objectives. The system may adapt the path and spray prescription using data obtained from onboard sensors and / or other connected devices to satisfy higher-level mission objectives, such as finding and spraying appropriate chemicals on plants with specific diseases. The system may perform one or more of the capabilities described above with or without user interaction with the system. Petition 870260066391, dated 06 / 07 / 2026, page 148 / 315 135 / 139
[381] In another embodiment, a drone (e.g., robot) may fly over a crop field for the first time (as a first stage) to collect data. Then, the drone may fly over the crop field a second time (in a second stage) to confirm the cause of an anomaly and / or apply a treatment. In one example, the path taken by the drone during the first stage may be the same as the path taken by the drone during the second stage. In another example, the path taken by the drone during the second stage may be different from the path taken by the drone during the first stage. In yet another example, the drone may be fitted with a sprayer attachment (or other similar device) and the confirmation may be based on some level of machine learning that was performed in the past (e.g., before the second stage).
[382] In another embodiment, a swarm (or group) of drones (e.g., robots) can work together. For example, one or more drones may first collect data on one (or more) first-stage paths (as discussed above). The collected data can then be analyzed (the analysis can be performed on board one or more of the drones and / or on a ground-based computing platform). Then, one or more different drones carrying the chemical(s) needed to treat the problem(s) in each of the identified areas can go to the area(s) to treat one or more problems.
[383] As described in this document, in one embodiment, the robot(s) that apply the treatment are the same ones that performed the identification (for example, the identification of an anomaly).
[384] As described in this document, in one embodiment, the robot(s) that apply the treatment are different from that(s) Petition 870260066391, dated 06 / 07 / 2026, p. 149 / 315 136 / 139 who made the identification (for example, the identification of an anomaly).
[385] In various contexts, the term drone can refer to any small robot, whether aerial or terrestrial.
[386] In one embodiment, a communication system may enable a device (e.g., a mobile robot) to communicate wirelessly (e.g., bidirectionally) with a remote computer from which the device may receive one or more instructions and / or send information,
[387] In one embodiment, the size and / or location of a window associated with foreground extraction can be changed. In one example, the size and / or location can be changed by an onboard computer or processor (e.g., onboard a mobile robot). In another example, the size and / or location can be communicated to the device (e.g., mobile robot) by an external computer or processor.
[388] In another embodiment, a support vector machine may comprise a smooth margin support vector machine that applies a slack variable to data points on or within a correct decision boundary.
[389] In another embodiment, width estimation may involve the use of foreground extraction and / or LIDAR depth estimation(s).
[390] In another embodiment, a distance for a line can be estimated using a single ranging device or a plurality of ranging devices. The ranging device or devices can emit electromagnetic, laser, infrared and / or acoustic waves.
[391] In another embodiment, a device (for example, a mobile robot) may have multiple cameras. In one example, the device Petition 870260066391, dated 06 / 07 / 2026, page 150 / 315 137 / 139 (e.g., mobile robot) can use an integrated processing system to combine information from multiple cameras to calculate the distance to the line(s).
[392] In introducing elements from various embodiments of this disclosure, the articles a, an, and said are intended to signify that there is one or more of the elements. The terms comprising, including, and having are intended to be inclusive and mean that there may be additional elements beyond those listed. Furthermore, any numerical examples in the discussion in this document are intended to be non-limiting and, therefore, additional numerical values, ranges, and percentages are within the scope of the embodiments disclosed.
[393] The Summary Disclosure is provided with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, in the preceding Detailed Description, it can be seen that several features are grouped into a single embodiment for the purpose of expediting disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than are expressly stated in each claim. Instead, as the following claims reflect, the inventive subject matter lies in less than all the features of a single disclosed embodiment. Thus, the following claims are incorporated herein in the Detailed Description, with each claim standing in itself as a separately claimed subject matter.
[394] Although the disclosure has been described in detail in connection with a limited number of embodiments, it should be readily understood that the invention is not limited to such disclosed embodiments. Instead, the disclosed may be Petition 870260066391, dated 06 / 07 / 2026, pp. 151 / 315 138 / 139 modified to incorporate any number of variations, alterations, substitutions or equivalent arrangements not described herein, but which are proportionate to the spirit and scope of the invention. Furthermore, although several embodiments of the invention have been described, it should be understood that the disclosed aspects may include only some of the described embodiments.
[395] One or more embodiments may utilize one or more features (for example, one or more of the systems, methods, and / or algorithms) described in U.S. Provisional Patent Application No. 62 / 688,885, filed June 22, 2018 (including each appendix attached thereto). The one or more features (for example, one or more of the systems, methods, and / or algorithms) described in U.S. Provisional Patent Application No. 62 / 688,885, filed June 22, 2018 (including each appendix attached thereto) may be used in place of and / or in addition to one or more features described in this document in relation to the various embodiments. The disclosure of U.S. Provisional Patent Application No. 62 / 688,885, filed June 22, 2018 (including each appendix attached thereto) is incorporated herein by reference in its entirety.
[396] One or more embodiments may utilize one or more features (for example, one or more of the systems, methods, and / or algorithms) described in U.S. Provisional Patent Application No. 62 / 596,506, filed December 8, 2017 (including each attached Appendix). The one or more features (for example, one or more of the systems, methods, and / or algorithms) described in U.S. Provisional Patent Application No. 62 / 596.506, filed on December 8, 2017 (including each (Annex attached to it) can be used in place of and / or in addition to Petition 870260066391, dated 06 / 07 / 2026, page 152 / 315 139 / 139 of one or more features described in this document in relation to the various embodiments. The disclosure of U.S. Provisional Patent Application No. 62 / 596,506, filed December 8, 2017 (including each appendix attached thereto) is incorporated by reference herein in its entirety.
[397] One or more embodiments may utilize one or more features (for example, one or more of the systems, methods, and / or algorithms) described in U.S. Provisional Patent Application No. 62 / 550,271, filed August 25, 2017 (including each appendix attached thereto). The one or more features (for example, one or more of the systems, methods, and / or algorithms) described in U.S. Provisional Patent Application No. 62 / 550,271, filed August 25, 2017 (including each appendix attached thereto) may be used in place of and / or in addition to one or more features described in this document with respect to the various embodiments. The disclosure of U.S. Provisional Patent Application No. 62 / 550,271, filed August 25, 2017 (including each attached Appendix), is incorporated by reference herein in its entirety.
Claims
1. Device characterized in that it comprises: a processing system (2102) including a processor; a communication system (2120) enabling the device to communicate wirelessly with a remote computer from which the device can receive instructions and send information;and a memory (2104) that stores executable instructions which, when executed by the processing system (2102), perform operations, the operations comprising: obtaining video data from a single monocular camera (404), wherein the video data comprises a plurality of frames, wherein the single monocular camera (404) is attached to a mobile ground robot (402) that is traveling along a track defined by a line of crops, wherein the mobile ground robot (402) comprises a plurality of wheels and a plurality of encoders, wherein each of the plurality of wheels has associated with it a respective encoder of the plurality of encoders, wherein the line of crops comprises a first plant stem (1102), and wherein the plurality of frames includes a representation of the first plant stem (1102);obtain robot speed data from a plurality of encoders, where robot speed data is determined based on an average of encoder values from all encoders in the plurality of encoders; perform foreground extraction on each of the plurality of frames from the video data, where foreground extraction results in a plurality of foreground images;and determine, based on the plurality of foreground images and based on robot speed data, an estimated width of the first plant stem (1102), wherein the determination comprises determining a ratio R, wherein R = Vr / Vx, wherein Vr is an instantaneous robot speed obtained from the robot speed data and Vx is an average horizontal foreground pixel speed obtained from camera motion estimation using optical flow calculated from the plurality of frames of the video data.; 2. Device according to claim 1, characterized in that the foreground extraction comprises processing, for each of the plurality of frames of the video data, only a fixed-size window (702) that is smaller than each of the plurality of frames wherein the fixed-size window (702) is a single window and is the only window on which the foreground extraction processing is applied; and / or wherein a window width of the window (702) is equal to a frame width of each of the plurality of frames divided by 4.
3. Device according to claim 2, characterized in that the fixed-size window (702) associated with each of the plurality of frames is located off-center in each of the plurality of frames.
4. Device according to claim 1, characterized in that the foreground extraction comprises, for each of the plurality of frames of the video data, a first function to perform edge detection.
5. Device, according to claim 4, characterized in that the foreground extraction further comprises, for each of the plurality of frames of the video data, a second function to perform morphological processing. Petition 870260066391, dated 06 / 07 / 2026, page 155 / 315 3 / 11 6. Device according to claim 5, characterized in that the foreground extraction further comprises, for each of the plurality of frames of the video data, a third function to perform connected component labeling.
7. Device according to claim 1, characterized in that determining the estimated width of the first plant stem (1102) comprises determining, based on the plurality of frames of the video data, an estimated camera movement using optical flow calculated from the plurality of frames of the video data.
8. Device according to claim 1, characterized in that the determination of the estimated width of the first plant stem (1102) further comprises: determining a first width, in pixels, at a first location of the first plant stem (1102), as represented in a first of the plurality of frames of the video data; and multiplying R times the first width, resulting in a first value.
9. Device according to claim 8, characterized in that determining the estimated width of the first plant stem (1102) further comprises: determining a second width, in pixels, at a second location of the first plant stem (1102), as represented in the first of the plurality of frames of the video data; multiplying R times the second width, resulting in a second value; and Petition 870260066391, dated 06 / 07 / 2026, page 156 / 315 4 / 11 calculating the average of the first value and the second value, resulting in the estimated width of the first plant stem (1102).
10. Device according to claim 8, characterized in that the size and location of a window (702) associated with foreground extraction can be changed, and in that the size and location can be changed by an onboard computer (2100) or communicated to the device by an external computer.
11. Device according to claim 1, characterized in that the operations further comprise: obtaining additional video data from the single monocular camera (404), wherein the additional video data comprises an additional plurality of frames, wherein the crop row comprises a second plant stem (1102), and wherein the additional plurality of frames includes another representation of the second plant stem (1102); obtaining additional robot speed data from the plurality of encoders, wherein the additional robot speed data is determined based on an average of encoder values from all encoders in the plurality of encoders; performing additional foreground extraction on each of the additional plurality of frames from the additional video data, wherein the additional foreground extraction results in an additional plurality of foreground images;and determine, based on the additional plurality of foreground images and based on the additional robot speed data, an additional estimated width of the second plant stem (1102).; 12. Device according to claim 1, Petition 870260066391, dated 06 / 07 / 2026, page 157 / 315 5 / 11 characterized in that the mobile ground robot (402) is an autonomous mobile robot, and in that the operations are performed without the use of global positioning system (GPS) data.
13. Device according to claim 1, characterized in that the operations further comprise: applying a convolutional neural network to the plurality of frames of the video data to extract feature vectors; applying a support vector machine to the feature vectors to classify the feature vectors, wherein the application of the support vector machine results in classified features; applying motion estimation based on a rigid transformation to the classified features, wherein the application of the motion estimation results in corrected count data; and determining a plant population based on the classified features and the corrected count data.
14. Device according to claim 1, characterized in that the mobile ground robot (402) further comprises: a plurality of wheels; one or more batteries; a plurality of electric motors, wherein each of the plurality of electric motors is electrically connected to one or more batteries, wherein each of the plurality of electric motors is connected to one of the plurality of wheels, wherein the plurality of electric motors allows the mobile ground robot (402) to rotate at least 180 degrees, wherein each of the plurality of encoders is connected to one of the plurality of wheels, and wherein each of the plurality of encoders transmits a respective speed feedback signal associated with a respective of each of the plurality of wheels to the processing system (2102);a global navigation satellite system (GNSS) comprising a GNSS antenna and GNSS computer coupled to the processing system (2102), wherein the GNSS receives input position information, processes the input position information, and transmits output position information to the processing system (2102); a gyroscope coupled to the processing system (2102), wherein the gyroscope transmits measurement signals including yaw rate measurements to the processing system, wherein the operations determine one or more control signals based on a desired angular velocity and a desired linear velocity, and wherein the one or more control signals are determined according to the velocity feedback signals, the position information, and the measurement signals by applying a real-time receding horizon estimate and control;and wherein the operations transmit one or more control signals to one or more of the plurality of electric motors, causing the mobile ground robot (402) to track a reference path.; 15. Apparatus characterized in that it comprises: a plurality of wheels; one or more batteries; a plurality of electric motors, wherein each of the plurality of electric motors is electrically connected to one or more batteries, wherein each of the plurality of electric motors is connected to one of the plurality of wheels, and wherein the plurality of electric motors allows the apparatus to rotate at least 180 degrees; one or more processors;a plurality of speed sensors coupled to one or more processors, wherein each of the plurality of speed sensors is connected to one of the plurality of wheels, wherein each of the plurality of speed sensors transmits a respective speed feedback signal associated with a respective of each of the plurality of wheels to the one or more processors, wherein the plurality of speed sensors comprises a plurality of encoders, wherein each of the plurality of wheels has associated with it a respective encoder of the plurality of encoders, and wherein robot speed data are determined based on an average of encoder values of all encoders of the plurality of encoders;a global navigation satellite system (GNSS) comprising a GNSS antenna and GNSS computer coupled to one or more processors, wherein the GNSS receives input position information, processes the input position information, and transmits output position information to the one or more processors; a gyroscope coupled to the one or more processors, wherein the gyroscope transmits measurement signals including yaw rate measurements to the one or more processors; and a plurality of sensors coupled to the one or more processors, wherein the plurality of sensors collects phenotypic data associated with crops;wherein one or more processors determine one or more control signals based on a desired angular velocity and a desired linear velocity, and wherein the one or more control signals are determined according to velocity feedback signals, position information, and measurement signals applying real-time receding horizon estimation and control; wherein the one or more processors transmit the one or more control signals to one or more of the plurality of electric motors, causing the apparatus to track a reference path; wherein the plurality of sensors comprises a camera (404) that is attached to the apparatus; wherein the apparatus comprises a mobile ground robot (402); and wherein the one or more processors facilitate operations, the operations comprising: applying a convolutional neural network to a plurality of frames from the camera (404) to extract feature vectors;Apply a support vector machine to the feature vectors to classify the feature vectors, wherein the application of the support vector machine results in classified features; apply motion estimation based on a rigid transformation to the classified features, wherein the application of motion estimation results in corrected count data; and determine a plant population based on the classified features and the corrected count data.
16. Apparatus according to claim 15, characterized in that the operations further comprise: obtaining video data from the camera (404), wherein the video data comprise a plurality of frames, in which the mobile ground robot (402) travels along a track defined by a line of crops, wherein the line of crops comprises a first plant stem (1102), and wherein the plurality of frames includes a representation of the first plant stem (1102); to obtain robot speed data from the plurality of encoders, wherein the robot speed data are determined based on an average of encoder values from all encoders in the plurality of encoders; performing foreground extraction on each of the plurality of frames of the video data, wherein the foreground extraction results in a plurality of foreground images;and determine, based on the plurality of foreground images and based on robot speed data, an estimated width of the first plant stem (1102), wherein the determination comprises determining a ratio R, wherein R = Vr / Vx, wherein Vr is an instantaneous robot speed obtained through robot speed data and Vx is an average horizontal foreground pixel speed obtained through camera motion estimation using optical flow calculated from the plurality of frames.; 17. Apparatus according to claim 16, characterized in that the camera is a single monocular camera (404).
18. Method characterized in that it comprises: Petition 870260066391, dated 06 / 07 / 2026, page 162 / 315 10 / 11 allowing, by a processing system (2102) including a processor, wireless communication with a remote computer from which the processing system (2102) can receive instructions and send information;to obtain, by the processing system (2102) comprising the processor, video data from a single monocular camera (404), wherein the video data comprise a plurality of frames, wherein the single monocular camera (404) is attached to a mobile ground robot (402) that is traveling along a track defined by a line of crops, wherein the mobile ground robot (402) comprises a plurality of wheels and a plurality of encoders, wherein each of the plurality of wheels has associated with it a respective encoder of the plurality of encoders, wherein the line of crops comprises a first plant stem (1102), and wherein the plurality of frames includes a representation of the first plant stem (1102);obtain, by processing system (2102), robot speed data from the plurality of encoders, wherein the robot speed data are determined based on an average of encoder values from all encoders in the plurality of encoders; perform, by processing system (2102), foreground extraction in each of the plurality of frames of the video data, wherein the foreground extraction results in a plurality of foreground images;and determine by the processing system (2102), based on the plurality of foreground images and based on robot speed data, an estimated width of the first plant stem (1102), wherein the determination comprises determining a ratio R, wherein R = Vr / Vx, wherein Vr is an instantaneous robot speed obtained through robot speed data and Vx is an average horizontal foreground pixel speed obtained through camera motion estimation using optical flow calculated from the plurality of frames of the video data.; 19. Method according to claim 18, characterized in that the foreground extraction comprises processing, for each of the plurality of frames of the video data, only a fixed-size window (702) that is smaller than each of the plurality of frames wherein the fixed-size window (702) is a single window and is the only window on which the foreground extraction processing is applied; and / or wherein a window width of the window (702) is equal to a frame width of each of the plurality of frames divided by 4.
20. Method according to claim 19, characterized in that the fixed-size window (702) associated with each of the plurality of frames is located off-center in each of the plurality of frames.