Object Detection Methods
Patent Information
- Application Number
- JP2024542173
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-04
- Filing Date
- 2023-01-18
- Publication Date
- 2026-01-28
AI Technical Summary
【0021】 本発明の新規な特徴は、添付の特許請求の範囲に詳細に記載される。本発明の特徴及び利点のより深い理解は、本発明の原理を利用した例示的な実施形態を説明する以下の詳細な説明及び添付の図面を参照することによって得られる。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 306,904, filed February 4, 2022, which is incorporated herein by reference.
[0002] As technology advances, tasks previously performed by humans are becoming increasingly automated. Tasks performed in highly controlled environments, such as factory assembly lines, can be automated by instructing machines to perform the task the same way every time, but tasks performed in unpredictable environments, such as agricultural environments, rely on dynamic feedback and adaptation to perform the task. Autonomous systems often struggle to identify and locate objects in unpredictable environments. Improvements in object tracking methods advance automation technology and improve the ability of autonomous systems to react and adapt to unpredictable environments. Summary of the Invention [Means for solving the problem]
[0003] In various aspects, the present disclosure provides a computer-implemented method for detecting a target plant, the computer-implemented method including: receiving an image of an area of a surface, the area including a target plant located on the surface; identifying one or more parameters of the target plant, the one or more parameters of the target plant including a point location of the target plant; and identifying the target plant in the image based on the one or more parameters of the target plant.
[0004] In some embodiments, the area of the surface further comprises one or more additional plants. In some embodiments, the target plant is a weed or an infested crop. In some embodiments, the point location corresponds to a feature of the target plant. In some embodiments, the feature is a center of the target plant, a meristem of the target plant, or a leaf of the target plant. In some embodiments, the one or more parameters further comprise a plant location, a plant size, a plant category, a plant type, a leaf shape, a phyllotaxis, a plant attitude, a plant health condition, or a combination thereof. In some embodiments, the surface is an agricultural surface.
[0005] In some embodiments, the point location includes a meristem location, a center of gravity location, or a leaf location of a plant. In some embodiments, the computer-implemented method further includes targeting the target plant at the point location with an instrument. In some embodiments, the instrument is a laser, a spray, or a grabber. In some embodiments, the laser is an infrared laser. In some embodiments, the computer-implemented method further includes operating the instrument at the point location for a period of time. In some embodiments, the period of time is sufficient to kill the target plant. In some embodiments, the period of time is based on one or more characteristics of the target plant. In some embodiments, the one or more characteristics include a plant size, a plant type, or both. In some embodiments, the period of time corresponds nonlinearly to a plant size. In some embodiments, the computer-implemented method includes killing the target plant with the instrument. In some embodiments, the computer-implemented method includes incinerating the characteristic of the target plant with the instrument.
[0006] In some embodiments, the computer-implemented method further comprises determining a size of the target plant. In some embodiments, the size of the plant comprises a size of one or more structures of the target plant. In some embodiments, the one or more structures are selected from the group consisting of leaves, stems, leaf blades, flowers, fruits, seeds, shoots, buds, and combinations thereof. In some embodiments, the size of the plant comprises a length, a radius, a diameter, an area, or any combination thereof.
[0007] In some embodiments, the computer-implemented method further comprises classifying the type of the target plant. In some embodiments, the type of the plant is based on the leaf shape of the target plant. In some embodiments, the type of the plant is selected from the group consisting of crop, weed, grass, broadleaf, purslane, or combinations thereof. In some embodiments, the computer-implemented method further comprises evaluating the condition of the target plant. In some embodiments, the condition comprises health, maturity, nutritional state, disease state, ripeness, crop yield, or any combination thereof. In some embodiments, the computer-implemented method further comprises determining a confidence score for the one or more parameters. In some embodiments, the computer-implemented method further comprises scheduling the target plant to be targeted based on the confidence score.
[0008] In some embodiments, the computer-implemented method further comprises obtaining labeled image data including parameterized objects corresponding to similar plants. In some embodiments, the computer-implemented method further comprises training a machine learning model to identify parameters corresponding to a target plant, the machine learning model being trained using the labeled image data. In some embodiments, the computer-implemented method further comprises generating a prediction of a plant corresponding to one or more parameters of the target plant, the one or more parameters of the target plant including point locations of the target plant, the one or more parameters being identified by using the image as input data to the machine learning model. In some embodiments, the computer-implemented method further comprises updating the machine learning model using the image, the one or more parameters, and information corresponding to the identification of the target plant, and when the machine learning model is updated, new object parameters are identified from new images using the machine learning model.
[0009] In some aspects, updating the machine learning model includes receiving further labeled image data that includes the image, and fine-tuning the machine learning model based on the further labeled image data. In some aspects, fine-tuning the machine learning model is performed using a subset of the further labeled image data and an image batch that includes the subset of the labeled image data. In some aspects, fine-tuning the machine learning model is performed using fewer batches, fewer epochs, or fewer batches and fewer epochs than training the machine learning model.
[0010] In some embodiments, the labeled image data includes an image of a plant. In some embodiments, the plant image includes a weed image, a crop image, a weed and crop image, or a combination thereof. In some embodiments, the weed image includes a purslane weed image, a broadleaf weed image, a side branch image, a grass image, or a combination thereof. In some embodiments, the crop image includes an onion image, a strawberry image, a carrot image, a corn image, a soybean image, a barley image, an oat image, a wheat image, an alfalfa image, a cotton image, a grass image, a tobacco image, a rice image, a sorghum image, a tomato image, a potato image, a grape image, a rice image, a lettuce image, a bean image, a pea image, a sugar beet image, or a combination thereof. In some embodiments, the plant image is labeled with a plant centroid location, a meristem location, a plant size, a plant category, a plant type, a leaf shape, a leaf number, a phyllotaxis, a plant posture, a plant health condition, or a combination thereof. In some aspects, the parameterized objects include data corresponding to point positions, shapes, sizes, categories, types, or combinations thereof.
[0011] In some embodiments, the computer-implemented method further comprises using a trained classifier to identify the target plant. In some embodiments, the computer-implemented method further comprises using a trained classifier to localize features of the target plant. In some embodiments, the trained classifier is trained using a training dataset that includes labeled images. In some embodiments, the labeled images are labeled with plant category, meristem location, plant size, plant condition, plant type, or any combination thereof.
[0012] In some aspects, the computer-implemented method further includes pre-training the machine learning model. In some aspects, pre-training the machine learning model is performed on a pre-training dataset including the labeled image data and pre-training labeled image data that share a common feature of the labeled image data. In some aspects, the common feature is an image of a plant.
[0013] In various aspects, the disclosure provides a computer-implemented method for detecting a target object, the computer-implemented method including: receiving an image of a region of a surface, the region including a target object located on the surface; obtaining labeled image data including parameterized objects corresponding to similarly positioned objects; training a machine learning model to identify object parameters corresponding to the target object, the machine learning model being trained using the labeled image data; generating a prediction of an object corresponding to one or more parameters of the target object, the one or more object parameters of the target object including a point location of the target object, the one or more object parameters being identified by using the image as input data to the machine learning model; identifying the target object in the image based on the one or more parameters; and updating the machine learning model using information corresponding to the image, the one or more parameters, and the identification of the target object, where if the machine learning model is updated, new object parameters are identified from a new image using the machine learning model.
[0014] In some embodiments, the target object is a target plant, a pest, a surface irregularity, or a piece of equipment. In some embodiments, the target plant is a weed or an infested crop. In some embodiments, the surface irregularity is a rock, a clump of soil, or a soil additive. In some embodiments, the piece of equipment is a sprinkler, a hose, or a marker. In some embodiments, the pest is an insect, a worm, an arthropod, a spider, a fungus, or a nematode.
[0015] In some embodiments, the labeled image data includes a plant image. In some embodiments, the plant image includes a weed image, a crop image, or a combination thereof. In some embodiments, the weed image includes a purslane weed image, a broadleaf weed image, a side branch image, a grass image, or a combination thereof. In some embodiments, the crop image includes an onion image, a strawberry image, a carrot image, a corn image, a soybean image, a barley image, an oat image, a wheat image, an alfalfa image, a cotton image, a grass image, a tobacco image, a rice image, a sorghum image, a tomato image, a potato image, a grape image, a rice image, a lettuce image, a bean image, a pea image, a sugar beet image, or a combination thereof. In some embodiments, the plant image is labeled with a plant centroid location, a meristem location, a plant size, a plant category, a plant type, a leaf shape, a leaf number, a phyllotaxis, a plant posture, a plant health condition, or a combination thereof.
[0016] In some embodiments, the one or more object parameters further comprise object location, object size, object category, plant type, leaf shape, phyllotaxis, plant posture, plant health, or a combination thereof. In some embodiments, the surface is an agricultural surface.
[0017] In some embodiments, the parameterized object includes data corresponding to a point location, a shape, a size, a category, a type, or a combination thereof. In some embodiments, the point location includes a plant meristem location, a centroid location, or a leaf location. In some embodiments, the computer-implemented method further includes using a trained classifier to identify the target object. In some embodiments, the computer-implemented method further includes using a trained classifier to locate features of the target object. In some embodiments, the trained classifier is trained using a training dataset including labeled images. In some embodiments, the labeled images are labeled with object category, meristem location, plant size, plant condition, plant type, or any combination thereof.
[0018] In some aspects, updating the machine learning model includes receiving further labeled image data that includes the image, and fine-tuning the machine learning model based on the further labeled image data. In some aspects, fine-tuning the machine learning model is performed using a subset of the further labeled image data and an image batch that includes the subset of the labeled image data. In some aspects, fine-tuning the machine learning model is performed using fewer batches, fewer epochs, or fewer batches and fewer epochs than training the machine learning model.
[0019] In some aspects, the computer-implemented method further includes pre-training the machine learning model. In some aspects, pre-training the machine learning model is performed on a pre-training dataset including the labeled image data and pre-training labeled image data that share a common feature of the labeled image data. In some aspects, the common feature is an image of a plant.
[0020] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0021] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Brief description of the drawings]
[0022] [Figure 1] 1 illustrates an isometric view of an autonomous laser weed eradication vehicle according to one or more embodiments herein. [Diagram 2] FIG. 1 shows a top view of an autonomous weed eradication vehicle moving through a field while implementing various techniques described herein. [Diagram 3] FIG. 1 illustrates a side view of a detection system disposed on an autonomous laser weed eradication vehicle according to one or more embodiments herein. [Figure 4] 1 shows an image of a plant with meristems located and leaf radius measured, according to one or more embodiments herein. [Diagram 5] 1 shows an image of a weed with meristems located and leaf radius measured, according to one or more embodiments herein. [Figure 6] 1 illustrates an architecture of a point detection system according to one or more embodiments herein. [Figure 7A] 1 illustrates a boundary region based plant detection method. [Figure 7B] 1 shows a mask-based plant detection method. [Figure 8] FIG. 2 is a block diagram illustrating components of a detection terminal according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is an example block diagram of a computing device architecture for a computing device capable of implementing various techniques described herein. [Figure 10] 1 is a flow diagram illustrating a method for training and using a point detection module according to an embodiment of the present disclosure. [Figure 11] FIG. 1 is a block diagram illustrating components of a prediction and targeting system for identifying, locating, targeting, and manipulating objects in accordance with one or more embodiments herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0023] Various exemplary embodiments of the present disclosure are described in detail below. Although specific embodiments are discussed, it should be understood that the description is for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the spirit and scope of the present disclosure. Therefore, the following description and drawings are illustrative and should not be construed as limiting. Numerous specific details are described to provide a thorough understanding of the present disclosure. However, in certain instances, well-known or conventional details are not described in order to avoid obscuring the description. References to one embodiment or an embodiment in the present disclosure may refer to the same embodiment or to any embodiment, and such references mean at least one exemplary embodiment.
[0024] Reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the disclosure. The appearances of the phrase "in one embodiment" in various places in this specification do not necessarily all refer to the same embodiment, nor are other exemplary embodiments mutually exclusive with other or alternative exemplary embodiments. Furthermore, various features are described that may be exhibited by some exemplary embodiments and not by others. Any feature of one example may be combined with or used in conjunction with any other feature of any other example.
[0025] The terms used herein generally have their ordinary meaning in the art, within the context of this disclosure and within the specific context in which each term is used. Alternative language and synonyms may be used for any one or more terms discussed herein, and no particular emphasis is to be placed on whether a term is elaborated or discussed herein. In some cases, synonyms for a particular term are provided. The description of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any term discussed herein, is for illustrative purposes only and is not intended to further limit the scope and meaning of the disclosure or any exemplary term. Similarly, the disclosure is not limited to the various exemplary embodiments provided herein.
[0026] Although not intended to limit the scope of the present disclosure, examples of devices, apparatus, methods and their related results according to exemplary embodiments of the present disclosure are given below. Please note that titles or subtitles may be used in the examples for the convenience of the reader, but do not limit the scope of the present disclosure in any way. Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs. In case of conflict, the present specification, including definitions, will take precedence.
[0027] Additional features and advantages of the present disclosure will be set forth in and in part will be apparent from the description which follows, or may be learned by the practice of the principles disclosed herein. The features and advantages of the present disclosure may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. These and other features of the present disclosure will become more fully apparent from the following description and the appended claims, or may be learned by the practice of the principles as described herein.
[0028] For clarity of explanation, in some cases, the technology may be presented as including individual functional blocks that represent devices, device components, steps or routines in a method implemented in software, or a combination of hardware and software.
[0029] In the figures, certain structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, such features may be arranged in a different manner and / or order than that shown in the illustrative figures. Furthermore, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and that in some embodiments it may not be included or may be combined with other features.
[0030] While the concepts of the present disclosure are susceptible to various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and are herein described in detail. It should be understood, however, that there is no intention to limit the concepts of the present disclosure to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives consistent with the scope of the present disclosure and the appended claims.
[0031] Described herein are systems and methods for identifying and locating objects, e.g., plants or pests, on a surface, e.g., a field. The systems and methods of the present disclosure can be used to identify and precisely target plants, e.g., weeds, for various crop management methods. For example, an autonomous weed eradication system implementing the point detection method can be used to identify weeds, locate the weed meristems, and precisely target the meristems for weed eradication. The point detection method described herein can be used to locate objects (e.g., plants, pests, equipment, surface irregularities, etc.) in an image (e.g., an image of a field), differentiate the objects by object category (e.g., as a weed or crop), identify the type and size of each object, and precisely locate features of the objects (e.g., plant meristems, plant leaves, or plant cores). Furthermore, the point detection method can be used to count objects of a particular category, locate crop rows, or assess plant health, nutrition, or maturity. A targeting system, such as an autonomous weed eradication system described herein, may target a feature of the object (e.g., the meristem of the plant or the center of the plant) for a period of time based on one or more parameters of the object (e.g., size, type, maturity, health, or nutrition), for example using an infrared laser.
[0032] As used herein, "image" may refer to a representation of an area or object. For example, an image may be a visual representation of an area or object formed by electromagnetic radiation (e.g., light, x-ray, microwave, or radio waves) scattered from the area or object. In another example, an image may be a point cloud model formed by a Light Detection and Ranging (LIDAR) or Radio Detection and Ranging (RADAR) sensor. In another example, an image may be a sonogram generated by detecting sound, infrasonic, or ultrasonic waves reflected from an area or object. As used herein, "imaging" may be used to describe the process of collecting or generating a representation (e.g., an image) of an area or object.
[0033] As used herein, a position, e.g., an object position or a sensor position, may be expressed relative to a reference frame. Exemplary reference frames include a surface reference frame, a vehicle reference frame, a sensor reference frame, or an actuator reference frame. Positions may be easily converted between reference frames, e.g., using a conversion factor or a calibration model. It should be understood that while a position, a change in position, or an offset may be expressed in one reference frame, the position, a change in position, or an offset may be easily converted between reference frames.
[0034] As used herein, a "sensor" may refer to a device capable of detecting or measuring an event, a change in an environment, or a physical property. For example, a sensor may detect light, such as visible light, ultraviolet light, or infrared light, and generate an image. Examples of sensors include a camera (e.g., a charge-coupled device (CCD) camera or a complementary metal-oxide semiconductor (CMOS) camera), a LIDAR detector, an infrared sensor, an ultraviolet sensor, or an x-ray detector.
[0035] As used herein, an "object" may refer to an item or distinct area that can be observed, tracked, manipulated, or targeted. For example, an object may be a plant, such as a crop or a weed. In another example, an object may be a piece of debris. In another example, an object may be a distinct area or point on a surface, such as a marking or surface irregularity.
[0036] As used herein, "targeting" or "aiming" may refer to pointing or directing a device or action to a specific location or object. For example, targeting an object may include pointing a sensor (e.g., a camera) or an instrument (e.g., a laser) to the object. Targeting or aiming may be dynamic, such that the device or action tracks an object that moves relative to the targeting system. For example, a device placed on a moving vehicle may dynamically target or aim at an object located on the ground by tracking the object as the vehicle moves relative to the ground.
[0037] As used herein, a "weed" may refer to an unwanted plant, such as an unwanted type of plant or a plant that grows in an undesirable place or at an undesirable time. For example, a weed may be a wild plant or an invasive plant. In another example, a weed may be a plant in a field of a cultivated crop that is not a cultivated species. In another example, a weed may be a plant that grows outside or between the rows of a cultivated crop.
[0038] As used herein, "manipulating" an object may refer to affecting, interacting with, or changing the state of an object. For example, manipulating may include irradiating, illuminating, heating, incinerating, killing, moving, lifting, grasping, spraying, or otherwise modifying an object.
[0039] As used herein, "electromagnetic radiation" may refer to radiation from the entire electromagnetic spectrum, which may include, but is not limited to, visible light, infrared light, ultraviolet light, radio waves, gamma rays, or microwaves.
[0040] Autonomous weed eradication system The detection methods described herein can be implemented by an autonomous weed eradication system to target and remove weeds. Such detection methods can facilitate object identification and tracking. For example, an autonomous weed eradication system can be used to detect and locate a weed of interest identified in images or representations collected over time for the autonomous weed eradication system by a first sensor, e.g., a predictive sensor. The detection information can be used to determine a predicted location of the weed relative to the system. The autonomous weed eradication system then uses the predicted location to locate the same weed in images or representations collected by a second sensor, e.g., a targeting sensor. In some embodiments, the first sensor is a predictive camera and the second sensor is a targeting camera. One or both of the first sensor and the second sensor can move relative to the weed. For example, the predictive camera can be coupled to the autonomous weed eradication system and move therewith.
[0041] The targeting of the weeds may include using the targeting sensor to precisely locate the weeds, targeting the weeds with a laser, and eradicating the weeds by burning them with a laser light, e.g., infrared light. The predictive sensor may be part of a predictive module configured to identify a predicted location of an object of interest, and the targeting sensor may be part of a targeting module configured to fine-tune the predicted location of the object of interest to identify a target location and target the object of interest at the target location with the laser. The predictive module may be configured to communicate with the targeting module and coordinate camera handover using point-to-point targeting as described herein. The targeting module may target the object at the predicted location. In some embodiments, the targeting module may dynamically target the object using the trajectory of the object during operation of the system, such that the position of the targeting sensor, the laser, or both is adjusted to keep the target.
[0042] The autonomous weed eradication system may identify, target, and remove weeds without human intervention. Optionally, the autonomous weed eradication system may be located in an autonomous vehicle or a manned vehicle, and may be towed by a vehicle such as a tractor. As shown in FIG. 1, the autonomous weed eradication system may be part of or coupled to a vehicle 100, such as a tractor or an autonomous vehicle. As shown in FIG. 2, the vehicle 100 may move through a field 200. As the vehicle 100 moves through the field 200, it may identify, target, and eradicate weeds in an unweeded portion 210 of the field, leaving a weeded field 220 behind. The detection methods described herein may be implemented by the autonomous weed eradication system to identify, target, and eradicate weeds while the vehicle 100 is operating. Such tracking methods with high accuracy allow for precise targeting of weeds, such as with a laser, to eradicate them without damaging adjacent crops.
[0043] Detection System In some embodiments, the detection method described herein may be performed by a detection system. The detection system may include a prediction system and, optionally, a targeting system. In some embodiments, the detection system may be disposed or coupled to a vehicle, such as an autonomous weeding vehicle or a laser weeding system towed by a tractor. The prediction system may include a prediction sensor configured to image an area of interest, and the targeting system may include a targeting sensor configured to image a portion of the area of interest. Imaging may include collecting a representation (e.g., an image) of the area of interest or a portion of the area of interest. In some embodiments, the prediction system may include multiple prediction sensors, allowing for a larger range of the area of interest. In some embodiments, the targeting system may include multiple targeting sensors.
[0044] The region of interest may correspond to an overlap region between the field of view of the targeting sensor and the field of view of the predictive sensor. Such overlap may be simultaneous or separated in time. For example, the field of view of the predictive sensor encompasses the region of interest at a first time point, and the field of view of the targeting sensor encompasses the region of interest at a second time point but not at the first time point. Optionally, the detection system may move relative to the region of interest between the first and second time points to facilitate a temporal separation of the overlap between the field of view of the predictive sensor and the field of view of the targeting sensor.
[0045] In some embodiments, the predictive sensor may have a wider field of view than the targeting sensor. The predictive system may further include an object identification module for identifying an object of interest in a predicted image or representation collected by the predictive sensor. The object identification module may distinguish the object of interest from other objects in the predicted image.
[0046] The prediction module may determine a predicted location of the object of interest and may transmit the predicted location to the targeting system. The predicted location of the object may be determined using the object tracking methods described herein.
[0047] The targeting system may point the targeting sensor to a desired portion of the region of interest that is predicted to contain the object based on the predicted position received from the prediction system. In some embodiments, the targeting module may direct an instrument to the object. In some embodiments, the instrument may act on or manipulate the object. In some embodiments, the targeting module may dynamically target the object using the trajectory of the object during operation of the system such that the position of the targeting sensor, the instrument, or both, is adjusted to hold the target.
[0048] An example of a detection system 300 is shown in FIG. 3. The detection system may be part of or coupled to a vehicle 100, e.g., an autonomous weeding vehicle or a laser weeding system towed by a tractor, moving along a surface, e.g., a field 200. The detection system 300 includes a prediction module 310 including a prediction sensor having a predicted field of view 315, and a targeting module 320 including a targeting sensor having a targeting field of view 325. The targeting module may further include an instrument such as a laser whose target area overlaps with the targeting field of view 325. In some embodiments, the prediction module 310 is positioned ahead of the targeting module 320 along the direction of travel of the vehicle 100 such that the targeting field of view 325 overlaps with the predicted field of view 315 with a time delay. For example, the predicted field of view 315 at a first time point may overlap with the targeting field of view 325 at a second time point. In some embodiments, the predicted field of view 315 at the first time point may not overlap with the targeting field of view 325 at the first time point.
[0049] The detection system of the present disclosure can be used to target objects on a surface, such as the ground, a soil surface, a floor, a wall, an agricultural surface (e.g., a field), a lawn, a road, a mound, a mountain, or a depression. In some embodiments, the surface can be non-flat, such as an uneven ground, an uneven terrain, or a rough floor. For example, the surface can be an uneven ground in a construction site, a field, or a mining tunnel, or the surface can be an uneven terrain including a field, a road, a forest, a hill, a mountain, a house, or a building. The detection system described herein can locate objects on a non-flat surface more accurately, faster, or in a large area than a single sensor system or a system lacking an object matching module.
[0050] Alternatively or additionally, the detection system may be used to target an object that may be floating above the surface on which it rests, such as a treetop away from its ground point, and / or to target an object that may be located relative to a surface, such as in the air or against the ground in the atmosphere. Further, the detection system may be used to target an object that may move relative to a surface, such as a vehicle, animal, human, or projectile.
[0051] FIG. 11 illustrates a detection system including a prediction system 400 and a targeting system 450 for tracking an object O during targeting relative to a moving body, e.g., the vehicle 100 shown in FIGS. 1-3. The prediction system 400, the targeting system 450, or both, may be located on or coupled to the moving body (e.g., a moving vehicle). The prediction system 400 may include a prediction sensor 410 configured to image an area, e.g., an area of a surface, including one or more objects, including the object O. Optionally, the prediction system 400 may include a speed tracking module 415. The speed tracking module may estimate the speed of the moving body relative to the area (e.g., the surface). In some embodiments, the speed tracking module 415 may include a device, e.g., a rotational encoder, for measuring the displacement of the moving body over time. Alternatively or in addition, the speed tracking module may use images collected by the prediction sensor 400 to estimate speed using optical flow.
[0052] The object identification module 420 may identify objects in images collected by the predictive sensor. For example, the object identification module 420 may identify weeds in an image and distinguish the weeds from other plants, e.g., crops, in the image. The object localization module 425 may locate objects identified by the object identification module 420 and compile a set of identified objects and their corresponding locations. Object identification and object localization may be performed on a series of images collected over time by the predictive sensor 410. The sets of identified objects and corresponding locations from two or more images from the object localization module 425 may be sent to the de-duplication module 430.
[0053] The de-duplication module 430 may use object locations in a first image collected at a first time and object locations in a second image collected at a second time to identify an object, e.g., object O, that appears in both the first image and the second image. The set of identified objects and corresponding locations may be de-duplicated by the de-duplication module 430 by assigning object locations that appear in both the first image and the second image to the same object O. In some embodiments, the de-duplication module 430 may use speed estimates from the velocity tracking module 415 to identify corresponding objects that appear in both images. The resulting de-duplicated set of identified objects may include unique objects, each having one or more corresponding locations identified at one or more time points. The adjustment module 435 may receive the de-duplicated set of objects from the de-duplication module 430 and may adjust the de-duplicated set by removing objects. In some embodiments, objects may be removed if they are no longer tracked. For example, an object may be removed if it has not been identified in a predetermined number of images in a sequence of images. In another example, an object may be removed if it has not been identified for a predetermined period of time. In some embodiments, an object may continue to be tracked that no longer appears in the images collected by the predictive sensor 410. For example, an object may continue to be tracked if, based on the predicted position of the object, the object is predicted to be within the predicted field of view. In another example, an object may continue to be tracked if, based on the predicted position of the object, the object is predicted to be within range of the targeting system. The adjustment module 435 may provide the adjusted set of objects to the position prediction module 440.
[0054] The position prediction module 440 may determine a predicted position of the object O at a future time from the adjusted set of objects. In some embodiments, the predicted position may be determined from two or more corresponding positions determined from images collected at two or more times, or from a single position in combination with speed information from the speed tracking module 415. The predicted position of the object O may be based on a vector velocity including the speed and direction of the object O relative to the moving body between the position of the object O in a first image collected at a first time and the position of the object O in a second image collected at a second time. Optionally, the vector velocity may take into account the distance of the object O from the moving body along the imaging axis (e.g., the height or elevation of the object relative to the surface). Alternatively or additionally, the predicted position of the object may be based on the position of the object O in the first image or the second image and the vector velocity of the vehicle determined from the speed tracking module 415.
[0055] The targeting system 450 may receive a predicted position of the object O at a future time from the prediction system 400 and may use the predicted position to precisely target the object with the instrument 475 at the future time. The targeting control module 460 of the targeting system 450 may receive the predicted position of the object O from the position prediction module 440 of the prediction system 435 and may direct the targeting sensor 465, the instrument 475, or both to point to the predicted position of the object. Optionally, the targeting sensor 465 may collect an image of the object O, and the position refinement module 470 may refine the predicted position of the object O based on the position of the object O determined from the image. In some embodiments, the position refinement module 470 may account for optical distortions in the images collected by the prediction sensor 410 or the targeting sensor 465, or distortions of the angular motion of the instrument 475 or the targeting sensor 465 due to nonlinearities in the angular motion relative to the object O. The targeting control module 460 may direct the tool 475, and optionally the targeting sensor 465, to point to the fine-tuned location of the object O. In some embodiments, the targeting control module 460 may adjust the position of the targeting sensor 465 or the tool 475 to follow the object to account for the movement of the vehicle during targeting. The tool 475, e.g., a laser, may then manipulate the object O. For example, the laser may direct infrared light to the predicted or fine-tuned location of the object O. The object O may be a weed, and directing infrared light to the location of the weed may eradicate the weed.
[0056] In some embodiments, the prediction system 400 may further include a scheduling module 445. The scheduling module 445 may select objects identified by the prediction module and schedule which ones to target with the targeting system. The scheduling module 445 may schedule objects for targeting based on parameters, such as object location, relative speed, tool activation time, confidence score, weed type, or combinations thereof. For example, the scheduling module 445 may prioritize targeting objects that are predicted to move out of the field of view of a prediction or targeting sensor or out of the range of a tool. Alternatively or in addition, the scheduling module 445 may prioritize targeting objects identified or located with high confidence. Alternatively or in addition, the scheduling module 445 may prioritize targeting objects with short activation times. In some embodiments, the scheduling module 445 may prioritize targeting objects based on a user's preferred parameters.
[0057] Prediction Module A prediction module of the present disclosure may be configured to detect objects using the detection methods described herein. In some embodiments, the prediction module is configured to capture an image or representation of an area of a surface using a predictive camera or sensor, identify an object of interest in the image, and determine a predicted location of the object.
[0058] The prediction module may include an object identification module configured to identify an object of interest and distinguish the object of interest from other objects in the predicted image. In some embodiments, the prediction module uses a machine learning model to identify and distinguish objects based on features extracted from a training dataset including images of labeled objects. For example, a machine learning model of or associated with the object identification module may be trained to identify weeds and distinguish weeds from other plants, e.g., crops. In another example, a machine learning model of or associated with the object identification module may be trained to identify debris and distinguish debris from other objects. The object identification module may be configured to identify plants and distinguish different plants, e.g., crops and weeds. In some embodiments, the machine learning model may be a deep learning model, e.g., a deep learning neural network.
[0059] The machine learning model may be trained using supervised, unsupervised, reinforcement, or other such training techniques. For example, a set of images, which may or may not contain various objects, may be analyzed using one of a variety of machine learning models to identify correlations between different elements of the images and specific objects without supervision and feedback (e.g., unsupervised training techniques). The machine learning model may also be trained to identify objects in predicted or target images using sample images, raw images, or labeled images. As an example of a supervised training technique, a set of labeled images may be selected for training the machine learning model to facilitate identification and classification of these objects. The machine learning model may be evaluated based on the sample images provided to the machine learning model to determine whether the machine learning model correctly identifies and classifies objects in these images. Based on this evaluation, the machine learning model may be modified to increase the likelihood that the machine learning model correctly identifies and classifies objects in predicted and / or target images. The machine learning model may further be dynamically trained by soliciting feedback (i.e., supervision) from a user regarding the accuracy of the machine learning model in identifying and classifying objects. The feedback may be used to further train the machine learning model to provide more accurate results over time.
[0060] In some embodiments, the object identification module includes using a discriminative machine learning model, e.g., a convolutional neural network. For example, the object identification module may include a point detection model (e.g., the point detection model shown in FIG. 6). The discriminative machine learning model may be trained using many images, e.g., high-resolution images of surfaces with or without objects of interest. For example, the machine learning model may be trained using images of fields with or without weeds. After training, the machine learning model may be configured to identify a region in the image that contains the object of interest. The region may be defined by a polygon, e.g., a rectangle. In some embodiments, the region is a bounding box. In some embodiments, the region is a polygonal mask that covers the identified region. In some embodiments, the discriminative machine learning model may be trained to identify the location of the object of interest, e.g., a pixel location in a predicted image.
[0061] The prediction module may further include a speed tracking module for determining the speed of the vehicle to which the prediction module is coupled. In some embodiments, a positioning system and a detection system may be located on the vehicle. Alternatively or in addition, the positioning system may be located on a vehicle spatially coupled to the detection system. For example, the positioning system may be located on a vehicle towing the detection system. The speed tracking module may include a positioning system, such as a wheel or rotational encoder, an inertial measurement unit (IMU), a global positioning system (GPS), a distance sensor (e.g., laser, SONAR, or RADAR), or an internal navigation system (INS). For example, a wheel encoder in communication with a wheel of the vehicle may estimate speed or distance traveled based on angular frequency, rotational frequency, rotation angle, or wheel rotation rate. In some embodiments, the speed tracking module may use images from the prediction sensor to determine the speed of the vehicle using optical flow.
[0062] The prediction module may include a system controller, such as a system computer having storage, random access memory (RAM), a central processing unit (CPU), and a graphics processing unit (GPU). The system computer may include a tensor processing unit (TPU). The system computer must have sufficient RAM, storage space, CPU power, and GPU power to perform operations to detect and identify targets. The prediction sensor must provide images of sufficient resolution to perform operations to detect and identify objects. In some embodiments, the prediction sensor may be a camera, such as a charge-coupled device (CCD) camera or a complementary metal oxide semiconductor (CMOS) camera, a LIDAR detector, an infrared sensor, an ultraviolet sensor, an x-ray detector, or any other sensor capable of generating an image.
[0063] Targeting Module The targeting module of the present disclosure may be configured to target an object detected by the prediction module. In some embodiments, the targeting module may direct an instrument to the object to manipulate the object. For example, the targeting module may be configured to direct a laser beam to a weed to incinerate the weed. In another example, the targeting module may direct a gripping tool to grab the object. In another example, the targeting module may direct a spraying tool to spray a fluid on the object. In some embodiments, the object may be a weed, a plant, an insect, a pest, a field, a piece of debris, an obstacle, an area of a surface, or any other object that can be manipulated. The targeting module may be configured to receive a predicted location of the object of interest from the prediction module and point a targeting camera or targeting sensor to the predicted location. In some embodiments, the targeting module may direct an instrument, such as a laser, to the predicted location. The location of the targeting sensor and the location of the instrument may be coupled. In some embodiments, multiple targeting modules are in communication with the prediction module.
[0064] The targeting module may include a targeting control module. In some embodiments, the targeting control module may control the targeting sensor, the instrument, or both. In some embodiments, the targeting control module may include an optical control system including optical components configured to control an optical path (e.g., a laser beam path or a camera imaging path). The targeting control module may include software-driven electrical components capable of controlling the activation and deactivation of the instrument. Activation or deactivation may depend on the presence or absence of an object detected by the targeting camera. Activation or deactivation may depend on the position of the instrument relative to the position of the target object. In some embodiments, the targeting control module may activate the instrument, e.g., a laser emitter, when an object is identified and located by the prediction system. In some embodiments, the targeting control module may activate the instrument when the instrument's range or targeting area is positioned to overlap with the target object position.
[0065] The targeting control module may shut off the tool after the object has been manipulated, e.g., grabbed, sprayed, burned, or irradiated, after an area containing the object has been targeted by the tool, at a point when the object is no longer identified by the target prediction module, after a designated period of time has elapsed, or any combination thereof. For example, the targeting control module may shut off the emitter after an area of a surface containing weeds has been scanned by the beam, after the weeds have been irradiated or burned, or after the beam has been activated for a predetermined period of time.
[0066] The prediction module and the targeting module described herein can be used in combination to locate, identify, and target an object with an instrument. The targeting control module can include an optical control system described herein. The prediction module and the targeting module can be in communication, for example, in electrical or digital communication. In some embodiments, the prediction module and the targeting module are directly or indirectly coupled. For example, the prediction module and the targeting module can be coupled to a support structure. In some embodiments, the prediction module and the targeting module are configured on or coupled to a vehicle, for example, the vehicle shown in FIG. 1 and FIG. 2. For example, the prediction module and the targeting module can be disposed in an autonomous vehicle. In another example, the prediction module and the targeting module can be towed by a vehicle, for example, a tractor.
[0067] The targeting module may include a system controller, such as a system computer having storage, random access memory (RAM), a central processing unit (CPU), and a graphics processing unit (GPU). The system computer may include a tensor processing unit (TPU). The system computer must have sufficient RAM, storage space, CPU power, and GPU power to perform the target detection and identification operations. The targeting sensor must provide an image with sufficient resolution to perform the object matching operations to the object identified in the predicted image.
[0068] Optical Control System The methods described herein can be implemented by an optical control system, such as a laser optical system, to target an object of interest.For example, an optical system can be used to target an object of interest identified in an image or display collected by a first sensor, such as a predictive sensor, and locate the same object in an image or display collected by a second sensor, such as a targeting sensor.In some embodiments, the first sensor is a predictive camera and the second sensor is a targeting camera.Targeting the object can include precisely locating the object using the targeting sensor and targeting the object with an instrument.
[0069] Described herein is an optical control system for directing a beam, e.g., a light beam, to a target location on a surface, e.g., the location of an object of interest. In some embodiments, the instrument is a laser. However, other instruments are also within the scope of the present disclosure, including, but not limited to, a gripping instrument, a spraying instrument, a planting instrument, a harvesting instrument, a pollination instrument, a marking instrument, a spraying instrument, or a deposition instrument.
[0070] In some embodiments, the emitter is configured to direct a beam along an optical path, e.g., a laser path. In some embodiments, the beam comprises electromagnetic radiation, e.g., light, radio waves, microwaves, or x-rays. In some embodiments, the light is visible light, infrared light, or ultraviolet light. The beam can be coherent. In one embodiment, the emitter is a laser, e.g., an infrared laser.
[0071] One or more optical elements may be arranged in the path of the beam. The optical elements may include a beam combiner, a lens, a reflective element, or any other optical element that may be configured to direct, focus, filter, or otherwise control light. These elements may be arranged in the following order in the direction of the beam path: the beam combiner, followed by a first reflective element, followed by a second reflective element. In another example, one or both of the first reflective element or the second reflective element may be arranged in front of the beam combiner in the direction of the beam path. In another example, the optical elements may be arranged in the following order in the direction of the beam path: the beam combiner, followed by the first reflective element. In another example, one or both of the first reflective element or the second reflective element may be arranged in front of the beam combiner in the direction of the beam path. Any number of additional reflective elements may be arranged in the beam path.
[0072] The beam combiner may also be referred to as a beam combining element. In some embodiments, the beam combiner may be a zinc selenide (ZnSe), zinc sulfide (ZnS), or germanium (Ge) beam combiner. For example, the beam combiner may be configured to transmit infrared light and reflect visible light. In some embodiments, the beam combiner may be a dichroic beam combiner. In some embodiments, the beam combiner may be configured to pass electromagnetic radiation having a wavelength longer than a cutoff wavelength and reflect electromagnetic radiation having a wavelength shorter than a cutoff wavelength. In some embodiments, the beam combiner may be configured to pass electromagnetic radiation having a wavelength shorter than a cutoff wavelength and reflect electromagnetic radiation having a wavelength longer than a cutoff wavelength. In other embodiments, the beam combiner may be a polarizing beam splitter, a long pass filter, a short pass filter, or a band pass filter.
[0073] The optical control system of the present disclosure may further include a lens disposed in the optical path. In some embodiments, the lens may be a focusing lens disposed such that the focusing lens focuses the beam, the scattered light, or both. For example, the focusing lens may be disposed in the visible optical path to focus the scattered light to the targeting camera. In some embodiments, the lens may be a defocusing lens disposed such that the defocusing lens defocuses the beam, the scattered light, or both. In some embodiments, the lens may be a collimating lens disposed such that the collimating lens collimates the beam, the scattered light, or both. In some embodiments, more than one lens may be disposed in the optical path. For example, two lenses may be disposed in series in the optical path to expand or contract the beam.
[0074] The position and orientation of one or both of the first and second reflective elements may be controlled by one or more actuators. In some embodiments, the actuators may be motors, solenoids, galvanometers, or servos. For example, the position of the first reflective element may be controlled by a first actuator, and the position and orientation of the second reflective element may be controlled by a second actuator. In some embodiments, a single reflective element may be controlled by multiple actuators. For example, the first reflective element may be controlled by a first actuator along a first axis and a second actuator along a second axis. Optionally, the mirror may be controlled by a first actuator, a second actuator, and a third actuator providing multi-axis control of the mirror. In some embodiments, a single actuator may control the reflective element along one or more axes. In some embodiments, a single reflective element may be controlled by a single actuator.
[0075] The actuator may rotate the reflective element to change the position of the reflective element, thus changing the angle of incidence of the beam that strikes the reflective element. Changing the angle of incidence may result in a translation of the location where the beam strikes the surface. In some embodiments, the angle of incidence may be adjusted to maintain the location where the beam strikes the surface while the optical system moves relative to the surface. In some embodiments, the first actuator rotates the first reflective element about a first rotation axis to translate the location where the beam strikes the surface along a first translation axis, and the second actuator rotates the second reflective element about a second rotation axis to translate the location where the beam strikes the surface along a second translation axis. In some embodiments, the first actuator and the second actuator rotate the first reflective element about the first rotation axis and the second rotation axis to translate the location where the beam strikes the surface of the first reflective element along the first translation axis and the second translation axis. For example, a single reflective element may be controlled by a first actuator and a second actuator to effect translation of the location where the beam strikes the surface along a first translational axis and a second translational axis with the single reflective element controlled by two actuators, hi another example, a single reflective element may be controlled by one, two, or three actuators.
[0076] The first translation axis and the second translation axis may be orthogonal. The coverage area of the surface may be defined by the maximum translational motion along the first translation axis and the maximum translational motion along the second translation axis. One or both of the first actuator and the second actuator may be servo-controlled, piezoelectrically driven, piezo-inertial driven, stepper motor controlled, galvanometer driven, linear actuator controlled, or any combination thereof. One or both of the first reflecting element and the second reflecting element may be a mirror, e.g., a dichroic or dielectric mirror, a prism, a beam splitter, or any combination thereof. In some embodiments, one or both of the first reflecting element and the second reflecting element may be any element capable of deflecting the beam.
[0077] The targeting camera can be arranged to capture light, e.g., visible light traveling along a visible light path in the opposite direction to the beam path, e.g., the laser path. The light can be scattered by a surface, e.g., a surface having an object of interest, or an object, e.g., the object of interest, and travel along a visible light path toward the targeting camera. In some embodiments, the targeting camera is arranged to capture light reflected from the beam combiner. In other embodiments, the targeting camera is arranged to capture light transmitted through the beam combiner. By capturing such light, the targeting camera can be configured to image a target field of view of a surface. The targeting camera can be coupled to the beam combiner, or the targeting camera can be coupled to a support structure that supports the beam combiner. In one embodiment, the targeting camera does not move relative to the beam combiner, so that the targeting camera maintains a fixed position relative to the beam combiner.
[0078] The optical control system of the present disclosure may further include an exit window disposed in the beam path. In some embodiments, the exit window may be the last optical element the beam hits before exiting the optical control system. The exit window may include a material that is substantially transparent to visible light, infrared light, ultraviolet light, or any combination thereof. For example, the exit window may include glass, quartz, fused silica, zinc selenide, zinc sulfide, transparent polymer, or any combination thereof. In some embodiments, the exit window may include a scratch-resistant coating, such as a diamond coating. The exit window may prevent dust, debris, water, or any combination thereof from reaching other optical elements of the optical control system. In some embodiments, the exit window may be part of a protective casing that surrounds the optical control system.
[0079] After exiting the optical control system, the beam can be directed along a beam path to a surface. In some embodiments, the surface includes an object of interest, such as a weed. Rotational movement of the reflective element can result in a laser sweep along a first translational axis and a laser sweep along a second translational axis. Rotational movement of the reflective element can control the location where the beam hits the surface. For example, rotational movement of the reflective element can move the location where the beam hits the surface to the location of the object of interest on the surface. In some embodiments, the beam is configured to damage the object of interest. For example, the beam can include electromagnetic radiation and the beam can irradiate the object. In another example, the beam can include infrared light and the beam can incinerate the object. In some embodiments, one or both of the reflective elements can rotate such that the beam scans an area surrounding and including the object.
[0080] The predictive camera or predictive sensor may be associated with an optical control system, such as an optical control system, for identifying and locating a target object. The predictive camera may have a field of view that encompasses the coverage area of the optical control system covered by the amiable laser sweep. The predictive camera may be configured to capture an image or representation of an area that includes the coverage area to identify and select a target object. The selected object may be assigned to the optical control system. In some embodiments, the field of view of the predictive camera and the coverage area of the optical control system may be separated in time such that the field of view of the predictive camera encompasses the target at a first time point and the coverage area of the optical control system encompasses the target at a second time point. Optionally, the predictive camera, the optical control system, or both may move relative to the target between the first time point and the second time point.
[0081] In some embodiments, multiple optical control systems may be combined to expand the coverage area of a surface. The multiple optical control systems may be configured such that a laser sweep along a translational axis of each optical control system overlaps with a laser sweep along a translational axis of an adjacent optical control system. The combined laser sweep defines a coverage area that at least one of the multiple beams from the multiple optical control systems may reach. One or more prediction cameras may be positioned such that a field of view of the prediction camera covered by the one or more prediction cameras completely encompasses the coverage area. In some embodiments, the detection system may include two or more prediction cameras, each having a field of view. The fields of view of the prediction cameras may be combined to form a predicted field of view that completely encompasses the coverage area. In some embodiments, the predicted field of view may not completely encompass the coverage area at a single time point, but may encompass the coverage area over two or more time points (e.g., image frames). Optionally, the prediction camera(s) may move relative to the coverage area during two or more time points to enable temporal coverage of the coverage area. The predictive camera or sensor may be configured to capture an image or representation of an area including a coverage area to identify and select an object for targeting. A selected object may be assigned to one of the optical control systems based on the object's location and the area covered by the laser sweep of an individual optical control system.
[0082] The optical control systems may be configured on a vehicle, such as the vehicle 100 shown in FIGS. 1-3. For example, the vehicle may be an unmanned vehicle. The unmanned vehicle may be a robot. In some embodiments, the vehicle may be controlled by a human. For example, the vehicle may be driven by a human driver. In some embodiments, the vehicle may be coupled to a second vehicle driven by a human driver, such as towed behind or pushed by the second vehicle. The vehicle may be remotely controlled by a human, such as by a remote control. In some embodiments, the vehicle may be remotely controlled via long wave signals, optical signals, satellite, or any other remote communication method. The optical control systems may be configured on the vehicle such that the coverage area overlaps with surfaces under, behind, in front of, or around the vehicle.
[0083] The vehicle may be configured to travel over a surface including a plurality of objects, including one or more objects of interest, such as a field including a plurality of plants and one or more weeds. The vehicle may include one or more of a plurality of wheels, a power source, a motor, a predictive camera, or any combination thereof. In some embodiments, the vehicle has sufficient clearance above the surface to travel over the plants, such as crops, without damaging the plants. In some embodiments, the space between the inner edge of the left wheel and the inner edge of the right wheel is wide enough to pass around a row of plants without damaging the plants. In some embodiments, the distance between the outer edge of the left wheel and the outer edge of the right wheel is narrow enough to allow the vehicle to pass between two rows of the plants, such as two rows of crops, without damaging the plants. In one embodiment, the vehicle including the plurality of wheels, the plurality of optical control systems, and the predictive camera may travel over a row of crops and emit beams of the plurality of beams at targets, such as weeds, to incinerate or irradiate the weeds.
[0084] Point Detection Described herein are point detection systems and methods for identifying and locating objects (e.g., plants, pests, pieces of equipment, surface irregularities, etc.) on a surface. These systems and methods may facilitate accurate localization of object features (e.g., object centers, object centroids, plant meristems, plant leaves, pest thoraxes, etc.) that may be targeted for autonomous surface maintenance, e.g., weed eradication, pest management, crop management, or soil management. Point detection may include using point-based localization to identify and locate objects (e.g., plants, pests, pieces of equipment, surface irregularities, etc.) in an image. In some embodiments, the points correspond to the plant meristems. Point-based localization may provide advantages over boundary region (FIG. 7A) or masking (FIG. 7B) based approaches by improving the ease of labeling objects and increasing the accuracy of localization and targeting.
[0085] Additionally, the point detection methods described herein may be used to assess various object parameters in addition to object location, including object size (e.g., radius, diameter, surface area, or a combination thereof), plant maturity (e.g., age, growth stage, ripeness, crop yield, or a combination thereof), object category (e.g., weed, crop, equipment, pest, or surface irregularity), weed type (e.g., grass, broadleaf, purslane, or side shoot), crop type (e.g., onion, strawberry, carrot, corn, soybean, These include, but are not limited to, the type of pest (e.g., barley, oats, wheat, alfalfa, cotton, grass, tobacco, rice, sorghum, tomato, potato, grapes, rice, lettuce, beans, peas, sugar beet, etc.), type of pest (e.g., spiders, insects, fungi, ants, grasshoppers, worms, beetles, caterpillars, etc.), plant health (e.g., nutrition, disease, or moisture), leaf shape, phyllotaxy (e.g., leaf number or leaf position), plant posture (e.g., upright, bent, or lying down), or a combination thereof.
[0086] The point detection model may be trained using training data that includes images of plants (e.g., images of weeds or images of crops) with labeled features. In some embodiments, images of plants may be labeled to indicate the center of the plant, the meristem of the plant, the leaves of the plant, the leaf outline, the radius of the plant, or a combination thereof.
[0087] Machine learning models The point detection method may be implemented by a point detection module configured to identify and locate objects in an image, e.g., a predicted image collected by a predictive sensor or a target image collected by a targeting sensor. In some embodiments, the point detection module may be part of or in communication with a predictive module. In some embodiments, the point detection module may be part of or in communication with a targeting module. The point detection module may implement one or more machine learning algorithms or networks that are implemented and dynamically trained to identify and locate objects in one or more images (e.g., predicted images, target images, etc.). The one or more machine learning algorithms or networks may include neural networks (e.g., convolutional neural networks (CNNs), deep neural networks (DNNs), etc.), geometric recognition algorithms, photometric recognition algorithms, principal component analysis using eigenvectors, linear discriminant analysis, You Only Look Once (YOLO) algorithm, hidden Markov modeling, multiple linear subspace learning using tensor representations, neuron-motivated dynamic link matching, support vector machines (SVMs), or any other suitable machine learning techniques. When the point detection module implements one or more neural networks for point detection, the one or more neural networks may include one or more convolutional layers, vision transformer layers, visual transformer layers, activation functions, pooling, batch normalization, other deep learning mechanisms, or combinations thereof. The point detection module may include a system controller, such as a system computer having storage, random access memory (RAM), a central processing unit (CPU), and a graphics processing unit (GPU). The system computer may include a tensor processing unit (TPU). The system computer should have sufficient RAM, storage space, CPU power, and GPU power to perform plant identification and location operations.
[0088] The point detection machine learning model may be trained using a sample training dataset of images, e.g., high-resolution images of surfaces with or without plants, pests, or other objects. The training images may include one or more object parameters, e.g., location (e.g., meristem location, thorax location, object center location, leaf outline, etc.), object size (e.g., radius, diameter, surface area, or combinations thereof), plant maturity (e.g., age, growth stage, ripeness, crop yield, or combinations thereof), object category (e.g., weed, crop, equipment, pest, or surface irregularity), weed type (e.g., grass, broadleaf, or purslane), crop type (e.g., onion, strawberry, corn, etc.), and / or other parameters. The points may be labeled with the following characteristics: type of pest (e.g., wheat, soybean, barley, oats, wheat, alfalfa, cotton, grass, tobacco, rice, sorghum, tomato, potato, grape, rice, lettuce, bean, pea, sugar beet, etc.), type of pest (e.g., spider, insect, fungus, ant, grasshopper, worm, beetle, caterpillar, etc.), plant health (e.g., nutrition, disease, moisture, or combinations thereof), leaf shape, phyllotaxy (e.g., leaf number or leaf position), plant posture (e.g., upright, bent, or lying down), or combinations thereof. In some embodiments, one or more machine learning algorithms implemented by the point detection module may be trained end-to-end (e.g., by training multiple parameters in combination). Alternatively, or in addition, sub-networks (e.g., point networks) may be trained independently. For example, these sub-networks may be trained using supervised, unsupervised, reinforcement, or other such training techniques as described above.
[0089] An example of a model architecture of the point detection module is provided in FIG. 6. An image (e.g., an image collected by a predictive or targeting sensor) may be received by a backbone network. The network may be a CNN including any number of nodes (e.g., neurons) organized into any number of layers, and the network may be constructed using a vision transformer or a visual transformer. In some embodiments, the convolutional neural network may include an input layer configured to receive an image, an identification layer configured to identify plants in the image, and an output layer configured to output data (e.g., a feature map, object count, object location, or other parameters). Each layer of the convolutional neural network may be connected by any number of additional hidden layers. For example, a convolutional network may include an input layer that receives an image, a series of hidden layers, and one or more output layers. Each of the hidden layers may perform a convolution on the image and output a feature map. The feature map may be passed from the hidden layer to the next convolutional layer. The output layer may output the results of the network, such as the size of the object, its location in the image, the category of the object (e.g., weed, crop, pest, equipment, or surface irregularity), and the type of object. In some embodiments, the output may be a multi-resolution output.
[0090] The backbone network may receive images via an input layer. In some embodiments, the backbone may include a pre-trained network (e.g., ResNet50, MobileNet, CBNetV2, etc.) or a custom-trained network. The backbone may include a series of convolutional layers, which may be organized into residual blocks, activation functions, pooling, batch normalization, vision transformers, visual transformers, other deep learning mechanisms, or combinations thereof. The backbone network may process images (e.g., predicted images, target images, etc.) as inputs and generate outputs, which may be fed to the rest of the machine learning network. For example, the outputs may include one or more feature maps that include features of the input image via an output layer.
[0091] The output of the backbone network (e.g., a feature map) may be received by one or more further networks or layers configured to identify one or more parameters, such as presence of an object, number of objects, location of an object, size of an object, type of object, maturity of a plant, category of a plant, type of weed, type of crop, health of a plant, or a combination thereof. In some embodiments, a network or layer may be configured to evaluate a single parameter. In some embodiments, a network or layer may be configured to evaluate two or more parameters. In some embodiments, the parameters may be evaluated by a single network or layer. The output of the network or layer may include a grid including one or more cells. A cell of the grid may represent an object (e.g., a plant, a pest, a piece of equipment, or a surface irregularity). The cell may further include a parameter of the object, such as location of an object, size of an object, maturity of a plant, category of a plant, type of weed, type of crop, health of a plant, or a combination thereof. In some embodiments, the location of the object may be represented as an offset relative to a reference point (e.g., relative to a corner of the grid cell). It should be noted that although grids and cells are broadly described and illustrated throughout this disclosure in accordance with a Cartesian coordinate system, the grids and cells may also be defined using other coordinate systems (e.g., polar coordinates, etc.).
[0092] In some examples, the output of the backbone network (e.g., feature maps) may be received by an atrous spatial pyramid pooling (ASPP) layer, which may apply a series of atrous convolutions to the output of the backbone network. In some embodiments, the outputs of the atrous convolutions may be pooled and provided to subsequent network layers of the point detection model.
[0093] The point detection model may further include one or more networks associated with the points configured to predict parameters of the object. The point networks may receive an output from the backbone network or the ASPP layer. Examples of networks associated with the points that may be implemented to predict parameters of an object may include a point hit network, a point category network, a point offset network, and a point size network. In some cases, the functions of the aforementioned networks may be combined such that a single network may be implemented to predict parameters of an object.
[0094] In an embodiment, a point hit network may be implemented to generate a grid of predictions with an output slice for each hit class. In some cases, a grid cell may be designated as including a hit (i.e., including an object). As an illustrative, but non-limiting example, a hit class may correspond to whether the object is a weed, a crop, a pest, or other class of predefined object. In some cases, a grid cell may be designated as not including a hit (i.e., not including an object). Another example of a hit class may include an infrastructure class, which may correspond to whether the object includes a drip tape or other watering mechanism. The point hit network may include a series of one or more CNNs running in parallel, each of which may include a series of convolutional layers, activation functions, batch normalization functions, skip connections, pooling, or other deep learning mechanisms (e.g., vision or visual transformers), or combinations thereof. The output of the point hit network may include an output slice for each hit class (e.g., weed, crop, equipment, pest, or surface irregularity). For example, the point hit network may include a first output slice corresponding to a weed category and a second output slice corresponding to a crop category. In some embodiments, an activation function (e.g., a sigmoid activation function, a softmax activation function, a staircase activation function, a linear activation function, a hyperbolic tangent activation function, a rectified linear unit (ReLU) activation function, a swish activation function, etc.) may be applied to the output of the point hit network. The output of the point hit network may include a set of predictions, optionally organized as a grid, including an output slice for each hit category (e.g., weed, crop, equipment, pest, or surface irregularity).
[0095] A point category network may be included to identify an object category or type for any object present in the image. The point category network may include a series of one or more CNNs running in parallel, each of which may include a series of convolutional layers, activation functions, batch normalization functions, skip connections, and pooling or other deep learning mechanisms (e.g., vision transformers, visual transformers, etc.), or combinations thereof. The output of the point category network may include output slices for each object category or type. For example, in the case of plant classification (e.g., weeds or crops), the point category network may identify a specific type of plant that corresponds to the identified plant classification (e.g., grass, broadleaf, purslane, side shoot, onion, strawberry, carrot, corn, soybean, barley, oats, wheat, alfalfa, cotton, grass, tobacco, rice, sorghum, tomato, potato, grape, rice, lettuce, bean, pea, sugar beet, etc.). The point category network, in this example, may include a first output slice corresponding to a type of grass, a second output slice corresponding to a type of broadleaf, a third output slice corresponding to a type of purslane, etc. In some embodiments, an activation function (e.g., sigmoid activation function, softmax activation function, staircase activation function, linear activation function, hyperbolic tangent activation function, ReLU activation function, swish activation function, etc.) may be applied to the output of the point category network. The output of the point category network may include a set of predictions, optionally organized as a grid, with an output slice for each hit category or type (e.g., grass, broadleaf, purslane, side branch, onion, strawberry, carrot, corn, soybean, barley, oats, wheat, alfalfa, cotton, grass, tobacco, rice, sorghum, tomato, potato, grape, rice, lettuce, bean, pea, sugar beet, spider, ant, grasshopper, worm, beetle, caterpillar, fungus, rock, etc.).
[0096] The point detection model may further include a point offset network, which may be included to identify point locations of objects present in the image. The point offset network may include a series of one or more CNNs running in parallel, each of which may include a series of convolutional layers, activation functions, batch normalization functions, skip connections, and pooling or other deep learning mechanisms (e.g., vision transformers, visual transformers, etc.), or combinations thereof. The output of the point offset network may include output slices for each coordinate dimension of each object category (e.g., an output slice of x-coordinates and an output slice of y-coordinates for each crop category, weed category, pest category, equipment category, surface roughness category, or combinations thereof). For example, the point offset network may include a first output slice corresponding to the x-coordinate of the weed category, a second output slice corresponding to the y-coordinate of the weed category, a third output slice corresponding to the x-coordinate of the crop category, and a fourth output slice corresponding to the y-coordinate of the crop category. In some embodiments, an activation function (e.g., sigmoid activation function, softmax activation function, staircase activation function, linear activation function, hyperbolic tangent activation function, ReLU activation function, swish activation function, etc.) may be applied to the output of the point offset network. The output of the point offset network may include a set of predictions, optionally organized as a grid, with an output slice for each coordinate dimension and type of category, as described above. In some embodiments, point locations may be represented as Cartesian coordinates (e.g., x-coordinate, y-coordinate, and / or z-coordinate) relative to a reference point in the image (e.g., an edge of the image, a center of the image, or a grid line in the image). In some embodiments, point locations may be represented as polar, spherical, or cylindrical coordinates (e.g., θ, r, and / or φ (spherical) or z (cylindrical) coordinates) relative to a reference point in the image (e.g., an edge of the image, a center of the image, or a polar grid line in the image).
[0097] The point detection model may further include a point size network, which may be included to identify the size of objects present in the image. The point size network may include a series of one or more CNNs running in parallel, each of which may include a series of convolutional layers, activation functions, batch normalization functions, skip connections, and pooling or other deep learning mechanisms (e.g., vision transformers, visual transformers, etc.), or combinations thereof. The output of the point size network may include an output slice for each hit class corresponding to the size of the item in a unit grid (in the case of a rectangular grid) having a hit class (e.g., weed size, crop size, equipment size, pest size, or surface irregularity size). For example, the point size network may include a first output slice corresponding to the size of a weed and a second output slice corresponding to the size of a crop. In some embodiments, an activation function (e.g., sigmoid activation function, softmax activation function, staircase activation function, linear activation function, hyperbolic tangent activation function, ReLU activation function, swish activation function, etc.) may be applied to the output of the point hit network. Optionally, the output may be scaled, for example using a multiplier or exponential modifier. The output of the point size network may include a set of predictions, optionally organized as a grid, including output slices for each hit category (e.g., weed size, crop size, equipment size, pest size, or surface irregularity size).
[0098] The point network predictions (e.g., category predictions from a point hit network, type or class predictions from a point category network, location predictions from a point offset network, size predictions from a point size network, or combinations thereof) may be further processed to reduce errors (e.g., remove false positives, remove points in error-prone image regions, and / or remove duplicates). For example, predictions of objects located within border regions of an image (e.g., within a predetermined distance from an edge of an image) may be discarded to remove objects that may not fit entirely within the image. Alternatively, or in addition, non-maximal suppression may be applied to the output points to remove duplicate predictions within the same region of the image.
[0099] The parameters identified by the point detection module may be provided to one or more systems configured to locate, track, target, or evaluate the identified plants. For example, the location of a plant's meristem may be provided to a targeting system to target the plant with an instrument (e.g., a laser) at the location of the plant's meristem. In some embodiments, the location of the plant's meristem may be a predicted location. In some embodiments, the location of the plant's meristem may be a target location. In another example, the parameters (e.g., plant size, plant type, or a combination thereof) may be provided to an operating time module configured to determine an operating time of an instrument (e.g., laser operating time) based on the provided parameters. In some embodiments, the parameters provided to the system may be separated based on one or more parameters. For example, weed parameters may be provided to a targeting module for weed eradication, and crop parameters may not be provided to the targeting module.
[0100] A machine learning model (e.g., a machine learning component of a point detection module or a prediction module) may be fine-tuned to update the model with additional training samples (e.g., additional training images). The fine-tuning process may be used to improve model performance without completely retraining the model, thereby reducing overall training time without compromising model performance. A machine learning model trained as described herein (e.g., using a standard number of training images, batches, and epochs) may be fine-tuned to incorporate additional samples. The trained model may be used as a parent or base model for the fine-tuning process. For example, weights determined for the trained model may be used as a starting point for updating the model with additional samples. The additional samples may be combined with samples used to train the parent model to form a training dataset. The samples may be referred to as "old" (e.g., images used to train the parent model) or "new" (e.g., additional images not used to train the parent model). Training batches may be formed using samples where a predetermined ratio of old and new samples are randomly selected from the training dataset for each batch. For example, each batch may contain 50% old samples and 50% new samples, or each batch may contain 70% old samples and 30% new samples. In some embodiments, the ratio of old and new data per batch may be selected based on the amount of data in each category, the similarity of the data between the two categories, or other parameters. By using batches that contain a mix of old and new samples, the model may be updated using fewer batches, fewer epochs, or fewer batches and fewer epochs than if the model were completely retrained. Furthermore, by using a mix of old and new samples, the model's performance on the old samples may be preserved while its performance on the new samples may be improved.
[0101] In some embodiments, a machine learning model (e.g., a machine learning component of a point detection module or a prediction module) may undergo a pre-training step before training. The pre-training step may improve model performance, reduce training time, or both. Pre-training may be performed using a large combined dataset of samples that share a common feature (e.g., images of plants). For example, the combined dataset may include images of weeds and images of crops, which share the common feature of being images of plants. Pre-training may use a larger number of epochs than the full model training (e.g., 80 epochs instead of 40 epochs) and a larger number of samples than the full model training (e.g., 15,000 images instead of 7,500 images). The pre-training process may be used to identify weights that better reflect the model data than common initial weights (e.g., initial weights of ResNet50, MobileNet, or CBNetV2). The weights identified from pre-training may be used as a starting point for the full model training. For example, pre-training may identify initial weights that better represent the plant image data, and the identified weights from pre-training may be used as initial weights for training a model to identify a particular type or category of plants (e.g., weeds, crops, weed types, or crop types) or to distinguish a particular type or category of plants (e.g., to distinguish onions from weeds or carrots from weeds). The pre-trained model may then be used as a starting point for training a full model on a subset of pre-training data specific to the full model. Training the full model may improve specialized performance compared to the pre-trained model. For example, a fully trained model may have improved performance for distinguishing weeds from other plants compared to a pre-trained model trained to identify unspecified plants. In some embodiments, the same pre-trained model may be used to train multiple specialized models. For example, the same pre-trained model may be used to train specialized models to identify weeds among crop types.For example, a specialized model may be trained to identify weeds in a field of onions. In another example, a specialized model may be trained to identify weeds in a field of carrots.
[0102] FIG. 10 illustrates an example of a method 1000 in which a point detection module may be trained and fine-tuned using the methods described herein. At step 1010, an untrained network may receive pre-trained image data. The pre-trained image data may include a combined dataset of images sharing common features, e.g., images of plants. The pre-trained image data may include labeled image data from multiple training datasets, e.g., labeled image data from a weed training set and labeled image data from a crop training set. The point detection module may be pre-trained at step 1020 using the pre-trained image data. Weights of a pre-trained model may be identified at step 1030 based on the pre-training. The weights of the pre-trained model may more faithfully represent an image dataset (e.g., a weed image dataset, a crop image dataset, a farm image dataset, a region image dataset, a company image dataset, a weed image dataset, or a seed image dataset) than weights from an untrained model. The pre-trained point detection module may receive labeled image data corresponding to a dataset of interest at step 1040. For example, the labeled image data may include labeled image data from a weed training set (e.g., labeled images of purslane weeds in a field, labeled images of broadleaf weeds in a field, labeled images of side shoots in a field, or labeled images of grasses in a field).In another example, the labeled image data may be a set of labeled image data from a crop training set (e.g., an image of an onion field with labeled onions and weeds, an image of a strawberry field with labeled strawberries and weeds, an image of a carrot field with labeled carrots and weeds, an image of a corn field with labeled corn plants and weeds, an image of a soybean field with labeled soybeans and weeds, an image of a barley field with labeled barley plants and weeds, an image of an oat field with labeled oats and weeds, an image of a wheat field with labeled wheat plants and weeds, an image of an alfalfa field with labeled alfalfa plants and weeds, an image of a cotton field with labeled cotton plants and weeds, an image of a wheat ... The labeled image data may include an image of a cotton field with plants and weeds, an image of a grass field with labeled grass plants and weeds, an image of a tobacco field with labeled tobacco plants and weeds, an image of a rice field with labeled rice and weeds, an image of a sorghum field with labeled sorghum plants and weeds, an image of a tomato field with labeled tomatoes and weeds, an image of a potato field with labeled potatoes and weeds, an image of a vineyard with labeled grapes and weeds, an image of a lettuce field with labeled lettuce plants and weeds, an image of a bean field with labeled beans and weeds, an image of a pea field with labeled peas and weeds, or an image of a sugar beet field with labeled sugar beets and weeds. In another example, the labeled image data may include labeled image data from a farm training set (e.g., an image of a particular farm field with labeled crops and weeds). In another example, the labeled image data may include labeled image data from a regional training set (e.g., images of fields in a particular agricultural region with labeled crops and weeds). The point detection module may be trained in step 1050 to identify object parameters (e.g., location, size, category, or type) for an object of interest (e.g., a plant, a weed, a type of weed, a crop, or a type of crop).Before, after, or during steps of process 1000, or simultaneously with process 1000, a training process of the point detection module (e.g., after any of steps 1050, 1060, or 1070), i.e., the trained, partially trained, or fine-tuned point detection module may be used for object detection 1051 to identify object parameters (e.g., location, size, category, or type) for an object of interest by receiving an image at step 1053, e.g., an image of a ground containing one or more objects. The point detection module (e.g., the trained, partially trained, or fine-tuned point detection module obtained from steps 1050, 1060, or 1070) may be used for object detection 1051. In some embodiments, object detection 1051 may be performed by a predictive system (e.g., predictive system 400 of FIG. 11) to perform point detection. An object may be detected in the image at step 1055. In step 1060, upon receipt of further labeled image data, e.g., a new image of the object of interest (e.g., a new image of a plant, a weed, a weed type, a crop, or a crop type), the point detection module may be pre-trained in step 1020, trained in step 1050, or fine-tuned in step 1070.
[0103] In some embodiments, labeled image data, such as the labeled image data received in step 1040 or further labeled image data received in step 1060, may be obtained from the images received in step 1053. The images received in step 1053 may be labeled and used to train or fine-tune a point detection model. In some embodiments, the object detection performed in step 1055 may be used to determine which images to further label and use for training or fine-tuning.
[0104] Operation of the device One or more parameters of the target object (e.g., target plant) evaluated by the point detection system may be used to determine the activation of the instrument (e.g., whether to activate the instrument, where on the object to activate the instrument, or duration of activation). In some embodiments, activation may be determined by an activation module based on one or more parameters. For example, whether to activate the instrument may be based on the object category (e.g., weed or crop). In another example, the location of activation on the object (e.g., the portion of the object targeted by the instrument) may be determined based on the object shape (e.g., center of gravity location, meristem location, leaf shape, or phyllotaxis) or the object posture (e.g., upright, bent, or lying down). In another example, the activation time of the instrument may be determined based on the object category (e.g., broadleaf, lateral branch, purslane, or grass), the object size (e.g., small, medium, or large), or a combination thereof. The activation module may be part of a prediction module, a location prediction module, a scheduling module, a targeting module, a targeting control module, or a combination thereof. The instrument (e.g., laser) of the targeting module controlled by the targeting control module can target the plant at the location of the plant's meristem for a time determined by the operation module. The operation time can be a time sufficient to manipulate (e.g., kill) the target plant. Targeting the plant's meristem with the instrument can facilitate accurate targeting of meristem cells. For example, irradiating the plant's meristem cells with an infrared laser instrument can incinerate the meristem cells, thus killing the plant. In addition to the location of the meristem, further parameters can be provided to the operation module to determine the operation time.
[0105] In one example, the plant size, the plant type, or both may be provided to the actuation module and used to determine the actuation time of a laser instrument configured to irradiate and incinerate a target plant. Larger plants or certain plant types may be more resistant to incineration and may require longer irradiation to kill the plant. Table 1 shows an example of a type factor multiplier that may be applied to the actuation time to account for the resistance of different plant types. [Table 1]
[0106] In some embodiments, an additional multiplier may be applied to the actuation time to account for the non-linear scaling of actuation time with plant size. An example of the size factor multiplication is shown in Table 2. [Table 2]
[0107] By way of example, an activation time sufficient to kill a plant may be determined as follows: Activation time (ms) = r (time factor) (type factor) (size factor) + reference time where r is the size of the plant. The time factor may account for system parameters or external conditions, such as laser intensity, temperature, altitude, or other factors. The type factor may account for the difference in kill time between different weed types. The size factor may account for the non-linear scaling of kill time with weed size, for example, as shown in Table 2. The reference time may be the minimum activation time and may be adjusted to account for system parameters or external conditions.
[0108] In some embodiments, the time multiplier may be about 50 ms. In some embodiments, the time multiplier may be about 10 ms, about 20 ms, about 30 ms, about 40 ms, about 50 ms, about 60 ms, about 70 ms, about 80 ms, about 90 ms, or about 100 ms. In some embodiments, the time multiplier may be about 10 ms to about 100 ms, about 20 ms to about 80 ms, about 30 ms to about 70 ms, or about 40 ms to about 60 ms. In some embodiments, the reference time may be about 50 ms. In some embodiments, the reference time may be about 10 ms, about 20 ms, about 30 ms, about 40 ms, about 50 ms, about 60 ms, about 70 ms, about 80 ms, about 90 ms, or about 100 ms. In some embodiments, the reference time can be about 10 ms to about 100 ms, about 20 ms to about 80 ms, about 30 ms to about 70 ms, or about 40 ms to about 60 ms. In some embodiments, the actuation time can be about 100 ms to about 10,000 ms, about 100 ms to about 5,000 ms, about 100 ms to about 2,000 ms, or about 200 ms to about 2,000 ms.
[0109] In another example, the activation time sufficient to kill a plant may be determined using a machine learning model. The machine learning model may be trained using a dataset of observed activation times sufficient to kill plants having various characteristics. For example, activation times may be measured for plants of various sizes and types, and the observed activation times may be used to train the machine learning model. For example, while using the point detection system to eradicate a plant or otherwise remove an object using a laser, the point detection system may record the activation time of the laser, and in some cases, further record image data that may be used to determine whether the removal of the plant or other object was successful. This data may be evaluated by a user or other entity to determine whether the activation time used for a particular plant or object was sufficient to successfully remove the plant or other object. Based on this evaluation, the dataset of observed activation times may be updated and used to iteratively train the machine learning model. For example, if the activation time of the laser is deemed insufficient to eradicate a particular type of plant having a particular size, the machine learning model may be updated such that activation times may be automatically extended for plants of similar types and sizes to ensure the eradication or otherwise removal of these plants is successful. Alternatively, if the activation time of the laser is deemed sufficient to eradicate a particular plant type having a particular size, the machine learning model may be enhanced so that this activation time may be used for plants of similar types and sizes. Thus, while using the point detection system to eradicate plants or otherwise remove objects using a laser, the machine learning model may be continually and iteratively updated to accurately identify appropriate activation times for different plants and objects.
[0110] The activation time may be used to determine whether to target an object. As described herein, the scheduling module may select an object identified by the predictive system and schedule the object to be targeted by the targeting system. In some embodiments, the scheduling module may prioritize targeting objects with shorter activation times over objects with longer activation times. For example, the scheduling module may schedule four weeds with shorter activation times to be targeted before one weed with a longer activation time so that more weeds can be targeted and killed in the available time.
[0111] In some embodiments, the operation of the tool may be based on the confidence score of the object. The confidence score may quantify the confidence that the object is identified, classified, located, or a combination thereof. For example, the confidence score may quantify the certainty that the plant is classified as a weed. In another example, the confidence score may quantify the certainty for classifying the plant as each of broadleaf, purslane, side shoot, or grass. In some embodiments, the confidence score may quantify the certainty that the object is not a particular class or type. For example, the confidence score may quantify the certainty that the object is not a crop and may be used to determine whether to destroy the crop with a laser. A confidence score may be assigned to each identified object in each collected image for each evaluated parameter (e.g., one or more of object location, weed classification, crop classification, purslane weed type, broadleaf weed type, side shoot weed type, grass weed type, onion crop type, strawberry crop type, carrot crop type, corn crop type, or soybean crop type). The confidence score may be used to determine how long to activate the instrument, for example, an object that is confidently identified as a large piece of grass may be targeted with a laser for longer than an object that is confidently identified as a small leaf.
[0112] In some embodiments, the confidence score may range from 0 to 1, with 0 corresponding to low confidence and 1 corresponding to high confidence. The threshold value for what is considered to be high confidence may be context sensitive and may be adjusted based on the desired outcome. In some embodiments, high confidence values may be considered to be 0.5 or more, 0.6 or more, 0.7 or more, 0.8 or more, or 0.9 or more. In some embodiments, low confidence values may be considered to be less than 0.3, less than 0.4, less than 0.5, less than 0.6, less than 0.7, or less than 0.8. Objects with higher confidence scores for a first object type and lower confidence scores for other object types may be identified as the first object type. For example, an object with a confidence score of 0.6 for a broadleaf type, a confidence score of 0.1 for a purslane type, a confidence score of 0.2 for a side branch type, and a confidence score of 0.1 for a grass type may be identified as a broadleaf weed.
[0113] The confidence score may be used to determine whether to activate the instrument on the object by evaluating the level of confidence that the object has the parameters selected for targeting with the instrument. For example, the confidence score may be used to determine whether to destroy the object with a laser by evaluating the level of confidence that the object is a weed. In some embodiments, determining whether to target the object with the instrument may include evaluating the confidence score over time (e.g., determining the confidence score for multiple observations of the object over a series of image frames). An object may be targeted if multiple high confidence observations are made. An object may not be targeted if a single high confidence observation and multiple low confidence or ambiguous observations are made. For example, an object may be targeted if it has weed confidence scores of 0.9, 0.8, 0.8, and 0.8 over four image frames. In another example, an object may be targeted if it has weed confidence scores of 0.9, 0.7, 0.5, and 0.8 over four image frames. In another example, an object may not be targeted if it has weed confidence scores of 0.4, 0.5, 0.8, and 0.4 across four image frames. A threshold for the confidence value, the number of observations, or both may be used to determine whether to target the object. The threshold may be situational and may be adjusted based on the desired outcome. In some embodiments, the threshold for the confidence value or the number of observations may be determined based on the number of opportunities for observation. For example, the threshold for the number of observations may be determined based on the number of frames in which the object is predicted to be in the camera's field of view. In some embodiments, the threshold may be determined empirically.
[0114] Computer system and method The detection and targeting methods described herein may be implemented using a computer system. In some embodiments, the detection system described herein includes a computer system. In some embodiments, the computer system may implement the object identification and targeting methods autonomously without human intervention. In some embodiments, the computer system may implement the object identification and targeting methods based on instructions provided by a human user via a detection terminal.
[0115] FIG. 8 illustrates components in a block diagram of a non-limiting exemplary embodiment of a detection terminal 1400 according to various aspects of the disclosure. In some embodiments, the detection terminal 1400 is a device that displays a user interface to provide access to a detection system. As shown, the detection terminal 1400 includes a detection interface 1420. The detection interface 1420 allows the detection terminal 1400 to communicate with the detection system. In some embodiments, the detection interface 1420 may include an antenna configured to communicate with the detection system, for example, by a remote control. In some embodiments, the detection terminal 1400 may also include a local communication interface, for example, an Ethernet interface, a Wi-Fi interface, or other interface that allows other devices associated with the detection system to connect to the detection system via the detection terminal 1400. For example, the detection terminal may be a portable device, for example, a mobile phone, that runs a graphical interface that allows a user to remotely operate or monitor the detection system via Bluetooth, Wi-Fi, or a mobile network.
[0116] The detection terminal 1400 further includes a detection engine 1410. The detection engine may receive information regarding the status of the detection system. The detection engine may receive information regarding the number of identified objects, the identity of the identified objects, the location of the identified objects, the trajectory and predicted location of the identified objects, the number of targeted objects, the identity of the targeted objects, the location of the targeted objects, the location of the detection system, the elapsed time of a task performed by the detection system, the area covered by the detection system, the charge of the detection system, or a combination thereof.
[0117] Actual embodiments of the illustrated devices will include many more components therein that will be known to those skilled in the art. For example, each of the illustrated devices will have a power source, one or more processors, computer-readable media for storing computer-executable instructions, etc. For clarity, these additional components are not illustrated herein.
[0118] In some examples, the procedures described herein (e.g., the procedures of FIG. 6, or other procedures described herein) may be performed by a computing device or apparatus, such as a computing device having the computing device architecture 1600 shown in FIG. 9. In one example, the procedures described herein may be performed by a computing device having the computing device architecture 1600. The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device, a server (e.g., in a Software as a Service (SaaS) system or other server-based system), and / or any other computing device having resource capabilities to perform the processes described herein, including the procedures of FIG. 6. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display (as an example of an output device or in addition to the output device), a network interface configured to communicate and / or receive data, any combination thereof, and / or other component(s), and the network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.
[0119] The components of the computing device may be implemented in circuitry, for example, the components may include and / or be implemented using electronic circuitry or other electronic hardware that may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry) and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.
[0120] A procedure is illustrated in FIG. 6, where the operations represent a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In terms of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed on one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to execute a process.
[0121] Additionally, the processes described herein may be executed under the control of one or more computer systems configured with executable instructions, implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors through hardware or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0122] 9 illustrates an example computing device architecture 1600 of an example computing device capable of implementing various techniques described herein. For example, the computing device architecture 1600 can implement the procedures illustrated in FIG. 6 or control the vehicle illustrated in FIGS. 1 and 2. The components of the computing device architecture 1600 are shown in electrical communication with each other using connections 1605, e.g., a bus. The example computing device architecture 1600 includes computing device connections 1605 coupling various computing device components to the processor 1610, including a processing unit (which may include a CPU and / or a GPU) 1610, and computing device memory 1615, e.g., read only memory (ROM) 1620 and random access memory (RAM) 1625. In some embodiments, the computing device may include a hardware accelerator.
[0123] The computing device architecture 1600 may include a cache of high-speed memory directly connected to, closely connected to, or integrated as part of, the processor 1610. The computing device architecture 1600 may copy data from the memory 1615 and / or storage device 1630 to the cache 1612 for quick access by the processor 1610. In this manner, the cache may provide a performance boost that avoids delays to the processor 1610 while waiting for data. These and other modules may control or be configured to control the processor 1610 to perform various actions. Other computing device memories 1615 may be used as well. The memory 1615 may include multiple different types of memories with different performance characteristics. The processor 1610 may include any general-purpose processor as well as hardware or software services, such as service 1 1632, service 2 1634, and service 3 1636 stored in the storage device 1630 configured to control the processor 1610, as well as dedicated processors whose software instructions are built into the processor design. Processor 1610 may be a self-contained system that includes multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0124] To enable user interaction with the computing device architecture 1610, the input device 1645 can represent any number of input mechanisms, such as a microphone for audio, a touch screen for gesture or graphic input, a keyboard, a mouse, motion input, voice, etc. The output device 1635 can also be one or more of several output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some cases, a multimodal computing device can enable a user to provide multiple types of input to communicate with the computing device architecture 1600. The communication interface 1640 can generally govern and manage user input and computing device output. There is no limitation to operation with any particular hardware configuration, so the basic features herein can be easily substituted for improved hardware or firmware configurations as they are developed.
[0125] The storage device 1630 may be a hard disk or other type of computer readable medium that is non-volatile memory and can store data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a random access memory (RAM) 1625, a read only memory (ROM) 1620, and hybrids thereof. The storage device 1630 may include services 1632, 1634, 1636 for controlling the processor 1610. Other hardware or software modules are also contemplated. The storage device 1630 may be connected to computer device connections 1605. In one aspect, a hardware module that performs a particular function may include software components stored on a computer readable medium in association with the hardware components necessary to perform that function, such as the processor 1610, the connections 1605, the output devices 1635, etc.
[0126] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or transporting instruction(s) and / or data. Computer-readable media may include non-transitory media, including carrier waves and / or transitory electronic signals capable of storing data and propagating wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media, such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memories, or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0127] In some embodiments, the computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referring to non-transitory computer readable storage media, media such as energy, carrier signals, electromagnetic waves, and the signals themselves are expressly excluded.
[0128] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For clarity of explanation, in some cases, the technology may be presented as including individual functional blocks, including functional blocks including devices, device components, steps or routines in a method implemented in software or a combination of hardware and software. Additional components other than those shown in the drawings and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other examples, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.
[0129] Particular embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of the operations may be rearranged. A process terminates when the operations are completed, but may include additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a call of the function or a return of the function to the main function.
[0130] The processes and methods according to the above examples may be implemented using computer-executable instructions stored on or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or configure a general purpose computer, a special purpose computer, or a processing device to perform a certain function or group of functions. Some of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binaries, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network storage devices, etc.
[0131] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may adopt any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program product) to perform the necessary tasks may be stored on a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. As a further example, such functionality may also be implemented on a circuit board between different chips or different processes executing within a single device.
[0132] The instructions, media for carrying such instructions, computing resources for executing them, and other structures for supporting such computing resources are examples of means for providing the functionality described in this disclosure.
[0133] The various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, the various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0134] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device with multiple uses, including applications in wireless communication device handsets and other devices. Any functions described as modules or components may be implemented together in an integrated logic device, or separately as separate but interoperable logic devices. When implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium that includes program code having instructions that, when executed, perform one or more of the above-described methods. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the techniques may be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.
[0135] The program code may be executed by one or more processors, such as a processor that may include one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein.
[0136] While illustrative embodiments have been illustrated and described, it will be understood that various changes can be made therein without departing from the spirit and scope of the disclosure.
[0137] In the foregoing description, aspects of the present application have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary embodiments of the present application have been described in detail herein, it will be understood that the inventive concepts may be embodied and used in various other ways, and that the appended claims are intended to be construed to include such variations except as limited by the prior art. The various features and aspects of the present application described above may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, the methods have been described in a particular order. It will be understood that in alternative embodiments, the methods may be performed in an order different from that described.
[0138] Those skilled in the art will recognize that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of the description.
[0139] When a component is described as being "configured to" perform a particular operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or by any combination thereof.
[0140] The phrase "coupled to" refers to any component that is directly or indirectly physically connected to another component and / or any component that is in direct or indirect communication with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).
[0141] Claim language or other language reciting "at least one" of a set and / or "one or more" of a set indicates that one element of the set, or multiple elements of the set (in any combination), satisfy the claim. For example, claim language reciting "at least one of A and B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" can mean A, B, or A and B, and can further include items not listed in the set of A and B.
[0142] As used herein, the terms "about" and "approximately" in reference to numerical values are used herein to include numerical values that fall within 10%, 5%, or 1% of that numerical value in either direction (greater than or less than), unless otherwise stated or clear from the context (except where such numerical value exceeds 100% of its possible values). EXAMPLES
[0143] The invention is further illustrated by the following non-limiting examples.
[0144] Example 1 Eradication of weeds in the field This example describes the eradication of weeds in a field using the detection method of the present disclosure. A vehicle with a prediction system, a targeting system, and an infrared laser, as shown in Figures 1 and 3, was placed in a field as shown in Figure 2. The vehicle moved through the crop rows at a speed of about 2 miles per hour, and a predictive camera collected images of the field. The predictive system identified weeds in the images and determined weed parameters, including leaf radius and weed type, as shown by the dashed circle in Figure 4. The predictive system determined the predicted location of the weed, which corresponds to the location of the weed meristem, as shown by the solid circle and center point in Figure 4. The predictive system sent the predicted location to the targeting system.
[0145] The targeting system was selected based on availability and proximity to the selected weeds. The targeting system included a targeting camera and an infrared laser, the direction of which was adjusted by mirrors controlled by actuators. The mirrors reflected visible light from the surface to the targeting camera and infrared light from the laser to the surface. The targeting system converted the predicted positions received from the prediction system into actuator positions. The targeting system adjusted the actuators to point the targeting camera and infrared laser beam toward the predicted positions of the selected weeds. The targeting camera imaged the field at the predicted positions of the weeds and modified its position to create a target position. The targeting system adjusted the position of the targeting camera and infrared laser beam based on the target positions of the weeds and activated the infrared beam toward the positions of the weeds. The beam irradiated the weeds with infrared light for a period of time based on the parameters of the weeds, killing them.
[0146] Example 2 Determination of laser operating time from weed parameters This example describes determining a sufficient laser activation time to kill a weed based on the weed's parameters. The weed was identified in an image and its parameters were determined. The parameters included leaf radius and weed type. The laser activation time in milliseconds (ms) was determined as follows: Activation time (ms) = r (time factor) (type factor) (size factor) + reference time where r is the leaf radius in millimeters (mm) measured from the meristem to the leaf tip furthest from the meristem. The time factor is a multiplier that can be adjusted to account for system parameters or external conditions, such as laser intensity, temperature, altitude, or other factors. The type factor is a multiplier that accounts for the difference in kill time between different weed types. The size factor is a size multiplier that accounts for the nonlinear scaling of kill time with weed size, where different size factors are applied to leaf radii in small, medium, or large size categories, and the multiplier increases with increasing size category. The reference time is the minimum activation time in milliseconds (ms) that is applied to each weed. In this example, the reference time is 50 ms, but this reference time can be adjusted to account for system parameters or external conditions.
[0147] Table 3 shows examples of weed control parameters, multipliers and activation times for the weeds shown in Figure 5. The weed meristem is marked with a solid circle containing a cross line and the leaf radius is shown with a dashed circle. [Table 3]
[0148] The determined laser activation time was provided to a targeting system including an infrared laser, the infrared laser was aimed at the meristem of the weed, and the laser was activated for the determined time, which was sufficient to incinerate the meristem of the plant, thereby killing the weed.
[0149] Example 3 Point detection model architecture This example describes a model architecture of a point detection system used to identify and locate weeds. Images of the ground surface are collected by a predictive camera and passed to a backbone network, as shown in FIG. 6. The backbone network is a convolutional neural network, or a network built on a vision transformer. The output of the backbone network is a set of feature maps that are fed into a series of further networks used to identify plant parameters. For example, the further networks include a point hit network to identify plant hits and differentiate the hits as plants or crops, a point category network to identify the type of weed or crop, a point size network to identify the size of the weed or crop, and a point offset network to identify the location of the weed or crop.
[0150] Each of the additional networks, including the point hit network, the point category network, the point size network, and the point offset network, generates a grid where each cell of the grid can represent a plant from which parameters (e.g., weed or crop, plant type, plant size, or plant offset / location) are identified.
[0151] Example 4 Selection and Scheduling of Targeted Weeds This example describes the selection and scheduling of weeds targeted for eradication. Objects are detected in images collected by the predictive camera of the autonomous weed eradication system. Each object is located and assigned a confidence score to a plant category and plant type, including a crop confidence score and a weed confidence score. The object confidence scores may be based on a single image or multiple images. Objects with a weed confidence score above a target threshold, a crop confidence score below a target threshold, or both, are identified as weeds. Objects with a crop confidence score above a target threshold, a weed confidence score below a target threshold, or both, are identified as crops. For objects identified as weeds, further parameters are identified, including weed type, confidence value for each weed type, weed size, and operation time.
[0152] Objects identified as weeds are scheduled for eradication based on parameters including weed location, plant and weed confidence scores, and eradication time. To ensure that weeds and not crops are targeted, objects with high weed confidence scores and / or low plant confidence scores are scheduled for eradication with a high priority, and objects with low weed confidence scores and / or high plant confidence scores are scheduled for eradication with a low priority. To eradicate as many weeds as possible within the available time, weeds with short operating times are scheduled for eradication with a higher priority than weeds with long operating times.
[0153] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art in the future without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. It is intended that the following claims define the scope of the invention, and that methods and structures within the scope of these claims, and their equivalents, are covered thereby.
Claims
1. 1. A computer-implemented method for detecting a target plant, the computer-implemented method comprising: receiving an image of an area of a surface, the area including a target plant located on the surface; Identifying one or more parameters of the target plant, wherein the one or more parameters of the target plant include a point location of the target plant; and Identifying the target plant in the image based on the one or more parameters of the target plant.
2. The computer-implemented method of claim 1 , wherein the area of the surface further comprises one or more additional plants.
3. The computer-implemented method of claim 1 or claim 2, wherein the target plant is a weed or a crop.
4. The computer-implemented method of claim 1 or claim 2, wherein the point locations correspond to features of the target plant.
5. A computer-implemented method as described in claim 4, wherein the feature is a meristem of the target plant.
6. The computer-implemented method of claim 4 , wherein the feature is a center of the target plant or a leaf of the target plant.
7. 3. The computer-implemented method of claim 1 or claim 2, wherein the one or more parameters further comprise plant location, plant size, plant category, plant type, leaf shape, phyllotaxy, plant posture, plant health, or a combination thereof.
8. The computer-implemented method of claim 1 or claim 2, further comprising targeting the target plant at the point location with an instrument.
9. The computer-implemented method of claim 8 , wherein the tool is a laser, a spray, or a grabber.
10. The computer-implemented method of claim 8 , further comprising actuating the instrument at the point location for a period of time.
11. The computer-implemented method of claim 10 , wherein the period of time is sufficient to kill the target plant.
12. The computer-implemented method of claim 10 , wherein the time period is based on one or more characteristics of the target plant.
13. The computer-implemented method of claim 12 , wherein the one or more characteristics include plant size, plant type, or both.
14. The computer-implemented method of claim 9 , further comprising killing the target plant with the instrument.
15. The computer-implemented method of claim 7 , comprising at least one of irradiating, irradiating with light, heating, or incinerating the features of the target plant using the instrument.
16. The computer-implemented method of claim 1 or claim 2, further comprising determining a size of the target plant.
17. The computer-implemented method of claim 16 , wherein the size of the plant comprises a size of one or more structures of the target plant.
18. 20. The computer-implemented method of claim 17, wherein the one or more structures are selected from the group consisting of leaves, stems, leaf blades, flowers, fruits, seeds, shoots, buds, and combinations thereof.
19. 17. The computer-implemented method of claim 16, wherein the size of the plant comprises a length, a radius, a diameter, an area, or any combination thereof.
20. The computer-implemented method of claim 1 or claim 2, further comprising classifying the target plant type.
21. 21. The computer-implemented method of claim 20, wherein the plant type is based on the leaf shape of the target plant.
22. 21. The computer-implemented method of claim 20, wherein the type of plant is selected from the group consisting of crops, weeds, grasses, broadleaves, purslane, or combinations thereof.
23. The computer-implemented method of claim 1 or claim 2, further comprising assessing the condition of the target plant.
24. 24. The computer-implemented method of claim 23, wherein the condition comprises health status, maturity, nutritional status, disease state, ripeness, crop yield, or any combination thereof.
25. The computer-implemented method of claim 1 or claim 2, further comprising determining a confidence score for the one or more parameters.
26. 26. The computer-implemented method of claim 25, further comprising scheduling the target plants to be targeted based on the confidence scores.
27. The computer-implemented method of claim 1 or claim 2, further comprising using a trained classifier to identify the target plant.
28. The computer-implemented method of claim 1 or claim 2, further comprising using a trained classifier to localize features of the target plant.
29. 30. The computer-implemented method of claim 28, wherein the trained classifier is trained using a training dataset comprising labeled images.
30. 30. The computer-implemented method of claim 29, wherein the labeled images are labeled with plant category, meristem location, plant size, plant condition, plant type, or any combination thereof.
31. The computer-implemented method of claim 1 or claim 2, further comprising pre-training the machine learning model.
32. 1. A computer-implemented method for detecting a target object, the computer-implemented method comprising: receiving an image of a region of a surface, the region including a target object located on the surface; obtaining labeled image data including parameterized objects corresponding to similarly positioned objects; training a machine learning model to identify object parameters corresponding to a target object, wherein the machine learning model is trained using the labeled image data; generating a prediction of an object corresponding to one or more parameters of the target object, the one or more object parameters of the target object including point positions of the target object, the one or more object parameters being identified by using the image as input data to the machine learning model; identifying the target object in the image based on the one or more parameters; and Updating the machine learning model using information corresponding to the image, the one or more parameters, and the identification of the target object, wherein when the machine learning model is updated, new object parameters are identified from new images using the machine learning model.
33. 33. The computer-implemented method of claim 32, wherein the target object is a target plant, a pest, a surface irregularity, or a piece of equipment.
34. 34. The computer-implemented method of claim 33, wherein the target plant is a weed or a crop.
35. 35. The computer-implemented method of claim 33 or claim 34, wherein the labeled image data comprises an image of a plant.
36. 36. The computer-implemented method of claim 35, wherein the plant images are labeled with plant centroid location, meristem location, plant size, plant category, plant type, leaf shape, leaf number, phyllotaxis, plant posture, plant health, or combinations thereof.
37. 35. The computer-implemented method of claim 33 or claim 34, wherein the one or more object parameters further comprise object position, object size, object category, plant type, leaf shape, phyllotaxis, plant posture, plant health, or a combination thereof.
38. 35. The computer-implemented method of claim 33 or claim 34, wherein the surface is an agricultural surface.
39. 35. The computer-implemented method of claim 33 or claim 34, wherein the parameterized object includes data corresponding to point position, shape, size, category, type, or combinations thereof.
40. 35. The computer-implemented method of claim 33 or claim 34, wherein the point locations comprise plant meristem locations, centroid locations, or leaf locations.
41. 35. The computer-implemented method of claim 33 or claim 34, further comprising using a trained classifier to identify the target object.
42. 35. The computer-implemented method of claim 33 or claim 34, further comprising using a trained classifier to localize features of the target object.
43. 43. The computer-implemented method of claim 42, wherein the trained classifier is trained using a training dataset comprising labeled images.
44. 44. The computer-implemented method of claim 43, wherein the labeled images are labeled by object category, meristem location, plant size, plant condition, plant type, or any combination thereof.
45. 35. The computer-implemented method of any one of claims 32 to 34, wherein updating the machine learning model comprises receiving further labeled image data that includes the image, and fine-tuning the machine learning model based on the further labeled image data.
46. 46. The computer-implemented method of claim 45, wherein fine-tuning the machine learning model is performed using the further subset of labeled image data and an image batch comprising the further subset of labeled image data.
47. 46. The computer-implemented method of claim 45, wherein fine-tuning the machine learning model is performed using fewer batches, fewer epochs, or fewer batches and fewer epochs than training the machine learning model.
48. The computer-implemented method of any one of claims 32 to 34, further comprising pre-training the machine learning model.
49. 49. The computer-implemented method of claim 48, wherein pre-training the machine learning model is performed on a pre-training dataset that includes the labeled image data and pre-training labeled image data that shares common features with the labeled image data.
50. 50. The computer-implemented method of claim 49, wherein the common feature is an image of a plant.