System and method for pre-harvest detection of latent infection in plants

Machine learning models predict latent plant infections using pre-harvest data to enable timely mitigation, addressing environmental issues and improving product quality by reducing waste and enhancing customer acceptance.

JP7772789B2Active Publication Date: 2025-11-18APEEL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023523563
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-05
Filing Date
2021-11-05
Publication Date
2025-11-18
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

Existing pre- and post-harvest mitigation methods for plant infections, such as fungicide applications, lead to environmental issues and the development of pathogen resistance, while latent infections in plants often go undetected until after harvest, causing quality loss and waste.

Method used

Utilizing machine learning models to predict latent infections in plants based on pre-harvest data, including geographic region, growth stage, and physiological characteristics, allowing for timely mitigation treatments and adjustments in harvest schedules.

Benefits of technology

Enables early detection of latent infections, reducing waste and improving product quality by allowing for targeted treatments and harvest adjustments, thus enhancing customer acceptance and reducing environmental impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007772789000008
    Figure 0007772789000008
  • Figure 0007772789000009
    Figure 0007772789000009
  • Figure 0007772789000010
    Figure 0007772789000010
Patent Text Reader

Abstract

[0009] Disclosed herein are systems and methods for identifying pre-harvest latent infection in plants. In one aspect, the method can include operations of obtaining data describing expression levels of one or more biomarkers present in a plant, encoding the obtained data into a data structure for input to a machine learning model, providing, by one or more computers, the encoded data structure as input to a machine learning model trained to generate output data indicative of the likelihood that the plant has a latent infection based on processing the encoded data structure, obtaining the generated output data indicative of the likelihood that the plant has a latent infection, determining that the plant has a latent infection based on the generated output data, and performing one or more actions to mitigate the infection in the plant.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 110,343, filed November 5, 2020, the disclosure of which is incorporated herein by reference in its entirety.

[0002] Technical Field This document describes devices, systems, and methods related to pre-harvest prediction of latent infection in various types of plants (eg, produce). [Background technology]

[0003] background

[0003] Many plant products, such as fruits, can have high infection rates. Preharvest infection of plant products can occur when the plant product is infected with a pathogen during growth. For example, approximately 20% to 40% of the avocado trees in an avocado orchard, and therefore the avocados they produce, can be infected with anthracnose before harvest. Unripe, preharvest plant products may appear free of symptoms of such infection. However, after the plant product is harvested, ripens, and undergoes senescence, the plant product may begin to exhibit symptoms of infection. Symptoms of infection can include stem rot, mildew, stem / heart dieback, and other features that can lead to a detrimental loss of plant product quality.

[0004]

[0004] Harvesting can involve removing fruit from a parent tree, for example, through plucking or cutting; both plucking and cutting can expose the xylem elements of the fruit's pedicel to infection and facilitate the transfer of fungal inoculants to the fruit's xylem tissue. While the fruit is immature, fungi can colonize the fruit's xylem but remain latent and actively spread to surrounding tissues as the fruit ripens, causing cellular damage and disease symptoms. This makes infections such as stem and stem rot (SER) challenging postharvest disorders: they originate on the farm but may still go undetected until the plant product (e.g., produce) reaches the consumer. Furthermore, infections such as SER can have temporally and spatially independent severities. This means that early harvests can have different SER incidence rates than later harvests from the same tree, and this incidence rate may not necessarily be comparable across neighboring trees. Summary of the Invention [Problem to be solved by the invention]

[0005]

[0005] Some pre- and post-harvest mitigation methods can involve the pre- and / or post-harvest application of fungicidal chemicals. Repeated fungicide applications can have negative environmental consequences, including the reduction of beneficial microflora on the parent tree. Repeated fungicide applications can also lead to the development of resistant strains of pathogens over time. [Means for solving the problem]

[0006] overview

[0006] The present disclosure generally relates to pre-harvest assessment of plants (e.g., agricultural crops) to predict the occurrence of latent infection in fruits, vegetables, seeds, and / or other plant products (e.g., leaves, stems, roots, etc.) grown from or otherwise derived from the evaluated plant. More specifically, machine learning models can be used to predict the likelihood of latent infection in plants using pre-harvest data about the plant and predictive characteristics. Predictive characteristics can include, but are not limited to, the plant's geographic region, growth stage, age, hormones, dry matter content, environmental conditions, etc.

[0007] In one exemplary embodiment of the disclosed techniques, by testing fruit from the same trees throughout the season, a dataset of physiological and transcriptomic information can be constructed that profiles trees that produce fruit with high or low incidence of one or more latent infections, such as stem rot (SER). While such latent infections in fruits, including but not limited to avocados, may not manifest until the fruit is fully developed on the tree and ripens after harvest, fruit may have markers of infection as early as six months before peak harvest. Thus, the disclosed techniques can provide a platform for using predictive assays for preharvest detection of latent infections in plants, allowing ample time for modifications to fungicide applications, plant harvest schedules, and / or plant marketing efforts (e.g., mitigation treatments) to mitigate the onset of symptoms of such infection in the plants.

[0008]

[0008] One or more embodiments described herein may include a method for identifying pre-harvest latent infection in a plant. In one aspect, the method may include operations of: acquiring, by one or more computers, data describing expression levels of one or more infection biomarkers present in above-ground portions of the plant; encoding, by the one or more computers, the acquired data into a data structure for input to a machine learning model; providing, by the one or more computers, the encoded data structure as input to a machine learning model trained to generate output data indicative of the likelihood that the plant has a latent infection based on processing of the encoded data structure; acquiring, by the one or more computers, the generated output data indicative of the likelihood that the plant has a latent infection; determining, by the one or more computers, that the plant has a latent infection based on the generated output data; and performing, by the one or more computers, one or more treatments to mitigate the latent infection in the plant.

[0009]

[0009] Other aspects may include corresponding systems, devices, and computer programs that perform the operations of the methods disclosed in the present disclosure, as defined by instructions encoded on a computer-readable storage device. These and other aspects may optionally include one or more of the following features. For example, performing, by one or more computers, one or more treatments to mitigate latent infection in the plant may include determining, by one or more computers, that a harvest schedule for the plant and / or one or more other plants of similar origin should be adjusted; and storing, by the one or more computers, data defining the adjusted harvest schedule in a memory device.

[0010] In some embodiments, performing, by the one or more computers, one or more treatments to mitigate latent infection in the plant can include generating, by the one or more computers, an alert message that, when processed by the user device, can cause the user device to output an alert notifying a user of the user device that the plant and one or more other plants of similar origin should be harvested, and transmitting, by the one or more computers, the generated alert to the user device. Performing, by the one or more computers, one or more treatments to mitigate latent infection in the plant can also include determining, by the one or more computers, based on the output data, that an antimicrobial treatment can be prescribed for the one or more other plants of similar origin.

[0011]

[0011] In some embodiments, the one or more other plants of similar origin can include plants in the same zone as the plant. Further, performing, by the one or more computers, one or more treatments to mitigate latent infection in the plants can include generating, by the one or more computers, one or more instructions that, when processed by an irrigation controller, can cause the irrigation controller to spray a liquid comprising the antimicrobial treatment, and transmitting, by the one or more computers, the one or more instructions to the irrigation controller using one or more networks. As yet another example, performing, by the one or more computers, one or more treatments to mitigate latent infection in the produce can include generating, by the one or more computers, one or more instructions that, when processed by a robotic device, can cause the robotic device to (i) navigate to a location associated with the plant and (ii) initiate an antimicrobial treatment that can cause the robotic device to spray a liquid comprising the antimicrobial treatment on the plant and one or more other plants native to the region, and transmitting, by the one or more computers, the one or more instructions to the robotic device using one or more networks.

[0012]

[0012] In some embodiments, performing one or more treatments to mitigate latent infection in the produce, by the one or more computers, may include generating, by the one or more computers, one or more instructions that, when processed by the robotic device, cause the robotic device to (i) navigate to a location associated with the plant and (ii) initiate a harvesting operation that causes the robotic device to harvest plant product from the plant and one or more other plants of similar origin, and transmitting, by the one or more computers, the one or more instructions to the robotic device using one or more networks.

[0013] In some embodiments, the acquired data may be generated based on output data generated by a nucleic acid sequencer based on sequencing (i) a bark sample extracted from the plant or (ii) a plant product sample extracted from the plant's plant products. Sometimes, the acquired data may be generated based on output data generated by a nucleic acid sequencer based on sequencing (i) a bark sample extracted from the plant or (ii) a plant product sample extracted from the plant's plant products. Furthermore, the data describing the expression levels of one or more infection biomarkers of the plant may include a list of one or more variants. The one or more variants may describe differences between the plant's read sequence and a reference genome of a healthy plant. In some embodiments, the machine learning model may include one or more of a binary logistic regression model, a logistic model tree, a random forest classifier, L2 regularization, partial least squares, or one or more neural networks. Sometimes, the plant product does not contain any visible signs of infection.

[0014]

[0014] In addition to the embodiments in the accompanying claims and the embodiments described above, the following numbered embodiments are also innovative.

[0015]

[0015] Embodiment 1 is a method for identifying pre-harvest latent infection in a plant, comprising: obtaining, by a processor, data describing expression levels of one or more infection biomarkers present in the plant, the infection biomarkers indicative of a likelihood of infection in the plant, the data indicative of differences between a read sequence of the plant and a reference genome of a healthy plant of the same species as the plant; and selecting, by the processor, one or more machine learning models based on the obtained data, the one or more machine learning models having been previously trained using data correlating other data with the identified one or more infection biomarkers of one or more other plants to generate an output indicative of the likelihood of the plant developing a latent infection before harvest, the one or more machine learning models being trained using a process that inputs, from a training dataset, (i) non-invasive measurements of the one or more other plants, (i a processor generating an output indicative of the likelihood that the plant will develop a latent infection before harvest based on applying the one or more machine learning models to the data; a processor determining, based on the output exceeding a predetermined threshold range, that the plant has a latent infection; a processor determining, based on the output exceeding a predetermined threshold range, one or more treatments to mitigate the latent infection in the plant; and a processor outputting, based on the output exceeding a predetermined threshold range, an indication that the plant has a latent infection and the one or more determined treatments.

[0016]

[0016] Embodiment 2 is a method as described in embodiment 1, in which the processor identifying one or more treatments to reduce latent infection in the plant includes identifying, for the plant and one or more other plants of similar origin, at least one of: (i) a harvest schedule change to earlier in the growing season, (ii) instructions to apply a predetermined amount of insecticide before harvest, and (iii) a time when the plant and one or more plants of similar origin are approaching the end of their respective healthy plant productive lives, and one or more machine learning models are pre-trained to identify (i) to (iii) using data from a training dataset.

[0017]

[0017] Embodiment 3 is a method according to embodiment 1 or 2, wherein the one or more other plants of similar origin include plants in the same zone as the plant.

[0018]

[0018] Embodiment 4 is a method described in any one of embodiments 1 to 3, which includes generating an alert message that, when processed by a user device, causes the processor to output an indication that the plant has a latent infection and one or more determined treatments, causing the user device to output an alert notifying a user of the user device to perform one or more of the determined treatments, and sending the generated alert message to the user device.

[0019]

[0019] Embodiment 5 is a method described in any one of embodiments 1 to 4, wherein determining one or more treatments to reduce latent infection in the plant by the processor includes determining, based on the generated output, that an antimicrobial treatment should be prescribed to one or more other plants of similar origin before harvest.

[0020]

[0020] Embodiment 6 is a method described in any one of embodiments 1 to 5, wherein the processor, when determining one or more treatments to reduce latent infection in the plant, generates instructions to the irrigation controller to automatically spray a liquid containing an antibacterial treatment when processed by the irrigation controller, and sends the instructions to the irrigation controller.

[0021]

[0021] Embodiment 7 is a method described in any one of embodiments 1 to 6, wherein when the processor determines one or more treatments to reduce latent infection in the plant, the determination is processed by a robotic device, and the processor generates instructions to the robotic device to (i) navigate to the location of the plant and (ii) spray an antimicrobial treatment on the plant and one or more other plants of similar origin, and transmits the instructions to the robotic device.

[0022]

[0022] Embodiment 8 is a method described in any one of embodiments 1 to 7, wherein the processor, when processing the robotic device to determine one or more treatments to reduce latent infection in the plant, generates instructions to the robotic device to (i) navigate to the location of the plant and (ii) harvest one or more plant products from the plant and one or more other plants of similar origin before the expected harvest timeframe, and transmits the instructions to the robotic device.

[0023]

[0023] Embodiment 9 is a method according to any one of embodiments 1 to 8, wherein the obtained data is generated by a nucleic acid sequencer based on sequencing of plant products extracted from the plant.

[0024]

[0024] Embodiment 10 is a method according to any one of embodiments 1 to 9, wherein the plant product comprises at least one of bark, leaf, flower, and fruit samples.

[0025]

[0025] Embodiment 11 is a method according to any one of embodiments 1 to 10, wherein the plant product is sampled non-destructively from the plant.

[0026]

[0026] Embodiment 12 is a method according to any one of embodiments 1 to 11, wherein the plant products include volatile components emitted from the plant.

[0027]

[0027] Embodiment 13 is a method described in any one of embodiments 1 to 12, wherein the one or more machine learning models include at least one of a binary logistic regression model, a logistic model tree, a random forest classifier, L2 regularization, partial least squares, and a convolutional neural network (CNN).

[0028]

[0028] Embodiment 14 is a method according to any one of embodiments 1 to 13, wherein the plant does not contain visible signs of infection.

[0029]

[0029] Embodiment 15 is a method described in any one of embodiments 1 to 14, further comprising: encoding, by a processor, the acquired data into a data structure for input to one or more machine learning models; and providing, by the processor, the encoded data structure as input to one or more machine learning models.

[0030]

[0030] Embodiment 16 is a method according to any one of embodiments 1 to 15, further comprising carrying out, by the processor, one or more of the determined treatments to reduce latent infection in the plant.

[0031]

[0031] Embodiment 17 is a method described in any one of embodiments 1 to 16, wherein the processor selects one or more machine learning models further based on one or more predictive features identified from the acquired data, and the one or more predictive features include at least one of plant growing region, environmental conditions, plant type, plant growth stage, non-invasive measurements of the plant, invasive measurements of the plant, dry matter content, gene expression, plant emitted volatile components, growth zone, and plant read sequence data.

[0032]

[0032] Embodiment 18 is a method described in any one of embodiments 1 to 17, wherein the known plant information includes at least one of growth history information for one or more other plants, length of growing season, soil conditions, precipitation level, amount of sunlight, environmental temperature, growing region, and plant type.

[0033]

[0033] Embodiment 19 is a system for predicting latent infection in plants, comprising one or more processors and one or more computer-readable storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform a method described in any one of embodiments 1 to 18.

[0034] The devices, systems, and techniques described herein may provide one or more of the following advantages. For example, early detection of infected plant products allows for treatment and separation of infected plant products from uninfected plant products. This can reduce the spread of infection and reduce plant product waste, as well as improve customer acceptance of the plant products. By reducing plant product waste and improving customer acceptance, providers and other related users in the supply chain can provide higher quality plant products, thereby enabling providers to differentiate themselves from other providers in the market. The disclosed techniques can be applied early in plant growth and can further provide for determining effective mitigation efforts to avoid the onset of infection and wasted plant products. Early establishment of latent infection onset can provide relevant stakeholders, such as farmers, ample time to determine and apply mitigation efforts. Furthermore, sampling plants in advance before harvesting is not a burden to farmers, even if the plant (e.g., tree) loses some plant product due to sampling, because mitigation efforts can be determined and applied to eliminate or reduce the possibility of losing plant product during the plant's maturation process.

[0035] As another example, the disclosed techniques provide for combining various data and / or characteristics about plant products that can be used to accurately and timely predict latent infections before harvest. Humans are not effective or accurate in collecting and assessing various types of data about plants during the plant's growing season, such as collecting invasive and non-invasive measurements, to predict the onset of latent infections. Thus, the disclosed techniques provide for the collection, synthesis, and analysis of various types of data that may be difficult, error-prone, and / or impossible for humans to perform without the present invention to accurately predict latent infections in plants before harvest.

[0036]

[0036] Furthermore, the present disclosure can reduce the amount of plant product that needs to be destroyed to detect latent infection in a zone of plants of like origin. In some embodiments, the present disclosure can reduce destruction of plant product by sampling only the above-ground portions of the plant, rather than sampling the plant product itself. In that case, the operations of the present disclosure can be performed using samples of the above-ground portions of the plant. In these or other embodiments, the present disclosure can also provide the benefit of preserving plant product by detecting latent infection in the plant and initiating one or more remedial actions to mitigate and potentially eliminate the latent infection before it manifests, thereby making the plant product salable and improving customer satisfaction. Such remedial actions can include, for example, harvesting plant product from the plant and instructing a user or robotic device to harvest plant product from one or more other plants in the zone of plants of like origin. In other embodiments, the harvest schedule can be updated to achieve the same or similar benefit. This can provide one or more economic benefits to users of the present disclosure.

[0037]

[0037] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0038] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1A]

[0038] A conceptual diagram of an example system for detecting latent infection in plants. [Figure 1B]

[0039] 1 is a flow chart of a process for detecting latent infection in plants. [Figure 2]

[0040] A conceptual diagram for predicting the likelihood of infection in plants. [Figure 3]

[0041] FIG. 1 is a conceptual diagram of the training of one or more models to predict infection likelihood in plants. [Figure 4]

[0042] 1 is a flowchart of a process for predicting the likelihood of infection in a plant. [Figure 5]

[0043] FIG. 1 is a system diagram of one or more components that can be used to implement the techniques described herein. [Figure 6A]

[0044] 1 illustrates an example process for collecting plant data for use in conjunction with the disclosed techniques. [Figure 6B]

[0044] An example process for collecting plant data is shown for use in conjunction with the disclosed techniques. [Figure 6C]

[0044] An example process for collecting plant data is shown for use in conjunction with the disclosed techniques. [Figure 7A]

[0045] 7 illustrates a graphical representation of example physical attributes of the plant of FIG. 6 that may be collected at different times. [Figure 7B] FIG. 7 shows a graphical representation of example physical attributes of the plant of FIG. 6 that may be collected at different times. [Figure 8A]

[0046] A biplot of principal component analysis of normalized gene expression across all time periods of data collection for the plants in Figures 6A-6C is shown. [Figure 8B]

[0046] Figures 6A-6C show biplots of principal component analysis of normalized gene expression across all time periods of data collection for the plants. [Figure 8C]

[0046] Figures 6A-6C show biplots of principal component analysis of normalized gene expression across all time periods of data collection for the plants. [Figure 9A]

[0047] Volcano plots of normalized gene expression for the plants in Figures 6A-6C are shown. [Figure 9B]

[0047] Volcano plots of normalized gene expression for the plants in Figures 6A-6C are shown. [Figure 9C]

[0047] Volcano plots of normalized gene expression for the plants in Figures 6A-6C are shown. [Figure 10A]

[0048] Clustered heat maps of normalized gene expression for genes analyzed at each collection time for plants in Figures 6A-6C are shown. [Figure 10B]

[0048] Figures 6A-6C show clustered heat maps of normalized gene expression for genes analyzed at each collection time for plants. [Figure 10C]

[0048] Figures 6A-6C show clustered heat maps of normalized gene expression for genes analyzed at each collection time for plants. [Figure 10D]

[0048] Figures 6A-6C show clustered heat maps of normalized gene expression for genes analyzed at each collection time for plants. [Figure 11]

[0049] FIG. 1 is a block diagram of system components that can be used to implement a system for detecting latent infection. DETAILED DESCRIPTION OF THE INVENTION

[0039]

[0050] Like reference symbols in the various drawings indicate like elements.

[0040] DETAILED DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS

[0051] The present disclosure generally relates to systems, methods, and computer programs for detecting latent infections in plants. Many plants fruit regularly and live productively for decades or even centuries. Throughout this long lifespan, plants can interact with and respond to their surrounding ecosystems, which can contain populations of beneficial and pathogenic microorganisms. Detection of and response to pathogenic microorganisms by plants can be observed through changes in the expression of genes involved in the plant's innate immune system. Thus, the disclosed techniques can provide for preharvest transcriptomic markers of signs of preharvest latent infections, such as stem rot (SER), by profiling the transcriptomes of various plant products from different types of plants, such as fruit from trees. While such latent infections may not become apparent until the fruit matures on the tree, is harvested, and / or ripens, fruits and other plant products can provide markers of infection as early as six months before peak harvest season. Utilizing this information toward the generation of detection assays allows prediction of high and low incidence trees, with ample time to modify fungicide applications, plant product harvest schedules, and / or plant product destination information to mitigate the manifestation of latent infection in various types of plants.

[0041]

[0052] The present disclosure may also provide a trained machine learning model capable of generating output data indicative of the likelihood that a plant has a latent infection based on machine learning model processing of input data representing plant characteristics. In some embodiments, the plant characteristics may include data representing expression levels of one or more biomarkers detected in the plant sample. If the plant is predicted to have a latent infection, the disclosed techniques may provide for determining one or more actions designed to mitigate the detected latent infection and, optionally, performing the one or more actions.

[0042]

[0053] Referring to the figures, FIG. 1A is a conceptual diagram of an example system 100 for detecting latent infection in plants. System 100 can include a user device 110, one or more robotic devices 120, 121, 122, 124, 125, 126, a nucleic acid sequencer 130, a computer 140 (e.g., a computer system), a network 150, one or more robotic device charging stations 160, 162, 164, an irrigation system 170, and / or any combination thereof. System 100 can be implemented with any subset, combination, or arrangement of these components to achieve the functionality described herein for detecting latent infection in plants, plant products, or both. In some embodiments, a nucleic acid sequencer may not be required. Instead, a qPCR machine may be used in place of sequencer 130. Furthermore, in some embodiments, sequencer 130 and computer 140 can be part of the same system.

[0043]

[0054] 1A , in some embodiments, an implementation of system 100 can involve user 105 extracting a sample 122 of plant product from plant 202-1 in zone 2, a sample of any above-ground portion of plant 202-1, or both. Sampling can occur prior to harvest, which can sometimes be six months or more before the plant product of the sampled plant is collected and harvested. Thus, the disclosed techniques can be performed while the sampled plants are undergoing their respective maturation processes and before they are harvested (e.g., before they reach full maturity). As a result, the disclosed techniques can be performed to predict the likelihood of infection of the sampled plant and other plants of similar origin prior to harvest, early enough so that efforts can be made to mitigate the potential onset of latent infection.

[0044]

[0055] Sample 122 can include genetic data of the plant products of the sampled plant. In some embodiments, sample 122 can be derived from fruit, leaves, bark, plant tissue, or other plant products of plant 202-1. Sample 122 can also include other information about the plant products of sampled plant 202-1. For example, sample 122 can include hardness and / or penetrometer measurements, volatile metabolites, and other compounds that are non-destructively and / or destructively detected, etc.

[0045]

[0056] In some embodiments, sample 122 can be extracted by one or more drones 120 or other types of collection devices or mechanisms. Examples of collection devices can include stationary probe devices and / or automated stationary samples. For example, a person can collect sample 122 using one or more types of collection devices described herein. In some embodiments, sample 122 can be pre-processed to prepare sample 122 for input into nucleic acid sequencer 130 for sequencing. Pre-processing can include, for example, processing sample 122 to obtain a nucleic acid sample from sample 122, applying a barcode to the sample, or any combination thereof. In some embodiments, sample 122 can include only the above-ground portion of a plant, such as plant 202-1. In such embodiments, the above-ground portion of the plant can be analyzed using techniques of the present disclosure to identify expression levels of one or more infection biomarkers in nucleic acids derived from sample 122 of plant 202-1. A prediction can then be made as to whether a plant product (shown in FIG. 1A using black dots on the plant, e.g., 102-1 through n02-z, shown in zones 1 through n, where n is any positive integer, x is any positive integer, y is any positive integer, and z is any positive integer) has developed or is likely to develop a latent infection based on treatment of only the above-ground portions of a particular plant, e.g., plant 202-1. In some embodiments, samples from both the above-ground portions of the plant and the plant product can be used.

[0046]

[0057] The nucleic acid sequencer 130 can process the received nucleic acids derived from the sample 122 and generate output data, which can be referred to as read sequence data 132, that represents the order of nucleotides within the nucleic acid (e.g., DNA) of the sample 122. In some embodiments, the read sequence data 132 can include nucleotides within the nucleic acid of the sample 122 using one or more nucleotide bases including guanine (G), cytosine (C), adenine (A), and thymine (T), in any combination. In some embodiments, the read sequence data 132 can be generated by transcribing the read sequence data generated by the nucleic acid sequencer 130 into cDNA to represent the RNA of the sample 122. In such embodiments, the read sequence data can be represented using bases including G, C, A, and uracil (U), in any combination.

[0047]

[0058] The read sequence data 132 can be provided to the computer 140 using one or more networks 150. The one or more networks 150 can include a wired Ethernet network, an optical network, a wireless network, a WAN, a LAN, a cellular network, a Wi-Fi network, the Internet, or any combination thereof. In some embodiments, the nucleic acid sequencer 130 can be directly connected to an application server using an Ethernet cable, a USB-C cable, or any other form of direct connection that facilitates data communication between the nucleic acid sequencer 130 and the computer 140. The computer 140 can provide the read sequence data 132 as input to a biomarker expression engine 151. The biomarker expression engine 151 can be part of the computer 140, as further described with reference to FIG. 5.

[0048]

[0059] In some embodiments, the data output by the nucleic acid sequencer 130 can be epigenetic data. The epigenetic data can include changes to the nucleic acids of the sample 122 (e.g., chemical modifications, changes to the structure of the DNA, etc.) that do not result in changes to the base calls (or nucleotides) of the read sequence data 132 generated by the nucleic acid sequencer 130 based on sequencing of the sample 122 by the nucleic acid sequencer 130. In some embodiments, such epigenetic data can be generated, for example, by a long-read sequencer. In some embodiments, if epigenetic data is not used, other nucleic acid sequencers can be used, including, but not limited to, short-read sequencers. These are merely examples of sequences that can be used to detect epigenetic data, but any other sequencer can also be used to generate the epigenetic data described herein.

[0049]

[0060] The biomarker expression engine 151 can be configured to receive the read sequence data 132 and identify expression levels of one or more infection biomarkers, levels of housekeeping biomarkers, or both, as described in International Publication No. WO 2021 / 092251 A1, PREDICTION OF INFECTION IN PLANT PRODUCTS, which is incorporated herein by reference in its entirety. The biomarker expression engine 151 can generate output data 151a based on processing the read sequence data 132 that represents expression levels of one or more infection biomarkers, levels of housekeeping biomarkers, or both. In some embodiments, the expression level of each biomarker can include the number of occurrences of each gene in the read sequence device 132. The number of occurrences can be expressed as copy numbers.

[0050]

[0061] In this example, nucleic acid sequencing device 130 is used to sequence sample 122 and generate read sequence data 132. Biomarker expression engine 151 can then determine expression levels of one or more infection biomarkers, levels of housekeeping biomarkers, or both. In such embodiments, this may include determining the number of occurrences of genes, presented as copy numbers, by biomarker expression engine 151. The disclosure is not so limited, and in some embodiments, nucleic acid sequencer 130 may not be required.

[0051]

[0062] For example, in some embodiments, a quantitative polymerase chain reaction (qPCR) machine and a microarray can be used to determine the occurrence of genes in sample 122. In such embodiments, primers can be configured in the qPCR machine to identify specific gene sequences within sample 122. The qPCR machine can be configured to amplify specific gene sequences identified by an assay for target genes. As a result, because the output of qPCR is a count of only a subset of the specific genes identified by the assay, qPCR machines can be faster than sequencing a complete genome or transcriptome using one or more types of sequencing devices. The counts output by the qPCR machine can be provided to computer 140, biomarker expression engine 151, and, in some embodiments, input generation engine 152 without first being provided to biomarker expression engine 151. The qPCR machine can be implemented in user device 110, computer 140, or any other computer, device, and / or system that can be configured to communicate with computer 140 using network 150.

[0052]

[0063] The input generation engine 152 can be part of the computer 140 and can receive output data 151a representing expression levels of one or more infection biomarkers, levels of housekeeping biomarkers, or both. Based on the received data 151a, the input generation engine 152 can generate an input data structure for input to the latent infection prediction engine 153. In some embodiments, this can include encoding the data representing the expression levels of the infection biomarkers, the housekeeping biomarkers, or both into an array data structure. For example, the encoded array data structure can include a binary string representing the levels of one or more infection biomarkers, the housekeeping biomarkers, or both. The input generation engine 152 can provide the encoded input data structure 152a as input to the latent infection prediction engine 153.

[0053]

[0064] Latent infection prediction engine 153 can be part of computer 140 and can receive input data structure 152a. Latent infection prediction engine 153 can be or otherwise include one or more machine learning models trained using the training process described in International Publication No. WO 2021 / 092251 A1, PREDICTION OF INFECTION IN PLANT PRODUCTS, which is incorporated herein by reference in its entirety. The one or more machine learning models can generate output data based on processing of input data structure 152a, which encodes data representing expression levels of one or more infection biomarkers, housekeeping biomarkers, or both. The one or more machine learning models may also be provided with additional input data, including, but not limited to, destructive (e.g., invasive) measurements taken from the sampled plant 202-1, non-destructive (e.g., non-invasive) measurements taken from the sampled plant 202-1, data about environmental conditions (e.g., weather patterns, rainfall levels, number of sunny days, other growing conditions), and / or data about the geographic region in which the sampled plant 202-1 is grown and harvested. Using the additional input data and / or input data structure 152a, the one or more machine learning models may predict the likelihood of a latent infection developing in the sampled plant 202-1 and other plants of similar origin in zone 2 or other zones within the growing area (e.g., field). The prediction may be in the form of output data.

[0054]

[0065] Thus, the output data can represent a likelihood 153a that the plant 202-1 from which the sample 122 was derived has, is developing, and / or may be developing one or more latent infections. The latent infection prediction engine 153 can perform the functions described by the predictive models described in WO 2021 / 092251 A1, PREDICTION OF INFECTION IN PLANT PRODUCTS, which is incorporated herein by reference in its entirety. The output data 153a generated by the latent infection prediction engine 153 can be provided to the latent infection detection module 154. In some embodiments, the latent infection prediction engine 153 can include one or more of a binary logistic regression model, a logistic model tree, a random forest classifier, L2 regularization, or one or more neural networks, and / or a partial least squares model.

[0055]

[0066] The latent infection detection module 154 can be part of the computer 140 and can evaluate the output data 153a to determine whether mitigation efforts may be necessary. By way of example, the output data 153a can be measured against one or more predetermined thresholds to determine whether the output data 153a indicates that the plant 202-1 from which the sample 122 was derived has developed, is developing, and / or may be developing one or more latent infections. In some embodiments, the latent infection detection module 154 can evaluate the output data 153a in the same or similar manner as the output data of the prediction module described in International Publication No. 2021 / 092251 A1, PREDICTION OF INFECTION IN PLANT PRODUCTS, which is incorporated herein by reference in its entirety. If the latent infection detection module 154 determines that the output data 153a does not indicate the presence of a latent infection, for example, techniques described herein can be transmitted at 156. Alternatively, if the latent infection detection module 154 determines that the output data 153a indicates the presence of a latent infection, the latent infection detection module 154 may provide instructions 154a to the mitigation engine 155 that determine, suggest, and / or perform one or more mitigation actions.

[0056]

[0067] The mitigation engine 155 can be part of the computer 140 and can initiate the determination, suggestion, and / or execution (e.g., implementation) of one or more mitigation actions. These actions can take a variety of forms. For example, in some embodiments, the mitigation engine 155 can determine that a harvest schedule for the sampled plant 202-1 and / or one or more other plants of similar origin can be adjusted based on the instructions 154a, the output data 153a, the expression level of one or more infection biomarkers, the expression level of one or more housekeeping biomarkers, or a combination thereof. In such embodiments, the mitigation engine 155 can access a database storing harvest schedules for the sampled plant 202-1, its plant products, plants and / or plant products of similar origin, or a combination thereof, and adjust the harvest schedule based on the output data 153a. In such embodiments, the output 155a of the mitigation engine 155 can include data written to the database to adjust the harvest schedule. In some embodiments, adjusting the harvest schedule can include adjusting post-harvest transportation routes for plant products harvested from sampled plant 202-1 and / or plants of similar origin. For example, plant products from sampled plant 202-1 and / or other plants of similar origin with a high predicted likelihood of latent infection (e.g., where the predicted likelihood of latent infection exceeds some predetermined threshold or range) can be scheduled for delivery to a geographic area closer to the location of harvest of the plant products, compared to other plants with a low predicted likelihood of latent infection or no latent infection, which can remain fresh longer and survive further distribution. In some embodiments, adjustments to the harvest schedule can be determined by mitigation engine 155 and transmitted to user device 110 or a computing device of another relevant stakeholder. The relevant stakeholder can then review the adjustments and determine whether to implement such adjustments. The adjustments can be implemented by the relevant stakeholder and / or semi-automatically by mitigation engine 155 or one or more other components described throughout this disclosure.

[0057]

[0068] For example, in some embodiments, the mitigation engine 155 can determine that a plant, e.g., 202-1, a plant product, a plant of similar origin, or a zone of plant products, should be harvested immediately to avoid product-based waste based on the instructions 154a, the output data 153a, the expression level of one or more infection biomarkers, the expression level of one or more housekeeping biomarkers, or a combination thereof. In such embodiments, the mitigation engine 155 can generate an alert message that, when processed by the user device, can cause the user device to output an alert notifying the user 105 of the user device 110 that the plant 202-1, the plant product, and / or one or more other plants or plant products of similar origin should be harvested. In such cases, the output 155a can include the generated alert. The computer 140 can then transmit the generated alert to the user device 110.

[0058]

[0069] Mitigation engine 155 may also determine one or more other actions that can be taken by user 105, another relevant stakeholder, and / or another device / robot (e.g., drone 120). For example, as described further below, mitigation engine 155 may determine that one or more products can be applied to plant 202-1, the plant product, and / or one or more other similarly aged plants or plant products to preserve such plants and / or products and prevent latent infection during the pre-harvest growing season.

[0059]

[0070] Abatement engine 155 may also cause plant 202-1 and other plants of similar origin within the same zone to be harvested in several ways. For example, in some embodiments, abatement engine 155 may deploy one or more robotic devices (e.g., drone 120) to initiate harvesting of plant products from plant 202-1 and other locally-sourced plants within the same zone, such as zone 2 of FIG. 1 . In such embodiments, abatement engine 155 may generate one or more instructions that, when processed by the robotic device, cause the robotic device to (i) navigate to a location associated with the plant and (ii) initiate a harvesting operation with the robotic device. Initiating harvesting with the robotic device may include causing the robotic device to begin performing one or more operations associated with harvesting at least one plant product from the plant and one or more other plants of similar origin within zone 2. In such embodiments, output 155a of abatement engine 155 may include instructions to cause the robotic device to navigate and initiate harvesting. In some implementations, mitigation engine 155 can use location data from user device 110 of user 105 or drone 120 that sampled plant 202-1 to identify the location of that plant and one or more other plants of similar origin within zone 2. Mitigation engine 155 can use network 150 to send one or more commands to robotic devices 124, 125, and / or 126 and / or robotic device 120, or any other type of automated harvesting and collecting device. While shown as quadcopter robotic devices, the disclosure is not so limited. Instead, robotic devices 124, 125, and / or 126 can include any robotic device, including a robotic device that moves as a rover on the ground, a robotic device that navigates via aerial flight, or any combination thereof.

[0060]

[0071] As yet another example, in some embodiments, mitigation engine 155 can determine that an antimicrobial treatment should be prescribed to one or more other plants of similar origin based on instructions 154a, output data 153a, expression levels of one or more infection biomarkers, expression levels of one or more housekeeping biomarkers, or a combination thereof. In some embodiments, the antimicrobial treatment can be formulated as described further below (see, e.g., Table 1) and in International Publication No. 2021 / 092251 A1, PREDICTION OF INFECTION IN PLANT PRODUCTS, which is incorporated herein by reference in its entirety. Mitigation engine 155 can cause the antimicrobial treatment to be administered in many different ways.

[0061]

[0072] For example, in some embodiments, mitigation engine 155 can generate one or more instructions that can be sent using network 150 to remote irrigation controller 171, causing irrigation controller 171 to activate irrigation system 170 and deploy an antimicrobial treatment. For example, irrigation controller 171 or mitigation engine 155 can instruct valve 172 to close and valve 174 to open. Irrigation controller 171 can then spray plant 202-1 and one or more locally grown plants in the same zone, such as zone 2, with the antimicrobial treatment to mitigate the detected or otherwise predicted latent infection. In other embodiments, mitigation engine 155 can determine, based on instructions 154a, output data 153a, the expression level of one or more infection biomarkers, the expression level of one or more housekeeping biomarkers, or a combination thereof, that plant 202-1 and other plants of similar origin in zone 2 need only be watered and need not be prescribed an antimicrobial solution treatment. In such an embodiment, the mitigation engine 155 can instruct the irrigation controller 171 (e.g., valves 174, 172) to be reconfigured to only water from a water source 178, which may be a water tank, a public water source, a well source, or a combination thereof.

[0062]

[0073] The abatement engine 155 may also spray the antimicrobial solution in other ways. For example, in some embodiments, the abatement engine 155 may command one or more robotic devices to initiate antimicrobial operations. In such embodiments, each robotic device, such as robotic device 120, may include a tank of antimicrobial solution and spraying devices 121 a, 122 a. In such embodiments, the robotic device 120 may receive instructions generated and transmitted by the abatement engine 155, process the instructions, navigate to a location associated with the plant and other plants of similar origin in the same zone, and administer the antimicrobial solution to the plant and other plants of similar origin in the same zone using sprayers 121 a, 122 a.

[0063]

[0074] The above examples may detect one or more infection biomarkers, one or more housekeeping biomarkers, or both, by analyzing read sequences to determine gene counts and / or using a qPCR machine to determine gene counts. The disclosure is not so limited. Instead, system 100 may be used to identify and count other types of biomarkers for input into prediction engine 153. For example, in some embodiments, system 100 may be used to detect and count genes that may be indicative of drought stress. Examples of types of genes that may be detected and counted as indicative of drought stress are described further below, such as in Table 1. Expression levels of these drought stress biomarkers may be analyzed using system 100 in the same or similar manner as expression levels of infection biomarkers, housekeeping biomarkers, or both. When system 100 detects a threshold expression level of drought stress, system 100 may use mitigation engine 155 to instruct irrigation system 171 to water the particular plant and / or one or more plants of similar origin in the same zone.

[0064]

[0075] 1B is a flowchart of a process 180 for detecting latent infection in a plant. Process 180 may be performed by computer system 140 shown and described throughout this disclosure (e.g., see FIGS. 1A and 5). Process 180 and / or one or more blocks of process 180 may be performed by other computing systems, devices, networks of computers / devices, cloud-based systems, and / or cloud-based services. For illustrative purposes, process 180 is described from the perspective of a computer system.

[0065]

[0076] Referring to process 180 of Figure 1B, a computer system can receive data describing the expression levels of infection biomarkers in plants within a particular zone (182). As described with reference to Figure 1A and throughout this specification, the data can be samples of one or more plants collected by a user (e.g., a human worker) in the field. Samples can also be collected by a device such as a drone or other automated collection device.

[0066]

[0077] The computer system can encode the data for input to one or more models (184). The data can be encoded into a data structure that is input to one or more machine learning models. See the description of the biomarker expression engine 151 and input generation engine 152 in FIG. 1A.

[0067]

[0078] The computer system can select at 186 at least one model trained to predict the likelihood of latent infection in plants. As further described below, different models can be selected for application based on plant type, plant developmental stage, and / or one or more environmental conditions. Models can be stratified in application so that they are applied sequentially. Models can also be applied in parallel.

[0068]

[0079] At 188, the computer system can input the coded data into the selected model. The computer system can then receive output from the model at 190. The output can include a value indicating the likelihood that the plant has or will have a latent infection. The value can be numeric (e.g., on a scale of 1 to 100, where 1 is the least likely to develop a latent infection and 100 is the most likely to develop a latent infection). The value can also be a string value and / or a Boolean value.

[0069]

[0080] At 192, the computer system can determine whether the predicted likelihood of latent infection for the plant (e.g., output from the model) exceeds a predetermined threshold range. If the predicted likelihood exceeds a predetermined threshold range, the computer system can determine one or more mitigation actions to mitigate latent infection for the plant and other plants of similar origin within the particular zone (194). In other words, the computer system can determine that the plant has a sufficiently high likelihood of developing latent infection that action should be taken early during the plant's growth to avoid wasting the plant at harvest and / or causing loss of the plant or its plant products.

[0070]

[0081] In some embodiments, the computer system can provide one or more mitigation actions as suggestions to the user device of the associated user. The user can then select which mitigation action to implement. The mitigation action can be performed manually by the user or another user. The mitigation action can also be performed automatically by the computer system. The mitigation action can also be performed semi-automatically by the computer system when the user provides input to the computer system indicating a selection of such mitigation action.

[0071]

[0082] If the predicted likelihood of latent infection for the plant does not exceed a predetermined threshold range, process 180 may terminate. In other words, the computer system may determine that a particular plant does not have a high enough likelihood of developing latent infection at this time in the plant's growing season to trigger any type of mitigation action.

[0072]

[0083] FIG. 2 is a conceptual diagram of predicting the likelihood of infection of a plant. The computer system 140, the user device 110, the plant data store 230, and the model data store 240 can communicate (e.g., wired, wireless) via a network 150. The sampler 105 can collect pre-harvest data from plants in a field / farm zone in step A. Pre-harvest collection can occur as early as six months before harvest. Pre-harvest collection can also occur at one or more time intervals during the plant's growing season. For example, pre-harvest collection and subsequent infection likelihood prediction can occur every two months until harvest. As another example, pre-harvest collection and infection likelihood prediction can be performed six months before harvest. At this point, mitigation actions can be determined and applied to the plant and / or plants of similar origin. At this point, it can also be determined whether additional pre-harvest collections and / or predictions may be necessary for the remainder of the growing season before harvest.

[0073]

[0084] Pre-harvest collection may include any of the techniques described herein (see, e.g., FIG. 1A). For example, a sampler 105 (e.g., a user) may walk around a plant and collect a subset of plant products from the plant. In some embodiments, the sampler 105 may collect four plant products from the plant being sampled and tested for likelihood of latent infection. The sampler 105 may collect any other quantity of plant products from the plant that may be optimal for predicting likelihood of latent infection. The quantity collected may depend on the type of plant, the particular area in which the plant is growing, and / or one or more other environmental conditions. Thus, collection of plant products may be used to determine whether a particular plant (e.g., tree) in a particular zone of a field / plot is at risk of being infected.

[0074]

[0085] As another example, collection can include detecting plant volatile metabolites using non-destructive techniques. Such data can also provide insight into how the plant responds to environmental conditions, which can be another indicator of whether the plant has a latent infection and / or is likely to develop a latent infection. The sampler 105 can hold a volatile collection device (e.g., an NIR device, an electronic nose, or other instrument) in proximity to the plant product to collect emitted volatiles. Other field-deployable instruments, including, but not limited to, electronic noses and / or NIR devices on semi-automated or automated robotic devices, can also be used to collect and measure volatiles emitted from the plant product.

[0075]

[0086] Any combination of pre-harvest collection can be performed. Pre-harvest collection can also depend on which model is used to predict the likelihood of latent infection of the plant products, which can further depend on the plant species, growing region, and / or environmental conditions. For example, in some embodiments, volatiles can be collected from the plant products, and a quantity of the plant products can be harvested from the plant and transported to a laboratory or facility for further analysis.

[0076]

[0087] Once the pre-harvest data is collected (Step A), the data can be transmitted to computer system 140 (Step B). As described herein, the data can include sequencing data, which can be transmitted to a nucleic acid sequencer or other sequencing device (see, e.g., FIG. 1A). The data can also include volatile measurements. In some embodiments, the data can include other non-destructively and / or destructively obtained measurements of plant products collected and / or harvested from the plants.

[0077]

[0088] The computer system 140 may receive plant information from the plant data store 230 (step C). The information may be specific to the type of plant being collected and tested for infection. The information may also be specific to the growing region of the collected plant and / or plant product. The information may also include current environmental conditions and / or future environmental conditions that may occur during the plant's growing season prior to harvest. Step C may be performed at any time before, during, and / or after one or more steps described in FIG. 2. For example, step C may be performed simultaneously with the performance of step D.

[0078]

[0089] The computer system 140 can process the data in step D. Processing the data can be performed as described in FIG. 1A. Processing the data can include preparing the collected harvest data to provide as input to one or more models, by encoding. Processing the data can also include preparing plant information for input to one or more models.

[0079]

[0090] The computer system may also receive models from the model data store 240 based on the plant information (step E). The computer system may determine how many models to apply to predict the likelihood of infection for the plant. The computer system may also determine which model to select based on plant information such as plant type and / or environmental conditions.

[0080]

[0091] The computer system can apply the selected model in step F. For example, the computer system can provide the collected pre-harvest data and / or plant information as input to the model (e.g., in a coding structure).

[0081]

[0092] Thus, the computer system can predict the likelihood of infection of the plant in step G. For example, the model can provide an output in the form of a value indicative of the likelihood of infection of the plant. The computer system can then determine whether the value exceeds some predetermined threshold range. Exceeding the threshold range can indicate a high likelihood that the plant will develop a plant infection. If the value does not exceed the threshold range, the likelihood that the plant will develop an infection is low.

[0082]

[0093] Based on the prediction in step G, computer system 140 may optionally determine a mitigating action in step H. For example, see Table 1 below for further discussion of the types of mitigating actions that may be determined, suggested, and / or implemented (e.g., manually, automatically, and / or semi-automatically).

[0083]

[0094] The computer system 140 can transmit the infection likelihood prediction and / or the determined mitigation actions to the user device 110 (step I). The user device 110 can output the prediction and / or mitigation actions to be presented to the associated user (step J). For example, the user device 110 can be a mobile phone, computer, laptop, and / or tablet with a display screen. The prediction and / or mitigation actions can be presented on the display screen in a graphical user interface (GUI). The associated user can view such information and take action in response to viewing the information. For example, the associated user can select one or more actions to perform from the presented mitigation actions.

[0084]

[0095] User device 110 and / or computer system 140 then, optionally, perform mitigating actions (step K). In some implementations, the associated user may select the mitigating actions to perform at user device 110. A notification of this selection may be sent to computer system 140. The computer system may then send a notification and instructions to one or more robotic devices to perform the user-selected mitigating actions. In some implementations, user device 110 may send instructions to one or more robotic devices if a mitigating action should be performed. Sometimes, instructions may be sent to one or more other user devices, notifying users of such data that one or more mitigating actions should be performed. Furthermore, in some implementations, user device 110 and / or computer system 140 may perform the mitigating actions in step K without receiving a selection and / or approval from the associated user at user device 110.

[0085] [Table 1]

[0086] [Table 2]

[0087] [Table 3]

[0088] [Table 4]

[0089] [Table 5]

[0090]

[0096] Exemplary gene categories in Table 1 are: PAL: phenylpropanoid ammonia lyase; CHS: chalcone synthase; PER: peroxidase; HSP: heat shock protein; WRKY: WRKY transcription factor; NAC: NAC transcription factor; ERF: ethylene response factor; ETR: ethylene transcription factor; EIN: ethylene-insensitive protein; ABA: abscisic acid; NCED: 9-cis-epoxycarotenoid dioxygenase; ZEP: zeaxanthin epoxidase; ABI: abscisic acid-insensitive transcription factor; JMT: jasmonate methyltransferase; SAMT: salicylate methyltransferase; and PIN: PIN-FORMED protein. One or more other gene categories can also be used in conjunction with the disclosed techniques.

[0091]

[0097] 3 is a conceptual diagram of training one or more models to predict the likelihood of infection in plants. The models can be trained using any of a variety of techniques, including, for example, any of the techniques described in this document and any of the techniques described in WO 2021 / 092251 A1, PREDICTION OF INFECTION IN PLANT PRODUCTS, which is incorporated herein by reference in its entirety and filed by one or more of the same inventors.

[0092]

[0098] Although the training of the one or more models is shown as being performed by computer system 140, the training may also be performed by one or more other computer systems, devices, networks, cloud-based systems, and / or cloud-based services.

[0093]

[0099] 3, computer system 140 may receive training data 302 from one or more sources (Step A). ​​The training data 302 (or portions thereof) may be received from a user device of an associated user, a volatile collection device, a robotic device, an NIR device, one or more other computing devices, and / or one or more data stores. In some implementations, computer system 140 may select one or more types of training data 302 to receive.

[0094]

[0100] The training data 302 can include one or more of non-invasive measurements 304, invasive measurements 306, known plant information 308, and positive infection identification information 310. The non-invasive measurements 304 can be collected using devices such as NIR devices, electronic noses, and other volatile component collection devices. The non-invasive measurements 304 can include, for example, data indicative of the plant's volatile composition as it is emitted by the plant. The non-invasive measurements 304 can be measured from plant products on the plant in the field (e.g., a farm). Thus, the plant product does not need to be harvested from the plant. In some embodiments, the non-invasive measurements 304 are measured from the plant product when it is harvested / picked from the plant and can be used for further analysis in a laboratory or other analytical setting. The non-invasive measurements 304 can also include a nucleic acid sample of the plant product being analyzed.

[0095]

[0101] Invasive measurements 306 can be collected using devices such as penetrometers and hardness meters. Invasive measurements 306 can be collected by an associated user from plant products that are still part of the plant and / or that have been harvested from the plant and used for further analysis. For example, an associated user can harvest a quantity of plant product (e.g., fruit) from a plant (e.g., a tree) and bring that quantity of plant product back to a laboratory or other analytical setting. In the laboratory, the associated user can capture measurements of the plant product, for example, by squeezing the plant product to identify a firmness measurement, harvesting some portion of the plant product's skin, measuring skin thickness by puncturing the plant product, etc. Invasive measurements can be used to determine dry matter percentage, fatty acid percentage, sugar content / sugar profiling, and / or acidity / acid profiling. Furthermore, such information can be determined using various other types of analytical methods, including, but not limited to, volatile analysis and metabolomics.

[0096]

[0102] Known plant information 308 may include growth history information for a particular plant, such as the length of the growing season, and environmental conditions that support the growth, maturation, and / or ripening of a particular plant. Environmental conditions may include rainfall levels, amount of sunlight, and / or temperatures that are favorable for the growth, maturation, and ripening of a particular plant.

[0097]

[0103] Positive infection identification 310 may indicate when a particular plant has been identified as having a latent infection or otherwise developing a latent infection. For example, identification 310 may indicate levels and expression of biomarkers that are positively associated with one or more types of latent infection in a particular plant. These associations may be made by an associated user. These associations may also be made by a computer, such as computer system 140. Positive infection identification 310 may also indicate a combination of biomarker levels or expression, environmental conditions, and one or more other information about the plants that are positively associated with one or more types of latent infection in a particular plant.

[0098]

[0104] Once the computer system 140 receives the training data (Step A), the computer system 140 can correlate the training data based on the predictive features (Step B). In other words, the computer system 140 can associate the non-invasive measurements 304, the invasive measurements 306, the known plant information 308, and / or the positive infection identification 310. The association can indicate the likelihood of a latent infection for a particular plant. For example, the combination of specific volatiles emitted by a particular plant product and unusual rainfall patterns and temperatures during the growing season can indicate the onset of a latent infection for that particular plant and plants of similar origin during the growing season before harvest. This combination of information can be correlated with the positive infection identification 310 for a particular plant.

[0099]

[0105] As mentioned, the training data can be correlated based on predictive features. Predictive features can include, but are not limited to, plant type, growing region, growing stage, ripening conditions, ripening stage, environmental conditions, etc. Thus, predictive features can represent different characteristics of a particular plant and plants of similar origin during the growing, maturation, and / or ripening period that may affect the onset of latent infection. In some embodiments, a model can be generated for each predictive feature. For example, one model can be generated to predict latent infection in avocados based on where the avocados were grown. As another example, a separate model can be generated to predict latent infection in a particular type of avocado (e.g., a particular type of latent infection, such as SER or any type of latent infection) based on volatiles emitted by such avocados during the growing stage. A separate model can also be generated to predict latent infection in a particular type of produce based on the timing of pre-harvest collection of plant gene expression. For example, one model can be used to predict latent infection in avocados during the first 0-3 months of growth, and another model can be used to predict latent infection in avocados during the first 3-6 months of growth. This can be beneficial because gene expression can change based on the age of the plant product and therefore provide a different indication of latent infection based on whether gene expression is collected and analyzed during the first 0-3 months of growth or the first 3-6 months of growth. One or more other time intervals can also be used in generating the model.

[0100]

[0106] In some embodiments, one or more of the predictive features can be combined to generate a model. For example, one model can be generated to predict latent infection in avocados based on a combination of volatiles emitted by avocados and environmental conditions. As another example, another model can be generated to predict latent infection in limes based on a combination of environmental conditions (e.g., temperature, rainfall), the growth stage of the lime, and the growing region. Other combinations of one or more of the predictive features are also possible. As yet another example, a model can be generated to predict latent infection in different types of plants based on each plant's type, growing region, emitted volatiles, and environmental conditions. Thus, the same model can be used to predict latent infection in different plants, including, but not limited to, avocados, limes, lemons, apples, citrus fruits, etc.

[0101]

[0107] Further, in some embodiments, a model can be generated for each growing region. Additional models can then be generated for each region based on age, growth stage, environmental conditions, and other information. Thus, multiple models within multiple models can be created to provide more accurate predictions of latent infection in plants early, before harvest time. Thus, mitigation actions can be determined and implemented early enough before harvest to reduce or otherwise eliminate wastage of plant systems at harvest time.

[0102]

[0108] The computer system 140 can train a model to predict the likelihood of infection using the correlated data (Step C). For example, the correlated data can be provided as input to the model. The model can generate an output as a value having the likelihood of infection for a particular plant associated with the input, correlated data. The computer system 140 can compare the model output with the positive infection identifications 310 to further refine and improve the accuracy of the predictive decisions made by the model. Once a desired level of accuracy is reached, the computer system 140 can output the model (Step D). The computer system 140 can also constantly improve the model in a feedback loop as the model is deployed and used during runtime.

[0103]

[0109] Overall, several biomarkers of the plant product can determine the model that is generated and / or should be generated during run-time. Biomarkers collected at different ages over different growing seasons can be used to validate the accuracy of the model in predicting the likelihood of developing latent infection before harvest. Invasive and / or non-invasive collection of hormones, gene expression, dry matter (e.g., oil content, maturity), etc. can also be used to determine the model that is generated and / or should be used during run-time.

[0104]

[0110] The model can be trained to generate an output indicating the likelihood that a particular plant and / or plants of similar origin will develop or are currently developing a particular type of latent infection, or any type of latent infection. In some embodiments, the model can also be trained to determine one or more mitigation actions that an associated user can take to prevent the onset of a predicted latent infection. For example, the model can be trained to select one or more mitigation actions from a repository of known mitigation actions typically performed by associated users. The model can be trained to select an optimal mitigation action to reduce or prevent wastage of plant systems at harvest. This can save the associated user time because the associated user does not have to think about which type of mitigation action may be preferable in a particular situation. Instead, the associated user can review the suggested mitigation actions and select one or more such mitigation actions to implement automatically or semi-automatically.

[0105]

[0111] In some embodiments, the model can be constructed using classification-based hierarchical cluster analysis. For example, k-means can be used to construct and train the model. Clustering analysis can be beneficial because gene expression may have similar trends and therefore provide clearer clusters for analysis and prediction decisions. Thus, the model can more accurately predict infection onset. One or more other techniques, including but not limited to random forest modeling, can also be used to construct and / or train the model. Decision trees can also be used to correlate / associate gene expression data, plant information, and other invasive and / or non-invasive plant measurements / data. Thus, any of the techniques described herein and incorporated by reference can be used to train the model to correlate non-invasive measurements 304 with invasive measurements 306 to determine pre-harvest onset of infection. In some embodiments, training and use of the model does not require invasive measurements 306 and can instead rely on predicting pre-harvest latent infection onset using non-invasive measurements 304, such as volatile composition, gene expression, and other biomarkers.

[0106]

[0112] 4 is a flowchart of a process 400 for predicting the likelihood of infection in a plant. Process 400 can be used to predict the onset of latent infection in a particular plant and / or plants of similar origin. Such prediction can be made prior to harvest, at any time up to six months prior to the time of harvest. Process 400 can also be performed at any other time prior to harvest of the plant and / or plants of similar origin, such that mitigation actions can be taken prior to harvest, sufficiently in advance to reduce or prevent wastage of the plant system.

[0107]

[0113] Process 400 may be performed by computer system 140 shown and described throughout this disclosure (e.g., with reference to FIGS. 1A and 5). Process 400 and / or one or more blocks of process 400 may be performed by other computing systems, devices, networks of computers / devices, cloud-based systems, and / or cloud-based services. For illustrative purposes, process 400 is described from the perspective of a computer system.

[0108]

[0114] Referring to process 400 of FIG. 4, a computer system can receive plant data at 402. See FIGS. 1 and 2 for further discussion of receiving plant data. At 404, the computer system can identify predictive features for the plant based on the plant data. In other words, analysis of the plant data can lead to the identification of one or more features that can be used to select an appropriate model for predicting the likelihood of infection for the corresponding plant. Predictive features can include, but are not limited to, plant type, growth stage, geographic region / location, soil type, age, hormones, gene expression, dry matter, and environmental conditions. As described herein, a different predictive model can be used for each plant type, growth stage, geographic region / location, soil type, age, hormones, gene expression, dry matter, and / or environmental condition. Thus, by identifying predictive features based on the plant data, the computer system can determine the number of and appropriate predictive models suitable for predicting latent infection for the corresponding plant. Sometimes, geographic region / location features identified from the plant data can be used to select a subset of models to be trained to predict latent infection for plants in that geographic region / location. The environmental conditions or one or more predictive features, also identified based on the plant data, can be used to select one or more models from the subset of models selected based on the geographic region / location features. Thus, at 406, the computer system can identify one or more models to apply based on the predictive features.

[0109]

[0115] At 408, the computer system can obtain the identified model. At 410, the computer system can then apply the model to the plant data. The plant data can be provided as input to the model. One or more different portions of the plant data can be input to each model. For example, one model can receive environmental conditions as inputs to predict the likelihood of latent infection and / or the likelihood of a particular type of latent infection. Another model can receive plant gene expression, volatile composition, and / or hardness measurements from the plant data as inputs to predict the likelihood of latent infection and / or the likelihood of a particular type of latent infection. Yet another model can receive a different combination of inputs, such as environmental conditions and gene expression, to predict the likelihood of latent infection in a plant. The model can be trained to receive one or more other inputs and / or combinations of inputs.

[0110]

[0116] The computer system can identify 412 the pre-harvest quality of the plant based on application of the models to the plant data. For example, the computer system can determine whether outputs from one or more models, alone and / or in combination, exceed some predetermined threshold range or value. If the outputs exceed the predetermined threshold range or value, the computer system can determine that the plant will and / or will develop a latent infection. On the other hand, if the outputs do not exceed the predetermined threshold range or value, the computer system can determine that the plant does not have any latent infection and / or is unlikely to develop a latent infection based on the current condition of the plant and / or the environment.

[0111]

[0117] As another example, the computer system can average the outputs from the applied models and determine whether the average output exceeds a predetermined threshold range or value. The computer system can also identify the arithmetic mean output of the models. In some embodiments, the computer system can implement a majority rule to identify outputs that are similarly expressed by the majority of the applied models. The computer system can then employ that output as the pre-harvest quality of the plant at 412. Further, at 412, the computer system can identify the pre-harvest quality of plants of similar origin. Plants of similar origin can be part of the same tree, the same zone of a field or other growing region, and / or the same area as the plant for which the plant data was received at 402. In some embodiments, plants of similar origin can be in a different growing region but have similar environmental conditions, soil, and / or other predictive characteristics as the plant for which the plant data was received at 402.

[0112]

[0118] At 414, the computer system may generate an output about the identified pre-harvest quality of the plant. The computer system may also store the identified pre-harvest quality of the plant. The computer system may also transmit the output to a user device, whereby the output is presented to an interested user, such as a farmer or other user in the supply chain. The output may indicate the pre-harvest quality of the plant and / or the quality the plant may have by the time the plant is harvested. The output may also indicate the pre-harvest quality of plants of similar origin. The computer system may also suggest one or more actions that may be performed to mitigate pre-harvest onset of latent infection. Optionally, the computer system may automatically perform one or more mitigation actions. The suggested and / or performed mitigation actions may include any of the actions listed in Table 1 above. One or more other mitigation actions may also be suggested and / or performed.

[0113]

[0119] 5 is a system diagram of one or more components that can be used to implement the techniques described herein. As described herein, computer system 140, user devices 110A-N, plant data store 230, and model data store 240 can communicate over network 150. Collection device 500 can also communicate with the components described herein over network 150. Collection device 500 can include any collection device described herein, particularly with reference to FIG. 1A , including, but not limited to, a handheld sampler of user 105, sample-collecting robotic device 120, harvesting robotic devices 160, 162, 164, electronic nose / sniffers, and / or other automated robotic devices capable of collecting samples from plants and plant products.

[0114]

[0120] Computer system 140 may include a model training module 502, a data processing engine 504, a latent infection prediction engine 153, a latent infection detection module 154, a mitigation engine 155, an output generator 506, and a communication interface 508. Computer system 140 may optionally include one or more additional, fewer, or different components that may be used to perform the techniques described throughout this disclosure.

[0115]

[0121] The model training engine 502 can be configured to train predictive models 510A-N. The engine 502 can receive training data, which can be stored in the plant data store 230 and / or the model data store 240. The engine 502 can use the training data to train the predictive models 510A-N to predict pre-harvest latent infection onset in plants. The engine 502 can train the predictive models 510A-N based on growing region, plant type, plant development stage, and / or environmental conditions. The engine 502 can also train one or more additional, fewer, or different predictive models 510A-N, as described throughout this disclosure. Additionally, the engine 502 can train multiple models within multiple models. The engine 502 can also continuously train and / or validate the predictive models 510A-N using predictive decisions made during runtime. The models generated and trained by the engine 502 can be stored as predictive models 510A-N in the model data store 240.

[0116]

[0122] The data processing engine 504 can be configured to process the plant data received from the collection device 500 as described throughout this disclosure. In some embodiments, the plant data collected by the collection device 500 can be stored in the plant data store 230 as plant information 512A-N. The engine 504 can then retrieve the plant information 512A-N from the plant data store 230 and process the information as described herein. The plant information 512A-N can include, but is not limited to, infestation predictions, growing regions, environmental conditions, group stages, non-invasive measurements, invasive measurements, growing zones in a field, farm, or other growing location, and / or read sequence data.

[0117]

[0123] The biomarker expression engine 151 can be configured to perform the techniques described in Figure 1 A. The input generation engine 152 can also be configured to perform the techniques described in Figure 1 A.

[0118]

[0124] The latent infection prediction engine 153 can be configured to predict the likelihood that the plant and / or plants of similar origin are, have developed, and / or are likely to develop a latent infection before harvest. See FIG. 1A for further discussion of the latent infection prediction engine 153. Briefly, the latent infection prediction engine 153 can receive plant data from the collection device 500, the plant data store 230, and / or the data processing engine 504. Using the plant data, the latent infection prediction engine 153 can select one or more of the prediction models 510A-N from the model data store 240 to apply to the plant data. Based on the application of the selected model 510A-N, the latent infection prediction engine 153 can predict the likelihood that the plant is developing, may develop, and / or has developed a latent infection before harvest. The latent infection prediction engine 153 can store this prediction in the plant data store 230 as part of the plant information 512A-N.

[0119]

[0125] The latent infection detection module 154 can be configured to determine one or more mitigation efforts that can be implemented in response to the predicted likelihood of pre-harvest latent infection, as described with reference to Figure 1 A. See Figure 1A for further discussion of the latent infection detection module 154.

[0120]

[0126] The mitigation engine 155 can be configured to determine whether one or more actions can and / or should be performed to mitigate the potential onset of latent infection in the plant and / or plants of similar origin prior to harvest. The engine 155 can also determine one or more suggestions for performing mitigation actions. The engine 155 can receive a prediction from the latent infection prediction engine 153 (or obtain a prediction from the plant information 512A-N in the plant data store 230) to determine one or more mitigation actions. In some embodiments, the mitigation engine 155 can optionally perform one or more of the mitigation actions with or without approval from a user at the user device 110A-N. For example, in a situation where the plant and / or plants of similar origin have already developed a latent infection and are likely to continue to develop the infection to the point where the plant system will be wasted by the time of harvest, the mitigation engine 155 can automatically perform one or more mitigation actions to attempt and reduce the continued development of the infection.

[0121]

[0127] The output generator 506 can be configured to generate an output for presentation to the user device 110A-N. The output generator 506 can receive a prediction decision from the latent infection prediction engine 153 and / or a mitigation action from the mitigation engine 155. Using this information, the output generator 506 can generate an output such as a message, a notification, an alert, and / or an alarm. The generated output can include, for example, a value indicative of a predicted likelihood of latent infection of the plant and / or plants of similar origin. The generated output can also include one or more mitigation actions that can be selected and performed (e.g., manually and / or automatically). The output can also include one or more other messages, notifications, alerts, and / or alarms as described herein. The output generator 506 can send the output to the user device 110A-N for presentation to the user. The output generator 506 can also send the output to the plant data store 230 for storage in plant information 512A-N.

[0122]

[0128] Finally, communication interface 508 may provide communication between one or more of the components described herein.

[0123]

[0129] User devices 110A-N may include any one or more of a computer, laptop, tablet, mobile device, smartphone, cell phone, etc. that may be used by an associated user. User devices 110A-N may include input devices, output devices, processors, and communication interfaces. Such components enable a user to provide information to computer system 140, display and / or access information from one or more of the components described herein, and / or act on information provided thereto.

[0124]

[0130] 6A-6C show an example process 600 for collecting plant data for use in conjunction with the disclosed techniques. Figures 6-10 provide illustrative examples of how the disclosed techniques can be applied to certain types of plants. Figures 6-10 are merely exemplary and do not limit the application of the disclosed techniques to other types of plants, including other perennial-bearing trees (e.g., mango, citrus, pear, apple, and stone fruit trees), grapes (e.g., grape vines), and shrubs (e.g., blackberry bushes, raspberry bushes, and / or strawberry bushes).

[0125]

[0131] With reference to Figures 6-10, avocado trees (Persea americana) can be long-lived perennial plants capable of producing large quantities of highly desirable fruit crops for decades. Throughout the lifespan of an avocado tree, the tree can constantly interact with and respond to pathogenic fungal populations. While many of these pathogenic relationships can be commensal or mutualistic, there are also many instances of parasites of avocado trees and / or fruit that can cause serious losses in tree health, fruit yield, and fruit quality. Stem rot (SER) can be a common quality defect in avocado fruit caused by parasitic infestation of the avocado fruit flesh near the base of the stem as the fruit ripens after harvest. The fungus that causes this disease can survive parasitically on the fruit's pedicels prior to harvest and ripening, typically without producing disease symptoms on the fruit or tree.

[0126]

[0132] Thus, using the disclosed techniques, physical and molecular attributes of early-, mid-, and late-season avocado fruit can be correlated with low or high incidence of SER at harvest and ripening. Using RNA sequencing, expression profiles can be generated that may indicate that fruit from trees with high SER incidence, at all time points prior to peak harvest maturity, are able to recognize and respond to increased fungal pressure through increased expression of genes involved in canonical pathogen response pathways. Leveraging this information toward the generation of detection assays allows prediction of high and low incidence trees in ample time to make modifications to fungicide applications, fruit harvest schedules, or fruit marketing destinations to mitigate the development of SER or other latent infections of avocados.

[0127]

[0133] In process 600, 602 is an aerial photograph of avocado trees available for sampling (see, e.g., FIG. 6A). Heat map 604 represents the average SER incidence for each tree at each collection time point (see, e.g., FIG. 6B). The incidence value for each tree at each time point is shown within each box in heat map 604. Additionally, blank boxes indicate that a particular tree was unavailable for sampling at that time point due to low fruit load or ongoing tree management / pruning. Image data 606 represents the difference between low (C2) and high (C1) incidence trees at t1 (days after flowering, DAF 330) and t3 (days after flowering, DAF 520) (see, e.g., FIG. 6C).

[0128]

[0134] Avocado fruit can be collected from trees in a field or other growing area. As shown in aerial photograph 602 in Figure 6A, samples were collected from 13 trees within block 7 of the field. Fruit can be collected periodically throughout the growing season. All fruit can be collected by gently picking the fruit with stainless steel scissors that are frequently sterilized with a spray of 70% EtOH throughout the harvest, leaving approximately 2-4 cm of pedicel. In the example process 600, at each collection time point, 30 fruits can be harvested from each tree and transported in crates to a research center (approximately 15 minutes of transport time). A total of 1,050 fruits can be collected from these trees. Weather data for the harvest date can also be obtained from a data store, as shown in graphs 702 and 704 in Figure 7A. The time points analyzed in the example process 600 can be from approximately 330 days DAF, approximately 430 days DAF, and approximately 520 days DAF.

[0129]

[0135] Upon arrival at the research center, the fruit are placed in thin produce crates and allowed to equilibrate to ambient temperature overnight. Approximately 24 hours after harvest, all samples (n = 30 per tree per time point) can be individually measured for fruit mass (g) and respiration rate (mL CO2 / kg-hr). Fruit respiration can be measured by sealing the fruit in custom airtight pods equipped with CO2 sensors and measuring CO2 concentrations every 5 seconds. The dry matter percentage of each fruit can be determined using a near-infrared handheld spectrometer that uses a custom dry matter prediction algorithm. Images of the fruit's external quality can also be taken. At this point, a subset of the fruit (n = 4 per tree per time point) can be destructively sampled for RNA extraction, while the remainder of the fruit (n = 24) can be allowed to ripen under ambient conditions. The length of time (in days) required for the fruit to reach 50 Shore (measured using a handheld hardness tester) can be considered the time to ripen. Once it reaches 50 Shore, the fruit can be cut in half, inspected for quality defects (including SER), and imaged as shown in image data 606.

[0130]

[0136] Once all collections are complete, the SER incidence of ripening fruit on individual trees at each time point can be determined. Trees that produced fruit marked by increasing SER incidence over the season can be designated "high incidence," while trees producing fruit with low SER incidence can be selected as "controls" or "low incidence" for comparison. Three "high incidence" (A1, C1, and D1) and three "low incidence" (A4, C2, and D2) trees can be selected for detailed transcriptional analysis. See heatmap 604 in Figure 6B.

[0131]

[0137] As described above, at each time point, 24 hours after harvest, four fruits can be destructively sampled from each tree by cutting cross sections of the fruit (which can include the seeds, endocarp, mesocarp, and exocarp) and grinding them to a fine powder under liquid nitrogen. These cross sections can include the skin, pulp, and seeds. Samples can be stored at -80°C until used for RNA extraction.

[0132]

[0138] RNA can then be extracted from each individual sample from the "high incidence" and "low incidence" trees collected at each time point using a cetyltrimethylammonium bromide (CTAB) buffer-based extraction protocol with bead cleanup. RNA-seq libraries can then be prepared and pooled by time point for sequencing. In some embodiments, libraries can be run for 75 cycles (high output). File quality can be assessed and aligned to a specific avocado reference genome. Bam file generation and indexing can also be performed.

[0133]

[0139] Sequence reads can also be aligned to genomic features. A separate read count table can be prepared for each time point. At each time point, differential expression between high and low incidence trees can be identified. Differentially expressed genes can be filtered for low counts (only genes with counts greater than 10 were retained) and p-values ​​less than 0.05.

[0134]

[0140] Principal component analysis can also be performed and visualized in the transformed normalized DeSeq matrix. In preparation for heatmap visualization, the fold-change matrix of the top, up-, and down-regulated genes (filtered for p-values ​​less than 0.05) at each time point can be normalized to each gene's value across samples. The R function hchust can be used to cluster genes by expression patterns within each time point (using method=complete). In some embodiments, the R function pheatmap can be used to visualize how individual samples cluster based on gene expression patterns.

[0135]

[0141] Clustering of differentially expressed genes according to time-course expression trends can be performed and visualized. Such functions can calculate gene similarity over time across samples within a condition. Relative gene expression can be normalized to a Z-score, where the value is centered around the arithmetic mean of each gene and scaled to the standard deviation of each gene. In some embodiments, genes analyzed for time-course trends can be limited to the top 500 up-regulated differentially expressed genes and the top 500 down-regulated differentially expressed genes with a p-value of less than 0.05. Gene clusters examined can be limited to those with a gene count greater than 20.

[0136]

[0142] Using the Arabidopsis homologs of each avocado transcript, gene ontology analysis can be performed on lists of the top differentially regulated avocado genes by time point and expression cluster. These lists can then be used as input to enrichment analysis tools that can map the provided gene lists to known sources of functional information to detect statistically significant enriched biological processes, pathways, regulatory motifs, and protein complexes. The regulated enrichment p-values ​​for each enrichment classification can be identified and output (e.g., reported).

[0137]

[0143] Overall, Figures 6A-6C show the location (e.g., area) of sampled trees and the incidence of SER in fruit collected from each tree at each time point. At the first collection date, 330 days after harvest, no SER was observed. Over the next two collections, SER incidence increased on average, with the most significant increase in fruit from trees A1, C1, and D1. Therefore, samples from these trees can be assigned to the "high incidence" group. Nearby trees A4, C2, and D2, which had minimal or no SER over the collections, can be assigned to the "low incidence" group as a control comparison.

[0138]

[0144] Figures 7A and 7B show graphical representations of example physical attributes of the plants of Figure 6 that may be collected at different times. Graph 702 in Figure 7A is a plot of maximum and minimum temperatures throughout the fruit collection described herein, with sample collection time points labeled 1, 2, and 3. Graph 704 in Figure 7A is a plot of rainfall measured over an area throughout the fruit collection described herein, with sample collection time points labeled 1, 2, and 3. Graph 706 in Figure 7B shows the average initial respiration per fruit collected at 330 days DAF, 430 days DAF, and 520 days DAF. Graph 708 in Figure 7B shows the average mass of fruit collected at each time point from selected high and low SER incidence trees, with standard deviations plotted as error bars.

[0139]

[0145] Referring to both Figures 7A and 7B, respiration rates 24 hours after harvest may be highest for fruit from all trees at the beginning of February, 330 days after flowering (see, e.g., graph 706 in Figure 7B). This may be due to lower average temperatures early in the year (see, e.g., graph 702 in Figure 7A) and / or the stage of fruit development at that time, as rapid growth and / or cell division may require increased metabolism and therefore increased respiration. While no significant differences in average dry matter content or respiration rate may be observed between high and low SER incidence trees, trends may be observed between the average mass of individual fruits and SER incidence. As days after flowering increase, fruit mass may also increase (see, e.g., graph 708 in Figure 7B), with average fruit mass being consistently higher for fruit from the high SER incidence group. This trend may be most pronounced in the third collection (520 days DAF), where the average fruit mass from A1, C1, and D1 is on average greater than 250 g, while the average fruit mass from A4, C2, and D2 are all less than 200 g (see, e.g., graph 708 in Figure 7B).

[0140]

[0146] 8A-8C show biplots of principal component analysis (PCA) of normalized gene expression across all times of data collection for the plants in FIGS. 6A-6C. In particular, biplot 800 shows the first two principal components from the PCA of gene expression data at 330 days DAF for both the low and high SER incidence groups. Biplot 802 shows the first two principal components from the PCA of gene expression data at 430 days DAF for both the low and high SER incidence groups. Biplot 804 shows the first two principal components from the PCA of gene expression data at 520 days DAF for both the low and high SER incidence groups.

[0141]

[0147] Referring to biplots 800-804 in Figures 8A-8C, across all three sequencing runs, 5-6 million reads were obtained per sample, with the percentage of identified reads (%RF) and percentage of Q-scores above 30 (%Q30) exceeding 90% and 95%, respectively. After alignment to the Persea americana cv. Hass genome, an average of 81.4%, 80.8%, and 80.9% of the reads could be successfully assigned to the annotated regions of the genome for the t1 (330 days DAF), t2 (430 days DAF), and t3 (520 days DAF) datasets, respectively.

[0142]

[0148] Visualization of PCA of normalized expression across all time points using pair plots can show that the strongest separation by PCA occurs between time points, particularly PC1-PC2. There may be some slight separation between low and high SER groups in PC3 and PC4. However, when PCA is performed within each individual time point and visualized using biplots, a slightly clearer separation between low and high SER samples can be observed across PC1 and PC2, as shown in Figures 8A-8C. This separation may be most clearly evident at the latest time point, 520 days DAF.

[0143]

[0149] Figures 9A-9C show volcano plots of normalized gene expression for the plants in Figures 6A-6C. Volcano plot 900 in Figure 9A shows differential gene expression between low and high incidence trees at 330 days DAF, volcano plot 902 in Figure 9B shows differential gene expression between low and high incidence trees at 430 days DAF, and volcano plot 904 in Figure 9C shows differential gene expression between low and high incidence trees at 520 days DAF. Thus, plots 900-904 show normalized gene expression (Log2 fold change) vs. statistical significance (-Log 10 Plots 900-904 show differentially expressed genes with a Log2 fold change greater than |2| and a p-value less than 0.05 falling into the pattern quadrant. The numbers within these quadrants represent the number of genes that meet these eligibility criteria. Thus, in Figure 9A, seven genes meet the eligibility criteria for quadrant 1, 11 genes meet the eligibility criteria for quadrant 2, 25 genes and 12 genes, respectively, in Figure 9B, and eight genes and 22 genes, respectively, in Figure 9C.

[0144]

[0150] 9A-9C, differences in gene expression can be seen by comparing the high-incidence tree group (A1, C1, and D1) with the low-incidence tree group (A4, C2, and D2) at all three time points. At t1, 863 genes were upregulated in the high-incidence tree (4% of identified genes) and 963 genes were downregulated (4.4% of identified genes). At t2, 1012 genes were upregulated in the high-incidence tree (4.8% of identified genes) and 786 genes were downregulated (3.7% of identified genes). At t3, 234 genes were upregulated in the high-incidence tree (1% of identified genes) and 166 genes were downregulated (0.08% of identified genes). When these DEGs are filtered for p-values ​​less than 0.05 and fold changes less than -2 or greater than 2, the up- and down-regulated genes in the high incidence trees at t1, t2, and t3 can be 11 up and 7 down, 12 up and 25 down, and 22 up and 8 down, respectively.

[0145]

[0151] In some embodiments, a heatmap visualization (not shown) of the top 30 up- and down-regulated genes at each time point (pre-filtered for p-values ​​less than 0.05 prior to sorting by fold change) may show that individual fruit expression patterns from trees marked as high incidence cluster together. This may suggest that these trees can be assigned high or low incidence based on the gene expression patterns of these genes in fruit as early as the first time point in February, 330 days DAF, and six months prior to the later harvest season. Notably, there may be little overlap of these genes across the three time points. Only four genes were conserved across these three groups: maker-Ctg1243-augustus-gene-0.16-mRNA-1, maker-Ctg0870-augustus-gene-0.10-mRNA-1, augustus_masked-Ctg0285-processed-gene-0.0-mRNA-1, and maker-Ctg0255-augustus-gene-4.16-mRNA-1. The predicted closest homologs of each of these genes in tomato ( S. lycopersicum ) and Arabidopsis ( A. thaliana ) could be identified using Blast analysis. These results may suggest that responses to latent fungal pathogenic mechanisms at the transcriptome level change throughout the fruit development phase on the tree.

[0146]

[0152] Statistical enrichment analysis of the top 30 up- and down-regulated genes at each time point can use the list of Arabidopsis homologs as input. Results can be filtered for Gene Ontology terms with a p-value of 0.05 or less and a false discovery rate (FDR) of less than 1. In the example shown in Figures 6A-6C, there can be 20, 35, and 79 Gene Ontology terms meeting these criteria at time points t1, t2, and t3, respectively. This analysis can provide the highest level of insight into the biological processes, molecular functions, and cellular component classifications most significantly represented in this subset of transcripts, which can show significant differences between low-SER and high-SER trees.

[0147]

[0153] As expected from the limited overlap of specific genes between time points in the top 30 up- and down-regulated gene sets, there may be limited overlap between the most significant Gene Ontology classifications, with t1 and t2 containing greater overlap than t1 and t3 and t2 and t3. For example, significant GO datasets for t1 and t2 may both contain the biological process terms response to other organisms (GO:0051707) and response to external biological stimuli (GO:0043207), along with the cellular component term cell margin (GO:0071944) and the molecular function term transporter activity (GO:0005215), while neither of these terms may overlap with t3. In addition to these terms, t2 may have additional terms related to oxidoreductase activity (GO:0016491), response to oxidative stress (GO:0006979), plant organ development (GO:0099402), response to osmotic stress (GO:0006970), defense responses to other organisms (GO:0098542), and aromatic compound biosynthesis (GO:0019438), which may not be present in the t1 dataset. The significant GO list for t3 may also include oxidoreductase activity terms (GO:0016491) and aromatic compound biosynthesis (GO:0019438), as well as terms related to plant organ development (GO:0099402), which may overlap with t2, such as postembryonic development (GO:0009791) and plant reproductive structure development (GO:0048608). Generally, t3 may not overlap with t1 or t2. Instead, the t3 GO list can include terms related to hormone signaling, including response to hormones (GO:0009725), cellular response to hormone stimulation (GO:0032870), and hormone-mediated signaling pathways (GO:0009755). Additionally, the t3 dataset can include many GO terms related to transcriptional regulation that are not found in the t1 or t2 datasets. These include, but are not limited to, transcription (GO:0006351), transcriptional regulation (GO:0006355), DNA binding (GO:0003677), RNA biogenesis (GO:0032774), and RNA biogenesis regulation (GO:2001141).

[0148]

[0154] Figures 10A-10D show normalized gene expression over time for selected gene clusters with strong time and condition trends in the plants in Figures 6A-6C. Referring to all of Figures 10A-10D, the numbers 1, 2, and 3 on the x-axis indicate the respective sample collection time points at 330 days DAF, 430 days DAF, and 520 days DAF. Clustering of differentially expressed genes by time expression trends can be performed and visualized using known computational techniques. Such techniques can calculate gene similarity over time across samples within a condition. Relative expression can be normalized to a Z-score, and values ​​can be centered around the arithmetic mean and scaled to the standard deviation for each gene. Genes analyzed for time trends can be limited to the top 500 up-regulated and top 500 down-regulated differentially expressed genes with p-values ​​less than 0.05. Visualized gene clusters can be limited to those with a gene count greater than 20. Sixteen gene clusters met these criteria and are shown in Figure 10A-D. Using g:GOSt and the list of Arabidopsis homologs to P. americana genes grouped into each cluster, statistical enrichment analysis can be performed. The top GO categories for each cluster are shown in Table 2.

[0149] [Table 6]

[0150] [Table 7]

[0151]

[0155] In particular, from the highly infected trees, there can be several groups (groups 9, 11, 12, and 18) that show an increasing trend over time and have relatively high gene content in fruit, as shown in Figures 10A-10D and Table 2. The genes and GO annotations of these groups can be examined for pathways specific to plant pathogen response, and group 11 can be found to have a significant number of terms associated with these pathways. For example, some of the top biological process GO categories with p-values ​​much smaller than 0.05 within group 11 can be response to stimuli (GO:0050896), response to chemicals (GO:0042221), and response to other organisms (GO:0051707), and examples of genes from these categories include, but are not limited to, predicted endochitinase 4 (maker-Ctg0022-augustus-gene-3.7), pathogenesis-related protein PR1 precursor (augustus_masked-Ctg0281-processed-gene-0.6), and osmotin-like pathogenesis-related protein R (augustus_masked-Ctg2009-processed-gene-0.2), which are plotted in Figures 10A-10D. Cluster 18 can further include GO clusters involved in responses to stress, such as reactive oxygen species metabolic processes (GO:0072593) and regulation of superoxide radical scavenging (GO:2000121). Additionally, cluster 12 can have an over-representation of genes involved in responses to jasmonic acid (GO:0009753), which can be established as involved in plant defense signaling. Thus, significant differences in gene expression over time can exist between high-SER and low-SER tree populations and can account for shifts in characteristic pathogen response pathways.

[0152]

[0156] Among the gene groups (Groups 1, 2, and 4) in which high-SER samples showed lower expression compared to low-SER groups, there may be a general overrepresentation of genes involved in cell development and differentiation. For example, the most significant biological process GO terms in Group 1 are floral organ identity specification (GO:0010093) and plant organ identity specification (GO:0090701), while the top two in Group 4 are xylem and phloem patterning (GO:0010051) and regionalization (GO:0003002). This may suggest that fruit from trees with higher SER may have differences in developmental patterns. Furthermore, Group 5 contains a high representation of genes involved in the phenylpropanoid biosynthesis process (GO:0009699). Phenylpropanoid compounds are known to be involved in plant defense. Therefore, higher expression of enzymes involved in phenylpropanoid biosynthesis may reduce the susceptibility of low-SER fruit to SER infection. Furthermore, the phenylpropanoid pathway can generate precursors for lignin biosynthesis and therefore may be related to the developmental gene expression differences seen in groups 1, 2, and 4.

[0153]

[0157] Figures 6-10 show that differential gene expression analysis can reveal significantly different up- and down-regulated genes in fruit samples from trees grouped by high and low SER incidence. These trees can be classified into high or low SER incidence groups based on the onset of SER symptoms in fruit ripening at harvest maturity. It can be observed that low-maturity fruit (less than approximately 30% dry matter) harvested and ripened at the first two time points may experience minimal, if any, SER onset, while more mature fruit at the final time point may develop significant levels of SER. This trend may be due to the increase in average temperature leading up to the final time point. Gene expression data can indicate measurable differences in the transcriptome landscape of fruit from high-incidence trees, even at these early time points before symptom onset, facilitating preharvest prediction of latent infection.

[0154]

[0158] Furthermore, as described herein in connection with Figures 6-10, the specific genes and processes most significantly altered in infected trees may shift over the time course of fruit development. This can be demonstrated through minimal overlap between the top 30 up-regulated and the top 30 down-regulated genes, which allows fruit expression trends to be clustered into expected high and low SER groups. Analysis of gene groups can be tailored differently over time, which may further indicate dynamic shifts in relative transcript abundance of genes across the three sampled time points. Nevertheless, gene ontology analysis of these grouped expression time trends can allow for the isolation of groups of known pathogen response genes, including osmotin-like proteins and chitinases, which may show consistent and / or increasing trends of differential expression over time in the high SER tree group. Furthermore, genes involved in jasmonate signaling and reactive oxygen species metabolism can be identified and up-regulated in fruit from high-incidence trees at early time points (see, e.g., Figures 10A-10D, groups 12 and 18 in Table 2). Reactive oxygen species and jasmonate signaling events may be known mechanisms deployed in the early host defense response to fungal pathogens, suggesting that fruit from trees with higher SER incidence may indeed exhibit a higher response to the presence of SER fungal pathogens, well before the onset of infectious symptoms.

[0155]

[0159] We can observe that the average postharvest SER incidence is up to 17% for tree C1, even for high-incidence trees. Using four individual fruit samples per tree, the likelihood that these particular fruits will develop SER symptoms is low. Therefore, the observed gene expression differences between fruit from low-incidence and high-incidence trees in the selected sample can be interpreted as a systematic response throughout the tree to the presence of SER fungi.

[0156]

[0160] Furthermore, analysis of grouped genes by differential expression over time (see, e.g., Figures 10A-10D, Table 2) may further suggest differences in fruit development between fruit from high and low SER incidence trees. Specifically, high SER fruit may show lower expression of genes involved in cellular development and differentiation. The measured physical difference between the high and low SER groups may be a small difference in the average mass of fruit harvested across the three time points (see, e.g., Figure 7B), with higher SER trees producing fruit with a higher average mass.

[0157]

[0161] FIG. 11 is a block diagram of system components that can be used to implement a system for detecting latent infection. Computing device 1100 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing system 1150 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. Additionally, computing device 1100 or 1150 may include a universal serial bus (USB) flash drive. A USB flash drive may store an operating system and other applications. A USB flash drive may include input / output components, such as a wireless transmitter or a USB connector that can be inserted into a USB port on another computing device. The components, their connections and relationships, and their functions shown herein are intended to be merely exemplary and are not intended to limit the embodiments of the invention described and / or claimed herein.

[0158]

[0162] Computing device 1100 includes processor 1102, memory 1104, storage device 1108, a high-speed interface 1108 connecting memory 1104 and a high-speed expansion port 1110, and a low-speed interface 1112 connecting a low-speed bus 1114 and storage device 1108. Each of components 1102, 1104, 1108, 1108, 1110, and 1112 are interconnected using various buses and may be mounted on a common motherboard or otherwise as appropriate. Processor 1102 may process instructions for execution within computing device 1100, including instructions stored in memory 1104 or storage device 1108 that display graphical information for a GUI on an external input / output device, such as a display 1116 coupled to high-speed interface 1108. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Additionally, multiple computing devices 1100 may be connected, for example, as a server bank, a cluster of blade servers, or a multi-processor system, with each device providing a portion of the required operations.

[0159]

[0163] The memory 1104 stores information within the computing device 1100. In one implementation, the memory 1104 is one or more volatile memory units. In another implementation, the memory 1104 is one or more non-volatile memory units. The memory 1104 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0160]

[0164] The storage device 1108 is capable of providing mass storage for the computing device 1100 .

[0161]

[0165] In one embodiment, storage device 1108 can be or include a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device or tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. A computer program product can be tangibly embodied on an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as memory 1104, storage device 1108, or memory 1102 on a processor.

[0162]

[0166] The high-speed controller 1108 manages bandwidth-intensive operations of the computing device 1100, while the low-speed controller 1112 manages less bandwidth-intensive operations. Such an allocation of functionality is merely exemplary. In one embodiment, the high-speed controller 1108 is coupled to the memory 1104, coupled to a display 1116, e.g., through a graphics processor or accelerator, and coupled to a high-speed expansion port 1110 that can accept various expansion cards (not shown). In that embodiment, the low-speed controller 1112 is coupled to the storage device 1108 and to a low-speed expansion port 1114. The low-speed expansion port can include various communication ports, e.g., USB, Bluetooth, Ethernet, wireless Ethernet, and can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a microphone / speaker pair, a scanner, or a networking device, e.g., through a network adapter, such as a switch or router. The computing device 1100 can be implemented in several different forms, as shown. For example, it can be implemented as a standard server 1120, or multiple times within a cluster of such servers. It may also be implemented as part of a rack server system 1124. Additionally, it may be implemented in a personal computer such as a laptop computer 1122.

[0163]

[0167] Alternatively, components from computing device 1100 may be combined with other components in a mobile device (not shown), such as device 1150. Each such device may include one or more of computing devices 1100, 1150, and the overall system may consist of multiple computing devices 1100, 1150 in communication with each other.

[0164]

[0168] The computing device 1100 can be implemented in several different forms, as shown in the figure. For example, it can be implemented as a standard server 1120, or multiple times within a cluster of such servers. It can also be implemented as part of a rack server system 1124. Additionally, it can be implemented in a personal computer, such as a laptop computer 1122.

[0165]

[0169] Alternatively, components from computing device 1100 may be combined with other components in a mobile device (not shown), such as device 1150. Each such device may include one or more of computing devices 1100, 1150, and the overall system may consist of multiple computing devices 1100, 1150 in communication with each other.

[0166]

[0170] Computing device 1150 includes, among other components, a processor 1152, memory 1164, an input / output device such as a display 1154, a communication interface 1166, and a transceiver 1168. Device 1150 may also be provided with a storage device such as a microdrive or other device to provide additional storage. Each of components 1150, 1152, 1164, 1166, and 1168 are interconnected using various buses, and some of the components may be mounted on a common motherboard or otherwise as appropriate.

[0167]

[0171] The processor 1152 can execute instructions within the computing device 1150, including instructions stored in the memory 1164. The processor can be implemented as a chipset of chips including multiple separate analog and digital processors. Furthermore, the processor can be implemented using any of several architectures. For example, the processor 1110 can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. The processor can provide, for example, a user interface, applications run by the device 1150, and coordination of other components of the device 1150, such as controlling wireless communication by the device 1150.

[0168]

[0172] The processor 1152 can communicate with a user through a control interface 1158 and a display interface 1156 coupled to a display 1154. The display 1154 can be, for example, a TFT (thin film transistor liquid crystal display) display, an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 1156 can include appropriate circuitry to drive the display 1154 to present graphical and other information to the user. The control interface 1158 can receive commands from the user and translate them for submission to the processor 1152. Additionally, an external interface 1162 can be provided in communication with the processor 1152, thereby enabling near-field communication of the device 1150 with other devices.

[0169]

[0173] External interface 1162 may, for example, provide for wired communication in some implementations or wireless communication in other implementations, and multiple interfaces may also be used.

[0170]

[0174] Memory 1164 stores information within computing system 1150. Memory 1164 may be embodied as one or more computer-readable media, one or more volatile units, or one or more non-volatile units. Expansion memory 1174 may also be provided and connected to device 1150 through expansion interface 1172, which may include, for example, a SIMM (single in-line memory module) card interface. Such expansion memory 1174 may provide additional storage space to device 1150 or may store applications or other information for device 1150. Specifically, expansion memory 1174 may include instructions for performing or supplementing the processes described above and may also include secure information. Thus, for example, expansion memory 1174 may be provided to device 1150 as a security module and programmed with instructions that enable secure use of device 1150. Additionally, secure applications may be provided along with additional information via a SIMM card, such as by placing identifying information on the SIMM card to prevent hacking.

[0171]

[0175] The memory may include, for example, flash memory and / or non-volatile random access memory (NVRAM) memory, as described below. In one embodiment, a computer program product may be tangibly embodied on an information carrier. The computer program product may also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier may be a computer- or machine-readable medium, such as memory 1164, expansion memory 1174, or memory 1152 on a processor, and may be received, for example, over transceiver 1168 or external interface 1162.

[0172]

[0176] Device 1150 can communicate wirelessly through communication interface 1166, which can include digital signal processing circuitry if necessary. Communication interface 1166 can provide communications under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communications can occur, for example, through radio frequency transceiver 1168. Additionally, short-range communications can occur using Bluetooth, Wi-Fi, or other such transceivers (not shown), etc. Additionally, a global positioning system (GPS) receiver module 1170 can provide additional navigation- and location-related wireless data to device 1150, which can be used as appropriate by applications running on device 1150.

[0173]

[0177] Device 1150 can also communicate audibly using audio codec 1160, which can receive voice information from a user and convert it into usable digital information. Audio codec 1160 can also generate audible sounds for the user, such as through a headset speaker of device 1150. Such sounds may include sounds from a voice telephone call, recorded sounds (e.g., voice messages, music files, etc.), and sounds created by applications running on device 1150.

[0174]

[0178] The computing device 1150 can be implemented in several different forms, as shown in the figure. For example, the computing device 1150 can be implemented as a cellular telephone 1180. The computing device 1150 can also be implemented as part of a smartphone 1182, a personal digital assistant, or other similar mobile device.

[0175]

[0179] Various implementations of the systems and methods described herein may be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations of such implementations. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0176]

[0180] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in a high-level procedural programming language, an object-oriented programming language, and / or an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, e.g., magnetic disk, optical disk, memory, programmable logic device (PLD), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0177]

[0181] To provide for user interaction, the systems and techniques described herein can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, for allowing the user to provide input to the computer. Other types of devices can be used to provide for user interaction as well; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0178]

[0182] The systems and techniques described herein can be implemented in a computing system that includes back-end components, such as data servers, or middleware components, such as application servers, or front-end components, such as client computers having graphical user interfaces or web browsers that allow users to interact with implementations of the systems and techniques described herein, or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0179]

[0183] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0180]

[0184] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of the disclosed technology or the scope of what may be claimed, but rather as descriptions of features that may be unique to particular embodiments of the disclosed technology. Certain features described herein in the context of separate embodiments may also be implemented in combination, either in part or in whole, in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described herein as acting in a particular combination and / or may initially be claimed as such, one or more features from the claimed combination may, in some cases, be deleted from that combination, and the claimed combination may be directed to a subcombination or variation of the subcombination. Similarly, although operations may be described in a particular order, this should not be understood as requiring such operations to be performed in that particular order or sequentially, or that all operations be performed, to achieve desirable results. Certain embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. In certain embodiments, for example, the following items are provided: (Item 1) 1. A method for identifying pre-harvest latent infection in a plant, comprising: obtaining, by a processor, data describing the expression levels of one or more infection biomarkers present in the plant, the infection biomarkers indicative of an infection likelihood of the plant, and the data indicative of a difference between a read sequence of the plant and a reference genome of a healthy plant of the same species as the plant; selecting, by the processor, one or more machine learning models based on the acquired data, the one or more machine learning models having been previously trained using data correlating other data with the identified one or more infection biomarkers of one or more other plants, and generating an output indicative of the likelihood of the plant developing a latent infection before harvest, the one or more machine learning models being trained using a process comprising: inputting into the one or more machine learning models from a training dataset: (i) non-invasive measurements of the one or more other plants; (ii) invasive measurements of the one or more other plants; (iii) known plant information of the one or more other plants; and (iv) positive infection identification information of the one or more other plants; determining a predicted likelihood that the one or more other plants will develop the latent infection before harvest based on inputting (i)-(iv) into the one or more machine learning models; and outputting the one or more machine learning models for runtime use. selecting, generating, by the processor, an output indicative of the likelihood that the plant will develop the latent infection before harvest based on applying the one or more machine learning models to the data; determining, by the processor, that the plant has a latent infection based on the output exceeding a predetermined threshold range; determining, by the processor, one or more treatments that mitigate the latent infection in the plant; outputting, by the processor, an indication that the plant has the latent infection and the one or more determined treatments; A method comprising: (Item 2) 2. The method of claim 1, wherein determining, by the processor, one or more treatments to mitigate the latent infection in the plant includes identifying, for the plant and one or more other plants of similar origin, at least one of: (i) a harvest schedule that is changed to earlier in the growing season; (ii) instructions to apply a predetermined amount of an insecticide before harvest; and (iii) a time when the plant and the one or more plants of similar origin are approaching the end of their respective healthy plant productive lives, wherein the one or more machine learning models are previously trained to identify (i)-(iii) using data from the training dataset. (Item 3) 3. The method according to item 2, wherein the one or more other plants of similar origin include plants in the same zone as the plant. (Item 4) outputting, by the processor, an indication that the plant has the latent infection and the one or more determined treatments; generating an alert message that, when processed by a user device, causes the user device to output an alert notifying a user of the user device to perform one or more of the determined actions; transmitting the generated alert message to the user device; Item 1. The method according to item 1, comprising: (Item 5) determining, by the processor, one or more treatments to mitigate the latent infection in the plant, determining, based on the generated output, that an antimicrobial treatment should be prescribed to one or more other plants of similar origin prior to harvest; Item 1. The method according to item 1, comprising: (Item 6) determining, by the processor, one or more treatments to mitigate the latent infection in the plant, generating instructions that, when treated by an irrigation controller, cause the irrigation controller to automatically spray a liquid containing the antimicrobial treatment; transmitting the instructions to the irrigation controller; Item 6. The method according to item 5, comprising: (Item 7) determining, by the processor, one or more treatments to mitigate the latent infection in the plant, generating instructions that, when processed by a robotic device, cause the robotic device to (i) navigate to a location of the plant and (ii) apply an antimicrobial treatment to the plant and one or more other plants of similar origin; transmitting the command to the robotic device; Item 1. The method according to item 1, comprising: (Item 8) determining, by the processor, one or more treatments to mitigate the latent infection in the plant, generating instructions that, when processed by a robotic device, cause the robotic device to (i) navigate to a location of the plant and (ii) harvest one or more plant products from the plant and one or more other plants of similar origin prior to an expected harvest timeframe; transmitting the command to the robotic device; Item 1. The method according to item 1, comprising: (Item 9) 2. The method according to item 1, wherein the obtained data is generated by a nucleic acid sequencer based on sequencing of plant products extracted from the plant. (Item 10) 10. The method of claim 9, wherein the plant product comprises at least one of a bark, leaf, flower, and fruit sample. (Item 11) 10. The method of claim 9, wherein the plant product is sampled non-destructively from the plant. (Item 12) 10. The method according to claim 9, wherein the plant products comprise volatile components emitted from the plant. (Item 13) Item 10. The method of item 1, wherein the one or more machine learning models include at least one of a binary logistic regression model, a logistic model tree, a random forest classifier, L2 regularization, partial least squares, and a convolutional neural network (CNN). (Item 14) 2. The method according to item 1, wherein the plant does not contain visible signs of infection. (Item 15) encoding, by the processor, the acquired data into a data structure for input to the one or more machine learning models; providing, by the processor, the encoded data structure as input to the one or more machine learning models; Item 1. The method according to item 1, further comprising: (Item 16) 10. The method of claim 1, further comprising carrying out, by the processor, one or more of the determined treatments to mitigate the latent infection in the plant. (Item 17) 2. The method of claim 1, wherein the processor's selection of one or more machine learning models is further based on one or more predictive features identified from the acquired data, wherein the one or more predictive features include at least one of a growing region of the plant, environmental conditions, a plant type, a growth stage of the plant, non-invasive measurements of the plant, invasive measurements of the plant, dry matter content, gene expression, emitted volatile components of the plant, a growing zone, and read sequence data of the plant. (Item 18) 2. The method of claim 1, wherein the known plant information includes at least one of growth history information, length of growing season, soil condition, precipitation level, amount of sunlight, environmental temperature, growing region, and plant type for the one or more other plants. (Item 19) A system for predicting latent infection in plants, comprising: one or more processors; one or more computer-readable storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations; and the operation comprises: obtaining data describing the expression levels of one or more infection biomarkers present in the plant, wherein the infection biomarkers indicate a likelihood of infection in the plant, and the data indicates a difference between a read sequence of the plant and a reference genome of a healthy plant of the same species as the plant; selecting one or more machine learning models based on the obtained data, the one or more machine learning models having been previously trained using data correlating other data with the identified one or more infection biomarkers of one or more other plants, and generating an output indicative of the likelihood of the plant developing a latent infection before harvest, the one or more machine learning models being trained using a process comprising: inputting into the one or more machine learning models from a training dataset: (i) non-invasive measurements of the one or more other plants; (ii) invasive measurements of the one or more other plants; (iii) known plant information of the one or more other plants; and (iv) positive infection identification information of the one or more other plants; determining a predicted likelihood that the one or more other plants will develop the latent infection before harvest based on inputting (i)-(iv) into the one or more machine learning models; and outputting the one or more machine learning models for runtime use. selecting, generating an output indicative of the likelihood of the plant developing the latent infection before harvest based on applying the one or more machine learning models to the data; determining that the plant has a latent infection based on the output exceeding a predetermined threshold range; determining one or more treatments that reduce the latent infection in the plant; outputting an indication that the plant has the latent infection and the one or more determined treatments; Including, the system. (Item 20) 20. The system of claim 19, wherein determining one or more treatments to mitigate the latent infection in the plant includes identifying, for the plant and one or more other plants of similar origin, at least one of: (i) a harvest schedule that is changed to earlier in the growing season; (ii) instructions to apply a predetermined amount of an insecticide before harvest; and (iii) a time when the plant and the one or more plants of similar origin are approaching the end of their respective healthy plant productive lives, wherein the one or more machine learning models are pre-trained to identify (i)-(iii) using data from the training dataset.

Claims

1. A method for identifying pre-harvest latent infection in a first plant, comprising: obtaining, by a processor, data describing the expression levels of one or more infection biomarkers present in the first plant, the data comprising a list of one or more variants, the infection biomarkers indicating the likelihood of infection of the first plant, the one or more variants describing differences between a read sequence of the first plant and a reference genome of a healthy plant of the same species as the first plant, the read sequence of the first plant representing an order of nucleotides within a nucleic acid of the first plant; selecting, by the processor, one or more machine learning models based on the acquired data, the one or more machine learning models having been previously trained using a training dataset associated with one or more second plants of the same species as the first plant, and generating an output indicative of a likelihood that the first plant will develop a latent infection before harvest, the one or more machine learning models being trained using a process comprising: inputting into the one or more machine learning models from the training dataset: (i) non-invasive measurements of the one or more second plants; (ii) invasive measurements of the one or more second plants; (iii) known plant information of the one or more second plants; and (iv) positive infection identification information of the one or more second plants, wherein the positive infection identification information comprises levels or expression of biomarkers positively associated with one or more types of latent infection; determining a predicted likelihood that the one or more second plants will develop the latent infection prior to harvest based on inputting (i)-(iv) into the one or more machine learning models; and outputting the one or more machine learning models for runtime use. selecting, generating, by the processor, an output indicative of the likelihood that the first plant will develop the latent infection before harvest based on applying the one or more machine learning models to the data; determining, by the processor, that the first plant has a latent infection based on the output exceeding a predetermined threshold range; determining, by the processor, one or more treatments that mitigate the latent infection in the first plant, wherein determining, by the processor, one or more treatments that mitigate the latent infection in the first plant; (a) identifying, for the first plant and one or more third plants of similar origin of the same species as the first plant, at least one of (i) a modified harvest schedule earlier in the growing season, (ii) instructions to apply a predetermined amount of pesticide before harvest, and (iii) when the first plant and the one or more third plants of similar origin are approaching the end of their respective healthy plant productive lives, wherein the one or more machine learning models were previously trained to determine (i)-(iii) using data from the training dataset, and the one or more third plants of similar origin include plants in the same zone as the first plant in a field; (b) determining, based on the generated output, that an antimicrobial treatment should be prescribed to one or more third plants of the same species and origin as the first plant before harvest, wherein the one or more third plants of similar origin include plants in the same zone as the first plant within a field or farm; (c) generating instructions that, when treated by an irrigation controller, cause the irrigation controller to automatically spray a liquid containing the antimicrobial treatment; transmitting said instructions to said irrigation controller; or (d) (α) generating instructions to the robotic device, when treated by the robotic device, to (i) navigate the robotic device to the location of the first plant, and (ii) apply an antimicrobial treatment to the first plant and one or more third plants of the same species and origin as the first plant; transmitting the command to the robotic device; or (β) generating instructions that, when processed by a robotic device, cause the robotic device to (i) navigate the robotic device to a location of the first plant, and (ii) harvest one or more plant products from the first plant and one or more third plants of the same species and similar origin as the first plant prior to an expected harvest timeframe; transmitting the instruction to the robotic device, wherein the one or more third plants of similar origin include plants in the same zone as the first plant in a farm or field. determining the number of times ... outputting, by the processor, an indication that the first plant has the latent infection and the one or more determined treatments; A method comprising:

2. determining, by the processor, the one or more treatments to mitigate the latent infection in the first plant includes identifying, for the first plant and one or more third plants of similar origin of the same species as the first plant, at least one of: (i) a modified harvest schedule earlier in the growing season; (ii) instructions to apply a predetermined amount of an insecticide before harvest; and (iii) a time when the first plant and the one or more third plants of similar origin are approaching the end of their respective healthy plant productive lives; wherein the one or more machine learning models were previously trained to identify (i)-(iii) using data from the training dataset; 10. The method of claim 1, wherein the one or more third plants of similar origin comprise plants in the same zone as the first plant in a field or farm.

3. outputting, by the processor, the indication that the plant has the latent infection and the one or more determined treatments; generating an alert message that, when processed by a user device, causes the user device to output an alert notifying a user of the user device to perform one or more of the determined actions; transmitting the generated alert message to the user device; The method of claim 1 , comprising:

4. Determining, by the processor, the one or more treatments that mitigate the latent infection in the first plant includes: determining, based on the generated output, that an antimicrobial treatment should be prescribed to one or more third plants of the same species and origin as the first plant prior to harvest, wherein the one or more third plants of similar origin include plants in the same zone as the first plant within a field or farm; The method of claim 1 , comprising:

5. Determining, by the processor, the one or more treatments that mitigate the latent infection in the first plant includes: generating instructions that, when treated by an irrigation controller, cause the irrigation controller to automatically spray a liquid containing the antimicrobial treatment; transmitting the instructions to the irrigation controller; The method of claim 4, comprising:

6. Determining, by the processor, the one or more treatments that mitigate the latent infection in the first plant includes: (a) generating instructions to the robotic device, when treated by the robotic device, to (i) navigate the robotic device to a location of the first plant, and (ii) apply an antimicrobial treatment to the first plant and one or more third plants of the same species and origin as the first plant; transmitting the command to the robotic device; or (b) generating instructions that, when processed by a robotic device, cause the robotic device to (i) navigate the robotic device to a location of the first plant, and (ii) harvest one or more plant products from the first plant and one or more third plants of the same species and similar origin as the first plant prior to an expected harvest timeframe; transmitting the instruction to the robotic device, wherein the one or more third plants of similar origin include plants in the same zone as the first plant in a farm or field; The method of claim 1 , comprising:

7. 10. The method of claim 1, wherein the obtained data is generated by a nucleic acid sequencer based on sequencing of plant products extracted from the first plant.

8. 8. The method of claim 7, wherein the plant product is sampled non-destructively from the first plant.

9. The method of claim 7 , wherein the plant products include volatile components emitted from the first plant.

10. 10. The method of claim 1, wherein the one or more machine learning models comprise at least one of a binary logistic regression model, a logistic model tree, a random forest classifier, L2 regularization, partial least squares, and a convolutional neural network (CNN).

11. encoding, by the processor, the acquired data into a data structure for input to the one or more machine learning models; providing, by the processor, the encoded data structure as input to the one or more machine learning models; The method of claim 1 further comprising:

12. 10. The method of claim 1, further comprising: performing, by the processor, one or more of the determined treatments to mitigate the latent infection in the first plant.

13. 2. The method of claim 1, wherein the selecting, by the processor, one or more machine learning models is further based on one or more predictive features identified from the acquired data, the one or more predictive features comprising at least one of a growing region of the first plant, environmental conditions, plant type, growth stage of the first plant, non-invasive measurements of the first plant, invasive measurements of the first plant, dry matter content, gene expression, emission volatiles of the first plant, a growing zone within a field, and read sequence data of the first plant.

14. 2. The method of claim 1, wherein the known plant information includes at least one of growth history information, growing season length, soil condition, precipitation level, sunlight amount, environmental temperature, growing region, and plant type for the one or more second plants.

15. A system for predicting latent infection in plants, comprising: one or more processors; one or more computer readable storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform actions in accordance with the method of any one of claims 1 to 14; A system comprising:

Citation Information

Patent Citations

  • Disease detection in plants

    US20140127672A1

  • Gene expression monitoring for risk assessment of apple and pear fruit storage stress and physiological disorders

    US20170260586A1