A method and system for differential crop measurement
Patent Information
- Application Number
- PCT/FI2026/050142
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-24
- Publication Date
- 2026-10-01
Smart Images

Figure FI2026050142_01102026_PF_FP_ABST
Abstract
Description
[0001] A METHOD AND SYSTEM FOR. DIFFERENTIAL CROP MEASUREMENT
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to agricultural research and plant breeding technologies, and more specifically, to a method and system for differential crop measurement.
[0004] BACKGROUND
[0005] The current state of the art in field-based plant breeding trials relies heavily on large, replicated plots to measure yield and other crop performance traits. Such plots are typically based on traditional field trial designs, where multiple small plots, typically ranging from 0.25 m2to 20 m2, are established, each containing plants of known genotypes and / or experimental treatments. Researchers use these plots to monitor agronomic performance and compare differences among genotypes, treatments, or environmental conditions across locations. Reference plots or checks containing well-known genotypes have commonly served as baselines for such comparisons. However, these baselines are derived from plot-level data, which limits the precision of comparisons due to spatial averaging and environmental variability within plots.
[0006] These traditional methods require a substantial volume of seed and space, particularly in multi-environment trials (METs), to account for the inherent spatial variability in field conditions, especially soil-related variation. Consequently, early-stage breeding generations, where only limited seed quantities are available, cannot be properly evaluated for yield, delaying the identification and promotion of promising genotypes.Attempts to control for environmental noise have typically included the use of reference plots or check varieties, but these are assessed at the plot level rather than the individual plant level. Soil heterogeneity, even within small plots, introduces confounding effects that reduce the reliability of such comparisons. Moreover, the physical and labor-intensive nature of individual plant-level assessments has rendered high-resolution phenotyping infeasible until recently.
[0007] A major shortcoming of these traditional approaches lies in their inability to control for micro-scale soil variation, which introduces substantial noise that can mask genuine genetic or treatment-related effects. Even within a single small plot, variability in soil properties, such as nutrient availability, moisture content, or compaction, can significantly influence plant growth and performance. As a result, plotwide comparisons either fail to detect genetic differences or require large numbers of replicates to statistically compensate for spatial noise.
[0008] The concept of accounting for such variability by placing reference plants immediately adjacent to test plants has been known, but hasn't been practically feasible. This close proximity allows both plants to grow under nearly identical environmental conditions, especially soil-related ones, thus reducing confounding factors. However, collecting accurate plant-level data in such paired designs has so far been considered impractical due to the high labor costs and narrow field observation windows. Manually recording traits for hundreds or thousands of plants and their references would require extensive manpower, making it economically unfeasible.
[0009] While recent advances in machine vision and automated data analysis have made it possible to efficiently extract plant-level traits, conventional trial plot layouts do not fully leverage this capability.Although imaging systems can now capture detailed phenotypic data across large areas, traditional spatial designs still do not support precise within-plot comparisons due to lack of reference.
[0010] The main limitation of current methods lies in their inability to isolate genetic or treatment effects from spatial soil variability within the plot, particularly when seed availability is limited. Given that the spatial resolution of soil-related noise often is less than 40 cm, different trial plots and even closely spaced plants may be subject to differing growth conditions, thereby undermining the accuracy of comparative analyses. Despite these technological advancements, a critical challenge remains unresolved in the context of modern plant breeding how to obtain accurate and reliable crop performance measurements at early stages of plant breeding when only a limited number of seeds are available, while minimizing the confounding effects of spatial soil variability in field trials that traditionally require large replicated plots and high seed volumes. In addition to this, within-plot references can be used for also other than breeding-related crop research and development. Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with conventional field trial designs and plot-level yield estimation.
[0011] SUMMARY
[0012] The aim of the present disclosure is to provide solution for accurate and reliable crop performance measurements at early stages of plant breeding when only a limited number of seeds are available, while minimizing the confounding effects of spatial soil variability in field trials that traditionally require large replicated plots and high seed volumes.The aim of the disclosure is achieved by a method and system for differential crop measurements as defined in the appended independent claims to which reference is made to. Advantageous features are set out in the appended dependent claims.
[0013] More specifically, the aim is achieved by a structured trial plot design where each test plant is placed in close proximity to a reference plant, ensuring they share similar environmental conditions, and the use of high-resolution imaging and automated analysis to extract and compare plant traits of the test plants with the plant traits of the reference plants efficiently.
[0014] The advantage of the present method and system is that it enables accurate and reliable plant-level performance comparisons by pairing test plants with closely positioned reference plants and leveraging high-resolution imaging and automated analysis.
[0015] The method and system for differential crop measurements provide significant advantages over traditional field trial methodologies, particularly in early-stage plant breeding where seed availability is limited. By combining a structured plot design, wherein each test plant is paired with one or more neighboring reference plants placed within a predefined comparison distance, with high-resolution imaging technologies, automated plant data collection, and computational trait analysis, the system enables accurate and reliable plant-level comparisons.
[0016] This close proximity between test and reference plants ensures that both experience nearly identical environmental conditions, thereby minimizing the confounding effects of spatial soil variability. The system's ability to perform automated extraction of crop performance indicators, such as yield components, phenology, disease symptoms,and nutrient status, allows for high-throughput, objective, and reproducible phenotyping.
[0017] The resulting differential crop measurements provide precise insights into genetic or treatment effects, even when derived from a limited number of plants. This increases statistical power, reduces the need for large replicated plots, and facilitates earlier and more informed selection decisions. By enhancing the signal-to-noise ratio through spatial pairing and automated analysis, the method overcomes fundamental limitations associated with conventional plot-level trial designs and supports more efficient and scalable breeding pipelines. The method and system may also be applied to other types of fieldbased plant research where environmental variability and limited seed availability limit the accuracy of performance comparisons.
[0018] Additional aspects, advantages, features and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative embodiments construed in conjunction with the appended claims that follow.
[0019] BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the embodiments of the disclosure are shown in the drawings, with references to the following diagrams wherein:
[0021] Fig. 1 is a schematic top-down view of structured plot design.DETAILED DESCRIPTION OF EMBODIMENTS
[0022] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. The present disclosure provides a method, system and a computer program for differential crop measurements. The method and system according to the present disclosure significantly improves the accuracy and reliability of crop performance measurements in early-stage plant breeding by enabling meaningful plant-level comparisons with minimal seed input, while effectively minimizing the confounding effects of spatial soil variability through the use of structured plot designs and high-resolution automated phenotyping.
[0023] The term " Confounding effects" throughout the description refers to environmental influences, particularly those arising from spatial soil variability such as differences in nutrient availability, moisture content, compaction, or microclimatic conditions, that hide or distort the relationship between a plant's genotype or treatment and its observed phenotypic traits. In the context of the present disclosure, confounding effects are minimized by placing each test plant in close proximity to one or more neighboring reference plants so that both experience similar local conditions, allowing for more accurate and reliable differential crop measurements.
[0024] The term " Meaningful", in the context of comparing plant traits or crop performance, refers to comparisons that are statistically and biologically relevant, such that the observed differences between test and reference plants can be attributed with confidence to genetic or treatment-related factors rather than to uncontrolled environmental variability. A meaningful comparison thus provides data that are useful for making selection or evaluation decisions in a breeding or research context.In an aspect, the present disclosure provides a method for differential crop measurements. The method comprises steps of: defining a comparison distance in a structured plot design for pairing one or more test plants with one or more neighboring reference plants that share the same micro-environmental conditions as the one or more test plants; sowing the one or more test plants and the one or more neighboring reference plants in the structured plot design where each test plant has one or more neighboring reference plants within the defined comparison distance to provide a reference comparison for the one or more test plants in the structured plot design; collecting plant data for the one or more test plants and the one or more neighboring reference plants; analyzing the collected plant data to extract plant traits related to crop performance indicators for the one or more test plants and for the one or more neighboring reference plants; computing differential crop measurements by comparing the extracted plant traits of each test plant of the one or more test plants with the extracted plant traits of its one or more paired neighboring reference plants.
[0025] The method enables to obtain accurate and reliable crop performance measurements at early stages of plant breeding, when seed quantities are highly limited, by introducing a trial design that enables precise plant-level comparisons while minimizing the confounding influence of spatial soil variability. By leveraging a structured trial layout, close proximity reference-based comparison, and automated data collection and analysis the method enables to achieve high-resolution measurement accuracy with significantly fewer seeds.
[0026] The first step, defining a comparison distance in a structured plot design for pairing one or more test plants with one or more neighboring reference plants that share the same micro-environmental conditionsas the one or more test plants, establishes the spatial configuration necessary to ensure that the test and reference plants grow in nearly identical environmental conditions. The comparison distance is defined based on crop type. For example, for cereals, this typically means the spacing between sowing rows plus a margin of approximately 3 cm, amounting to a total of 10 – 30 cm. This corresponds to one sowing row distance, a range within which the soil conditions, such as nutrient availability and moisture, can be assumed to be highly similar. This proximity reduces the variance caused by micro-scale soil heterogeneity and ensures that environmental differences between the test and reference plants are negligible. Optionally, also the distance of more than one sowing row and a margin of approximately 3 cm may be used the comparison distance.
[0027] The second step, sowing the one or more test plants and the one or more neighboring reference plants in the structured plot design, where each test plant has one or more neighboring reference plants within the defined comparison distance to provide a reference comparison for the one or more test plants in the structured plot design, implements the defined comparison distance by physically arranging the plants such that every test plant is neighbored by one or more neighboring reference plants within the comparison distance. The structured plot design ensures consistent spacing and allows for the reduction of border effects by including padding rows if necessary. This setup facilitates paired comparisons that block out the influence of local environmental noise, especially from soil heterogeneity, thereby increasing the accuracy of phenotypic comparisons even with a limited number of test plants.
[0028] In the third step, collecting plant data for the one or more test plants and the one or more neighboring reference plants, high-resolutionimage or video data is captured using cameras, hyperspectral sensors, or other imaging devices. These data streams record visible plant traits at high accuracy across all plants in the structured plot. Because the spatial layout is known, each plant can be automatically identified and assigned to its respective category, test plant or reference plant.
[0029] The fourth step, analyzing the collected plant data to extract plant traits related to crop performance indicators, involves the use of computational tools and machine vision algorithms to segment and quantify specific plant features. Crop performance indicators may include for example number and size of yield components (such as spike count), biomass, phenology (e.g., flowering time), disease infection (e.g., lesion area), and nutrient deficiency (e.g., through reflectance data). Importantly, spatial estimates of these traits are also derived, enabling not just raw trait comparisons but also distributional analyses across the plot.
[0030] The fifth step, computing differential crop measurements by comparing the extracted plant traits of each test plant of the one or more test plants with the extracted plant traits of its one or more paired neighboring reference plants, enables to isolate the genetic or treatment-related performance differences from the background noise introduced by soil heterogeneity. This results in a differential crop measurement that more accurately reflects the performance of the test plant, such as its genetic potential or treatment response, by minimizing the influence of environmental variability, particularly from soil differences. In the case of yield component comparisons, while processing aims at obtaining plant-specific estimation, mapping each measured yield component to a known plant individual is not necessary as comparisons can be made directly between the estimated yieldcomponents or their spatial distributions without knowledge of the exact plant individual that they relate to.
[0031] This method thus enables significant reduction in the number of seeds required for yield estimation; ability to conduct meaningful performance evaluations at early breeding stages; enhanced precision in trait measurement by controlling for environmental variation; scalability of automated phenotyping through high-resolution video intelligence; facilitation of early and more accurate selection decisions in breeding pipelines. This method allows small-scale, high-resolution trials to serve as effective proxies for large, replicated plots, thereby transforming early-stage phenotyping into a precise and resource-efficient process.
[0032] The plant data comprises high-resolution video of the sown one or more test plants and one or more neighboring reference plants, one or more images of the sown one or more test plants and one or more neighboring reference plants or hyperspectral measurements of the sown one or more test plants and one or more neighboring reference plants.
[0033] Using a high-resolution video, one or more still images, or hyperspectral measurements of the sown test and reference plants enable accurate and efficient extraction of plant traits necessary for differential crop measurements. By using any one of these imaging methods, the method remains flexible and adaptable to various operational contexts and resource constraints. For example, high- resolution video provides a rapid, continuous way to capture spatial and morphological details of plants along sowing lines and supports automated processing for dynamic traits such as tiller development or canopy coverage. Alternatively, still images offer precise, high-quality snapshots that are particularly effective for analyzing specific traits atkey growth stages, such as spike count or disease symptoms. Hyperspectral measurements, on the other hand, enable the detection of subtle physiological changes such as chlorophyll concentration, water stress, or early disease symptoms, which are not visible in standard RGB imagery.
[0034] Each of these plant data capturing methods supports automated, objective, and high-resolution plant-level phenotyping, making it possible to derive meaningful trait data from small numbers of plants with minimal manual labor. This is especially critical at early breeding stages when seed availability is limited and traditional large-plot evaluations are not feasible.
[0035] The technical effect of this method is that reliable crop performance measurements are obtained as differential trait estimates in field conditions despite the variability of soil. The ability to compare test and reference plants at high resolution, regardless of the specific imaging method used, helps minimize the confounding effects of spatial soil variability and increases the reliability of the performance estimates. Thus, it ensures that even minimal field trials with limited seed input can yield precise and useful crop performance data.
[0036] Optionally, the structured plot design may include padding on the sides of the structured plot. The inclusion of padding on the sides of the structured plot design refers to additional rows of plants, it may comprise the same genotype as the reference plants, surrounding the structured plot that contains the test and reference plants intended for measurement. This enables to mitigate the border effects, which are sources of bias in field trials. Border effects occur when plants located at the border of a plot experience different environmental conditions, such as increased light exposure, wind, or reduced competition for resources, compared to those in the interior. These uncontrolledvariables can distort plant growth and performance, introducing additional variability unrelated to the genetic or treatment effects under study.
[0037] In small-scale structured plots used in early-stage breeding, where only a limited number of seeds are available and accurate plant-level measurements are essential, minimizing such non-genetic sources of variation becomes critically important. By including padding rows, the test and reference plants are buffered from these border effects, ensuring that their growth conditions more accurately reflect the intended production environment.
[0038] The padding thus improves the similarity between the central measurement area of the plot and production use conditions. This increases the reliability of the differential crop measurements between test and reference plants as an indicator for production performance in conditions where border effects are not present. In production conditions, each plant grows surrounded by other plants and thus needs to compete for resources. Near to the borders of a trial plot there are more resources available than in the canopy further away from the borders. By ensuring that there is no border close to the region of the plot where the one or more test plants and the one or more neighboring reference plants are growing, the comparative crop performance measurements will resemble production conditions to a further extent. The fundamental objective is to estimate production performance and this is hence very desirable.
[0039] The inclusion of padding in the structured plot design enhances the accuracy and consistency of crop performance measurements by minimizing the impact of border effects. This supports reliable phenotyping under conditions where traditional replicated large plots are impractical due to seed or space limitations.The one or more neighboring reference plants may be selected based to match the expected development speed (phenology) and height of the test plants where the expected height and development speed for the test plants are obtained by computational methods such as genomic prediction or pedigree-based prediction based on a statistical framework such as but not limited to best linear unbiased prediction. The method may further comprise using large conventional test plots for measuring trait data, and predicting crop yield of the test plant based on the differential crop measurement and trait data obtained from large conventional plots of the one or more neighboring reference plants. Allowing the integration of large conventional test plots for measuring trait data, particularly traits such as yield, for the varieties and / or treatments used as reference plants provides an additional data layer that enhances prediction capability. Specifically, this approach leverages yield or other trait measurements collected from large, fully replicated plots, where reference genotypes are grown under conventional conditions and using these data as a calibration baseline. The trait data from these large plots, such as absolute yield values, are used to scale or interpret the differential crop measurements observed in the structured plot design. For example, if a test plant is observed to have 15% higher biomass than its paired reference plant, and the reference plant's yield is known from the large plot, a yield estimate for the test plant can be inferred using that differential and a calibrated relationship between biomass and yield.
[0040] This creates a bridge between relative plant-level comparisons and absolute trait values, allowing more robust estimation of performance traits even for genotypes that are only grown in the structured small plots. Importantly, this enables yield estimation and trait prediction for test plants from the differential crop measurements without requiringlarge replicated plots for each individual genotype, which would otherwise be impossible at early breeding stages due to limited seed availability.
[0041] This provides a significant improvement in the accuracy and interpretability of the differential crop measurements. It enhances the method's ability to produce reliable trait predictions by anchoring relative plant-level differences to empirically validated absolute values. This enables early-stage genotypes, tested only in the structured plots with a minimal number of seeds, to be evaluated with a level of confidence approaching that of full-scale yield trials.
[0042] The optional use of large conventional test plots enhances thus the method by enabling trait calibration and yield prediction based on the differential measurements. This supports obtaining accurate and reliable crop performance data early in the breeding process, while maintaining minimal seed use and controlling for spatial soil variability. In another aspect, the present disclosure provides a system for differential crop measurements. The system comprises a plant data collecting means configured to collect the plant data from one or more test plants and one or more neighboring reference plants sown in a structured plot design, wherein each test plant is paired with one or more neighboring reference plants that share the same micro- environmental conditions as the one or more test plants within a comparison distance; a computing device configured to receive and analyze the collected plant data to extract plant traits related to crop performance indicators for the one or more test plants and the one or more neighboring reference plants, and to pair each test plant with its one or more neighboring reference plants, and to compute differential crop measurements by comparing the extracted plant traits of each test plant of the one or more test plants with the extracted plant traitsof its paired one or more neighboring reference plants. The system for implementing the method of differential crop measurements enables to obtain accurate and reliable crop performance measurements at early stages of plant breeding when only limited seed quantities are available, while minimizing the confounding effects of spatial soil variability in field trials.
[0043] The plant data collecting means configured to capture data from one or more test plants and one or more neighboring reference plants grown in a structured plot design may be high-resolution video camera, imaging rigs, handheld devices, drones, or hyperspectral sensors, depending on the type of plant data being collected. It enables efficient, high-throughput, and non-destructive acquisition of plant-level data, capturing visual and spectral information necessary to quantify crop performance indicators such as yield components, phenology, disease symptoms, or nutrient status. This automated and scalable data collection is crucial when working with small plots and limited seeds, where manual data acquisition would be impractical or prohibitively labor-intensive.
[0044] For example, if the type of plant data is high-resolution video, the plant data collecting means may be a handheld video camera, a rig-mounted recording system, or a mobile platform (like a rover or drone) equipped with video capture. If the type of plant data is one or more still images, the collecting means may be a digital camera system mounted on a frame or carried manually to capture the images of specific plant sections. If hyperspectral measurements are used, the collecting means may be specialized sensors capable of detecting reflectance across a broad range of wavelengths, mounted on a platform that can scan plots with high spatial and spectral precision.The computing device is configured to receive the collected plant data and perform automated analysis to extract relevant plant traits from both the test and reference plants. This comprises the identification and segmentation of individual plants and the calculation of trait values using machine vision or statistical algorithms. The computing device is also responsible for computing differential crop measurements by comparing the traits of the test plants to those of their neighboring reference plants within the structured plot. This comparison is conducted in a way that controls for local environmental variability, especially soil-related differences, by relying on the close proximity of test and reference plants.
[0045] The system enables a precise, automated, and scalable framework for obtaining high-quality crop performance data from minimal field trials. It ensures that comparisons are made at the plant level under nearly identical environmental conditions, thereby significantly reducing the impact of spatial noise. Moreover, the system allows this process to be executed efficiently across hundreds or thousands of plants, making it feasible even under the time constraints typical of narrow field observation windows.
[0046] The system provides the technical infrastructure to support accurate, high-resolution, and resource-efficient plant-level phenotyping. It enables early-stage crop performance evaluation with minimal seed input and high reliability, directly addressing the challenges posed by traditional field trials that depend on large replicated plots and high seed volumes.
[0047] In another aspect, the present disclosure provides a computer program differential crop measurements, wherein the computer program comprises instructions which, when executed by a computing device, cause the computing device to receive plant data collected from one ormore test plants and one or more neighboring reference plants sown in a structured plot design in which each test plant is assigned one or more neighboring reference plants within a comparison distance, extract plant traits, compute differential measurements by comparing extracted plant traits of each test plant with extracted plant traits of its assigned one or more neighboring reference plants.
[0048] This computer program comprises instructions that, when executed by a computing device (such as a laptop, field computer, smartphone or server), automate the core analytical steps of the method. First, the program receives plant data that have been collected from a structured plot design, where each test plant is paired with one or more closely spaced reference plants. This data may consist of video frames, still images, or hyperspectral measurements as described in earlier claims. The software then processes this data to extract plant traits related to crop performance indicators, including yield components, phenology, disease symptoms, and nutrient status. These traits are derived using image processing, machine vision, or statistical algorithms embedded in the program. This automated extraction is critical for enabling scalable, objective, and reproducible phenotyping across a large number of plants within a short observation window.
[0049] Next, the program performs computational comparisons between each test plant and its corresponding reference plant(s), producing differential crop measurements. These comparisons are based on relative differences or ratios of trait values, and they are designed to minimize the impact of local soil heterogeneity by ensuring that both test and reference plants share the same micro-environmental conditions.The computer program enables the automation of high-resolution trait analysis, reduction in human error and labor costs, and the ability to generate statistically robust crop performance estimates from small, spatially sensitive datasets. This is essential for enabling the method and system to be applied efficiently and reliably in practical breeding programs, even under conditions where traditional large-scale field trials are not feasible.
[0050] The computer program ensures the method's scalability, repeatability, and computational precision, thereby supporting accurate crop performance measurement with minimal seed input and improved resistance to environmental noise.
[0051] Additional advantages of the method and system for differential crop measurements disclosed herein include the integration of synthetic data-enabled video intelligence, which improves trait measurement accuracy at scale. The approach enables clear visibility of the plant canopy while capturing high-resolution video footage, even using simple devices such as a smartphone.
[0052] The system addresses a longstanding bottleneck in acquiring annotated training data for machine vision algorithms by introducing a synthetic data pipeline. This pipeline enables training of accurate plant segmentation models using only tens of manually annotated images. While this significantly reduces cost and development effort, all plant data used in differential measurement is always collected from real-world field conditions.
[0053] The method allows the extraction of more accurate plant trait data for each individual sowing row. The resulting dataset provides a surprisingly detailed and reliable description of the canopy structureand trait variation, sufficient to answer complex agronomic questions using minimal seed quantities.
[0054] The present method helps to address for example a question, what is the spatial resolution of soil-related environmental noise? In standard field trials, it is common to use approximately 2400 seeds in an 8 m2plot. Yet from a statistical perspective, only about 900 seeds should be sufficient to identify high-performing lines in a single environment, assuming realistic noise levels. This discrepancy points to unaccounted spatial variability.
[0055] Using a trait with low to medium heritability, such as yield, simulations show that after large-scale spatial correction, the genetic effect (G) typically explains only 40–60% of total variance. The rest is attributed to noise, largely due to spatial soil variation.
[0056] To better understand the spatial scale of this variability, we simulated yield noise effects under different soil block sizes and examined the resulting correlation between genotype and yield performance. When noise blocks were simulated at 10×10 cm, 20×20 cm, or even 30×30 cm, the number of blocks per plot was large enough for noise to cancel out, and yield data remained well-correlated with genetic effects.
[0057] However, problems consistent with real-world data appear when the spatial scale of soil noise reaches ~50 cm. At that point, each plot accommodates so few noise blocks that the noise no longer cancels out effectively. This explains why simply reshaping plots (e.g., into wide rectangles) would not help: such layouts would capture only one or two noise blocks, rendering the data unreliable. In contrast, long and narrow plots intersect multiple noise zones, enabling averaging and better noise control.The proposed solution is thus based on a within-plot reference design, where each test plant is grown in close proximity (e.g., one or two sowing rows apart) from one or more neighboring reference plants. For each unit of row length, the method allows direct comparison of plant traits between the test plant and the reference plant. These comparisons are made within the comparison distance, which is carefully selected to ensure that both plants are subject to highly similar soil conditions. This effectively blocks the influence of soil variability, allowing trait differences to be attributed more confidently to genetic or treatment effects.
[0058] Statistically, this design aligns with the logic of paired test setups (e.g., paired t-tests), which benefit from high correlation in the noise affecting both units of a pair. If we normalize yield variance to 1 after spatial correction, and assume that G explains 40% of this variance, then, given a soil-related micro-scale correlation of 0.80 between paired plants spaced 20 cm apart, the effective residual noise drops from 0.60 to 0.12. This improves the signal-to-noise ratio from 0.40 to 0.76, nearly doubling measurement efficiency.
[0059] In practical terms, this means that instead of needing 900 seeds per variety to achieve statistical confidence, the method could achieve the same with approximately 170 seeds per variety. This has profound implications for breeding programs. Potential applications for plant breeding may be for examples as follows: obtaining yield-related trait data already at the F5 generation for every genotype; enabling multienvironment testing (MET) at the F6 generation, when seed volumes are still limited; gaining deeper understanding of soil heterogeneity and its impact on plant performance.
[0060] From a genomics and GxE modeling perspective, one of the common frustrations is that yield data is typically collected from only a limitedsubset of genotypes. Expanding coverage to 10 times more genotypes using this method would drastically improve the quality of prediction models. It would also enable a new trial strategy, planting in more locations than are ultimately phenotyped, and then choosing to collect data only from plots and environments that yield the highest-quality differential measurements.
[0061] Interestingly, areas of the field previously dismissed as "noise zones" could become the most informative regions. With high-resolution differential data, it becomes feasible to ask why a plant grows well or poorly at specific points, and to conduct targeted soil sampling at very fine spatial scales.
[0062] To ensure the validity of plant-to-plant comparisons, it is important that the reference plant and the test plant have similar phenology and plant height, minimizing shading and growth-stage mismatches. These traits typically have high heritability, meaning that suitable reference genotypes can be identified using computational prediction methods such as A-BLUP. The differential crop measurements focus primarily on traits such as biomass and yield components, rather than total yield. However, because harvest index is often more heritable than yield itself, it can serve as a reliable basis for scaling differential traits to yield predictions. Genomic prediction can further enhance this process. Border effects, although still possible, can be mitigated effectively by including additional rows of reference plants on both sides of the testreference pairing. Harvesting logistics may require extra care for genotypes of interest, but this is a manageable challenge.
[0063] Thus, using reference plants within structured plots to quantify and control for spatial variation enables precise, early-stage, and scalablecrop performance evaluation, even when seed availability is very limited.
[0064] EXAMPLES
[0065] Example 1
[0066] Crop: Wheat
[0067] Structured plot design: One or more test plants are sown in a sowing row located approximately 20 cm apart from one or more neighboring reference plants sown in an adjacent row.
[0068] Plant data collection: A plant data collecting means, comprising a camera, captures high-resolution video (4K) while being moved at approximately 1.5 km / h along the structured plot.
[0069] Padding: Two additional sowing rows containing filler plants are included on both sides of the structured plot to reduce border effects. Data analysis: A computing device executing a neural-network-based computer program processes the video stream to segment individual plants, extract plant traits such as plant height and number of spikes, and compute differential crop measurements by comparing each test plant with its neighboring reference plant(s).
[0070] Result: The relative differences in total biomass between test and reference plants could be reliably measured using data from only 200 seeds, whereas standard plot-level comparisons typically require more than 1000 seeds to achieve comparable statistical power.
[0071] Example 2
[0072] Crop: Corn or Soybean
[0073] Structured plot design: Row spacing is adjusted according to the crop, 70–75 cm for corn and 45–50 cm for soybeans, with test plants and reference plants placed in adjacent rows within the defined comparison distance.Plant data collection: Hyperspectral measurements are used as the plant data, captured via a hyperspectral sensor-based plant data collecting means.
[0074] Data analysis: The computing device extracts spectral traits from the test and reference plants. Early signs of nutrient deficiency are detected using differences in near-infrared reflectance.
[0075] Result: The differential crop measurements showed strong correlation (r=0.85) with SPAD meter readings for chlorophyll content, confirming the system's ability to detect physiological differences at early growth stages.
[0076] Example 3
[0077] General implementation: The method and system for differential crop measurements are applied to both small grain cereals (e.g., wheat, barley) and row crops with wider row spacing (e.g., corn, soybean). The structured plot design, use of high-resolution imaging or hyperspectral sensing as plant data collecting means, and automated analysis via a computing device enabled accurate crop performance estimation across varying crop types and plot geometries, demonstrating the method's adaptability and scalability in real-world breeding trials.
[0078] The method and system leverages high-resolution imagery (e.g., RGB video, hyperspectral images) and computational methods to extract traits from each plant in the trial. Automated image analysis identifies each plant, quantifies relevant traits (e.g., plant height, leaf area, disease symptoms, ear / grain metrics, yield components, other morphological features), and compares the test plant's trait values to the corresponding one or more neighboring reference plants. This plant-by-plant comparison reduces the confounding effect of micro-environmental variability that is common in agricultural fields. Insteadof comparing one plant to another, spatial distributions of the traits can be estimated and comparison can be done between the spatial distributions.
[0079] The reference plant genotype is chosen, or predicted, based on predicted growth parameters, such as height and phenology to ensure that it does not overshadow the test plant or diverge excessively in growth timing. The resulting combined dataset enables differential crop measurements to be calculated, facilitating genotype or treatment performance assessment in early-generation testing, advanced plant breeding triage, or product-level evaluations (e.g., biostimulants or pesticides). Development speed (phenology) and height have typically high heritabilities and such traits based on a measured genotype of progeny information with methods such as G-BLUP or A-BLUP, respectively can be routinely predict. Also other machine learning approaches can be used.
[0080] In one embodiment, test plants are interspersed with reference plants at carefully controlled distances (e.g., on the order of centimeters rather than meters). In this way, each test plant is "paired" with at least one reference plant that experiences very similar environmental conditions (e.g., soil composition, localized moisture, microclimate, etc.) for direct comparison.
[0081] In another embodiment, the test plants are grown in one sowing line whereas the reference plants are grown in the neighboring sowing line (parallel to the sowing line of the test plants) located approximately 10-20 cm away.
[0082] Beyond the test- reference comparison at the plant level, the method and system contemplates integrating traditional large-field plot data to further predict yield or agronomic performance. By correlating theplant-specific differential measurements (or comparisons of spatial distributions) and the conventionally measured large-plot data collected for the reference plants, and integrating information such as the dependency between biomass and yield the method allows e.g the prediction of yield (or other agregate measurement) also for the test plants: for example the yield of a test plant can be predicted after observing relative increase of 15% in biomass as compared to the reference plant, predicted relationship between biomass and yield of the test plant and the observed yield (from large plot) of the reference variety.
[0083] More specifically, in one embodiment, large-plot yield data are first obtained for the reference genotype(s) or treatment(s) grown under conventional conditions (e.g., in multiple rows or larger trial areas). The trait estimates, such as biomass, plant height, leaf area, or other morphological and physiological indicators, are then compared between the test plant and its local reference plant to yield a relative difference or ratio (i.e. the differential crop measurement). This relative difference is converted into a predicted yield value for the test plant by applying a scaling factor or function derived from either historical data or concurrent calibration measurements. For instance, if the test plant's biomass is observed to be 15% higher than the reference under similar management, a 15% yield increment can be inferred if the harvest index or yield-forming physiology is known or assumed to be similar. Where genetic factors (e.g., a known genotypespecific harvest index) influence the ultimate yield response, the system can incorporate these factors into a linear, polynomial, or machine-learning-based model (such as partial least-squares regression, ridge regression, random forests, or neural networks). In some cases, data from multiple growth stages may be used to refine the yield prediction progressively through the season. By integratingmulti-environment data or repeated measurements, the method and system supports robust predictive modeling of yield or other aggregated traits for the test plants even when seed availability or replicated trial space is limited.
[0084] In one embodiment, agricultural cereal crops (e.g., wheat, barley, rice) are used in a field trial. Test seeds are sown in rows no more than ~20 cm away from a reference plant row, effectively ensuring that each target test plant is within a distance (including row spacing + an additional margin of about 3 cm) that statistically justifies the assumption of nearly identical soil conditions.
[0085] To collect sufficiently high-resolution plant-level data, a person or automated platform (e.g., a rover, tractor-mounted camera, or drone) traverses the field, capturing video or images of each plant. The system may employ RGB, multispectral, or hyperspectral sensors to measure reflectance or color cues relevant to plant health, disease, or development. Simple manual tools such as imaging rigs may be used to improve visibility to all above ground parts of the canopy.
[0086] The captured data are then transferred (wirelessly or via physical media) to a remote or on-site computing system. Machine vision algorithms segment and identify each plant and plant component individually, storing various morphological or spectral features. The system then matches each test plant with its nearest reference plant(s), depending on the margin set. A suite of plant traits (e.g., height, biomass, number and size of yield components, leaf color, disease lesion area) is calculated for both sets of plants. The differential measurement is computed as the trait value for the test plant minus or divided by the trait value for its paired reference plant(s). Multiple reference plants for a single test plant may be averaged or otherwise combined statistically to form a robust reference trait value.By comparing test plants to these local references, the noise induced by field heterogeneity is substantially reduced. Hence, relatively small differences in performance or morphology may be detected more reliably. The method also contemplates "padding" plants to reduce border effects. For instance, if the test and reference pairs are near the plot boundary, dummy or filler plants may be sown around the perimeter of the plot to create uniform micro-environment conditions (e.g., minimizing unrealistic resource availability at the edges).
[0087] In certain embodiments, the method and system can be applied to plant breeding programs, wherein the reference plants correspond to well-known or check varieties grown under identical management conditions. By comparing test plants (e.g., newly created lines or hybrids) to these reference varieties in very close proximity, breeding decisions can be made more accurately based on genetic differences rather than environmental noise.
[0088] In other embodiments, especially for evaluating products such as biostimulants, the test plant and its reference plant may be of the same genotype but receive different treatments, for example, one is treated with a specific biostimulant while the other is left untreated. Such a design isolates treatment effects on plant performance by minimizing confounding variables in soil and microclimate.
[0089] In an extended approach, large-scale yield data can be collected from conventional harvest plots in the same trial or from related fields. Statistical models (e.g., best linear unbiased prediction or genomic prediction) can then correlate the differential measurements with end-of-season yields to refine or predict performance at scale and use computational methods to convert yield for reference plants measured from large plots into yield estimates of test plants. For example, biomass correlates with yield. If a test plant has 15% higher biomassthan its reference plant (measured from the large plot), a 15% higher yield can be expected given a similar harvest index.
[0090] In practice, the correlation can involve combining multiple plant-trait measurements, e.g., leaf area index, canopy temperature, chlorophyll content, and mapping them against yield outcomes observed in a fully replicated large plot for the same reference genotype(s). By modeling the relationship (e.g., via a linear regression, mixed-model regression, or advanced machine-learning algorithm), differing trait values in the test plant can be used to infer adjusted yield expectations. Furthermore, genotype-specific or environment-specific calibration constants (such as harvest index ranges, stress tolerance factors, or phenological rates) can be input into the model to address biological or climatic variations. Repeated trait measurements at different growth stages may be integrated for dynamic updating of the yield prediction. Multi-location data can also be pooled to account for genotype-by-environment interactions, further refining yield estimates for varied field conditions. Thus, even with minimal seed availability or limited replication, the method and system enable robust yield predictions for test genotypes or treatments, greatly enhancing the utility of early-phase trials.
[0091] The present method and system is valuable in particular in the early stages of the breeding program where the availability of seeds is very low for a new genotype (variety candidate). In such a situation, obtaining a useful yield estimate is not viable. This method and system allow obtaining differential estimates that are more reliable and allow evaluating also yield traits earlier. Also in the second year of the breeding program the quantity of seed available for each genotype will still remain so limited that testing only at one location is feasible. The method and system enable multi environment testing at stages whereconventionally the availability of seed has enabled making a single test plot (large-enough for yield estimation) at a single location.
[0092] A variety of approaches can be employed to recognize and quantify each plant in a high-throughput manner, ranging from classical computer vision (CV) methods to deep learning techniques.
[0093] In one scenario, classical CV algorithms (for example, those using color segmentation, morphological operations, feature extraction, and template matching) can detect and segment plant structures in well-illuminated images. Such methods might incorporate image filtering, thresholding, and edge detection (e.g., Canny edge or Sobel filters) to isolate plant regions from soil background. Alternatively, convolutional neural networks (CNNs) and other deep learning architectures (such as U-Net for segmentation or YOLO for object detection) can be used to directly segment or locate each plant in RGB or multispectral frames. A well-trained CNN can also classify various morphological features, such as leaf count or disease lesions, at the pixel or bounding-box level. Hyperspectral data can be processed through modern machinelearning pipelines, e.g., 3D CNNs or specialized architectures that fuse spectral and spatial information, enabling more nuanced trait inference for individual plants. These algorithms can run on specialized hardware accelerated by graphics processing units (GPUs) or tensor processing units (TPUs) for real-time or near-real-time data analysis. For portable field applications, a light-weight inference model can be used on embedded systems, while the bulk of training and large-scale data analytics may occur on remote servers. Depending on the farm or research station's connectivity, data can be streamed locally or uploaded to a cloud-computing environment for further analysis. Details such as calibration, lens distortion correction, and sensor fusion (e.g., combining RGB and near-infrared channels) can be included toimprove accuracy. The particular choice of algorithm typically depends on factors such as the resolution and type of imaging sensor, computational resources available, and the complexity of the traits under investigation.
[0094] To implement the close-proximity test-and-reference design, specialized hardware may be employed for precisely placing different seeds in adjacent rows or lines. For instance, a "cassette planter" or a multi-bin precision seeder can hold multiple seed types and deposit them in separate lines with minimal row-to-row variation. These multicompartment sowing machines typically have independent metering systems or "hoppers" for the test seeds and for the reference seeds, respectively. By carefully calibrating the seed distribution rate, the operator can ensure that each sowing line remains uniform in seed density and spacing.
[0095] Additionally, physical guides or rods can be used to fix the distance between rows, ensuring that the test row is exactly 10-20 cm apart from the reference row. In some implementations, mechanical adaptors or extension arms are attached to conventional plot sowers, so a single pass can place test seeds in one furrow and reference seeds in a neighboring furrow. Also twin-row arrangements where each pass places two parallel lines may be used. In certain cases, a cassettebased planter is installed on a tractor-mounted rig that simultaneously manages the depth control, row spacing, and metering of at least two different seed types. This approach minimizes manual intervention and ensures replicable row spacing across the entire field. Where extremely high precision is needed, such as in advanced breeding research, small-scale manual or semi-automatic planters may be used, allowing the operator to feed test seeds and reference seeds individually into adjacent drills. These specialized devices, in combination with theadjustable row markers, guarantee that the distance remains stable, optimizing the local micro-environment for subsequent plant-level comparisons.
[0096] As in the early stages of plant breeding, the availability of sowing seed is a significant limiting factor for achieving genetic gains, this limitation appears in two ways. The number of seeds available is insufficient to enable conventional yield measurement for early-generation plants (e.g., F4 generation). By the time seed availability is sufficient for yield estimation, particularly in multi-environment trials (METs), only a small fraction of the initial germplasm remains, thereby limiting the diversity represented in the yield data. Both challenges can be overcome enabling reliable yield estimation using a minimal number of seeds. The present method and system enables yield estimation with limited seed input by combining accurate plant-specific phenotyping with a structured trial design that uses reference plants placed in close proximity to test plants. By modeling spatial variation, especially soil-related variation, this design allows plant-level comparisons that are not confounded by micro-environmental differences.
[0097] Since soil characteristics such as nutrient availability and moisture distribution are major determinants of yield, yet are difficult or impractical to measure directly at fine spatial scales, the system instead compares the performance of a test plant to a nearby reference plant that shares the same micro-environmental conditions. Simulations have shown that the spatial scale of soil variation often exceeds 40 cm. Therefore, by placing test and reference plants within 10-30 cm of each other, such as in adjacent sowing rows, reliable comparisons can be achieved.Furthermore, localized effects (e.g., from tractor tracks or minor soil anomalies) can be effectively neutralized if they impact both test and reference plants similarly, which is achieved through this spatial pairing.
[0098] To assess plant growth, the system employs visible yield component–based comparisons. These include traits such as spike number, tiller count, and biomass. Although exact grain number and grain size may not be directly observable, these visible traits offer a high correlation with biomass, and biomass in turn can be combined with a harvest index, often a highly heritable trait, to estimate total yield.
[0099] This method can also be used to quantify treatment effects or detect disease symptoms, provided appropriate management tools (e.g., localized sprayers) are available.
[0100] Another challenge is shading or competition effects due to differing development speeds or plant heights between test and reference plants. This is mitigated by selecting reference plants that are predicted to match the test plants in phenology and height. Traits such as flowering time and plant height are highly heritable and can be predicted using computational methods such as genomic prediction (e.g., A-BLUP or G-BLUP). Where performance comparisons are made between plants with very different developmental profiles, the results can still be adjusted statistically.
[0101] At early breeding stages, selection is typically intense, and yield data are only gathered for a narrow subset of the original genotypes. The present method significantly enhances genomic selection by enabling yield-related trait data collection for a broader range of genotypes. This increases the diversity and representativeness of the genomic training set. Additionally, this approach allows yield estimation and evenpreliminary multi-environment testing (albeit with slightly higher noise) at stages where traditionally only one large test plot could be allocated due to seed limitations.
[0102] The present method and system allow replacement of the small plot phase, where yield is typically not measured, with a structured, reference-based design supported by automated phenotyping. This can reduce the amount of seed needed for statistically significant yield comparisons by more than 90%, enabling early-stage genotypes to be assessed for yield potential with high resolution and minimal input. In an example of representative trial setup a 10-meter plot with 6 sowing rows is planted for each of 50 test varieties. One row in each plot contains the variety of interest; the adjacent rows contain reference plants. The goal is to reach visible spike development by demonstration day. Reference plants are selected to match the test varieties in phenology and plant height. Large replicated plots of the reference varieties are included elsewhere in the field to support calibration and yield prediction.
[0103] The purpose of such a trial is to assess whether the rank ordering of variety candidates relative to the reference variety, as determined using the proposed structured plot design, corresponds with the ordering observed in conventional large replicated plots. A strong correspondence would validate the predictive power and practical utility of the method.
[0104] In traditional field research, reference plots have been used for decades to create performance baselines. However, these baselines are typically generated at the plot level and do not capture plant-to-plant variation. The present method and system build upon this concept by placing reference plants immediately adjacent to test plants, thuscontrolling for soil variation and enabling meaningful plant-level comparisons.
[0105] The reference and test plants may differ in genotype or be genetically identical but subject to different treatments (e.g., with or without a biostimulant). Historically, collecting accurate plant-level data in such designs was not feasible due to labor constraints. Measuring even a few traits across hundreds of plots would require a massive workforce. With the advent of automated machine vision systems, such high-resolution phenotyping is now possible.
[0106] DETAILED DESCRIPTION OF THE DRAWINGS
[0107] Referring to Fig. 1, there is shown a schematic top-down view illustrating an example of a structured plot design for differential crop measurements. The figure shows multiple parallel sowing rows, where test plants (marked with asterisks *) are sown in a central row (Row 4), and reference plants (marked with circles o) are sown in neighboring rows (Rows 3 and 5) within a predefined comparison distance. Additional rows (Rows 1-2 and 6) are included as padding to minimize border effects and ensure uniform environmental conditions within the interior of the plot.
[0108] The imaging and computing device is positioned above the sowing rows and configured to collect plant data from the structured plot design. The system may include one or more cameras, sensors, or computing units capable of capturing high-resolution video, images, or hyperspectral measurements of the test and reference plants. The collected plant data are processed by the computing device to extract crop performance indicators and compute differential crop measurements by comparing each test plant with its neighboring reference plants).
Claims
1. CLAIMS1. A method for differential crop measurements, the method comprises steps of:- defining a comparison distance in a structured plot design for pairing one or more test plants with one or more neighboring reference plants that share the same micro-environmental conditions as the one or more test plants;- sowing the one or more test plants and the one or more neighboring reference plants in the structured plot design, where each test plant has one or more neighboring reference plants within the defined comparison distance to provide a reference comparison for the one or more test plants in the structured plot design;- collecting plant data for the one or more test plants and the one or more neighboring reference plants;- analyzing the collected plant data to extract plant traits related to crop performance indicators for the one or more test plants and for the one or more neighboring reference plants;- computing differential crop measurements by comparing the extracted plant traits of each test plant of the one or more test plants with the extracted plant traits of its one or more paired neighboring reference plants.
2. The method according to claim 1, wherein the plant data comprises high-resolution video of the sown one or more test plants and one or more neighboring reference plants, one or more images of the sown one or more test plants and one or more neighboring reference plants or hyperspectral measurements of the sown one or more test plants and one or more neighboring reference plants.
3. The method according to claim 1 or 2, wherein the structured plot design includes padding on the sides of the structured plot.
4. The method according to any one of the preceding claims, wherein the one or more neighboring reference plants are selected to match the expected development speed and height of the test plants where the expected height and development speed for the test plants are obtained by computational methods such as genomic prediction or pedigree-based prediction based on a statistical framework such as but not limited to best linear unbiased prediction.
5. The method according to any one of the preceding claims, wherein the method further comprises using large conventional test plots for measuring trait data, and predicting crop yield of the test plant based on the differential crop measurement and trait data obtained from large conventional plots of the one or more neighboring reference plants.
6. A system for differential crop measurements, the system comprises - a plant data collecting means configured to collect the plant data from one or more test plants and one or more neighboring reference plants sown in a structured plot design, wherein each test plant is paired with one or more neighboring reference plants that share the same micro- environmental conditions as the one or more test plants within a comparison distance;- a computing device configured to receive and analyze the collected plant data to extract plant traits related to crop performance indicators for the one or more test plants and the one or more neighboring reference plants, and to pair each test plant with its one or more neighboring reference plants, and to compute differential crop measurements by comparing the extracted plant traits of each test plant of the one or more test plants with the extracted plant traits of its paired one or more neighboring reference plants.
7. A computer program for differential crop measurements, wherein the computer program comprises instructions which, when executedby a computing device, cause the computing device to receive plant data collected from one or more test plants and one or more neighboring reference plants sown in a structured plot design in which each test plant is assigned one or more neighboring reference plants within a comparison distance, extract plant traits, compute differential measurements by comparing extracted plant traits of each test plant with extracted plant traits of its assigned one or more neighboring reference plants.