Seedling phenotype high-throughput analysis method and system based on time-series growth images
By acquiring images and performing time-frequency analysis multiple times during the seedling growth cycle, operational disturbances are identified and corrected, solving the problem of data distortion caused by disturbances in seedling cultivation and improving the accuracy and screening precision of seedling phenotypic data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI AGRICULTURAL UNIVERSITY
- Filing Date
- 2026-01-16
- Publication Date
- 2026-06-02
AI Technical Summary
During seed and seedling cultivation, repeated movement and imaging operations cause physical disturbance to the seedlings, leading to physiological responses, which affects the accuracy and reliability of time-series image data and reduces the accuracy of seedling vigor assessment and screening.
By acquiring images of the same batch of seedlings at multiple time points during the seedling growth cycle, a time-series growth image sequence is constructed, phenotypic parameters are extracted, time-frequency transformation and high-frequency energy ratio analysis are performed, operational disturbances are identified and corrected, and the accuracy of the data is ensured.
It effectively separates operational disturbance noise from real growth signals, improves the accuracy and reliability of seedling phenotypic data, and enhances the precision of seedling growth dynamic tracking and early screening.
Smart Images

Figure CN122134618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of seedling cultivation and phenotypic analysis technology, and more specifically, to a high-throughput analysis method and system for seedling phenotypic analysis based on time-series growth images. Background Technology
[0002] In the field of seed and seedling cultivation, especially in intensive seedling raising and variety selection, efficient and objective monitoring and analysis of seedling growth phenotypes are crucial. High-throughput analysis methods based on time-series growth images can periodically and automatically collect image sequences of seedling populations and utilize image processing and analysis techniques to non-destructively extract morphological and physiological phenotypic parameters such as plant height, leaf area, and color index. This enables dynamic growth tracking, early screening of superior individual plants, and assessment of biotic and abiotic stress responses. This technology provides important data support for modern precision agriculture and efficient breeding.
[0003] However, in the aforementioned high-throughput analysis process, to obtain time-series images, the seedling units (such as plug trays) need to be periodically moved to the imaging site. This repeated movement and imaging operation itself causes physical disturbance and environmental changes to the seedlings in sensitive growth stages, thereby triggering their transient physiological responses. This transient stress signal introduced by the measurement operation is captured by the imaging system and mixed with the real growth signal reflecting genetic characteristics or treatment effects in the same time-series image data. Because this disturbance noise overlaps with the real biological signal in the time domain, and its influence may vary from individual to individual, it leads to distortion of the growth curves resolved from the time-series images, misjudgment of early stress response characteristics, and ultimately reduces the accuracy and reliability of seedling vigor assessment and screening based on phenotypic data. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a high-throughput analysis method and system for seedling phenotype based on time-series growth images to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A high-throughput analysis method for seedling phenotypes based on time-series growth images includes the following steps: S1. During the seedling growth cycle, images of the same batch of seedlings are collected at multiple time points to obtain a time-series growth image sequence. S2. Extract at least one phenotypic parameter from each image in the time-series growth image sequence for all individual plants in the seedling population. S3. Based on the numerical values of the same phenotypic parameter of all individual plants at each time point, the population-level time series of the corresponding phenotypic parameter is calculated, and the time change curve of the phenotypic parameter of each individual plant is constructed. S4. Perform time-frequency transformation on the time series of the population and the time change curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and calculate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. S5. Identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters at the population level and for each individual plant. S6. The time point where the peak fluctuation of the high-frequency energy proportion of the population at the same time is consistent with the peak fluctuation of the high-frequency energy proportion of a single plant exceeding a preset number is determined to be an operational disturbance, and the corresponding phenotypic parameter values at the same time are corrected.
[0006] Furthermore, S1 includes: During the seedling growth cycle, determine the operation time for periodic movement and imaging of the seedling nursery unit; Based on the operation time, images of the same batch of seedlings were collected at the first preset time point before the operation time, the second preset time point after the operation time, and the third preset time point. Images collected at multiple time points are arranged in chronological order to obtain a time-series grown image sequence.
[0007] Furthermore, S2 includes: Each image in the time-series growth image sequence is segmented to obtain the independent region of each individual plant within the seedling population in each image; Extract at least one geometric parameter that characterizes the morphological features of an individual plant from its independent regions. The geometric parameters extracted from all images and all individual plants in the time-series growth image sequence are organized to obtain at least one phenotypic parameter for each individual plant in the seedling population.
[0008] Furthermore, S3 includes: According to the time point order of the time-series growth image sequence, the values of the same phenotypic parameter of all individual plants in the seedling population corresponding to each time point are obtained sequentially; For each time point, the values of the same phenotypic parameter of all individual plants corresponding to that time point are aggregated and calculated to obtain the population level value of that phenotypic parameter at that time point. The population level values at each time point are connected in chronological order to form the population level time series of the corresponding phenotypic parameters; For each individual plant in the seedling population, the values of the same phenotypic parameter of that individual plant at each time point in the time-series growth image sequence are connected in chronological order to form the time change curve of the phenotypic parameter of that individual plant.
[0009] Furthermore, S4 includes: Time-frequency analysis was performed on the time series of the population and the time variation curves of the phenotypic parameters of each individual plant to obtain the time-frequency distribution of each time series containing frequency components and time information. Based on the periodic characteristics of the seedling growth process, a high-frequency sub-band representing instantaneous disturbances is determined in the frequency domain of the time-frequency distribution. For the time-frequency distribution of the population-level time series and the time-frequency distribution of the phenotypic parameter time variation curves of each individual plant, the energy value contained in the high-frequency subband at each time point is calculated respectively. For the time-frequency distribution of the population-level time series and the time-frequency distribution of the phenotypic parameter time variation curve of each individual plant, the ratio of the energy value contained in the high-frequency subband to the total energy value corresponding to that time point is calculated respectively. The calculated ratios are arranged in chronological order to generate the high-frequency energy proportion time series of the population-level time series and the high-frequency energy proportion time series corresponding to the time change curve of the phenotypic parameters of each individual plant.
[0010] Furthermore, based on the periodic characteristics of the seedling growth process, a high-frequency sub-band characterizing instantaneous disturbances is determined in the frequency domain of the time-frequency distribution, specifically including: The frequency band formed by the frequency components in the time-frequency distribution whose corresponding period is shorter than the preset growth period threshold is determined as the high-frequency sub-band.
[0011] Furthermore, S5 includes: Local extrema detection is performed on the high-frequency energy proportion time series of the population level time series, and the locations of all local maxima in the time series that are greater than their adjacent time points are marked. For each individual plant, local extremum detection is performed on the time series corresponding to the high-frequency energy proportion of the phenotypic parameter change curve. The positions of all local maxima in the time series of each individual plant that are greater than the values of its adjacent time points are marked. Set a threshold for the energy percentage that represents significant fluctuations, and filter out the marked local maxima in the high-frequency energy percentage time series of the population-level time series and the high-frequency energy percentage time series of individual plants respectively. The locations of local maxima that remain after filtering and whose corresponding energy percentage values are greater than the energy percentage threshold are recorded as peak fluctuations in the corresponding time series.
[0012] Furthermore, S6 includes: On the timeline, examine the time points where the peak fluctuations of the high-frequency energy proportion in the population-level time series occur; For each detected time point, count the number of individual plants that exhibited high-frequency energy percentage time-series peak fluctuations at that time point; When the number of individual plants exhibiting high-frequency energy proportion time series peak fluctuations at a given time point exceeds a preset number, it is determined that the time point meets the condition that the high-frequency energy proportion time series peak fluctuations of the population-level time series are consistent with the high-frequency energy proportion time series peak fluctuations of the individual plants exceeding the preset number. For time points where operational disturbances are identified, the corrected phenotypic parameter values for that time point are obtained through interpolation by using the phenotypic parameter values corresponding to the adjacent time points before and after that time point.
[0013] Furthermore, using the phenotypic parameter values corresponding to adjacent time points before and after this time point, the corrected phenotypic parameter values for this time point are obtained through interpolation calculation, specifically including: Using a linear interpolation method, the phenotypic parameter values corrected for the time point with the operational disturbance are calculated based on the phenotypic parameter values corresponding to the nearest time point before the time point with the operational disturbance and the phenotypic parameter values corresponding to the nearest time point after the time point with the operational disturbance.
[0014] On the other hand, the present invention provides a high-throughput analysis system for seedling phenotypes based on time-series growth images, comprising the following modules: The sequence acquisition module is used to acquire images of the same batch of seedlings at multiple time points during the seedling growth cycle, and obtain a time-series growth image sequence. Attached Figure Description
[0015] Figure 1 This is a flowchart of the high-throughput analysis method for seedling phenotypes based on time-series growth images according to the present invention; Figure 2 This is a schematic diagram of the structure of the seedling phenotypic high-throughput analysis system based on time-series growth images according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1: Figure 1 The present invention provides a high-throughput analysis method for seedling phenotypes based on time-series growth images, which includes the following steps: S1. During the seedling growth cycle, images of the same batch of seedlings are collected at multiple time points to obtain a time-series growth image sequence. S2. Extract at least one phenotypic parameter from each image in the time-series growth image sequence for all individual plants in the seedling population. S3. Based on the numerical values of the same phenotypic parameter of all individual plants at each time point, the population-level time series of the corresponding phenotypic parameter is calculated, and the time change curve of the phenotypic parameter of each individual plant is constructed. S4. Perform time-frequency transformation on the time series of the population and the time change curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and calculate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. S5. Identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters at the population level and for each individual plant. S6. The time point where the peak fluctuation of the high-frequency energy proportion of the population at the same time is consistent with the peak fluctuation of the high-frequency energy proportion of a single plant exceeding a preset number is determined to be an operational disturbance, and the corresponding phenotypic parameter values at the same time are corrected.
[0018] S1. During the seedling growth cycle, images of the same batch of seedlings are acquired at multiple time points to obtain a time-series growth image sequence. The specific implementation is as follows: Throughout the entire growth cycle of seedlings, it is necessary to determine the operation time for the periodic movement and imaging of the seedling cultivation unit. This operation time specifically refers to the starting point or time range of the entire process of moving the seedling-carrying unit, such as a standard seed tray, from the growth area to the dedicated imaging area for image acquisition, completing the imaging operation, and then moving it back to the growth area. In actual high-throughput cultivation facilities, this operation time is usually preset and fixed by the operating program built into the automated transport and control equipment. The operation time can be obtained by directly reading a pre-arranged time schedule from the operation log file of the automated transport and control equipment, or by having the equipment operator manually set and save a fixed time point in the control interface according to the experimental design. For example, a fixed time each day, such as 9:00 AM, can be set as the start time for the daily moving imaging operation. In this way, the operation time becomes a known and repeatable time reference, ensuring that all subsequent image acquisition steps are precisely synchronized with the moving operation event, providing an accurate time anchor for analyzing operational disturbances.
[0019] After determining the operation time, specific image acquisition is planned and executed based on this operation time. Specifically, a first preset time point needs to be selected before the operation time, and a second and third preset time points need to be selected after the operation time. The setting of these preset time points needs to be based on the growth physiological characteristics of the seedlings and their response patterns to physical disturbances. The core logic is to systematically capture the complete process of changes in the seedling phenotypic state before, immediately after, and for a period of time after the operation disturbance. The first preset time point is located before the operation time, and its purpose is to obtain a baseline image of the seedlings when they are not affected by the current movement operation disturbance. The interval between this time point and the operation time should be long enough to ensure that the seedlings are in a stable state, but not so long that natural growth changes would mask the differences before and after the disturbance. For example, for most herbaceous crop seedlings, the first preset time point can be set 30 minutes before the start of the operation time. The determination of this 30-minute interval is based on preliminary experimental observations, that is, in the current growth stage of this type of seedling, the phenotypic changes caused by natural growth within 30 minutes are much smaller than the instantaneous changes that may be caused by the movement operation, thus effectively separating the two signals.
[0020] The second preset time point is located after the operation time and immediately adjacent to the operation completion time. Its purpose is to capture the transient physiological or morphological responses of the seedlings that the movement operation may cause as promptly as possible. The setting of this time point needs to consider the actual duration of the movement operation and the shortest time it takes for the seedlings to produce a detectable response to the disturbance. For example, the entire movement operation, including transmission, positioning, imaging, and return, may take approximately 4 minutes. After the seedling unit has re-stabilized in its growth position, the second preset time point is set 5 minutes after the operation is completed. This 5-minute interval is determined by measuring the average response time of the seedlings after experiencing slight physical disturbances, resulting in subtle changes in leaf angle or plant height that can be detected by the image.
[0021] The third preset time point is also located after the operation time, but with a longer time interval compared to the second preset time point. Its purpose is to observe whether the seedlings exhibit a sustained effect or recovery trend after experiencing a transient perturbation, thereby helping to distinguish between transient spurious changes and genuine, persistent growth anomalies. This time interval is set based on the time required for the seedlings to recover from the transient perturbation to a stable growth state. For example, for most vegetable seedlings, the third preset time point is set 60 minutes after the operation is completed. This 60-minute interval is determined based on extensive observational data showing that, after being subjected to similar levels of operational perturbation, the phenotypic parameters of this type of seedling typically recover to a level with no statistical difference compared to before the perturbation within 60 minutes. The specific values of all the above time intervals, including 30 minutes, 5 minutes, and 60 minutes, are illustrative examples. In practical applications, they need to be determined after calibration through limited preliminary experiments based on the specific seedling species being cultivated, such as tomato, cucumber, or corn, its growth stage (e.g., cotyledon stage, true leaf stage), and known physiological response kinetic data to the perturbation. The preliminary experiment method can be to simulate the movement operation on a small batch of seedlings and intensively sample images at different time points. By analyzing the change curves of phenotypic parameters, the start time, peak time and recovery time of the disturbance response can be identified, so as to reasonably set the first preset time point, the second preset time point and the third preset time point.
[0022] At each preset time point, images of the same batch of seedlings are acquired. The acquisition is conducted in a dedicated imaging chamber or station under controlled lighting conditions with constant light intensity and a fixed spectral composition to eliminate the impact of ambient light variations on image color and contrast, ensuring consistency and comparability of image quality across time points. A high-resolution industrial digital camera is used for image acquisition, fixedly mounted directly above the plane of the seedling unit, with the optical axis of the camera lens perpendicular to the plane of the seedling unit. Before acquisition begins, the camera undergoes white balance and focus calibration to ensure accurate and clear image colors. Key camera parameters are set, including a resolution of at least 20 megapixels, an exposure time that is automatically adjusted based on light intensity or fixed at a value such as 10 milliseconds, and an ISO sensitivity set to the lowest native value such as 100 to reduce image noise. During acquisition, the seedling unit is precisely positioned at the center of the camera's field of view by the conveyor system, ensuring that the position and proportion of the seedlings in the image remain consistent with each shot. The camera is then triggered to capture a single color image, or, as needed, images of specific bands within a multispectral image. Each time an image is captured, the system automatically records a precise image acquisition timestamp, derived from a high-precision clock synchronized with the control system. For example, with an operating time base of 9:00 AM, the system will automatically execute the aforementioned image acquisition process at 8:30 AM (the first preset time point), 9:05 AM (the second preset time point), and 10:00 AM (the third preset time point), thereby obtaining seedling images at three different time points. Each image acquired at a specific time point corresponds to the visual state of the same batch of seedlings at that particular moment. The images are stored in lossless compression formats such as RAW or PNG, and include metadata containing time point identifiers and seedling batch numbers.
[0023] All the images acquired from different time points are arranged and organized according to their acquisition timestamps to obtain a time-series grown image sequence. Specifically, by reading the timestamp information from the metadata accompanying each image file, the image file paths or the image data themselves are sorted into an ordered list according to the timestamps from earliest to latest. For example, images from the first preset time point (8:30 AM), the second preset time point (9:05 AM), and the third preset time point (10:00 AM) are arranged according to their corresponding time sequences of 8:30 AM, 9:05 AM, and 10:00 AM, forming a sequence containing three image elements. If the growth cycle includes multiple moving imaging operations, the process of acquiring images at three time points around the operation time is repeated for each operation, and all images are arranged in global time order to form a longer time-series grown image sequence covering the entire growth cycle. This sequence is an ordered collection of images, each with a unique time stamp, comprehensively recording the visual changes in the phenotypic state of the seedlings within a time window surrounding one or more movement events. This time-series growth image sequence serves as the sole data foundation for all subsequent phenotypic parameter extraction and analysis. The image sequence constructed in this way not only includes natural growth information of the seedlings but also intentionally incorporates instantaneous changes that may arise due to operational disturbances, providing the necessary data structure for subsequent identification and separation of such disturbances. The entire acquisition process, from determining the time baseline and planning acquisition points to image acquisition and sorting, ensures the continuity, relevance, and traceability of the data over time, laying a solid foundation for subsequent time-series analysis.
[0024] S2. From each image in the time-series growth image sequence, extract at least one phenotypic parameter from all individual plants in the seedling population. Specifically, this is implemented as follows: Each image in the time-series growth image sequence is segmented. The segmentation aims to separate the closely packed seedling population into independent connected regions representing each individual plant. Since the seedling trays have a regular grid, an effective segmentation strategy combines prior location information with image color features. Specifically, for each image in the time-series growth image sequence, a regular grid mask is first defined on the image based on the known number of rows and columns of the seedling trays, initially dividing the image into several candidate regions that may contain individual plants. However, since seedlings may grow beyond the grid boundaries or be tilted, further refined segmentation is needed. Within each candidate region, the color image is converted from the red-green-blue color space to a color space more suitable for vegetation identification, such as the green-red difference index color space. In this color space, the green-red difference index value for each pixel is calculated, specifically the green channel pixel value minus the red channel pixel value. Subsequently, the distribution of the green-red difference index values within the candidate region is statistically analyzed, and a segmentation threshold is automatically calculated using Otsu's method. The candidate region is binarized using the segmentation threshold. The pixels with values greater than the segmentation threshold are initially identified as seedling pixels, forming an initial binary mask.
[0025] Morphological post-processing is performed on the initial binary mask to optimize the segmentation results. First, morphological opening operations are used to remove small noise points; the structuring element can be a square with a side length of 3 pixels. Then, since a single seedling should be a connected region in the binary mask, it is necessary to handle the potential adhesion of multiple regions due to leaf contact. At this point, a watershed segmentation algorithm is used. Specifically, the distance transform of the initial binary mask is calculated, and then local maxima in the distance transform map are found as markers. The watershed algorithm is then applied to separate the adhered regions. For each connected region obtained after segmentation, its pixel area is calculated. A minimum pixel area threshold is set, for example, 200 pixels. Connected regions with a pixel area smaller than this threshold are considered noise or residual matrix and are removed. The minimum pixel area threshold is pre-set based on statistical measurements of the minimum projected area of a single plant in a large number of seedling images. After these steps, a series of non-overlapping connected regions are finally obtained from an image, and each connected region is identified as an independent region of an individual plant within the seedling population in the image. This independent region is defined by a set of pixels, containing all visible pixels of that individual plant in the current image. For each image in the time-series image growth sequence, the complete segmentation process described above is independently repeated, resulting in a set of independent regions containing all individual plants for each image. Key parameters of the segmentation process, such as the green-red difference index segmentation threshold, are automatically calculated for each image using the Otsu method, while the minimum pixel area threshold is a globally preset value based on empirical data, ensuring the method's robustness to changes in image brightness and differences in individual plant size.
[0026] After successfully obtaining the independent region of each individual plant in each image, the feature extraction stage begins. This involves extracting at least one geometric parameter characterizing the morphological features of each individual plant from its independent region. Here, the geometric parameter refers to a quantified index calculated from the projected shape of the plant in the two-dimensional image. The extraction process is performed separately for each identified independent region within each image. First, the pixel set defining the independent region of the individual plant is obtained. Based on this pixel set, several fundamental geometric parameters can be calculated. A core geometric parameter is the projected area. The projected area is calculated by counting the total number of pixels constituting the independent region and then multiplying the number of pixels by the area represented by a single pixel in actual physical space. The actual area of a single pixel is obtained through camera calibration. For example, given the camera resolution and imaging height, each pixel corresponds to 0.1 mm multiplied by 0.1 mm, or 0.01 square millimeters. Therefore, the projected area equals the number of pixels multiplied by 0.01, in square millimeters.
[0027] Another key geometric parameter is the principal axis length, also known as plant height or maximum Feret diameter. It is calculated as follows: first, calculate the minimum bounding rectangle or fitted ellipse of the pixel set of the independent region; then, determine the maximum length of the region in a certain direction. A more precise approach is to calculate the Euclidean distances between all pixel pairs in the binary region and find the maximum value. This maximum value is the principal axis length, which also needs to be multiplied by the pixel equivalent, for example, 0.1 mm per pixel, to convert it to physical length in millimeters.
[0028] Density or roundness can also be calculated as geometric parameters. Density reflects the compactness of a region's shape and is calculated as: 4 times the region's projected area multiplied by π, divided by the square of the region's perimeter. The perimeter of a region can be obtained by calculating the chain code length of the boundary pixels of that independent region. The closer the roundness is to 1, the closer the shape is to a circle. Which geometric parameter(s) to extract depends on the needs of subsequent analysis; for example, if the focus is on biomass accumulation, the projected area is used, and if the focus is on plant height growth, the principal axis length is used. In practice, the projected area and principal axis length can be extracted as representative geometric parameters. The extraction process is purely geometric calculation, applying the same formula to each independent region. The input is the set of pixel coordinates of the region, and the output is one or more numerical geometric parameters corresponding to that region. For each image in the time-series growth image sequence, and for each independent region of each individual plant in that image, the above geometric parameters are calculated, thus generating a set of geometric parameter values for each individual plant at each imaging time point.
[0029] The scattered geometric parameters extracted from all images and individual plants in the time-series growth image sequence need to be systematically organized to obtain at least one phenotypic parameter for each plant in the seedling population. This organization refers to constructing a structured data table or array such that each phenotypic parameter value can be uniquely identified by three-dimensional indices: which image in the time-series growth image sequence corresponds to the time point, which specific plant in the seedling population corresponds to the individual plant, and which geometric parameter type corresponds to the parameter type. A specific implementation can be as follows: establish a three-dimensional data matrix or a table containing multiple records. For example, a relational data table can be established where each record contains the following fields: a time point identifier referencing the timestamp of the time-series growth image sequence, a plant identifier assigned during the initial segmentation and maintained consistent across subsequent images through a tracking algorithm, a parameter name such as projected area or principal axis length, and a parameter value, which is the calculated floating-point number. For each plant in each image, a record is generated for each extracted geometric parameter.
[0030] In this way, the originally scattered data is organized into an ordered multidimensional dataset. At least one phenotypic parameter of each individual plant in the seedling population is represented in this structured dataset. For example, if the geometric parameter of projected area is extracted, then the projected area phenotypic parameter of each individual plant in the seedling population is reflected in the set of records in the data table whose parameter name is projected area. Furthermore, the projected area value of any individual plant at any given time can be retrieved based on the time point identifier and the individual plant identifier. If multiple geometric parameters are extracted, they exist in parallel within this data structure. This organizational process forms the direct data foundation for the subsequent time series construction in step S3. It ensures that subsequent analysis can efficiently and accurately access the specific phenotypic parameter value of any individual plant at any given time, thereby supporting the construction of population-level time series and individual phenotypic parameter time change curves. The entire step S2, from image segmentation to feature extraction to data organization, forms a complete and repeatable computational pipeline from the original image to structured phenotypic parameters. The calculation methods and organizational logic of all parameters are clearly defined, ensuring the interpretability and verifiability of the results.
[0031] S3. Based on the numerical values of the same phenotypic parameter of all individual plants at each time point, the population-level time series of the corresponding phenotypic parameter is calculated, and the time change curve of the phenotypic parameter of each individual plant is constructed. The specific implementation is as follows: The input to this step is the structured phenotypic parameter data obtained in step S2. This data specifies the values of specific geometric parameters, such as the projected area, for each individual plant in the seedling population at each time point in the time-series growth image sequence. The goal of step S3 is to generate a time series reflecting the overall average state of the population and an independent development trajectory curve for each individual plant. First, following the time point order of the time-series growth image sequence, the values of the same phenotypic parameter for all plants in the seedling population corresponding to each time point are obtained sequentially. Specifically, the time point order of the time-series growth image sequence is known, for example, the time point sequence is time point A, time point B, and time point C. For time point A, records with time point identifiers equal to time point A and parameter names equal to a specified parameter, such as projected area, need to be filtered from the structured data table generated in step S2. Then, the parameter values corresponding to each individual plant identifier are extracted from these records. This results in a dataset containing the projected area values of all surviving and successfully detected individual plants at time point A. This dataset may be a list or array containing multiple values. The exact same operation was repeated for time points B and C: records with time point identifiers of B and C and parameter name "projected area" were selected from the data table and their numerical lists were extracted. This sequential access and filtering mechanism ensured the integrity of the data acquisition and the temporal correspondence, preparing a set of numerical values for the same phenotypic parameter for all individuals in the population at each time point.
[0032] For each time point, the values of the same phenotypic parameter for all individuals at that time point need to be aggregated to obtain a single value that represents the overall level of the population at that time—the population level value for that phenotypic parameter at that time point. The purpose of aggregation is to integrate individual differences and reflect the central tendency of the population. The most commonly used and robust aggregation method is to calculate the median. Specifically, for N individual parameter values obtained at a certain time point, these N values are first sorted in ascending order. If N is odd, the population level value is the value in the middle position after sorting; if N is even, the population level value is the arithmetic mean of the two values in the middle position after sorting. The median is chosen over the arithmetic mean because the median is not sensitive to extreme outliers in the data, such as maxima or minima caused by segmentation errors, and more robustly represents the typical state of the population. For example, at time point A, if the projected area values of 100 individual plants are obtained, after sorting them, the average of the 50th and 51st values is calculated as the population-level projected area value at time point A, in square millimeters. Besides the median, assuming a stable population size and high data quality, the arithmetic mean can also be used as the aggregation method, which involves summing all N values at that time point and then dividing by N. The calculated population-level value needs to be stored in conjunction with its corresponding time point.
[0033] The population-level values at each time point are concatenated in chronological order to form a population-level time series for the corresponding phenotypic parameter. The chronological order is the inherent order of the time-series growth image sequence, such as time point A, time point B, and time point C. In practice, the concatenation operation involves creating a new sequence data structure, such as a list or a dedicated time series object. The first element of this data structure is the population-level value at time point A, the second element is the population-level value at time point B, and the third element is the population-level value at time point C. Each element not only contains the value itself but also implicitly or explicitly associates it with its corresponding time point information. Thus, these discrete, time-point-calculated population-level values are chained together into a continuous, time-varying sequence—the population-level time series. This sequence forms the basis for subsequent population-level time-frequency analysis, characterizing the pattern and trend of the evolution of a specified phenotypic parameter at the population average level over time.
[0034] For each individual plant in the seedling population, its independent growth trajectory needs to be constructed. This involves connecting the values of the same phenotypic parameter at each time point in the time-series growth image sequence of that individual plant in chronological order, forming a phenotypic parameter time-varying curve for that individual plant. Each individual plant is defined by its unique plant identifier. The implementation process requires traversing all plant identifiers in the seedling population. For a specific plant identifier, such as plant number five, all records whose plant identifier is equal to plant number five and whose parameter name is equal to a specified parameter, such as projected area, are selected from the structured data table generated in step S2. These records already contain parameter values at different time points. Then, strictly following the time point order of the time-series growth image sequence, such as time point A, time point B, and time point C, the parameter values of that individual plant at each time point are extracted sequentially. If the data for that individual plant is missing at a certain time point, for example due to segmentation failure, it needs to be processed according to a strategy. One strategy is to mark the value at that time point as a missing value; another strategy is to use the values from adjacent time points for linear interpolation to fill the missing value. The extracted or filled series of values are arranged in chronological order to form the time-varying curve of the projected area phenotypic parameters for individual plant number five. This curve is also represented in the data structure as an ordered list of values, with each position in the list corresponding to a parameter value at a given time point. This process is repeated for each individual plant identifier within the population, generating an independent time-varying curve for the phenotypic parameters for each plant. These curves form the foundation for subsequent individual-level time-frequency analysis and population-individual comparative analysis. Through the series of operations in step S3, the original discrete phenotypic observation data are systematically constructed into two different perspectives of time-series data: a macroscopic population-level time series and a microscopic individual time-varying curve. This lays a structured data foundation for in-depth analysis of seedling growth dynamics and identification of operational perturbations. The entire process relies on explicit sorting, filtering, aggregation, and joining rules; all operations are deterministic, ensuring the repeatability of the results.
[0035] S4. Perform time-frequency transformation on the time series of the population and the time variation curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and calculate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. The specific implementation is as follows: Time-frequency analysis was performed on the population-level time series and the temporal variation curves of phenotypic parameters for each individual plant. The purpose of time-frequency analysis is to transform one-dimensional time series data into a two-dimensional time-frequency distribution, which can reveal the patterns of change of different frequency components in the signal over time. For each time series to be analyzed, continuous wavelet transform was used as the specific method for time-frequency analysis. Continuous wavelet transform requires a mother wavelet function. In this embodiment, the Morlet wavelet is chosen because it has good localization characteristics in both the time and frequency domains, making it suitable for analyzing non-stationary biological signals. For a given time series, such as the temporal variation curve of the projected area of an individual plant, the curve consists of a series of values sampled at equal time intervals. The calculation process of continuous wavelet transform involves translating the mother wavelet function on the time axis while scaling it at different scales. The convolution or inner product of the scaled and translated wavelet function with the original time series at each time point is calculated, resulting in a series of wavelet coefficients. Each wavelet coefficient corresponds to a specific time point and a specific scale, with the scale parameter inversely proportional to the frequency. By calculating the wavelet coefficients at all scales of interest, covering the entire time range, a wavelet scale map corresponding to the time series is obtained, i.e., a time-frequency distribution. This time-frequency distribution is a two-dimensional matrix, where the rows represent different scales (i.e., frequencies), and the columns represent different time points. Each element in the matrix is the square of the modulus of the wavelet coefficient at the corresponding time and scale, which characterizes the energy intensity of that frequency component at that time point. The above continuous wavelet transform calculation is performed independently on the population-level time series and the temporal variation curves of the phenotypic parameters of each individual plant, thereby generating their respective time-frequency distributions for the population time series and each individual curve. Each time-frequency distribution contains the correlation between frequency components and temporal information.
[0036] Based on the cyclical characteristics of seedling growth, a high-frequency sub-band characterizing transient disturbances is determined within the frequency domain of the time-frequency distribution. The core of this step is dividing the frequency domain into a low-frequency component reflecting slow growth changes and a high-frequency component reflecting rapid transient disturbances. Specific implementation relies on prior knowledge of the typical timescales of seedling growth. Seedling growth changes, such as biomass accumulation or increased plant height, are relatively slow processes, with the smallest period of significant change typically measured in hours or days. Conversely, transient physiological disturbances caused by movement or other manipulations have very short durations, usually within minutes to an hour. A preset growth cycle threshold is used to distinguish between these two different timescales. This preset growth cycle threshold needs to be determined through observation based on the specific seedling species and growth stage being cultivated. One method is to analyze the fastest reliably detectable fluctuation period in the time-varying curves of seedling phenotypic parameters during a known stable growth period without external disturbances. For example, for most vegetable seedlings in their seedling stage, the natural growth curve of their projected area is smooth over several consecutive hours. By calculating their autocorrelation function or power spectrum, it can be determined that the periods corresponding to their effective frequency components are all greater than 2 hours. Therefore, the preset growth period threshold can be set to 2 hours. Then, in the time-frequency distribution, the frequency band formed by the frequency components with periods shorter than 2 hours is determined as the high-frequency subband. Since there is a clear conversion relationship between scale and period in wavelet transform, based on the characteristics of the selected mother wavelet, the scale range with a scale value less than a certain critical value (Scrit) can be calculated to correspond to frequency components with periods less than 2 hours. This set consisting of all scales with scale values less than the critical value (Scrit) is the determined high-frequency subband. This high-frequency subband is a continuous or discrete scale interval, which is pre-calculated and stored for subsequent unified energy extraction operations on all time-frequency distributions.
[0037] For both the time-frequency distribution of the population-level time series and the time-frequency distribution of the phenotypic parameter time variation curves of each individual plant, the energy value contained in the high-frequency subband at each time point is calculated. Specifically, for a given time-frequency distribution matrix, at a given time point column index, the squared values of the wavelet coefficient moduli corresponding to all rows belonging to a predetermined high-frequency subband scale interval are extracted. These extracted squared values of wavelet coefficient moduli are then summed. This summation represents the total energy of all high-frequency components in the signal at that specific time point—that is, components with periods shorter than a preset growth period threshold—and is called the high-frequency subband energy value. This extraction and summation operation is repeated for each time point column index, resulting in a one-dimensional array of the same length as the original time series. Each element in the array is the high-frequency subband energy value for the corresponding time point. This calculation is independently applied to the time-frequency distribution of the population-level time series and the time-frequency distribution of each individual plant curve, generating a high-frequency subband energy value sequence for each analysis object.
[0038] For the time-frequency distribution of the population-level time series and the time-frequency distribution of the phenotypic parameter time variation curves of each individual plant, the ratio of the energy value contained in the high-frequency subband to the total energy value corresponding to that time point is calculated at each time point. The method for calculating the total energy value at that time point is as follows: for the same time-frequency distribution matrix and the same time point column index, the squares of the wavelet coefficient moduli corresponding to all scales in that column, including both low and high frequencies, are extracted, and all these values are summed to obtain the total energy value at that time point. Then, the previously calculated high-frequency subband energy value at that time point is divided by the total energy value at that time point; the quotient is the high-frequency energy ratio. This ratio is a dimensionless number between 0 and 1, quantifying the proportion of the energy of the high-frequency perturbation component relative to the total signal energy at that time point. The closer the ratio is to 1, the more the signal fluctuation at that moment mainly comes from the instantaneous high-frequency component; the closer the ratio is to 0, the more the signal is dominated by the low-frequency gradual variation component. Similarly, this division operation is performed at each time point, thereby generating a high-frequency energy ratio sequence for each group or individual being analyzed.
[0039] The calculated high-frequency energy ratios are arranged in chronological order to generate high-frequency energy ratio time series for the population-level time series and high-frequency energy ratio time series corresponding to the phenotypic parameter time change curves of each individual plant. The chronological order is the time point order of the original input time series. In practice, the one-dimensional ratio array calculated in the previous step is mapped to the original time point order according to its natural index order to form an ordered ratio sequence. This sequence is the high-frequency energy ratio time series; it is a new time series, where each data point represents the relative intensity of high-frequency energy in the original phenotypic parameter sequence at the corresponding time. The high-frequency energy ratio time series for the population-level time series is calculated based on the population's time-frequency distribution, reflecting the change in the intensity of high-frequency fluctuations in the overall population signal over time. The high-frequency energy ratio time series corresponding to the phenotypic parameter time change curves of each individual plant is calculated based on its own time-frequency distribution, reflecting the change in the intensity of high-frequency fluctuations in the individual signal over time. These newly generated high-frequency energy ratio time series are the direct input data for identifying peak fluctuations, i.e., potential disturbance events, in the subsequent step S5. The entire step S4 decomposes the time-domain mixed signal into the frequency domain through time-frequency transformation, and constructs a new characteristic time series that is highly sensitive to instantaneous disturbances by quantizing and normalizing the energy of specific high-frequency sub-bands, providing key data conversion and feature extraction methods for accurate detection of operational disturbance events.
[0040] S5. Identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters at the population level and for each individual plant. Specifically, this is implemented as follows: Local extremum detection is performed on the high-frequency energy proportion time series of a population-level time series to identify all local maxima locations where the value is greater than that of the adjacent time points. In practice, the high-frequency energy proportion time series is treated as a discrete one-dimensional signal. A local maximum is defined as follows: for a specific time point in the time series, it is necessary to determine whether the time series value corresponding to that time point is simultaneously greater than the time series values of both the preceding and following time points. For the first data point in the time series, it is only necessary to determine whether the time series value corresponding to that data point is greater than the time series value of the second adjacent data point; for the last data point in the time series, it is only necessary to determine whether the time series value corresponding to that data point is greater than the time series value of the second-to-last adjacent data point. The detection algorithm sequentially traverses every internal point in the time series except for the first and last points, comparing its value with the values of the preceding and following points. If the condition of being greater than both the preceding and following values is met, the index position of that point is recorded as a candidate local maximum location. For the first and last points, the above one-sided comparison is performed. By iterating and comparing in this way, a list containing all candidate position indices can be obtained. Each position in this list corresponds to a point in time in the original time series, and the value of that point is convex within a local range.
[0041] For each individual plant, local extremum detection is performed on the time series corresponding to the phenotypic parameter change curve of the high-frequency energy proportion. Local maxima positions for each individual plant within that time series that are greater than their adjacent time points are marked. This process is completely independent of the processing of the population time series and uses the same algorithmic logic, but it requires traversing each individual plant in the seedling population one by one. For a given individual plant, its high-frequency energy proportion time series is obtained; this is an independent one-dimensional array with the same length as the population time series but a different numerical sequence. Then, the exact same local maximum detection algorithm is applied: traversing the array, comparing the values of each internal point with its preceding and following adjacent points, recording the index of the position that satisfies the condition of being greater than the preceding and following values, and also handling one-sided comparisons of the first and last points. This detection process generates a list of candidate local maximum positions for each individual plant. The detection of all individual plants is performed in parallel or sequentially, with consistent algorithm core and judgment criteria, ensuring the comparability of detection results between individuals.
[0042] A threshold for the energy percentage is set to characterize significant fluctuations. This threshold distinguishes meaningful significant fluctuations from local maxima that may be caused by measurement noise or minor fluctuations. The setting of the energy percentage threshold must be based on the analysis of the distribution characteristics of high-frequency energy percentage time-series values to ensure its rationality and effectiveness. A specific and feasible method is based on statistical analysis of historical or training data. A batch of seedling high-frequency energy percentage time-series data acquired during known periods without significant operational disturbances can be collected, and the distribution characteristics of the values corresponding to all local maxima in this batch of data can be calculated. For example, the mean of these values can be calculated plus twice the standard deviation, and the result can be used as the initial value of the energy percentage threshold. Another method is to use the percentile method, for example, setting the threshold as the 95th percentile of these historical local maxima values. Suppose that analysis of historical data shows that the high-frequency energy percentage values of local maxima during periods without disturbances are generally below 0.25, then the energy percentage threshold can be initially set at 0.25. This threshold is a dimensionless value. Physically, it means that only when the proportion of high-frequency energy at a local maximum exceeds 0.25 is the fluctuation considered significant enough to potentially correspond to a real transient disturbance event. The final determination of this threshold may require fine-tuning with a small amount of validation data, but the core principle is to ensure it filters out fluctuations in background noise levels and retains truly prominent peak signals. This pre-defined energy proportion threshold is a global parameter that will be applied to subsequent filtering operations on all time-series local maxima.
[0043] The local maxima locations in the high-frequency energy proportion time series of the population-level time series and the high-frequency energy proportion time series of individual plants are filtered separately. The specific filtering process is as follows: For the population-level time series, firstly, read the list of all previously detected candidate local maxima locations. For each candidate location index in the list, find the corresponding value in the original population-level high-frequency energy proportion time series. Then, compare this value with a pre-set energy proportion threshold. If the value is less than or equal to the energy proportion threshold, remove the candidate location from the list; if the value is greater than the energy proportion threshold, keep the candidate location in the list. After this round of comparison and removal, the remaining list of locations represents the significant local maxima locations in the filtered population-level time series. For each individual plant, the same filtering operation is performed. For each individual plant's list of candidate local maxima locations, sequentially check the individual's high-frequency energy proportion time series value corresponding to each location index and compare it with the same energy proportion threshold, removing candidate locations whose values are not greater than the threshold. Finally, a filtered list of locations containing only significant local maxima is obtained for each individual plant. This filtering step is crucial because it uses a unified, data-statistic-based quantitative standard to eliminate small-amplitude local fluctuations that may be caused by random factors, and focuses on potential disturbance signals with higher energy proportions that are more likely to represent real events.
[0044] The local maxima positions remaining after filtering, whose corresponding energy percentage values are greater than the energy percentage threshold, are recorded as peak fluctuations in the corresponding time series. This step is a natural result and formal record of the previous filtering operation. For the population-level time series, each of the local maxima positions retained after filtering satisfies the condition that its corresponding value is greater than the energy percentage threshold. These position indices and their corresponding time point information, such as the number of hours calculated from the start of the time series, along with their specific high-frequency energy percentage values, are packaged and recorded into a data structure, such as a list of tuples containing time points and peak intensity values. This list is formally defined as the set of peak fluctuations in the high-frequency energy percentage time series of the population-level time series. Similarly, for each individual plant, its filtered local maxima positions and related information are recorded as the set of peak fluctuations in the high-frequency energy percentage time series of that individual plant. At this point, step S5 is completed, and its output is clear: for the population time series, a set of peak fluctuation records is obtained; for each individual time series, its own set of peak fluctuation records is also obtained. These peak fluctuations represent the points in the corresponding time series where the proportion of high-frequency energy is significantly higher. They are the set of events that form the basis for determining the existence of operational disturbances and conducting group-individual comparative analysis in subsequent step S6. The entire step S5, through rigorous local extremum detection combined with statistically based threshold filtering, achieves the goal of automatically and objectively extracting potential significant event points from continuous time series signals. Its algorithm steps are clear, the parameter settings are well-defined, and it has high operability and repeatability.
[0045] S6. For time points where the peak fluctuation of the high-frequency energy proportion in the population-level time series at the same time point coincides with the peak fluctuation of the high-frequency energy proportion in a single plant exceeding a preset number, it is determined that an operational disturbance exists, and the corresponding phenotypic parameter values at the same time point are corrected. Specifically, this is implemented as follows: On the timeline, examine the time points where the peak fluctuations of the high-frequency energy proportion in the population-level time series occur. Specifically, extract the time point information associated with each peak fluctuation record from the population peak fluctuation set output in step S5. These time points are typically represented as offsets relative to the time series start point, such as timestamps in hours or days. Compile all extracted time points into a list containing all moments when significant high-frequency energy peaks were detected at the population level. Since multiple peak fluctuations may exist, this list may contain multiple time points. The examination operation involves sequentially or iteratively accessing this list of time points to prepare for further population-individual consistency analysis for each candidate time point.
[0046] For each detected time point, the number of individual plants exhibiting high-frequency energy percentage temporal peak fluctuations at that time point is counted. This requires comparing the time points detected at the population level with the peak fluctuation records of each individual plant. For a specific time point currently being processed, all individual plants in the seedling population need to be traversed. For each individual plant, its individual peak fluctuation set output in step S5 is retrieved, and it is checked whether there is a peak fluctuation in the set whose associated time point is sufficiently close to the specific time point currently being examined, i.e., satisfying the time consistency condition. The time consistency condition requires a clearly defined maximum allowable time deviation range, i.e., a time window threshold. The setting of this time window threshold is based on the consideration of possible small time differences in physiological responses between individuals and the temporal sampling interval. For example, if the sampling interval of the time series image is 1 hour, then the time window threshold can be set to half of the sampling interval, i.e., 0.5 hours. This means that if a peak fluctuation time point of an individual plant satisfies that the absolute time difference between it and the specific time point currently being examined is less than or equal to 0.5 hours, then it is considered that the individual exhibited a high-frequency energy percentage temporal peak fluctuation at the specific time point currently being examined. During implementation, for each individual plant, the absolute time difference between all its peak fluctuation time points and the specific time point currently being inspected is calculated. If any absolute time difference is less than or equal to 0.5 hours, the individual is included in the statistics. After traversing all individual plants, the counts of those meeting the condition are accumulated to obtain the number of individual plants exhibiting high-frequency energy peak fluctuations at the specific time point currently being inspected. This number is an integer, and its maximum possible value is the total number of individual plants in the entire seedling population.
[0047] It is necessary to determine whether the number exceeds a preset number. The preset number is a crucial threshold that defines the proportion of the population that must be affected by an operational disturbance event to be considered valid. The preset number should be set based on an understanding of the expected consistency of the seedling population's response to a common external operational disturbance. A reasonable method is based on statistical significance principles or empirical proportions. For example, the preset number can be set as an integer representing a certain percentage of the total number of individual seedlings, such as 70%. If the total number of individual seedlings is 100, then the preset number would be 70. Another method is to analyze known operational disturbance events in historical data, statistically analyze the distribution of the number of individuals exhibiting synchronous high-frequency energy peaks during these events, and take the lower bound of this distribution as the preset number. For example, analyzing 10 known movement operations and recording that the minimum number of individuals exhibiting synchronous peaks after each operation is 65, the preset number can be conservatively set to 60. The preset number is a predefined and stored integer parameter. During the judgment process, the statistically obtained number is compared with the preset number. When the number of individual plants exhibiting high-frequency energy percentage time-series peak fluctuations at a given time point exceeds a preset number, that time point is deemed to meet the condition that the high-frequency energy percentage time-series peak fluctuations in the population-level time series are consistent with the high-frequency energy percentage time-series peak fluctuations of more than the preset number of individual plants. This condition satisfies both temporal consistency and population response scale consistency, and is a sufficient condition for determining the existence of an operational disturbance event. Once met, this time point is officially marked as a time point with an operational disturbance.
[0048] For time points identified as having operational perturbations, the original phenotypic parameter values corresponding to these time points need to be corrected, as these values are considered to be contaminated by the operational perturbations. The goal of the correction is to use the phenotypic parameter values corresponding to adjacent time points before and after the time point that are unperturbed or minimally affected by the perturbations, and estimate the phenotypic parameter values that the time point should have had no perturbations through interpolation calculations. The specific implementation includes several sub-steps. First, two reference time points need to be determined for interpolation: the nearest time point before the time point with operational perturbations and the nearest time point after the time point with operational perturbations. Here, "nearest" refers to the time point that is earlier than and closest to the time point with operational perturbations on the time axis, and the time point that is later than and closest to the time point with operational perturbations, among all sampling time points of the original time-series grown image sequence. The phenotypic parameter values corresponding to these two time points are known; they are derived from the original data extracted in step S2 and organized through step S3. Next, linear interpolation is used for calculation. Linear interpolation assumes that the phenotypic parameters change linearly over short time intervals. Let the phenotypic parameter value corresponding to the nearest time point before the operational disturbance be the first reference value, and the phenotypic parameter value corresponding to the nearest time point after the operational disturbance be the second reference value. The time interval between the first and second reference values is the total time interval. The time interval between the operational disturbance and the nearest previous time point is the partial time interval. The corrected phenotypic parameter value is calculated using the following relationship: the corrected phenotypic parameter value equals the first reference value plus the difference between the second and first reference values multiplied by the partial time interval divided by the total time interval. This calculation process requires specifying the source of the first and second reference values; that is, they must be values of the same phenotypic parameter at the same individual or population level at the corresponding time point, and have the same physical unit, such as square millimeters. The calculated corrected phenotypic parameter value is also a value with the same physical unit. For the correction of population-level time series, the first and second reference values are derived from the values of the population-level time series at the nearest previous and nearest subsequent time points. To correct the temporal variation curves of individual plant phenotypic parameters, the aforementioned interpolation calculation needs to be performed independently for each individual plant at each time point identified as a disturbance, using the parameter values of that individual at the most recent previous and most recent subsequent time points. After correction, the parameter values corresponding to the time points in the original data where the operational disturbance occurred are replaced with the corrected phenotypic parameter values. In this way, all data points identified as affected by operational disturbances are replaced by an estimated value based on their normal growth trend before and after, effectively eliminating the transient anomalies introduced by operational disturbances from the data, making subsequent growth analysis, model fitting, or variety selection based on the corrected data more accurate and reliable.
[0049] Example 2: Figure 2A schematic diagram of the high-throughput seedling phenotyping system based on time-series growth images of the present invention is provided. The high-throughput seedling phenotyping system based on time-series growth images includes the following modules: The sequence acquisition module is used to acquire images of the same batch of seedlings at multiple time points during the seedling growth cycle, and obtain a time-series growth image sequence. The parameter extraction module is used to extract at least one phenotypic parameter of all individual plants in the seedling population from each image of the time-series growth image sequence. The curve construction module is used to calculate the population-level time series of the corresponding phenotypic parameter based on the numerical values of the same phenotypic parameter of all individual plants at each time point, and to construct the time change curve of the phenotypic parameter of each individual plant. The time series calculation module is used to perform time-frequency transformation on the time series of the population and the time change curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and to generate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. The fluctuation identification module is used to identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters of the population level and the individual plants, respectively. The numerical correction module is used to determine the existence of operational disturbances at time points where the peak fluctuation of the high-frequency energy proportion in the population-level time series at the same time coincides with the peak fluctuation of the high-frequency energy proportion in a single plant exceeding a preset number of individuals, and to correct the corresponding phenotypic parameter values at the same time.
[0050] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.
[0051] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0052] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0053] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0054] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0055] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0056] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0057] The parameter extraction module is used to extract at least one phenotypic parameter of all individual plants in the seedling population from each image of the time-series growth image sequence. The curve construction module is used to calculate the population-level time series of the corresponding phenotypic parameter based on the numerical values of the same phenotypic parameter of all individual plants at each time point, and to construct the time change curve of the phenotypic parameter of each individual plant. The time series calculation module is used to perform time-frequency transformation on the time series of the population and the time change curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and to generate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. The fluctuation identification module is used to identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters of the population level and the individual plants, respectively. The numerical correction module is used to determine the existence of operational disturbances at time points where the peak fluctuation of the high-frequency energy proportion in the population-level time series at the same time coincides with the peak fluctuation of the high-frequency energy proportion in a single plant exceeding a preset number of individuals, and to correct the corresponding phenotypic parameter values at the same time.
[0058] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention introduces time-frequency transformation analysis into seedling phenotypic time-series data processing, enabling the decomposition of mixed growth signals into different frequency dimensions. This clearly separates the instantaneous high-frequency components caused by operational disturbances from the low-frequency trends representing true growth. This frequency-domain-based feature extraction method makes transient stress signals, which are difficult to identify in the time domain, explicit, providing a direct basis for accurate identification of operational disturbances. By constructing high-frequency energy time series at the population level and individual plant level and detecting their peak fluctuations, and utilizing the synchronous response of the population under external disturbances, a discrimination mechanism based on consistency statistics is established. This effectively overcomes the problem of disturbance signal submersion caused by individual physiological differences, and achieves objective and reliable automated detection of operational disturbance events.
[0059] 2. Based on accurately pinpointing the moment of operational disturbance, interpolation correction of phenotypic parameters at the time of contamination based on adjacent normal data can effectively repair abnormal fluctuations in the growth curve, restoring a time-series trajectory that better conforms to natural growth patterns. This significantly improves the accuracy and reliability of data relied upon for subsequent growth dynamic analysis, early stress response assessment, and selection of superior individual plants. It directly enhances the practical value of high-throughput phenotypic analysis technology in seedling cultivation, providing more solid data support for variety selection and refined cultivation management, and helping to improve breeding efficiency and overall seedling quality.
Claims
1. A high-throughput analysis method for seedling phenotypic analysis based on time-series growth images, characterized in that, Includes the following steps: S1. During the seedling growth cycle, images of the same batch of seedlings are collected at multiple time points to obtain a time-series growth image sequence. S2. Extract at least one phenotypic parameter from each image in the time-series growth image sequence for all individual plants in the seedling population. S3. Based on the numerical values of the same phenotypic parameter of all individual plants at each time point, the population-level time series of the corresponding phenotypic parameter is calculated, and the time change curve of the phenotypic parameter of each individual plant is constructed. S4. Perform time-frequency transformation on the time series of the population and the time change curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and calculate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. S5. Identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters at the population level and for each individual plant. S6. The time point where the peak fluctuation of the high-frequency energy proportion of the population at the same time is consistent with the peak fluctuation of the high-frequency energy proportion of a single plant exceeding a preset number is determined to be an operational disturbance, and the corresponding phenotypic parameter values at the same time are corrected.
2. The high-throughput analysis method for seedling phenotypic analysis based on time-series growth images according to claim 1, characterized in that, S1 includes: During the seedling growth cycle, determine the operation time for periodic movement and imaging of the seedling nursery unit; Based on the operation time, images of the same batch of seedlings were collected at the first preset time point before the operation time, the second preset time point after the operation time, and the third preset time point. Images collected at multiple time points are arranged in chronological order to obtain a time-series grown image sequence.
3. The high-throughput analysis method for seedling phenotypes based on time-series growth images according to claim 1, characterized in that, S2 include: Each image in the time-series growth image sequence is segmented to obtain the independent region of each individual plant within the seedling population in each image; Extract at least one geometric parameter that characterizes the morphological features of an individual plant from its independent regions. The geometric parameters extracted from all images and all individual plants in the time-series growth image sequence are organized to obtain at least one phenotypic parameter for each individual plant in the seedling population.
4. The high-throughput analysis method for seedling phenotypes based on time-series growth images according to claim 1, characterized in that, S3 includes: According to the time point order of the time-series growth image sequence, the values of the same phenotypic parameter of all individual plants in the seedling population corresponding to each time point are obtained sequentially; For each time point, the values of the same phenotypic parameter of all individual plants corresponding to that time point are aggregated and calculated to obtain the population level value of that phenotypic parameter at that time point. The population level values at each time point are connected in chronological order to form the population level time series of the corresponding phenotypic parameters; For each individual plant in the seedling population, the values of the same phenotypic parameter of that individual plant at each time point in the time-series growth image sequence are connected in chronological order to form the time change curve of the phenotypic parameter of that individual plant.
5. The high-throughput analysis method for seedling phenotypic analysis based on time-series growth images according to claim 1, characterized in that, S4 include: Time-frequency analysis was performed on the time series of the population and the time variation curves of the phenotypic parameters of each individual plant to obtain the time-frequency distribution of each time series containing frequency components and time information. Based on the periodic characteristics of the seedling growth process, a high-frequency sub-band representing instantaneous disturbances is determined in the frequency domain of the time-frequency distribution. For the time-frequency distribution of the population-level time series and the time-frequency distribution of the phenotypic parameter time variation curves of each individual plant, the energy value contained in the high-frequency subband at each time point is calculated respectively. For the time-frequency distribution of the population-level time series and the time-frequency distribution of the phenotypic parameter time variation curve of each individual plant, the ratio of the energy value contained in the high-frequency subband to the total energy value corresponding to that time point is calculated respectively. The calculated ratios are arranged in chronological order to generate the high-frequency energy proportion time series of the population-level time series and the high-frequency energy proportion time series corresponding to the time change curve of the phenotypic parameters of each individual plant.
6. The high-throughput analysis method for seedling phenotypic analysis based on time-series growth images according to claim 5, characterized in that, Based on the periodic characteristics of seedling growth, a high-frequency sub-band characterizing instantaneous disturbances is determined in the frequency domain of the time-frequency distribution, specifically including: The frequency band formed by the frequency components in the time-frequency distribution whose corresponding period is shorter than the preset growth period threshold is determined as the high-frequency sub-band.
7. The high-throughput analysis method for seedling phenotypic analysis based on time-series growth images according to claim 1, characterized in that, S5 include: Local extrema detection is performed on the high-frequency energy proportion time series of the population level time series, and the locations of all local maxima in the time series that are greater than their adjacent time points are marked. For each individual plant, local extremum detection is performed on the time series corresponding to the high-frequency energy proportion of the phenotypic parameter change curve. The positions of all local maxima in the time series of each individual plant that are greater than the values of its adjacent time points are marked. Set a threshold for the energy percentage that represents significant fluctuations, and filter out the marked local maxima in the high-frequency energy percentage time series of the population-level time series and the high-frequency energy percentage time series of individual plants respectively. The locations of local maxima that remain after filtering and whose corresponding energy percentage values are greater than the energy percentage threshold are recorded as peak fluctuations in the corresponding time series.
8. The high-throughput analysis method for seedling phenotypes based on time-series growth images according to claim 1, characterized in that, S6 include: On the timeline, examine the time points where the peak fluctuations of the high-frequency energy proportion in the population-level time series occur; For each detected time point, count the number of individual plants that exhibited high-frequency energy percentage time-series peak fluctuations at that time point; When the number of individual plants exhibiting high-frequency energy proportion time series peak fluctuations at a given time point exceeds a preset number, it is determined that the time point meets the condition that the high-frequency energy proportion time series peak fluctuations of the population-level time series are consistent with the high-frequency energy proportion time series peak fluctuations of the individual plants exceeding the preset number. For time points where operational disturbances are identified, the corrected phenotypic parameter values for that time point are obtained through interpolation by using the phenotypic parameter values corresponding to the adjacent time points before and after that time point.
9. The high-throughput analysis method for seedling phenotypic analysis based on time-series growth images according to claim 8, characterized in that, Using the phenotypic parameter values corresponding to adjacent time points before and after this time point, the corrected phenotypic parameter values for this time point are obtained through interpolation calculation, specifically including: Using a linear interpolation method, the phenotypic parameter values corrected for the time point with the operational disturbance are calculated based on the phenotypic parameter values corresponding to the nearest time point before the time point with the operational disturbance and the phenotypic parameter values corresponding to the nearest time point after the time point with the operational disturbance.
10. A high-throughput analysis system for seedling phenotypic analysis based on time-series growth images, used to implement the high-throughput analysis method for seedling phenotypic analysis based on time-series growth images as described in any one of claims 1-9, characterized in that, Includes the following modules: The sequence acquisition module is used to acquire images of the same batch of seedlings at multiple time points during the seedling growth cycle, and obtain a time-series growth image sequence. The parameter extraction module is used to extract at least one phenotypic parameter of all individual plants in the seedling population from each image of the time-series growth image sequence. The curve construction module is used to calculate the population-level time series of the corresponding phenotypic parameter based on the numerical values of the same phenotypic parameter of all individual plants at each time point, and to construct the time change curve of the phenotypic parameter of each individual plant. The time series calculation module is used to perform time-frequency transformation on the time series of the population and the time change curve of the phenotypic parameters of each individual plant to obtain the corresponding time-frequency distribution, and to generate the high-frequency energy proportion time series based on the energy of the preset high-frequency subband in the time-frequency distribution. The fluctuation identification module is used to identify the peak fluctuations in the high-frequency energy proportion time series corresponding to the time change curves of phenotypic parameters of the population level and the individual plants, respectively. The numerical correction module is used to determine the existence of operational disturbances at time points where the peak fluctuation of the high-frequency energy proportion in the population-level time series at the same time coincides with the peak fluctuation of the high-frequency energy proportion in a single plant exceeding a preset number of individuals, and to correct the corresponding phenotypic parameter values at the same time.