Data labeling method and determination system
By constructing a multi-geographical examination site collaborative calibration framework and iterative update method, the problems of insufficient parameter adaptability and cross-environment consistency in existing technologies are solved, and high-precision simulation and parameter optimization of crop digital twins in multiple environments are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INNER MONGOLIA ZHONGFU MINGFENG AGRI TECH CO LTD
- Filing Date
- 2026-06-15
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies lack a collaborative calibration mechanism for multiple geographical test sites when constructing crop digital twins, making it difficult to adapt parameters to different environmental conditions. Furthermore, they lack the ability to integrate real observed total photosynthetic values with multi-environment simulation results into the same evaluation framework, resulting in insufficient cross-environment consistency. Additionally, they lack an iterative calibration process for updating parameter distribution.
By constructing a multi-geographical examination site collaborative calibration framework, combining environmental time series, actual observed yield, multi-environmental joint loss value, prior distribution and posterior distribution, the dynamic physiological parameter set is iteratively updated to achieve crop digital twin parameter calibration and digital germplasm configuration file generation under multi-environmental collaborative constraints.
The target parameter set has been adapted to various environmental conditions, which improves the simulation accuracy of crop digital twins and ensures that parameter updates are within a physiologically reasonable range. The generated configuration file can be directly used for subsequent simulation and optimization.
Smart Images

Figure CN122491682A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural information technology, specifically to a data labeling method and judgment system. Background Technology
[0002] With the development of agricultural informatization, smart agriculture, and crop growth modeling technologies, constructing crop digital twins using environmental data, crop physiological parameters, and observation results to simulate and analyze crop growth status and yield performance has become an important technical direction in agricultural digitalization research. The simulation accuracy of crop digital twins largely depends on the accuracy of parameter calibration results. Therefore, how to calibrate key parameters in crop models based on measured environmental data and observation results has become a crucial problem to be solved in this field. In existing technologies, parameter calibration for crop models typically employs statistical inference-based methods or optimization-based methods to complete parameter search and updates, thereby improving the model's adaptability to specific application scenarios.
[0003] Chinese invention patent application CN109472320A, published on March 15, 2019, discloses an automatic correction framework for crop growth period model variety parameters under uncertain conditions. By selecting meteorological data with large differences in temperature and light attributes over multiple years as correction sample data, key phenological variables that can constrain the parameters to be adjusted are screened. The adaptive differential evolution algorithm is used to automatically correct the crop growth period model variety parameters, and cluster analysis is performed on the obtained multiple sets of parameter results to select representative variety parameters with good robustness, thereby reducing simulation errors and improving parameter stability.
[0004] While the aforementioned schemes can achieve automatic calibration of crop model parameters to some extent, their focus remains on optimizing variety parameters within crop growth stage models, primarily revolving around constraints on key phenological variables, evolutionary search, and the selection of representative parameters. For application scenarios requiring the construction of crop digital twins, relying solely on single-type model parameter calibration methods still has the following shortcomings: First, there is a lack of a collaborative calibration mechanism for multiple geographical testing sites, making it difficult to adapt the obtained parameters to different environmental conditions simultaneously; second, there is a lack of a processing method to integrate the actual observed total photosynthetic rate with multi-environment simulation results into the same evaluation framework, resulting in insufficient cross-environment consistency; third, there is a lack of an iterative calibration process based on parameter distribution updates, making it difficult to continuously shrink the search space and output directly reusable digital germplasm configuration files within the constraints of feasible parameter ranges.
[0005] Therefore, the present invention provides a data annotation method and a judgment system. Summary of the Invention
[0006] (a) Technical problems to be solved
[0007] To address the shortcomings of existing technologies, this invention provides a data annotation method and judgment system. By constructing a multi-geographical examination site collaborative calibration framework, it iteratively updates the dynamic physiological parameter set by combining environmental time series, actual observed yield, multi-environment joint loss value, as well as prior distribution, likelihood term, and posterior distribution, thereby realizing crop digital twin parameter calibration and digital germplasm configuration file generation oriented towards multi-environment collaborative constraints.
[0008] (II) Technical Solution
[0009] To achieve the above objectives, the present invention provides the following technical solution: a data annotation method, comprising:
[0010] S1. Construct an expert consensus parameter set and a dynamic physiological parameter set, and obtain an environmental time series set and a real observed yield set;
[0011] S2. Calculate the simulated yield values for each geographical examination site based on the crop growth mechanism simulation kernel;
[0012] S3. Calculate the target parameter set, update the dynamic physiological parameter set, and complete the adaptive calibration of the digital twin of the variety to be tested.
[0013] Preferably, the input parameters of the variety to be tested are obtained; the input parameters include effective accumulated temperature, developmental base temperature, extinction coefficient, carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential; the input parameters are divided to construct an expert consensus parameter set and a dynamic physiological parameter set, the expert consensus parameter set including effective accumulated temperature, developmental base temperature, and extinction coefficient; the dynamic physiological parameter set including carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential;
[0014] The developmental base temperature is the minimum temperature required for the growth and development of the tested variety; the effective accumulated temperature represents the sum of the effective temperatures during the entire growth period of the tested variety; the extinction coefficient is used to represent the degree of shading of radiation by the canopy; the carbon conversion rate is used to represent the ability of the tested variety to convert intercepted energy into dry matter; the specific leaf area is used to represent the leaf area that a unit of dry matter of the tested variety can form; and the upper limit of daily assimilation potential is used to represent the maximum value of the total photosynthetic output that the tested variety can produce in a single day.
[0015] Obtain the environmental time series set of N different geographical test sites and the actual observed yield set corresponding to each geographical test site; the actual observed yield refers to the total photosynthetic amount measured by the test variety in the corresponding geographical test site during the growth experiment.
[0016] Preferably, for any set of dynamic physiological parameters, the crop growth mechanism simulation kernel is invoked to calculate the simulated yield value for each geographical examination site. Specifically:
[0017] Calculate the effective accumulated temperature increment of the variety to be tested on day t by obtaining the average daily temperature of the geography examination room on day t; t is a positive integer.
[0018] The cumulative effective accumulated temperature of the tested variety is calculated based on the effective accumulated temperature increment; when the cumulative effective accumulated temperature is not less than the effective accumulated temperature, it indicates that the tested variety has completed its growth period, and data observation is stopped for subsequent calculations.
[0019] The leaf area index and extinction coefficient of the tested variety on day t were obtained to calculate the light interception rate on day t.
[0020] Obtain the water stress index and photoactive radiation on day t, and calculate the simulated daily total photosynthesis on day t.
[0021] The simulated yield value of the tested variety was calculated based on the simulated daily total photosynthesis.
[0022] Preferably, for any set of dynamic physiological parameters, the crop growth mechanism simulation kernel is invoked to calculate the simulated yield value for each geographical examination site. Specifically:
[0023] The red band, near-infrared band, and green band reflectance intensities were obtained for each geographical examination site on each collection date. These intensities were then combined to form a daily canopy reflectance vector. Multiple daily canopy reflectance vectors from the same geographical examination site during the growing season were arranged chronologically according to the collection date to obtain the canopy reflectance trajectory. This trajectory represents the change in canopy reflectance status of the tested variety over the collection date. Multiple collection dates from the same geographical examination site were then sorted chronologically and numbered sequentially starting from 1 to obtain the collection date sequence number.
[0024] The reflection intensities of the red band, near-infrared band, or green band in each canopy reflection trajectory are normalized to obtain normalized reflection intensities, which are then used to form normalized canopy reflection trajectories. Normalization refers to the process where, for the same geographical examination site, the minimum and maximum values of the red band, near-infrared band, or green band reflection intensities across all collection dates are used. The normalized reflection intensities corresponding to the minimum value are 0, and those corresponding to the maximum value are 1. The normalized reflection intensities of other red band, near-infrared band, or green band reflection intensities are converted to values between 0 and 1 based on their position between the minimum and maximum values.
[0025] A current keypoint set is generated based on the normalized canopy reflection trajectory of the current geographical examination site. This set includes multiple current keypoints, including a starting keypoint, a peak keypoint, a descending keypoint, and an ending keypoint. The starting keypoint is the single-day canopy reflection vector corresponding to the first acquisition date; the ending keypoint is the single-day canopy reflection vector corresponding to the last acquisition date. The canopy activity value is obtained by adding the normalized reflection intensity of the near-infrared band and the normalized reflection intensity of the green band on the same acquisition date. The peak date is the acquisition date on which the canopy activity value reaches its maximum value. If the canopy activity values of two or more acquisition dates are both at their maximum values, the acquisition date with the earliest date sequence is taken as the peak date. The peak keypoint is the single-day canopy reflection vector corresponding to the peak date. The descending date is the acquisition date after the peak date on which the canopy activity value is less than the canopy activity value of the previous acquisition date. If there are no acquisition dates after the peak date on which the canopy activity value is less than the canopy activity value of the previous acquisition date, the last acquisition date is taken as the descending date. The descending keypoint is the single-day canopy reflection vector corresponding to the descending date.
[0026] Obtain a historical geography exam site sample set; the historical geography exam site sample set includes M historical geography exam site samples; each historical geography exam site sample includes the normalized canopy reflectance trajectory, dynamic physiological parameter set, expert consensus parameter set, and actual observed yield of that historical geography exam site; generate a historical key point set for each historical geography exam site sample; the historical key point set includes multiple historical key points, including historical starting key point, historical peak key point, historical decline key point, and historical ending key point; M is a positive integer not less than 2;
[0027] Calculate the keypoint difference value between the current keypoint set and each historical keypoint set. The keypoint difference value represents the difference between the current geography exam room and a historical geography exam room sample at key locations of canopy change. The keypoint difference value is calculated as follows: Pair the starting keypoint, peak keypoint, descending keypoint, and ending keypoint in the current keypoint set with the historical starting keypoint, historical peak keypoint, historical descending keypoint, and historical ending keypoint in the historical keypoint set, respectively. For each pair of current and historical keypoints, calculate the red light band values for both the current and historical keypoints. The absolute values of the normalized reflectance intensity differences, the near-infrared band, the green band, and the acquisition date sequence number are calculated. These values are then summed to obtain the single-point difference value between the current key point and the historical key point. Finally, the single-point difference values of the four current key points and the historical key points are summed to obtain the key point difference value between the current geography exam room and the historical geography exam room sample.
[0028] The historical geography exam room samples are sorted in ascending order of key point difference values, and the top P historical geography exam room samples are selected as target reference samples; P is the square root of M rounded down; if P is less than 2, then P is 2; if P is greater than M, then P is M; if two or more historical geography exam room samples have the same key point difference value, then the sample with the earlier historical geography exam room sample number is ranked first; the historical geography exam room sample number refers to a consecutive number generated according to the time sequence in which the historical geography exam room samples entered the historical geography exam room sample set.
[0029] The actual observed yield of each target reference sample corresponding to the geography examination site is allocated into multiple historical single-day contribution reference values. Specifically, for any target reference sample, the historical canopy activity value of each collection date is calculated first, and the historical canopy activity values of all collection dates are added together to obtain the total historical activity value. The historical canopy activity value of a certain collection date is divided by the total historical activity value to obtain the yield allocation ratio of that collection date. The actual observed yield is multiplied by the yield allocation ratio to obtain the historical single-day contribution reference value of that collection date.
[0030] A daily photosynthetic contribution prediction model is trained based on the target reference sample. For any collection date of any target reference sample, the daily training input vector is composed of the collection date number, normalized reflectance intensity of the red band, normalized reflectance intensity of the near-infrared band, normalized reflectance intensity of the green band, carbon conversion rate, specific leaf area, upper limit of daily assimilation potential, effective accumulated temperature, developmental base temperature, and extinction coefficient. The historical daily contribution reference value corresponding to the collection date is used as the training label for the daily training input vector. The daily contribution training sample set is composed of all daily training input vectors and the corresponding historical daily contribution reference values.
[0031] A gradient boosting regression tree model was used to train the daily contribution training sample set. During training, the daily training input vector was input into the gradient boosting regression tree model to obtain the daily photosynthetic contribution value output by the model. The squared difference between the daily photosynthetic contribution value output by the model and the corresponding historical daily contribution reference value was calculated. The training objective was to minimize the sum of the squared differences between the daily photosynthetic contribution values corresponding to all daily training input vectors and the corresponding historical daily contribution reference values to complete the training and obtain the daily photosynthetic contribution prediction model.
[0032] The daily prediction input vector corresponding to each collection date of the current geography examination site is input into the daily photosynthetic contribution prediction model to obtain the daily photosynthetic contribution value of each collection date of the current geography examination site. The daily prediction input vector consists of the collection date number of the current geography examination site on the corresponding collection date, the normalized reflectance intensity of the red light band, the normalized reflectance intensity of the near-infrared band, the normalized reflectance intensity of the green light band, carbon conversion rate, specific leaf area, upper limit of daily assimilation potential, effective accumulated temperature, developmental base temperature and extinction coefficient.
[0033] If the daily photosynthetic contribution value on a certain collection date is greater than the daily assimilation potential limit, then the daily photosynthetic contribution value on that collection date is adjusted to the daily assimilation potential limit; if the daily photosynthetic contribution value on a certain collection date is not greater than the daily assimilation potential limit, then the daily photosynthetic contribution value on that collection date is retained; finally, the daily photosynthetic contribution values of all collection dates in the current geography examination room are added together to obtain the simulated yield value for the current geography examination room. .
[0034] Preferably, the simulated yield value and the actual observed yield corresponding to the geography examination room are obtained, and the yield error is calculated;
[0035] Calculate the combined multi-environmental loss value based on simulated output values and actual observed output values using a dynamic physiological parameter set. ;
[0036] The multi-environment joint loss value is used to characterize the overall fit of the dynamic physiological parameter set to the actual observed output of all geographical examination sites.
[0037] The multi-environment joint loss value is converted into a likelihood term, which represents the degree of matching between a certain dynamic physiological parameter set and the current real observation results of multiple environments. Specifically, each dynamic physiological parameter set is input into the crop growth mechanism simulation kernel to obtain the multi-environment joint loss value corresponding to each dynamic physiological parameter set; then the dispersion of the joint loss value of all dynamic physiological parameter sets in the current round is calculated, where the dispersion is the variance of the yield error; the multi-environment joint loss value corresponding to a certain dynamic physiological parameter set is substituted into the formula to obtain the likelihood term corresponding to that dynamic physiological parameter set.
[0038] For each set of dynamic physiological parameters, a prior distribution is constructed to represent the initial confidence level of different values of each dynamic physiological parameter before the introduction of current multi-environment real observation results. Specifically, the common value ranges of each dynamic physiological parameter in publicly available agronomic data are first retrieved, and the physiologically feasible intervals given by experts are determined. Then, the overlapping part of the common value ranges and the physiologically feasible intervals is used to determine the lower and upper search boundaries. The lower and upper search boundaries are used as the lower and upper limits of the allowable value range of the corresponding dynamic physiological parameter. A uniform distribution is constructed between the lower and upper search boundaries to form the prior distribution of the dynamic physiological parameter.
[0039] The posterior distribution is obtained by combining the prior distribution and the likelihood term. The posterior distribution is used to represent the updated confidence level of each dynamic physiological parameter set after integrating the initial confidence level and the current observation matching results. Specifically, the prior probability corresponding to each dynamic physiological parameter set is multiplied by its corresponding likelihood term to obtain the unnormalized updated value of each dynamic physiological parameter set; the posterior distribution is then calculated based on the unnormalized updated value.
[0040] The system obtains the confidence interval of each parameter in the dynamic physiological parameter set, divides the interval into multiple segments, takes the center value of each segment as the new parameter value, and then combines them into a new set of parameters to form a new set of dynamic physiological parameters for the next round. The system then repeatedly calculates the simulated yield value and joint loss value of each geographical test site through the crop growth mechanism simulation kernel, and iterates until the target parameter set with the minimum joint loss value is obtained.
[0041] Preferably, the optimal joint loss value for each iteration is calculated based on the simulated output value and the actual observed output value;
[0042] Calculate the rate of change of the optimal joint loss value between two consecutive iterations based on the optimal joint loss value in each iteration.
[0043] If at least one parameter in the target parameter set reaches its lower or upper search boundary, and the average relative yield error is greater than the median of the difference between the simulated yield value and the actual observed yield in the current round, it is determined that there is a risk of boundary conflict between the expert consensus parameter set in the current digital germplasm configuration file and the actual growth law. At this time, an abnormal diagnosis prompt message is output, and the R&D personnel are prompted to re-check the expert consensus parameters such as effective accumulated temperature, developmental base temperature or extinction coefficient.
[0044] Finally, the target parameter set and the expert consensus parameter set are merged to generate an updated digital germplasm configuration file, and the confidence intervals of each parameter in each dynamic physiological parameter set are written synchronously to complete the adaptive calibration of the digital twin of the variety under test under the current multi-environment observation conditions.
[0045] A data annotation system includes a parameter management module, a multi-environment data management module, a simulation calculation module, a loss assessment module, a distributed update module, and a configuration file output module.
[0046] The parameter management module includes an expert consensus parameter storage unit, a dynamic physiological parameter storage unit, a parameter range management unit, and a parameter retrieval unit;
[0047] The multi-environment data management module includes a geographical examination room index unit, a time series storage unit, a temperature data storage unit, a radiation data storage unit, a soil moisture data storage unit, and an observed total photosynthetic amount storage unit;
[0048] The simulation calculation module includes an effective accumulated temperature increment calculation unit, a cumulative effective accumulated temperature calculation unit, a simulation termination date determination unit, a leaf area index generation unit, a light interception rate calculation unit, a water stress index calculation unit, a daily total photosynthesis calculation unit, and a simulated total photosynthesis output unit.
[0049] The loss assessment module includes a difference calculation unit and a multi-environment joint loss value generation unit;
[0050] The distribution update module includes a prior distribution generation unit, a joint loss receiving unit, a likelihood term generation unit, a posterior distribution generation unit, a posterior result normalization unit, a confidence interval extraction unit, a candidate parameter set recombination unit, and an iteration state determination unit.
[0051] The configuration file output module includes a target parameter group writing unit and a digital germplasm configuration file generation unit.
[0052] Preferably, the parameter management module includes an expert consensus parameter storage unit, a dynamic physiological parameter storage unit, a parameter range management unit, and a parameter calling unit. The expert consensus parameter storage unit is used to store effective accumulated temperature, developmental base temperature, and extinction coefficient. The dynamic physiological parameter storage unit is used to store carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential. The parameter range management unit is used to store the lower and upper limits corresponding to each dynamic physiological parameter. The parameter calling unit is used to provide parameter data to the simulation calculation module and the distribution update module. The multi-environment data management module includes a geographical examination room index unit, a time series storage unit, a temperature data storage unit, a radiation data storage unit, a soil moisture data storage unit, and an observed total photosynthetic amount storage unit.
[0053] Preferably, the distribution update module includes a prior distribution generation unit, a joint loss receiving unit, a likelihood term generation unit, a posterior distribution generation unit, a posterior result normalization unit, a confidence interval extraction unit, a candidate parameter set recombination unit, and an iteration state determination unit. The prior distribution generation unit generates a prior distribution based on the lower and upper limits corresponding to each dynamic physiological parameter. The joint loss receiving unit receives the multi-environment joint loss value output by the loss assessment module. The likelihood term generation unit generates a likelihood term corresponding to the dynamic physiological parameter set based on the multi-environment joint loss value. The posterior distribution generation unit generates a posterior distribution based on the prior distribution and the likelihood term. The posterior result normalization unit normalizes the posterior distribution result. The confidence interval extraction unit extracts the confidence interval of each dynamic physiological parameter based on the normalized posterior distribution. The candidate parameter set recombination unit generates a new dynamic physiological parameter set based on the confidence interval. The iteration state determination unit determines whether the iteration termination condition is met based on the distribution update results of the current and previous rounds.
[0054] (III) Beneficial Effects
[0055] This invention provides a data annotation method and judgment system, which have the following beneficial effects:
[0056] 1. By combining environmental time series data and actual observed yield data from multiple geographical examination sites, the dynamic physiological parameter set is updated collaboratively. This allows the target parameter set to adapt to multiple environmental conditions simultaneously, reducing the problem of distortion in single-environment calibration results when applied across regions.
[0057] 2. By constructing a multi-environment joint loss value, the simulated yield and the actual observed yield values under each geographical test site are uniformly incorporated into the same evaluation framework. Then, the dynamic physiological parameter set is updated by combining the prior distribution, likelihood term and posterior distribution. Therefore, the selected target parameter set can be closer to the actual observation results, thereby improving the simulation accuracy of crop digital twins.
[0058] 3. This invention first determines the lower and upper limits of each dynamic physiological parameter, then generates a prior distribution within the allowable value range, and extracts the parameter confidence interval based on the posterior distribution to form a new set of dynamic physiological parameters. This ensures that the parameter update process is always constrained by the reasonable physiological range, avoids disordered drift of parameter values, and at the same time, the basis for updating each parameter is clear, making it easy to interpret and verify.
[0059] 4. After the iteration termination condition is met, the set of dynamic physiological parameters with the minimum joint loss value is determined as the target parameter set, and it is combined with the expert consensus parameter set to form an updated digital germplasm configuration file. Therefore, the generated configuration results can be directly called by subsequent crop digital twins, which facilitates repeated simulation, parameter reuse and continuous optimization in different geographical test sites. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of a data annotation method according to the present invention;
[0061] Figure 2 This is a schematic diagram of the structure of a data annotation system according to the present invention. Detailed Implementation
[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0063] Please see Figure 1 This invention provides a data annotation method and judgment system, comprising the following steps:
[0064] S1. Construct an expert consensus parameter set and a dynamic physiological parameter set, and obtain an environmental time series set and a set of actual observed yields.
[0065] The input parameters of the variety to be tested are obtained; the input parameters include effective accumulated temperature, developmental base temperature, extinction coefficient, carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential; the input parameters are divided to construct an expert consensus parameter set and a dynamic physiological parameter set, the expert consensus parameter set including effective accumulated temperature, developmental base temperature, and extinction coefficient; the dynamic physiological parameter set including carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential;
[0066] Among them, the expert consensus parameter set is used This indicates that the dynamic physiological parameter set is used It indicates that the effective accumulated temperature is used Indication, developmental base temperature The extinction coefficient is expressed as... Indicates carbon conversion rate To represent, specifically leaf area is used It indicates that the upper limit of Japan's assimilation potential is used express;
[0067] The developmental base temperature is the minimum temperature required for the growth and development of the tested variety; the effective accumulated temperature represents the sum of the effective temperatures during the entire growth period of the tested variety; the extinction coefficient is used to represent the degree of shading of radiation by the canopy; the carbon conversion rate is used to represent the ability of the tested variety to convert intercepted energy into dry matter; the specific leaf area is used to represent the leaf area that a unit of dry matter of the tested variety can form; and the upper limit of daily assimilation potential is used to represent the maximum value of the total photosynthetic output that the tested variety can produce in a single day.
[0068] Obtain the environmental time series set of N different geography test sites and the set of actual observed yields corresponding to each geography examination room The actual observed yield refers to the total photosynthetic amount measured by the test variety in the corresponding geographical test site during growth experiments.
[0069] S2. Calculate the simulated yield value of each geographical examination site based on the crop growth mechanism simulation kernel.
[0070] In an optional embodiment, for any set of dynamic physiological parameters The system invokes the crop growth mechanism simulation kernel to calculate the simulated yield values for each geographical examination site. Specifically:
[0071] Get the average daily temperature of the geography exam room on day t. Calculate the effective accumulated temperature increment of the variety to be tested on day t. :
[0072]
[0073] Where t = 1, 2, 3, ..., m, and m is a positive integer;
[0074] Calculate the cumulative effective accumulated temperature of the variety to be tested. :
[0075]
[0076] When the cumulative effective temperature ≥Effective accumulated temperature When the test variety has completed its growth period, data observation is stopped and subsequent calculations are performed.
[0077] Obtain the leaf area index of the tested variety on day t. and extinction coefficient Calculate the light interception rate on day t. :
[0078]
[0079] Obtain the water stress index on day t. and effective radiation Calculate the simulated total daily photosynthesis on day t. :
[0080]
[0081] Calculate the simulated yield value of the tested variety. :
[0082]
[0083] Where i represents the geographical examination room number of the variety to be tested, i = 1, 2, 3, ..., N, and N is a positive integer.
[0084] In an optional embodiment, for any set of dynamic physiological parameters The system invokes the crop growth mechanism simulation kernel to calculate the simulated yield values for each geographical examination site. Specifically:
[0085] The red band, near-infrared band, and green band reflectance intensities were obtained for each geographical examination site on each collection date. These intensities were then combined to form a daily canopy reflectance vector. Multiple daily canopy reflectance vectors from the same geographical examination site during the growing season were arranged chronologically according to the collection date to obtain the canopy reflectance trajectory. This trajectory represents the change in canopy reflectance status of the tested variety over the collection date. Multiple collection dates from the same geographical examination site were then sorted chronologically and numbered sequentially starting from 1 to obtain the collection date sequence number.
[0086] The reflection intensities of the red band, near-infrared band, or green band in each canopy reflection trajectory are normalized to obtain normalized reflection intensities, which are then used to form normalized canopy reflection trajectories. Normalization refers to the process where, for the same geographical examination site, the minimum and maximum values of the red band, near-infrared band, or green band reflection intensities across all collection dates are used. The normalized reflection intensities corresponding to the minimum value are 0, and those corresponding to the maximum value are 1. The normalized reflection intensities of other red band, near-infrared band, or green band reflection intensities are converted to values between 0 and 1 based on their position between the minimum and maximum values.
[0087] A current keypoint set is generated based on the normalized canopy reflection trajectory of the current geographical examination site. This set includes multiple current keypoints, including a starting keypoint, a peak keypoint, a descending keypoint, and an ending keypoint. The starting keypoint is the single-day canopy reflection vector corresponding to the first acquisition date; the ending keypoint is the single-day canopy reflection vector corresponding to the last acquisition date. The canopy activity value is obtained by adding the normalized reflection intensity of the near-infrared band and the normalized reflection intensity of the green band on the same acquisition date. The peak date is the acquisition date on which the canopy activity value reaches its maximum value. If the canopy activity values of two or more acquisition dates are both at their maximum values, the acquisition date with the earliest date sequence is taken as the peak date. The peak keypoint is the single-day canopy reflection vector corresponding to the peak date. The descending date is the acquisition date after the peak date on which the canopy activity value is less than the canopy activity value of the previous acquisition date. If there are no acquisition dates after the peak date on which the canopy activity value is less than the canopy activity value of the previous acquisition date, the last acquisition date is taken as the descending date. The descending keypoint is the single-day canopy reflection vector corresponding to the descending date.
[0088] Obtain a historical geography exam site sample set; the historical geography exam site sample set includes M historical geography exam site samples; each historical geography exam site sample includes the normalized canopy reflectance trajectory, dynamic physiological parameter set, expert consensus parameter set, and actual observed yield of that historical geography exam site; generate a historical key point set for each historical geography exam site sample; the historical key point set includes multiple historical key points, including historical starting key point, historical peak key point, historical decline key point, and historical ending key point; M is a positive integer not less than 2;
[0089] Calculate the keypoint difference value between the current keypoint set and each historical keypoint set. The keypoint difference value represents the difference between the current geography exam site and a historical geography exam site sample at key locations of canopy change. The keypoint difference value is calculated as follows: the starting keypoint, peak keypoint, descending keypoint, and ending keypoint in the current keypoint set are paired one-to-one with the historical starting keypoint, historical peak keypoint, historical descending keypoint, and historical ending keypoint in the historical keypoint set. For each pair of current and historical keypoints, calculate the absolute value of the normalized reflectance intensity difference in the red band and the absolute value of the normalized reflectance intensity difference in the near-infrared band between the current and historical keypoints. The absolute values of the normalized reflectance intensity difference in the green band and the absolute values of the difference in the collection date sequence number are calculated. These are then summed to obtain the single-point difference value between the current key point and the historical key point. Finally, the single-point difference values of the four current key points and the historical key points are summed to obtain the key point difference value between the current geography exam room and the historical geography exam room sample. A smaller key point difference value indicates a lower cumulative difference between the current geography exam room and the historical geography exam room sample in the initial state, peak state, initial decline state, and final state.
[0090] The historical geography exam room samples are sorted in ascending order of key point difference values, and the top P historical geography exam room samples are selected as target reference samples; P is the square root of M rounded down; if P is less than 2, then P is 2; if P is greater than M, then P is M; if two or more historical geography exam room samples have the same key point difference value, then the sample with the earlier historical geography exam room sample number is ranked first; the historical geography exam room sample number refers to a consecutive number generated according to the time sequence in which the historical geography exam room samples entered the historical geography exam room sample set.
[0091] The actual observed yield of each target reference sample corresponding to the geography examination site is allocated into multiple historical single-day contribution reference values. Specifically, for any target reference sample, the historical canopy activity value of each collection date is calculated first, and the historical canopy activity values of all collection dates are added together to obtain the total historical activity value. The historical canopy activity value of a certain collection date is divided by the total historical activity value to obtain the yield allocation ratio of that collection date. The actual observed yield is multiplied by the yield allocation ratio to obtain the historical single-day contribution reference value of that collection date.
[0092] A daily photosynthetic contribution prediction model is trained based on the target reference sample. For any collection date of any target reference sample, the daily training input vector is composed of the collection date number, normalized reflectance intensity of the red band, normalized reflectance intensity of the near-infrared band, normalized reflectance intensity of the green band, carbon conversion rate, specific leaf area, upper limit of daily assimilation potential, effective accumulated temperature, developmental base temperature, and extinction coefficient. The historical daily contribution reference value corresponding to the collection date is used as the training label for the daily training input vector. The daily contribution training sample set is composed of all daily training input vectors and the corresponding historical daily contribution reference values.
[0093] A gradient boosting regression tree model was used to train the daily contribution training sample set. During training, the daily training input vector was input into the gradient boosting regression tree model to obtain the daily photosynthetic contribution value output by the model. The squared difference between the daily photosynthetic contribution value output by the model and the corresponding historical daily contribution reference value was calculated. The training objective was to minimize the sum of the squared differences between the daily photosynthetic contribution values corresponding to all daily training input vectors and the corresponding historical daily contribution reference values to complete the training and obtain the daily photosynthetic contribution prediction model.
[0094] The daily prediction input vector corresponding to each collection date of the current geography examination site is input into the daily photosynthetic contribution prediction model to obtain the daily photosynthetic contribution value of each collection date of the current geography examination site. The daily prediction input vector consists of the collection date number of the current geography examination site on the corresponding collection date, the normalized reflectance intensity of the red light band, the normalized reflectance intensity of the near-infrared band, the normalized reflectance intensity of the green light band, carbon conversion rate, specific leaf area, upper limit of daily assimilation potential, effective accumulated temperature, developmental base temperature and extinction coefficient.
[0095] If the daily photosynthetic contribution value on a certain collection date is greater than the daily assimilation potential limit, then the daily photosynthetic contribution value on that collection date is adjusted to the daily assimilation potential limit; if the daily photosynthetic contribution value on a certain collection date is not greater than the daily assimilation potential limit, then the daily photosynthetic contribution value on that collection date is retained; finally, the daily photosynthetic contribution values of all collection dates in the current geography examination room are added together to obtain the simulated yield value for the current geography examination room. .
[0096] S3. Calculate the target parameter set, update the dynamic physiological parameter set, and complete the adaptive calibration of the digital twin of the variety to be tested.
[0097] Obtain the simulated production value corresponding to the geography exam room and actual observed yield Calculate production error :
[0098]
[0099] Calculate dynamic physiological parameter set Multi-environment joint loss value :
[0100]
[0101] The multi-environment joint loss value This is used to characterize the overall fit of the dynamic physiological parameter set to the actual observed yields of all geographical examination sites. The smaller the value, the better the fit.
[0102] Combined loss value of multiple environments Convert to likelihood term The likelihood term is used to represent the degree of matching between a set of dynamic physiological parameters and current real-world observations across multiple environments. Specifically, it involves comparing the sets of dynamic physiological parameters... By inputting the crop growth mechanism simulation kernel into each set of dynamic physiological parameters, the multi-environment joint loss values corresponding to each set of parameters are obtained. Next, calculate the dispersion of the joint loss value of the entire set of dynamic physiological parameters for the current round, where the dispersion is the production error. The variance; substituting the multi-environment joint loss value corresponding to a certain dynamic physiological parameter set into the formula, we obtain the likelihood term corresponding to that dynamic physiological parameter set:
[0103]
[0104] in, Indicates production error The variance; since the numerator in the exponential function is the joint loss value, and it is preceded by a negative sign, the smaller the joint loss value, the larger the calculated likelihood term, indicating that the dynamic physiological parameter set matches the actual observation results; the larger the joint loss value, the smaller the calculated likelihood term, indicating that the dynamic physiological parameter set matches the actual observation results.
[0105] For each set of dynamic physiological parameters, a prior distribution is constructed to represent the initial confidence level of different values of each dynamic physiological parameter before incorporating current multi-environmental real-world observation results. Specifically, the common value ranges of each dynamic physiological parameter in publicly available agronomic data are first retrieved, and the physiologically feasible intervals given by experts are determined. Then, the overlapping portion of the common value ranges and the physiologically feasible intervals is used to determine the lower and upper search boundaries. These lower and upper search boundaries are used as the lower and upper limits of the allowable value range for the corresponding dynamic physiological parameter. A uniform distribution is constructed between the lower and upper search boundaries to form the prior distribution of that dynamic physiological parameter. ;
[0106] Combining prior distribution and likelihood term Obtain the posterior distribution posterior distribution This is used to represent the updated confidence level of each dynamic physiological parameter set after integrating the initial confidence level and the current observation matching results. Specifically, the prior probability corresponding to each dynamic physiological parameter set is multiplied by its corresponding likelihood term to obtain the unnormalized updated value of each dynamic physiological parameter set; then, the unnormalized updated value is substituted into the formula:
[0107]
[0108] The system obtains the confidence interval for each parameter in the dynamic physiological parameter set, divides the interval into multiple equal segments, takes the center value of each segment as the new parameter value, and combines them to form a new parameter set. This forms the next round of a new dynamic physiological parameter set. The simulated yield value and joint loss value for each geographical test site are calculated repeatedly through the crop growth mechanism simulation kernel. This process is iterated until the target parameter set with the minimum joint loss value is obtained. :
[0109]
[0110] target parameter set The target simulated yield value for each geographical examination site is calculated using a crop growth mechanism simulation kernel. Based on the target simulated output value Calculate the average relative production error :
[0111]
[0112] Calculate the optimal joint loss value for each iteration. :
[0113]
[0114] Where k represents the iteration round;
[0115] Based on the optimal joint loss value in each iteration Calculate the rate of change of the optimal joint loss value for two consecutive iterations. :
[0116]
[0117] If the target parameter set At least one parameter reaches its lower or upper search boundary, and the average relative yield error is simultaneously... If the value is greater than the median of the difference between the simulated yield value and the actual observed yield in the current round, it is determined that there is a risk of boundary conflict between the expert consensus parameter set in the current digital germplasm configuration file and the actual growth law. At this time, an abnormal diagnosis prompt message is output, and the R&D personnel are prompted to re-check the expert consensus parameters such as effective accumulated temperature Tsum, developmental base temperature Tbase, or extinction coefficient Kext.
[0118] Finally, the target parameter set By merging with the expert consensus parameter set, an updated digital germplasm configuration file is generated and synchronously written into the confidence interval of each parameter in each dynamic physiological parameter set, so as to complete the adaptive calibration of the digital twin of the variety under test under the current multi-environment observation conditions.
[0119] Please see Figure 2 This invention provides a data annotation method and judgment system, including a parameter management module, a multi-environment data management module, a simulation calculation module, a loss assessment module, a distribution update module, and a configuration file output module.
[0120] The parameter management module includes an expert consensus parameter storage unit, a dynamic physiological parameter storage unit, a parameter range management unit, and a parameter retrieval unit;
[0121] The multi-environment data management module includes a geographical examination room index unit, a time series storage unit, a temperature data storage unit, a radiation data storage unit, a soil moisture data storage unit, and an observed total photosynthetic amount storage unit;
[0122] The simulation calculation module includes an effective accumulated temperature increment calculation unit, a cumulative effective accumulated temperature calculation unit, a simulation termination date determination unit, a leaf area index generation unit, a light interception rate calculation unit, a water stress index calculation unit, a daily total photosynthesis calculation unit, and a simulated total photosynthesis output unit.
[0123] The loss assessment module includes a difference calculation unit and a multi-environment joint loss value generation unit;
[0124] The distribution update module includes a prior distribution generation unit, a joint loss receiving unit, a likelihood term generation unit, a posterior distribution generation unit, a posterior result normalization unit, a confidence interval extraction unit, a candidate parameter set recombination unit, and an iteration state determination unit.
[0125] The configuration file output module includes a target parameter group writing unit and a digital germplasm configuration file generation unit.
[0126] The parameter management module includes an expert consensus parameter storage unit, a dynamic physiological parameter storage unit, a parameter range management unit, and a parameter calling unit. The expert consensus parameter storage unit is used to store effective accumulated temperature, developmental base temperature, and extinction coefficient. The dynamic physiological parameter storage unit is used to store carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential. The parameter range management unit is used to store the lower and upper limits corresponding to each dynamic physiological parameter. The parameter calling unit is used to provide parameter data to the simulation calculation module and the distribution update module. The multi-environment data management module includes a geographical examination room index unit, a time series storage unit, a temperature data storage unit, a radiation data storage unit, a soil moisture data storage unit, and an observed total photosynthetic amount storage unit.
[0127] The distribution update module includes a prior distribution generation unit, a joint loss receiving unit, a likelihood term generation unit, a posterior distribution generation unit, a posterior result normalization unit, a confidence interval extraction unit, a candidate parameter set recombination unit, and an iteration state determination unit. Specifically, the prior distribution generation unit generates a prior distribution based on the lower and upper limits corresponding to each dynamic physiological parameter; the joint loss receiving unit receives the multi-environment joint loss value output by the loss assessment module; the likelihood term generation unit generates a likelihood term corresponding to the dynamic physiological parameter set based on the multi-environment joint loss value; the posterior distribution generation unit generates a posterior distribution based on the prior distribution and the likelihood term; the posterior result normalization unit normalizes the posterior distribution result; the confidence interval extraction unit extracts the parameter confidence interval for each dynamic physiological parameter based on the normalized posterior distribution; the candidate parameter set recombination unit generates a new round of dynamic physiological parameter sets based on the parameter confidence intervals; and the iteration state determination unit determines whether the iteration termination condition is met based on the distribution update results of the current and previous rounds.
[0128] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0129] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0130] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A data annotation method, characterized in that, include: S1. Construct an expert consensus parameter set and a dynamic physiological parameter set, and obtain an environmental time series set and a real observed yield set; S2. Calculate the simulated yield values for each geographical examination site based on the crop growth mechanism simulation kernel; S3. Calculate the target parameter set, update the dynamic physiological parameter set, and complete the adaptive calibration of the digital twin of the variety to be tested.
2. The data annotation method according to claim 1, characterized in that: In S1, an expert consensus parameter set and a dynamic physiological parameter set are constructed, and an environmental time series set and a set of actual observed yields are obtained, specifically as follows: Obtain the input parameters of the variety to be tested; the input parameters include effective accumulated temperature, developmental base temperature, extinction coefficient, carbon conversion rate, specific leaf area and upper limit of daily assimilation potential; The input parameters are divided into an expert consensus parameter set and a dynamic physiological parameter set. The expert consensus parameter set includes effective accumulated temperature, developmental base temperature, and extinction coefficient; the dynamic physiological parameter set includes carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential. The developmental base temperature is the minimum temperature required for the growth and development of the tested variety; the effective accumulated temperature represents the sum of the effective temperatures during the entire growth period of the tested variety; the extinction coefficient is used to represent the degree of shading of radiation by the canopy; the carbon conversion rate is used to represent the ability of the tested variety to convert intercepted energy into dry matter; the specific leaf area is used to represent the leaf area that a unit of dry matter of the tested variety can form; and the upper limit of daily assimilation potential is used to represent the maximum value of the total photosynthetic output that the tested variety can produce in a single day. Obtain the environmental time series set of N different geographical test sites and the actual observed yield set corresponding to each geographical test site; the actual observed yield refers to the total photosynthetic amount measured by the test variety in the corresponding geographical test site during the growth experiment.
3. The data annotation method according to claim 1, characterized in that: In S2, the simulated yield values for each geographical examination site are calculated based on the crop growth mechanism simulation kernel, as follows: For any set of dynamic physiological parameters, the crop growth mechanism simulation kernel is invoked to calculate the simulated yield values for each geographical examination site. Specifically: Calculate the effective accumulated temperature increment of the variety to be tested on day t by obtaining the average daily temperature of the geography examination room on day t; t is a positive integer. The cumulative effective accumulated temperature of the tested variety is calculated based on the effective accumulated temperature increment; when the cumulative effective accumulated temperature is not less than the effective accumulated temperature, it indicates that the tested variety has completed its growth period, and data observation is stopped for subsequent calculations. The leaf area index and extinction coefficient of the tested variety on day t were obtained to calculate the light interception rate on day t. Obtain the water stress index and photoactive radiation on day t, and calculate the simulated daily total photosynthesis on day t. The simulated yield value of the tested variety was calculated based on the simulated daily total photosynthesis.
4. The data annotation method according to claim 1, characterized in that: In S2, the simulated yield values for each geographical examination site are calculated based on the crop growth mechanism simulation kernel, as follows: For any set of dynamic physiological parameters, the crop growth mechanism simulation kernel is invoked to calculate the simulated yield values for each geographical examination site. Specifically: The red band, near-infrared band, and green band reflectance intensities were obtained for each geographical examination site on each collection date. These intensities were then combined to form a daily canopy reflectance vector. Multiple daily canopy reflectance vectors from the same geographical examination site during the growing season were arranged chronologically according to the collection date to obtain the canopy reflectance trajectory. This trajectory represents the change in canopy reflectance status of the tested variety over the collection date. Multiple collection dates from the same geographical examination site were then sorted chronologically and numbered sequentially starting from 1 to obtain the collection date sequence number. The reflection intensities of the red band, near-infrared band, or green band in each canopy reflection trajectory are normalized to obtain normalized reflection intensities, which are then used to form normalized canopy reflection trajectories. Normalization refers to the process where, for the same geographical examination site, the minimum and maximum values of the red band, near-infrared band, or green band reflection intensities across all collection dates are used. The normalized reflection intensities corresponding to the minimum value are 0, and those corresponding to the maximum value are 1. The normalized reflection intensities of other red band, near-infrared band, or green band reflection intensities are converted to values between 0 and 1 based on their position between the minimum and maximum values. Generate the current set of key points based on the normalized canopy reflection trajectory of the current geography exam room; The current keypoint set includes multiple current keypoints, including a starting keypoint, a peak keypoint, a falling keypoint, and an ending keypoint. The starting keypoint is the daily canopy reflection vector corresponding to the first acquisition date; the ending keypoint is the daily canopy reflection vector corresponding to the last acquisition date; the canopy activity value is obtained by adding the normalized reflection intensity of the near-infrared band and the normalized reflection intensity of the green band on the same acquisition date; the peak date is the acquisition date on which the canopy activity value reaches its maximum value. If the canopy activity value is the maximum value on two or more acquisition dates, the acquisition date with the earliest date sequence is taken as the peak date; the peak keypoint is the daily canopy reflection vector corresponding to the peak date; the falling date is the acquisition date after the peak date on which the canopy activity value is less than the canopy activity value of the previous acquisition date. If there is no acquisition date after the peak date on which the canopy activity value is less than the canopy activity value of the previous acquisition date, the last acquisition date is taken as the falling date; the falling keypoint is the daily canopy reflection vector corresponding to the falling date. Obtain a historical geography exam site sample set; the historical geography exam site sample set includes M historical geography exam site samples; each historical geography exam site sample includes the normalized canopy reflectance trajectory, dynamic physiological parameter set, expert consensus parameter set, and actual observed yield of that historical geography exam site; generate a historical key point set for each historical geography exam site sample; the historical key point set includes multiple historical key points, including historical starting key point, historical peak key point, historical decline key point, and historical ending key point; M is a positive integer not less than 2; Calculate the keypoint difference value between the current keypoint set and each historical keypoint set. The keypoint difference value represents the difference between the current geography exam room and a historical geography exam room sample at key locations of canopy change. The keypoint difference value is calculated as follows: Pair the starting keypoint, peak keypoint, descending keypoint, and ending keypoint in the current keypoint set with the historical starting keypoint, historical peak keypoint, historical descending keypoint, and historical ending keypoint in the historical keypoint set, respectively. For each pair of current and historical keypoints, calculate the red light band values for both the current and historical keypoints. The absolute values of the normalized reflectance intensity differences, the near-infrared band, the green band, and the acquisition date sequence number are calculated. These values are then summed to obtain the single-point difference value between the current key point and the historical key point. Finally, the single-point difference values of the four current key points and the historical key points are summed to obtain the key point difference value between the current geography exam room and the historical geography exam room sample. The historical geography exam room samples were sorted in ascending order of key point difference values, and the top P historical geography exam room samples were selected as target reference samples; P is the square root of M rounded down. If P is less than 2, then P is 2; if P is greater than M, then P is M; if the key point difference values of two or more historical geography exam room samples are the same, then the sample with the earlier historical geography exam room sample number will be ranked first; the historical geography exam room sample number refers to the consecutive number generated according to the order in which the historical geography exam room samples entered the historical geography exam room sample set. The actual observed yield of the geographical examination room corresponding to each target reference sample is allocated into multiple historical single-day contribution reference values. Specifically, for any target reference sample, the historical canopy activity value of each collection date of the target reference sample is calculated first, and the historical canopy activity values of all collection dates are added together to obtain the sum of historical activity values. Divide the historical canopy activity value of a certain collection date by the sum of historical activity values to obtain the production allocation ratio for that collection date. Multiply the actual observed production by this production allocation ratio to obtain the historical daily contribution reference value for that collection date. A daily photosynthetic contribution prediction model is trained based on the target reference sample. For any collection date of any target reference sample, the daily training input vector is composed of the collection date number, normalized reflectance intensity of the red band, normalized reflectance intensity of the near-infrared band, normalized reflectance intensity of the green band, carbon conversion rate, specific leaf area, upper limit of daily assimilation potential, effective accumulated temperature, developmental base temperature, and extinction coefficient. The historical daily contribution reference value corresponding to the collection date is used as the training label for the daily training input vector. The daily contribution training sample set is composed of all daily training input vectors and the corresponding historical daily contribution reference values. A gradient boosting regression tree model was used to train the daily contribution training sample set. During training, the daily training input vector is input into the gradient boosting regression tree model to obtain the daily photosynthetic contribution value output by the model; the squared difference between the daily photosynthetic contribution value output by the model and the corresponding historical daily contribution reference value is calculated; the training objective is to minimize the sum of the squared differences between the daily photosynthetic contribution values corresponding to all daily training input vectors and the corresponding historical daily contribution reference values to complete the training and obtain the daily photosynthetic contribution prediction model. The daily prediction input vector corresponding to each collection date of the current geography examination site is input into the daily photosynthetic contribution prediction model to obtain the daily photosynthetic contribution value of each collection date of the current geography examination site. The daily prediction input vector consists of the collection date number of the current geography examination site on the corresponding collection date, the normalized reflectance intensity of the red light band, the normalized reflectance intensity of the near-infrared band, the normalized reflectance intensity of the green light band, carbon conversion rate, specific leaf area, upper limit of daily assimilation potential, effective accumulated temperature, developmental base temperature and extinction coefficient. If the daily photosynthetic contribution value on a certain collection date is greater than the daily assimilation potential limit, then the daily photosynthetic contribution value on that collection date will be corrected to the daily assimilation potential limit. If the daily photosynthetic contribution value on a certain collection date is not greater than the daily assimilation potential limit, then the daily photosynthetic contribution value for that collection date is retained. Finally, the daily photosynthetic contribution values for all collection dates in the current geography exam room are summed to obtain the simulated yield value for the current geography exam room. .
5. The data annotation method according to claim 1, characterized in that: In S3, the target parameter set is calculated, the dynamic physiological parameter set is updated, and the adaptive calibration of the digital twin of the test variety is completed, as follows: Obtain the simulated yield value and the actual observed yield corresponding to the geography examination room, and calculate the yield error; Calculate the combined multi-environmental loss value based on simulated output values and actual observed output values using a dynamic physiological parameter set. ; The multi-environment joint loss value is converted into a likelihood term, which represents the degree of matching between a certain dynamic physiological parameter set and the current real observation results of multiple environments. Specifically, each dynamic physiological parameter set is input into the crop growth mechanism simulation kernel to obtain the multi-environment joint loss value corresponding to each dynamic physiological parameter set; then the dispersion of the joint loss value of all dynamic physiological parameter sets in the current round is calculated, where the dispersion is the variance of the yield error; the multi-environment joint loss value corresponding to a certain dynamic physiological parameter set is substituted into the formula to obtain the likelihood term corresponding to that dynamic physiological parameter set. For each set of dynamic physiological parameters, a prior distribution is constructed to represent the initial confidence level of different values of each dynamic physiological parameter before the current real-world observation results of multiple environments are introduced; Specifically, first, the common value ranges of each dynamic physiological parameter in publicly available agronomic data are retrieved, and the physiologically feasible intervals given by experts are determined; then, the overlapping part of the common value ranges and the physiologically feasible intervals is used to determine the lower and upper limits of the search, and the lower and upper limits of the search are used as the lower and upper limits of the allowable value range of the corresponding dynamic physiological parameter. A uniform distribution is constructed between the lower and upper limits of the search to form the prior distribution of the dynamic physiological parameter. The posterior distribution is obtained by combining the prior distribution and the likelihood term. The posterior distribution is used to represent the updated confidence level of each dynamic physiological parameter set after integrating the initial confidence level and the current observation matching results. Specifically, the prior probability corresponding to each dynamic physiological parameter set is multiplied by its corresponding likelihood term to obtain the unnormalized updated value of each dynamic physiological parameter set. The posterior distribution is calculated based on the unnormalized update value. The system obtains the confidence interval of each parameter in the dynamic physiological parameter set, divides the interval into multiple segments, takes the center value of each segment as the new parameter value, and then combines them into a new set of parameters to form a new set of dynamic physiological parameters for the next round. The system then repeatedly calculates the simulated yield value and joint loss value of each geographical test site through the crop growth mechanism simulation kernel, and iterates until the target parameter set with the minimum joint loss value is obtained.
6. The data annotation method according to claim 1, characterized in that: In S3, the target parameter set is calculated, the dynamic physiological parameter set is updated, and the adaptive calibration of the digital twin of the test variety is completed, as follows: The target parameter set is used to calculate the target simulated yield value for each geographical test site through the crop growth mechanism simulation kernel, and the average relative yield error is calculated based on the target simulated yield value. Calculate the optimal joint loss value for each iteration based on simulated output values and actual observed output values; Calculate the rate of change of the optimal joint loss value between two consecutive iterations based on the optimal joint loss value in each iteration. If at least one parameter in the target parameter set reaches its lower or upper search boundary, and the average relative yield error is greater than the median of the difference between the simulated yield value and the actual observed yield in the current round, it is determined that there is a risk of boundary conflict between the expert consensus parameter set in the current digital germplasm configuration file and the actual growth law. At this time, an abnormal diagnosis prompt message is output, and the R&D personnel are prompted to re-check the expert consensus parameters such as effective accumulated temperature, developmental base temperature or extinction coefficient. Finally, the target parameter set and the expert consensus parameter set are merged to generate an updated digital germplasm configuration file, and the confidence intervals of each parameter in each dynamic physiological parameter set are written synchronously to complete the adaptive calibration of the digital twin of the variety under test under the current multi-environment observation conditions.
7. A data annotation system, comprising a parameter management module, a multi-environment data management module, a simulation calculation module, a loss assessment module, a distributed update module, and a configuration file output module: The parameter management module includes an expert consensus parameter storage unit, a dynamic physiological parameter storage unit, a parameter range management unit, and a parameter retrieval unit; The multi-environment data management module includes a geographical examination room index unit, a time series storage unit, a temperature data storage unit, a radiation data storage unit, a soil moisture data storage unit, and an observed total photosynthetic amount storage unit; The simulation calculation module includes an effective accumulated temperature increment calculation unit, a cumulative effective accumulated temperature calculation unit, a simulation termination date determination unit, a leaf area index generation unit, a light interception rate calculation unit, a water stress index calculation unit, a daily total photosynthesis calculation unit, and a simulated total photosynthesis output unit. The loss assessment module includes a difference calculation unit and a multi-environment joint loss value generation unit; The distribution update module includes a prior distribution generation unit, a joint loss receiving unit, a likelihood term generation unit, a posterior distribution generation unit, a posterior result normalization unit, a confidence interval extraction unit, a candidate parameter set recombination unit, and an iteration state determination unit. The configuration file output module includes a target parameter group writing unit and a digital germplasm configuration file generation unit.
8. The data determination system according to claim 7, characterized in that: The parameter management module includes an expert consensus parameter storage unit, a dynamic physiological parameter storage unit, a parameter range management unit, and a parameter calling unit. The expert consensus parameter storage unit is used to store effective accumulated temperature, developmental base temperature, and extinction coefficient. The dynamic physiological parameter storage unit is used to store carbon conversion rate, specific leaf area, and upper limit of daily assimilation potential. The parameter range management unit is used to store the lower and upper limits corresponding to each dynamic physiological parameter. The parameter calling unit is used to provide parameter data to the simulation calculation module and the distribution update module. The multi-environment data management module includes a geographical examination room index unit, a time series storage unit, a temperature data storage unit, a radiation data storage unit, a soil moisture data storage unit, and an observed total photosynthetic amount storage unit. The distribution update module includes a prior distribution generation unit, a joint loss receiving unit, a likelihood term generation unit, a posterior distribution generation unit, a posterior result normalization unit, a confidence interval extraction unit, a candidate parameter set recombination unit, and an iteration state determination unit. Specifically, the prior distribution generation unit generates a prior distribution based on the lower and upper limits corresponding to each dynamic physiological parameter; the joint loss receiving unit receives the multi-environment joint loss value output by the loss assessment module; the likelihood term generation unit generates a likelihood term corresponding to the dynamic physiological parameter set based on the multi-environment joint loss value; the posterior distribution generation unit generates a posterior distribution based on the prior distribution and the likelihood term; the posterior result normalization unit normalizes the posterior distribution result; the confidence interval extraction unit extracts the parameter confidence interval for each dynamic physiological parameter based on the normalized posterior distribution; the candidate parameter set recombination unit generates a new round of dynamic physiological parameter set based on the parameter confidence interval; and the iteration state determination unit determines whether the iteration termination condition is met based on the distribution update results of the current and previous rounds.