Crop yield prediction system and method based on automatic modeling agent

By constructing a crop yield prediction system using automated modeling intelligent agents, the problems of poor regional generalization ability and difficulty in linking prediction results in existing technologies have been solved. This has enabled efficient and dynamic crop yield prediction and supply chain linkage, thereby improving the decision-making efficiency of grain supply and dispatch.

CN122022012APending Publication Date: 2026-05-12XIAMEN YUZHI FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN YUZHI FUTURE TECHNOLOGY CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing crop yield forecasting methods have limited ability to respond to extreme weather and sudden factors, poor regional generalization ability, and difficulty in linking forecast results with the supply chain. This results in low efficiency and high engineering costs of manual modeling, slow updates and strong subjectivity, making it difficult to achieve high-frequency dynamic updates and lacking a closed-loop mechanism to support food supply and dispatch.

Method used

A crop yield forecasting system based on automated modeling agents is adopted. Through multi-source data collection and standardized processing, a regional adaptive yield forecasting model is constructed. Combined with supply chain data, supply and demand assessment is carried out to generate strategy recommendations to support the scheduling of food supply.

Benefits of technology

It has improved the efficiency of modeling and forecasting updates, enhanced the adaptability of cross-regional forecasts, improved the timeliness of supply scheduling decisions, and realized intelligent management of the entire process of crop production, processing, storage, transportation and supply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122022012A_ABST
    Figure CN122022012A_ABST
Patent Text Reader

Abstract

The invention provides a crop yield prediction system and method based on an automatic modeling agent. Comprising a data acquisition module used for performing standardization processing and space-time alignment on multi-source data to generate a unified data set; the feature construction module is used for constructing a feature set for prediction according to a stage rule and a time window rule related to the growth process of the target crop; the automatic modeling module is used for generating region self-adaptive yield prediction models corresponding to different regions; the multi-scale yield prediction module is used for executing scale summarization on the yield prediction results of the regional scale to generate a yield prediction result of a higher scale; the supply chain safety evaluation module is used for generating a supply and demand evaluation result and a risk identifier according to a preset supply and demand evaluation rule; and the strategy output module is used for generating strategy suggestions or instruction information for grain insurance supply scheduling. According to the method, the modeling and prediction updating efficiency can be improved, the cross-region prediction adaptability is enhanced, and the timeliness of an insurance supply scheduling decision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent analysis technology for agricultural data, and in particular to a crop yield prediction system and method based on an automated modeling intelligent agent. Background Technology

[0002] Crops (such as soybeans) are important sources of grain, oil, and feed. Changes in their production directly affect supply chain decisions such as import arrangements, reserve plans, crushing schedules, port and inland logistics capacity allocation, and feed cost control. Countries, enterprises, and trading entities typically need to understand the production trends and risks in major producing areas several months in advance.

[0003] Existing crop yield forecasting methods mainly fall into two categories: one is statistical models such as linear regression and stepwise regression, which are usually based on meteorological averages, historical average yields, etc., to build predictive relationships; the other is subjective judgment methods that rely on surveys and experience, such as estimates formed by observing sown area, pod formation, and growth.

[0004] However, the above methods still have the following shortcomings: statistical models have limited ability to respond to sudden factors such as extreme weather, drought, and pests and diseases, and the hydrothermal structure and characteristic sensitivity of different regions vary significantly, leading to a high engineering cost due to the reliance on extensive manual intervention for model transfer and generalization; experience-based judgment methods are slow to update, costly, and highly subjective, making it difficult to achieve high-frequency dynamic updates at the weekly or monthly level. Furthermore, existing forecasts mostly focus on outputting production figures, lacking a closed-loop mechanism that links forecast results with supply chain operation information such as inventory, crushing, and port arrival schedules, making it difficult to directly form an actionable basis for supply guarantee scheduling. Summary of the Invention

[0005] In view of this, embodiments of this application provide a crop yield prediction system and method based on an automated modeling agent to solve the problems of low efficiency of manual modeling, poor regional generalization ability, and difficulty in linking prediction results with supply guarantee scheduling in the prior art.

[0006] The first aspect of this application provides a crop yield prediction system based on an automated modeling agent, comprising: a data acquisition module for acquiring multi-source data related to the growth and supply chain operation of a target crop, and performing standardization and spatiotemporal alignment on the multi-source data to generate a unified dataset; a feature construction module for constructing a feature set for prediction based on the unified dataset, according to stage rules and time window rules related to the growth process of the target crop; an automated modeling module for the automated modeling agent to perform candidate model selection and parameter optimization based on the feature set, generating a regional adaptive yield prediction model corresponding to different regions, and iteratively updating the regional adaptive yield prediction model when new data arrives; a multi-scale yield prediction module for outputting regional-scale yield prediction results using the regional adaptive yield prediction model, and performing scale aggregation on the regional-scale yield prediction results to generate higher-scale yield prediction results; a supply chain security assessment module for associating the yield prediction results with supply chain operation data, and generating supply and demand assessment results and risk labels according to preset supply and demand assessment rules; and a strategy output module for generating strategy suggestions or instruction information for grain supply security scheduling based on the supply and demand assessment results and risk labels.

[0007] The second aspect of this application provides a method for predicting crop yield based on an automated modeling agent using the system of the first aspect. The method includes: acquiring multi-source data related to the growth and supply chain operation of a target crop; performing standardization and spatiotemporal alignment on the multi-source data to generate a unified dataset; constructing a feature set for prediction based on the unified dataset according to stage rules and time window rules related to the growth process of the target crop; having the automated modeling agent perform candidate model selection and parameter optimization based on the feature set to generate a regional adaptive yield prediction model corresponding to different regions, and iteratively updating the regional adaptive yield prediction model when new data arrives; outputting regional-scale yield prediction results using the regional adaptive yield prediction model, and performing scale aggregation on the regional-scale yield prediction results to generate higher-scale yield prediction results; associating the yield prediction results with supply chain operation data, generating supply and demand assessment results and risk indicators according to preset supply and demand assessment rules; and generating strategy suggestions or instruction information for grain supply scheduling based on the supply and demand assessment results and risk indicators.

[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The system employs a data acquisition module to acquire multi-source data related to the growth of the target crop and the operation of its supply chain. This data is then standardized and spatiotemporally aligned to generate a unified dataset. A feature construction module, based on this unified dataset, constructs a feature set for prediction according to stage rules and time window rules related to the target crop's growth process. An automated modeling module, powered by the feature set, selects candidate models and optimizes parameters to generate regionally adaptive yield prediction models for different regions. These models are iteratively updated as new data arrives. A multi-scale yield prediction module outputs regional-scale yield predictions using the regionally adaptive yield prediction models and performs scale aggregation to generate higher-scale yield predictions. A supply chain security assessment module correlates yield predictions with supply chain operation data, generating supply and demand assessment results and risk indicators based on preset supply and demand assessment rules. Finally, a strategy output module generates strategy suggestions or instructions for grain supply security scheduling based on the supply and demand assessment results and risk indicators. This application improves modeling and prediction update efficiency, enhances cross-regional forecast adaptability, and improves the timeliness of supply security scheduling decisions. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of the structural composition of the crop yield prediction system based on automated modeling intelligent agents provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the crop yield prediction method based on automated modeling intelligent agents provided in this application embodiment. Detailed Implementation

[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0012] Crops (such as soybeans) are important sources of grain, oil, and feed. Countries, businesses, and traders need to predict production levels months in advance to plan imports, reserves, crushing capacity, and logistics. If forecasts are delayed or inaccurate, it will directly affect crushing schedules, feed costs, inventory security, and even local food supply plans.

[0013] Existing methods for predicting crop yields mainly fall into the following two categories: 1. Traditional statistical models (linear regression, stepwise regression, etc.) are often based on meteorological averages and historical average yields, making it difficult to capture sudden risks such as extreme weather, drought, and pests and diseases.

[0014] 2. Experience-based subjective judgment, such as surveying farmers' planting area and pod formation. This method is costly, slow to update, and highly subjective, making it impossible to achieve high-frequency (weekly / monthly) updates.

[0015] However, the aforementioned existing technologies still have the following problems: For supply chain scheduling departments, what they truly need is not "a static year-end production figure," but rather: dynamic knowledge at each key stage—sowing, seedling, flowering and grain-filling, and pre-harvest—regarding: 1) which state or production region will experience a production reduction; 2) by how much; 3) when this reduction will affect crushing plants, ports, and feed mills; and 4) whether it's necessary to pre-schedule import shipments or utilize reserves. Existing methods almost never create a closed-loop digital system for "forecasting → supply chain decision-making → supply guarantee execution," let alone achieve automation, reusability, and self-evolution.

[0016] Meanwhile, agricultural forecasting modeling faces two typical problems: 1) The features are extremely complex (rainfall, temperature, radiation, vegetation index, sowing time, popular varieties, diseases, differences in farmer management, etc.); 2) The sensitivity of features varies greatly across different regions (for example, the states of Mato Grosso and Bahia in Brazil have completely different hydrothermal structures). Therefore, traditional methods require a large amount of manual parameter tuning / model replacement, resulting in extremely high engineering costs.

[0017] To address the problems of complex data processing, low efficiency of manual modeling, poor model generalization ability, and ineffective guidance of food supply chain security decisions in existing crop yield forecasting methods, this application proposes a crop yield forecasting system based on an automated modeling agent. The system utilizes multi-source agricultural data, integrating remote sensing imagery, meteorological observations, sown area, historical yields, and supply chain operation indicators. Through an automated modeling agent with adaptive learning and optimization capabilities, it achieves automatic modeling and dynamic forecasting of soybean yields under different regions and climatic conditions. Furthermore, the system integrates the forecasting results with a supply chain model, constructing a closed-loop mechanism of yield forecasting—supply and demand assessment—security strategy—feedback update. This allows forecasting information to be directly transformed into actionable decision-making suggestions for food scheduling and security, thereby achieving intelligent management of the entire process of crop production, processing, storage, transportation, and supply security.

[0018] The specific framework and functions of the crop yield prediction system based on automated modeling intelligent agents provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments. Figure 1 This is a schematic diagram of the structural composition of the crop yield prediction system based on automated modeling intelligent agents provided in the embodiments of this application, as shown below. Figure 1 As shown, the crop yield prediction system based on automated modeling agents may specifically include the following modules: The data acquisition module 101 is used to acquire multi-source data related to the growth of the target crop and the operation of the supply chain, and to perform standardization processing and spatiotemporal alignment on the multi-source data to generate a unified dataset. Feature construction module 102 is used to construct a feature set for prediction based on a unified dataset, according to stage rules and time window rules related to the growth process of the target crop. The automated modeling module 103 is used by the automated modeling agent to perform candidate model selection and parameter optimization based on the feature set, generate regional adaptive yield prediction models corresponding to different regions, and iteratively update the regional adaptive yield prediction models when new data arrives. The multi-scale yield forecasting module 104 is used to output regional-scale yield forecasting results using a regional adaptive yield forecasting model, and to perform scale aggregation on the regional-scale yield forecasting results to generate higher-scale yield forecasting results. The supply chain security assessment module 105 is used to correlate production forecast results with supply chain operation data and generate supply and demand assessment results and risk labels according to preset supply and demand assessment rules. The strategy output module 106 is used to generate strategy suggestions or instruction information for grain supply scheduling based on supply and demand assessment results and risk identification.

[0019] In some embodiments, multi-source data related to the growth of the target crop and the operation of the supply chain are acquired, and the multi-source data are standardized and spatiotemporally aligned to generate a unified dataset, including: Acquire remote sensing observation data related to the growth of target crops, and generate remote sensing index data characterizing crop growth based on the remote sensing observation data; Acquire meteorological observation data related to the growth of the target crop, and perform spatial aggregation on the meteorological observation data according to the preset regional granularity, and perform temporal aggregation according to the preset time slice granularity to generate meteorological element data; Obtain area data and historical yield data corresponding to the target crop, and obtain supply chain operation data related to supply chain operation; Standardization processing is performed on multi-source data, including unifying field definitions, converting units, handling missing data, and handling abnormal data. Based on region identifiers and time slice identifiers, the multi-source data that has undergone standardization is associated and aligned to form a data record set with a preset structure, which serves as a unified dataset.

[0020] Specifically, this embodiment uses soybeans as the target crop and organizes data using "regional identifier + time slice identifier" as the primary key. A multi-source data acquisition and fusion process is established. By standardizing and spatiotemporally aligning remote sensing observation data, meteorological observation data, area data, historical yield data, and supply chain operation data, a unified dataset is formed for subsequent feature construction and automated modeling. This unified dataset is expressed using a pre-structured data record set. Each data record corresponds to a set of indicators for a region within a time slice. The indicator set includes at least one or more of the following: crop growth indicators, meteorological elements, area and yield-related indicators, and supply chain operation indicators.

[0021] In some examples, region identifiers are used to characterize the spatial granularity of prediction and evaluation, and can be state / province, production area grid, or port service area; time slice identifiers are used to characterize the temporal granularity of data alignment, and can be daily, weekly, or monthly. This embodiment uses state-level regions and monthly time slices as examples. Region identifiers can use major production area identifiers such as Mato Grosso and Bahia states in Brazil, and time slice identifiers can be in the form of "year-month" to ensure that multi-source data from different sources can be correlated and aligned under a unified spatial and temporal coordinate system.

[0022] Furthermore, the data acquisition module obtains remote sensing observation data covering the target area through a satellite remote sensing data interface, such as Sentinel-2 imagery or multispectral imagery data of equivalent resolution. To ensure that the remote sensing data reflects crop growth and is suitable for subsequent modeling, the data acquisition module performs regionalization processing on the remote sensing observation data: first, the image is cropped according to the geographical boundaries of the target area, and invalid pixels such as clouds and shadows are removed or marked; then, multiple images of the same area are time-aggregated within a time slice to form a remote sensing observation sequence for that area within that time slice; furthermore, vegetation index and leaf area correlation indicators are calculated based on the remote sensing observation sequence to generate remote sensing index data characterizing crop growth. The remote sensing index data may include one or more of NDVI, EVI, and LAI.

[0023] For example, within the "2025-06" time slice, for the state of Mato Grosso, the system collects multiple images from that month, removes cloud cover, calculates NDVI and EVI separately, and then performs temporal aggregation on the NDVI and EVI for that month to obtain the crop growth index value for that month in the state, which is used as a remote sensing index field in the subsequent unified dataset.

[0024] Furthermore, the data acquisition module obtains meteorological observation data related to the growth of the target crop from meteorological observation or reanalysis data sources. This meteorological observation data may include daily or hourly records of precipitation, maximum temperature, minimum temperature, and average temperature. Since meteorological observations typically involve uneven station distribution or differences in grid resolution, this embodiment performs spatial aggregation on the meteorological observation data according to a preset regional granularity to obtain a regional-level meteorological sequence consistent with the regional identifier; then, it performs temporal aggregation on the regional-level meteorological sequence according to a preset time-slice granularity to generate meteorological element data.

[0025] For example, the system performs regional aggregation on meteorological data from stations within the boundary of Mato Grosso state to obtain daily precipitation and daily temperature sequences for the state. Subsequently, within the "2025-06" time slice, it performs cumulative aggregation on daily precipitation and average or extreme value aggregation on daily temperature to generate meteorological data such as cumulative precipitation, average temperature, maximum temperature, and minimum temperature for the state in that month, allowing it to be directly aligned with remote sensing index data on the same time slice.

[0026] Furthermore, the data acquisition module obtains area data and historical yield data corresponding to the target crop. The area data can be the sown area or the harvested area, and the historical yield data can be the historical yield per unit area or the historical total output of the region. The area data and historical yield data can come from agricultural statistics, planting structure databases, or data sources maintained by the business side, and are indexed by regional identifiers and time slice identifiers.

[0027] Simultaneously, the data acquisition module obtains supply chain operation data related to supply chain operations. This data includes at least one or more of the following: inventory data, crushing operation data, import arrival schedule data, consumption data, and transportation capacity data. Supply chain operation data can originate from data sources such as port inventory systems, oil mill operation systems, trade arrival plans and customs clearance records, feed consumption statistics, and inland logistics capacity monitoring. To ensure its usability for subsequent supply and demand assessment, this embodiment also maps supply chain operation data to regional and time-slice identifiers. For example, port inventory is grouped by port service area, oil mill crushing rate is grouped by oil mill location, and import arrival schedule is grouped by destination port and its service area, thus maintaining consistency with the regional scale of production forecast output.

[0028] In some examples, due to differences in field naming, unit definitions, missing data patterns, and outlier patterns among multi-source data, this embodiment performs standardization processing on the multi-source data before performing spatiotemporal alignment to ensure consistent dataset definitions. Standardization processing includes at least field definition unification, unit conversion, missing data handling, and outlier data handling.

[0029] Unified field definitions include standardizing the naming and statistical definitions of indicators with the same meaning. For example, mapping "port inventory," "available inventory," and "in-port inventory" to a unified field and labeling the definitions. Unit conversion includes standardizing units of precipitation, temperature, inventory, production, and area to make subsequent model inputs comparable in terms of unit dimension. Missing data processing includes filling or imputing missing remote sensing indicators caused by cloud cover, missing meteorological station data, and missing supply chain indicators caused by intermittent reporting, and retaining missing data markers for quality control. Anomaly data processing includes identifying, correcting, or removing temperature, precipitation, inventory, or crushing rate data that significantly deviates from reasonable ranges, and retaining anomaly marker fields in the data records for traceability.

[0030] Furthermore, after standardization, this embodiment performs association and alignment on multi-source data based on region identifiers and time slice identifiers to form a data record set with a preset structure, serving as a unified dataset. During alignment, the "region identifier + time slice identifier" is used as the primary key, merging remote sensing index data, meteorological element data, area data, historical production data, and supply chain operation data under the same primary key to obtain a unified data record oriented towards modeling.

[0031] For example, for a primary key like "Region ID = Mato Grosso State, Time Slice ID = 2025-06", the unified data record should include at least the following data: remote sensing indicators such as NDVI and EVI for that month; meteorological data such as cumulative precipitation and average temperature; planting or harvesting area data for that region; historical yield data; and supply chain operation data such as port inventory, oil mill crushing rate, or import arrival schedule associated with that region. For a primary key like "Region ID = Bahia State, Time Slice ID = 2025-06", the system generates corresponding data records according to the same rules, thereby ensuring that the data structure of different regions is consistent under the same time slice, providing a unified entry point for subsequent feature construction based on stage rules and time window rules.

[0032] Through the above embodiments, the system standardizes, integrates, and aligns heterogeneous data from multiple sources, such as remote sensing, meteorology, agricultural statistics, and supply chain operations, at a unified regional and time-slice granularity, forming a structured and traceable unified dataset. This provides a consistent data foundation for subsequent regional adaptive modeling, rolling forecasting, supply and demand assessment, and strategy output.

[0033] In some embodiments, based on a unified dataset, a feature set for prediction is constructed according to stage rules and time window rules related to the growth process of the target crop, including: Based on the stage rules, multiple growth stages corresponding to the growth process of the target crop are determined in a unified dataset, and the unified dataset is segmented according to multiple growth stages; Remote sensing index data and meteorological element data at each growth stage are aggregated to generate stage characteristics that represent crop growth and meteorological conditions at each growth stage. A rolling aggregation process is performed on the unified dataset according to the time window rule to generate window features that characterize the changes of the unified dataset within the rolling time window. The stage features and window features are combined to obtain the feature set used for prediction.

[0034] Specifically, in this embodiment, the stage rule is used to divide the soybean growth process into multiple physiologically significant stages, so as to characterize the impact of weather and growth changes on yield at different stages. The stage rule can be determined based on a preset soybean growth calendar, which includes at least several stages such as sowing period, emergence period, vegetative growth period, flowering and grain filling period, and maturity and harvesting period. For situations where there are differences in sowing time in different regions, this embodiment uses sowing time information associated with regional identifiers or regional growth calendar parameters to adaptively correct the stage boundaries, so that the same stage corresponds to different start and end time slices in different regions.

[0035] For example, taking the Brazilian states of Mato Grosso and Bahia as examples, the two states have significant differences in hydrothermal structure and different sowing windows. The system maintains or infers the corresponding sowing start time slice for each region and determines the stage boundaries of the sowing period, vegetative growth period, and flowering and grain filling period for that region accordingly. When entering the "2025-06" prediction time slice, Mato Grosso may be in the flowering and grain filling period, while Bahia may still be in the vegetative growth period. The system segments the unified dataset according to the rules of each stage to avoid different regions being forcibly mapped to the same growth stage, which would cause feature distortion.

[0036] When performing segmentation, the feature construction module reads the time-series data records of the region using the "region identifier + time slice identifier" as the primary key in the unified dataset. Based on the stage rules, the time-series data of the region is divided into multiple stage data segments. Each stage data segment contains multiple time slice records covered by that stage, and the records contain at least remote sensing index data and meteorological element data.

[0037] Furthermore, in order to transform the growth status and meteorological conditions within each stage into modelable numerical representations, the feature construction module performs stage aggregation processing on the remote sensing index data and meteorological element data within each growth stage to generate stage features. Stage aggregation processing includes at least two types of operations: stage statistics calculation and stage change characterization.

[0038] Stage statistics calculations are used to extract stage averages, stage cumulatives, stage extremes, or stage fluctuation characteristics. For example, stage cumulative precipitation can be calculated for precipitation during the flowering and grain-filling period, stage average temperature can be calculated for the average temperature within the stage, stage extremes can be calculated for the highest temperature, and stage averages or fluctuation ranges can be calculated for the diurnal temperature range.

[0039] Stage change characterization is used to extract trend features or inflection point features of remote sensing indicators over time. For example, during the flowering and grain-filling stage, trend fitting of the remote sensing indicator sequence can yield the intensity of changes in crop growth intensity during the stage, thus characterizing the stage-specific changes in crop growth status. For instance, "stage NDVI slope" and "accumulated heat temperature" can be constructed. Here, the stage slope of the remote sensing indicator sequence can be used as a stage feature, and the accumulated temperature within the stage can be used to form a heat accumulation-type stage feature.

[0040] For example, in Mato Grosso state during the flowering and grain-filling stage, the system extracts NDVI sequences, precipitation, and temperature sequences from multiple time slices covering this stage from a unified dataset, and calculates stage characteristics such as average NDVI, NDVI slope, cumulative precipitation, average temperature, and heat accumulation. For Bahia state during the vegetative growth stage, the system calculates stage characteristics such as cumulative precipitation, average temperature, average NDVI, and slope for the corresponding stage data segments. Because the stage boundaries adjust with the region, the stage characteristics can more accurately reflect the environmental pressure and growth status of different regions within the corresponding growth stage.

[0041] Furthermore, time window rules are used to capture short- and medium-term dynamic changes with a fixed-length sliding window without relying on the boundaries of the reproductive stage, supporting weekly or monthly rolling forecasting needs. Time window rules can consist of multiple window lengths, such as 30 days, 60 days, and 90 days. The end point of the window is aligned with the forecast time slice identifier, ensuring that the window features are all derived from observable data records before or up to the forecast time point.

[0042] The feature construction module performs rolling aggregation on a unified dataset according to time window rules to generate window features. Rolling aggregation includes at least cumulative window features, mean window features, and fluctuation window features.

[0043] For example, within the window corresponding to the "2025-06" prediction time slice, the system calculates the cumulative precipitation over the past thirty days, the average temperature over the past sixty days, and the diurnal temperature range fluctuation over the past ninety days. It can further calculate "abnormal fluctuation" characteristics, that is, to differentiate the precipitation or temperature within the window from the historical distribution or historical average, so as to characterize the window-level impact of abnormal events such as drought and heat waves.

[0044] Furthermore, after obtaining the stage features and window features, the feature construction module combines the two types of features to form a feature set for prediction. The combination method can be to merge them using a unified key, that is, using "region identifier + prediction time slice identifier" as the index, merging the stage feature fields and window feature fields corresponding to the prediction time point of the region into a single sample record.

[0045] For example, for "Region ID = Mato Grosso State, Prediction Time Slice ID = 2025-06", the system-generated sample records can simultaneously include stage features such as cumulative precipitation during the flowering and grain-filling stage, the slope of NDVI change during the stage, and cumulative heat accumulation during the stage, as well as window features such as cumulative precipitation over the past 30 days, average temperature over the past 60 days, and temperature fluctuation over the past 90 days. For "Region ID = Bahia State, Prediction Time Slice ID = 2025-06", the system-generated sample records include stage features corresponding to its vegetative growth stage and window features under the same window length rule. Through this combination, the feature set retains both the sensitivity information of the reproductive stage and the short-to-medium-term dynamic change information, which can support the needs of automated modeling agents for differentiated modeling and rolling updates for different regions.

[0046] In some examples, when combining phase features and window features, the feature construction module can also perform consistency checks and availability checks on feature fields to ensure that the feature dimensions and field meanings of each sample record remain consistent across different regions. This allows the automated modeling module to perform unified data input and model evaluation during candidate model selection and parameter optimization. Through the above embodiments, the system completes the transformation from a unified dataset to a set of modelable features, providing direct input for the subsequent construction and iterative updates of regional adaptive yield prediction models.

[0047] In some embodiments, an automated modeling agent performs candidate model selection and parameter optimization based on a feature set to generate a regional adaptive yield prediction model corresponding to different regions, including: Based on the feature set, historical sample sets are constructed for different regions, and the historical sample sets are rolled and divided in chronological order to form training datasets and validation datasets. The candidate models are trained in the preset candidate model set, and the parameter configuration corresponding to the candidate models is determined by the preset parameter optimization strategy. The evaluation results of different candidate models and parameter configurations on the validation dataset are compared according to the preset evaluation rules. The candidate models and corresponding parameter configurations that meet the preferred conditions are selected as the regional adaptive yield prediction models associated with the corresponding regions.

[0048] Specifically, in this embodiment, the automated modeling agent first extracts historical sample sequences of the corresponding regions from the feature set output by the feature construction module, forming a historical sample set for that region. Each sample in the historical sample set is indexed by "regional identifier + time slice identifier," containing observable stage features and window features of that time slice, and is associated with the supervised target value corresponding to that region in that time slice. The supervised target value can be the historical statistical value of regional yield or regional total output. To meet the temporal constraints of agricultural forecasting, the automated modeling agent does not use a random partitioning method, but instead performs a rolling partitioning of the historical sample set according to chronological order to form a training dataset and a validation dataset.

[0049] For example, for Mato Grosso, the agent organizes multi-year sample sequences along an annual timeline. In each rolling partition, earlier years and their corresponding monthly samples are used as the training dataset, while samples from one or more subsequent years are used as the validation dataset. For Bahia, the agent employs the same rolling partitioning strategy, but due to differences in planting windows, rainfall seasonality, and crop growth rhythms, its sample distribution and feature sensitivity differ significantly from Mato Grosso. Therefore, it is necessary to independently create training and validation datasets and independently perform model search and parameter optimization. Through this method, the agent ensures that the validation dataset is later than the training dataset, making the model evaluation more closely resemble the real-world application scenario of "predicting the future from history."

[0050] Furthermore, the automated modeling agent pre-maintains a candidate model set, which covers a family of models from different modeling paradigms to adapt to the complex characteristics of agriculture, significant nonlinear relationships, and strong regional heterogeneity. The candidate model set may include one or more of the following: gradient boosting tree models, categorical feature processing models, multilayer perceptron models, and temporal hybrid models.

[0051] The agent trains candidate models one by one or in parallel from the candidate model set for each region's training dataset. During training, the agent takes stage features and window features from the feature set as input and historical yield per unit area or historical total yield as the supervision target, forming a regression training task. For candidate models containing time-series fusion mechanisms, the agent can use feature sequences from multiple time slices as input windows, enabling the model to learn the cumulative impact of crop growth changes and weather fluctuations on yield. For tree models or shallow network-type candidate models, the agent takes the comprehensive feature vector corresponding to the current prediction time slice as input to achieve regional-scale yield regression prediction.

[0052] Furthermore, after setting up the candidate model training framework, the automated modeling agent uses a preset parameter optimization strategy to determine the parameter configuration of each candidate model. The parameter optimization strategy can be Bayesian optimization, evolutionary search, or an iterative optimization strategy based on the search space. In each optimization iteration, the agent generates a set of parameter candidates, applies these parameter candidates to the training process of the corresponding candidate model, obtains evaluation results on the validation dataset, and then updates the parameter candidates for the next round according to the optimization strategy.

[0053] For example, for Mato Grosso, the agent might iterate multiple times on parameters such as the learning rate, tree depth, and number of leaf nodes in the tree model; for Bahia, the agent might iterate on parameters such as sequence length, fusion layer width, and regularization parameters in the temporal hybrid model. Because the feature sensitivities of the two states differ significantly, the optimal parameter configurations obtained are often different, thus achieving regional adaptation.

[0054] Furthermore, the automated modeling agent compares the evaluation results of different candidate models and their parameter configurations on the validation dataset according to preset evaluation rules, and selects the candidate model and corresponding parameter configuration that meet the optimization criteria as the regional adaptive yield prediction model associated with the corresponding region. The preset evaluation rules can adopt a multi-index comprehensive evaluation method, and the evaluation indicators can include one or more of the following: mean absolute percentage error, root mean square error, and coefficient of determination. The optimization criteria can be set as "optimal comprehensive score" or "preferential selection of models with higher stability under the condition that the error index meets the threshold".

[0055] For example, for Mato Grosso state, the agent compares the yield prediction error and interannual stability of different models in the validation year. If a candidate model maintains a low error even in the main years of reduced yield, it is selected as the regional adaptive yield prediction model for that state. For Bahia state, the agent focuses more on the generalization performance under seasonal variations in rainfall, prioritizing candidate models that perform stably under different climatic conditions. After determining the optimal model, the agent saves the model's structural identifier, parameter configuration, and evaluation result summary for that region, allowing it to be directly accessed by the multi-scale yield prediction module and providing a baseline configuration for iterative updates when new data arrives.

[0056] Through the above embodiments, the system utilizes an automated modeling agent to organize training and validation samples in different regions in a rolling manner over time, and automatically trains, optimizes, and selects the best candidate model to form a regional adaptive yield prediction model. This reduces the reliance on manual model selection and parameter tuning for cross-regional modeling and improves the availability and update efficiency of the model under different climatic conditions and regional differences.

[0057] In some embodiments, the regional adaptive yield prediction model is iteratively updated as new data arrives, including: New data corresponding to a unified dataset is accessed, and features of the new data are updated based on stage rules and time window rules to form incremental samples; Whether to initiate a model update is determined based on preset update trigger conditions. The update trigger conditions include at least one of the following: the time span of newly added data coverage meets a threshold condition and the change in model prediction error meets a threshold condition. When determining whether to initiate a model update, the regional adaptive yield prediction model is retrained or incrementally updated based on incremental samples to obtain the updated regional adaptive yield prediction model.

[0058] Specifically, new data refers to the latest data records that continuously arrive at a preset time slice granularity after the model has been deployed and is running. New data includes at least one or more of the following: remote sensing index data generated from the latest remote sensing observation data, meteorological element data generated from the latest meteorological observation data, and supply chain operation data updated within the same time slice. The automated modeling module accesses the new data through the update interface of the data acquisition module and, following the standardized processing and spatiotemporal alignment rules in the corresponding embodiment of claim 2, maps the new data into a "regional identifier + time slice identifier" structure record consistent with the unified dataset, and writes it into the unified dataset to form an expanded unified dataset.

[0059] For example, using "2025-07" as a new time slice, the system accesses remote sensing imagery of Mato Grosso state for that month and generates remote sensing index data such as NDVI and EVI. It also accesses the precipitation and temperature series for that month and aggregates them to form meteorological element data. Simultaneously, it accesses supply chain operation data corresponding to that month, such as port inventory, oil mill crushing rates, or arrival schedules. After standardization and spatiotemporal alignment, a new data record is generated with "Region ID = Mato Grosso state, Time slice ID = 2025-07" and merged into a unified dataset. The same rules are applied to other regions, such as Bahia state, to merge corresponding new data records.

[0060] Furthermore, after the new data is incorporated into a unified dataset, the feature construction module, under the scheduling of the automated modeling module, updates the features of the data corresponding to the new time slices based on existing stage rules and time window rules to form incremental samples. Incremental samples refer to the set of observable features and their index information for each region in the new prediction time slice, used for subsequent model updates and inference.

[0061] In some examples, for the phase rule, the system first determines the phase affiliation of the new time slice in the growth process of the region, and updates the phase aggregation results within that phase, such as updating the phase characteristics such as cumulative precipitation, average temperature, average remote sensing index and trend within the phase; for the time window rule, the system takes the new time slice as the end point of the window and updates the rolling cumulative, rolling average and fluctuation characteristics of the windows of the past thirty days, sixty days or ninety days to form window characteristics.

[0062] For example, after Mato Grosso state enters the flowering and grain-filling stage, when new data for "2025-07" arrives, the system incorporates this time slice data into the flowering and grain-filling stage data segment and updates the stage's cumulative precipitation and NDVI slope. Simultaneously, it updates window characteristics such as the cumulative precipitation over the past sixty days ending in "2025-07" and the temperature fluctuation over the past ninety days; thus forming an incremental sample with "Region Identifier = Mato Grosso State, Forecast Time Slice Identifier = 2025-07". For Bahia state, if it is still in the vegetative growth stage in this time slice, the stage affiliation and stage aggregation range will differ accordingly, and the generated incremental sample reflects the stage differences and window dynamics of the region.

[0063] Furthermore, to avoid the computational overhead caused by indiscriminately updating the model every time new data arrives, the automated modeling module sets preset update trigger conditions to determine whether to initiate a model update. The update trigger conditions must include at least one of the following: the time span covered by the new data meets a threshold condition, and the change in model prediction error meets a threshold condition.

[0064] The threshold condition for the time span of newly added data coverage is used to determine whether the cumulative amount of newly added data has reached the minimum scale to support stable updates. For example, when the newly added data of several consecutive time slices has been merged into a unified dataset, the time span is considered to meet the threshold condition. The threshold condition for the change in model prediction error is used to determine whether the error of the model in the most recent rolling prediction has changed significantly. For example, the prediction error of the model for the most recently revealed historical output is compared with the historical error baseline. When the error increases or the fluctuation expands beyond the threshold, the error change is considered to meet the threshold condition.

[0065] For example, the system performs monthly rolling forecasts for the Mato Grosso state model and records the forecast error. When new data has been available for three consecutive months but the error remains stable, the system can trigger only a minor update. When a drought occurs in a certain month, causing remote sensing indicators and meteorological elements to deviate significantly from their historical distributions, and the model's forecast error increases significantly in that month or subsequent months, the system determines that the error change meets the threshold condition and initiates a model update so that the model can be re-adapted to the new climate and crop distribution.

[0066] Furthermore, when determining to initiate a model update, the automated modeling module performs retraining or incremental updates on the regional adaptive yield prediction model based on incremental samples, resulting in an updated regional adaptive yield prediction model.

[0067] Retraining and updating is suitable for scenarios where the newly added data covers a long time span, the data distribution has shifted significantly, or the model structure needs to be re-optimized. In this case, the automated modeling agent can reorganize the historical sample set of the region, re-incorporate the feature set and supervision target corresponding to the expanded unified dataset into the training and validation datasets, and use the time-rolling partitioning mechanism and candidate model set search mechanism to retrain, optimize parameters, and optimize the model, outputting a new regional adaptive yield prediction model.

[0068] Incremental updates are suitable for scenarios where the model structure is stable and the new data is mainly used to refine recent parameters. In this case, the automated modeling module incorporates incremental samples into the training process while keeping the model structure unchanged, and updates the model parameters using a preset incremental learning strategy, so that the model can absorb the latest growth and weather change information carried by the new time slices.

[0069] For example, for tree-based regional adaptive yield prediction models, model parameters can be retrained using a preset strategy after adding samples or adjusted incrementally. For models containing a time-series fusion substructure, some parameters can be fine-tuned using new samples under a fixed structure to better adapt to the impact of recent changes in remote sensing indicators and meteorological fluctuations on yield. After the update is completed, the system saves the updated model version, its corresponding training time range, and evaluation summary, enabling the multi-scale yield prediction module to automatically call the latest version model in subsequent predictions.

[0070] Through the above embodiments, as new remote sensing, meteorological, and supply chain operation data continue to arrive, the system can synchronously update features based on predetermined stage rules and time window rules, and trigger model retraining or incremental updates through time span thresholds or prediction error change thresholds, thereby realizing the rolling iterative update of the regional adaptive yield prediction model, improving the prediction model's adaptability to the latest data distribution changes, and reducing the need for repetitive manual modeling.

[0071] In some embodiments, the regional adaptive yield prediction model includes: The feature input substructure is used to receive a set of features and generate corresponding feature representations. The temporal fusion substructure is used to fuse the feature representations corresponding to different time slices to obtain fused feature representations; The regression prediction substructure is used to output the yield prediction results associated with the corresponding region based on the fused feature representation.

[0072] Specifically, the feature input substructure receives the feature set and generates the corresponding feature representation. The feature set originates from the feature construction module and includes stage features and window features, indexed by "region identifier + prediction time slice identifier". The feature input substructure first performs field mapping and vectorization processing on the feature set, mapping feature fields from different sources and with different statistical calibers into fixed-dimensional input vectors. Among them, stage features are used to characterize the growth status and meteorological conditions during stages such as sowing, vegetative growth, flowering and grain filling, and maturity and harvesting, while window features are used to characterize the short- and medium-term dynamic changes obtained by rolling aggregation with a preset window length.

[0073] In some implementations, to ensure the consistency of inputs for models in different regions, the feature input substructure can align feature fields based on a preset feature dictionary, and use a preset imputation strategy and add missing markers for missing fields, thereby ensuring that the input vector has consistent dimensions and semantics in different regions and time slices.

[0074] For example, for the feature set "Region ID = Mato Grosso State, Predicted Time Slot ID = 2025-07", the feature input substructure maps the window features corresponding to this time slot, such as the stage cumulative precipitation, stage average temperature, stage remote sensing index change trend, and the cumulative precipitation and average temperature of the past thirty days and the past sixty days, into an input vector, and outputs the feature representation for subsequent fusion. For "Region ID = Bahia State, Predicted Time Slot ID = 2025-07", the feature input substructure completes the mapping with the same feature dictionary, so that the models of the two states are consistent in terms of input interface, but the feature value distribution reflects regional differences.

[0075] Furthermore, the temporal fusion substructure is used to fuse the feature representations corresponding to different time slices to obtain a fused feature representation. Unlike using only single-time-slice features, this embodiment, to enhance the expression of the cumulative effect of crop growth, allows the temporal fusion substructure to receive feature representations from multiple adjacent time slices as input and perform fusion calculations along the time dimension. Multiple time slices can correspond to consecutive time slices within the same growth stage, or to multiple time slices covered by a time window rule, thus enabling the fusion process to reflect the continuous temporal changes of growth indicators and meteorological elements.

[0076] Temporal fusion can be implemented using either weighted fusion or sequence fusion. In weighted fusion, the temporal fusion substructure assigns weights to the feature representations of each time slice and performs weighted summation or weighted combination of the feature representations based on these weights. The weights can be determined based on time distance, stage position, or preset rules to highlight the contribution of time slices that are more sensitive to prediction. In sequence fusion, the temporal fusion substructure performs sequence modeling processing on the feature representation sequence to learn the correlation between temporal patterns such as meteorological anomalies and inflection points of remote sensing indicators and yield changes.

[0077] For example, for Mato Grosso state, the temporal fusion substructure inputs the feature representations corresponding to "2025-05, 2025-06, 2025-07" into the fusion process to generate fused feature representations, reflecting the cumulative impact of precipitation changes, temperature fluctuations, and remotely sensed growth changes during the key growth stages in the state. For Bahia state, the fused time slice set and stage affiliation may be different, but the temporal fusion substructure still generates fused feature representations according to the same fusion mechanism, thereby achieving adaptive expression of regional differences.

[0078] Furthermore, the regression prediction substructure is used to output yield prediction results associated with the corresponding region based on the fused feature representation. The regression prediction substructure receives the fused feature representation output by the time-series fusion substructure and performs regression mapping to obtain the yield prediction results. The yield prediction results can be expressed in tons / hectare or equivalent, and serve as input for the multi-scale yield prediction module to further fuse with area data to calculate the regional total output prediction results.

[0079] In some implementations, the regression prediction substructure can be set as a one-layer or multi-layer regression mapping structure, and its structural configuration can be determined based on the candidate model selection and parameter optimization results of the automated modeling agent, so that it can adapt to nonlinear feature relationships and maintain scalability in different regions.

[0080] For example, in the Mato Grosso state model, the regression prediction substructure outputs the state's yield forecast for the "2025-07" forecast time slot based on the fused feature representation; in the Bahia state model, the regression prediction substructure outputs the state's yield forecast for the same forecast time slot. Both provide yield forecasts through the same output radial multi-scale yield forecast module, facilitating subsequent regional total output forecasts and higher-scale aggregation.

[0081] Through the above embodiments, the regional adaptive yield prediction model achieves unified access to stage features and window features, fusion expression of cross-time slice features, and output of yield prediction results by combining feature input substructure, temporal fusion substructure and regression prediction substructure. This enables the model to not only depict the temporal changes in the crop growth process, but also maintain a consistent structural interface under the feature sensitivity differences in different regions, supporting the automated modeling agent to perform regional adaptive modeling and rolling updates in the candidate model set.

[0082] In some embodiments, a regional adaptive yield forecasting model is used to output regional-scale yield forecasts, and scale aggregation is performed on the regional-scale yield forecasts to generate higher-scale yield forecasts, including: Input the feature set corresponding to the target region into the regional adaptive yield prediction model, and output the yield prediction result of the target region; The total output forecast for the target region is obtained by fusing the yield forecast results with the corresponding area data in a unified dataset, and this serves as the regional-scale output forecast result. The total output forecast results from multiple regions are weighted or cumulatively aggregated according to preset aggregation rules to generate higher-scale output forecast results.

[0083] Specifically, in this embodiment, for each target region, the multi-scale yield prediction module first extracts the feature set entries corresponding to that target region from the feature set, and inputs the feature set entries into the regional adaptive yield prediction model associated with that target region, outputting the yield prediction result for that target region in the prediction time slice. The yield prediction result can be expressed in tons / hectare or equivalent caliber, and should be consistent with the statistical caliber of historical yield data so that the same supervised target definition can be used in the training and inference phases.

[0084] For example, within the "2025-07" prediction time slot, the system extracts feature sets corresponding to Mato Grosso and Bahia states respectively, inputs them into their respective regional adaptive yield prediction models, and obtains the yield prediction results for each state within that time slot. Due to differences in reproductive stage, hydrothermal structure, and crop growth variations between the two states, their yield prediction results reflect regional differences, providing a basis for subsequent calculations of regional total output and national aggregation.

[0085] Furthermore, after obtaining the yield per unit area forecast, the multi-scale yield forecast module reads area data matching the target region and forecast time slice from a unified dataset, and performs a fusion operation based on the yield per unit area forecast and the area data to obtain the total yield forecast for the target region, which serves as the regional-scale yield forecast. The area data can be at least one of the sown area data or the harvested area data. The system can select the area type according to a preset caliber during fusion to ensure that the total yield forecast is consistent with the supply caliber of the supply chain assessment.

[0086] For example, for the Mato Grosso state in the "2025-07" forecast time slot, the system reads the state's sown area or harvested area data from a unified dataset, and then merges the yield forecast with the area data to obtain the state's total yield forecast. The same operation is performed for the Bahia state. Since both the area data and the yield forecast are indexed and aligned using "regional identifier + time slot identifier," the fusion process has a clear reference basis and avoids calculation errors caused by cross-regional and cross-time slot mismatches.

[0087] Furthermore, after obtaining total output forecasts for multiple regions, the multi-scale output forecasting module performs weighted or cumulative summation of the total output forecasts for multiple regions according to preset summation rules, generating higher-scale output forecasts. The preset summation rules can be selected as cumulative summation or summation by weight based on statistical caliber, where the weights can be related to area data, regional contribution, or preset zoning weights.

[0088] For example, the system aims to aggregate the total production forecasts of all states in Brazil at the national level to obtain the total production forecast for the whole country. In some implementations, the system can also set different aggregation weights or regional weights for major producing areas and non-major producing areas to form regional aggregation results and national aggregation results, which will help the subsequent supply chain security assessment module to identify the impact path of production reduction in major producing areas on national supply.

[0089] Through the above embodiments, the multi-scale production forecasting module realizes a continuous calculation link from regional feature set to regional yield forecast, then to regional total production forecast and higher-scale summary forecast. This enables the forecast results to support the judgment of "which region will reduce production and by how much" at the state or production area level, and to form a total forecast at the national level, providing consistent input data for subsequent supply and demand assessment, gap calculation and supply guarantee strategy output.

[0090] In some embodiments, production forecast results are correlated with supply chain operation data, and supply and demand assessment results and risk labels are generated according to preset supply and demand assessment rules, including: Production forecast results are matched and correlated with supply chain operation data based on regional and time slice identifiers. According to the preset supply and demand assessment rules, the available supply represented by the production forecast results and the inventory, port replenishment and processing consumption represented by the supply chain operation data are used to calculate the supply and demand balance, and the supply and demand gap information including gap size information and gap time window information is obtained as the supply and demand assessment result. Based on supply and demand gap information, determine the risk category or risk level corresponding to the regional identifier and time slice identifier, and generate risk identifiers.

[0091] Specifically, the supply chain security assessment module uses "regional identifier + time slice identifier" as the matching primary key to match and correlate production forecast results with supply chain operation data. Production forecast results include total production forecasts at the regional scale and optional higher-scale aggregated forecasts. Supply chain operation data includes at least one or more of the following: inventory data, crushing operation data, import arrival schedule data, consumption data, and transportation capacity data. To ensure consistency, the supply chain operation data has already been mapped and aligned according to preset regional and time slice granularities during the data acquisition module phase. Therefore, in this embodiment, matching can be directly completed using the same primary key within a unified dataset or its derived index.

[0092] For example, under the primary key "Region ID = Port Service Region A, Time Slice ID = 2025-07", the system can simultaneously read the corresponding total production forecast, port inventory, planned arrival replenishment, oil mill crushing operations, and feed consumption data for that region. Under the primary key "Region ID = Mato Grosso State, Time Slice ID = 2025-07", the system can read the total production forecast for that state and associate it with its corresponding inventory, processing, and arrival schedule data. For cross-regional transmission scenarios, the region ID can also be selected with spatial granularity consistent with the supply chain nodes, such as "Port Service Region" or "Oil Mill Supply Radius Region", so that the forecast results can be directly mapped to the scheduling object.

[0093] After the matching and association are completed, the supply chain security assessment module calculates the supply and demand balance by combining the available supply represented by the production forecast results with the inventory, arrival replenishment and processing consumption represented by the supply chain operation data, according to the preset supply and demand assessment rules. The result is the supply and demand gap information.

[0094] In this embodiment, the available supply is determined by production forecast results and can be expanded into a regional supply sequence by time slice; inventory is used to represent the initial available inventory or the available inventory at port; port arrivals are used to represent the arrival of imported supplies in each time slice; processing consumption is used to represent the demand-side consumption formed by crushing consumption, feed consumption, or other processing consumption. Pre-defined supply and demand assessment rules at least define the statistical scope and time alignment of each quantity, ensuring that supply and demand balance calculations are completed under the same criteria.

[0095] For example, within the continuous time slice from "July 2025 to September 2025", the system can allocate the total output forecast for a certain region into available supply according to the harvest schedule, and then overlay it with the region's initial inventory and planned port arrival replenishment to form an available supply sequence. Next, the system converts the oil mill's crushing operations and feed consumption into a processing consumption sequence and subtracts from this, obtaining the supply-demand difference sequence for each time slice. When the supply-demand difference is negative, a gap is formed. Thus, the system obtains supply-demand gap information that includes gap size information and gap time window information. The gap size information is used to characterize the cumulative or peak amount of the gap, and the gap time window information is used to characterize the start and end time slices and duration of the gap.

[0096] Furthermore, after obtaining supply and demand gap information, the supply chain security assessment module determines the risk category or risk level corresponding to the regional identifier and time slice identifier based on the supply and demand gap information, and generates a risk identifier. The risk category or risk level can be classified according to the gap size, duration, time and location of the gap occurrence, and supply chain operation constraints, forming a risk input that can be directly used by the strategy output module.

[0097] For example, when the gap size exceeds a preset threshold and its duration covers the key peak crushing season, the system marks the area as high-risk; when the gap size is small but concentrated within the time window of insufficient port replenishment, the system can mark it as medium-risk and prompt attention to the import pace; when the supply-demand difference is positive and inventory is within a safe range, the system can mark it as low-risk. In cases where there is a significant production reduction forecast in major producing areas and the gap time window overlaps with the time window of rapid decline in port inventory, the system can generate risk labels for the corresponding areas and aggregate them spatially to form a "production reduction risk distribution" result, facilitating dispatchers to identify key areas and key time periods.

[0098] Through the above embodiments, the system can associate the production forecast results with supply chain operation data such as inventory, port replenishment, and processing consumption using a unified primary key, and complete the supply and demand balance calculation according to the preset supply and demand assessment rules. It outputs the supply and demand assessment results including the gap size and gap time window, as well as the corresponding risk indicators, so that the crop production forecast results can be directly transformed into quantifiable supply chain security risk inputs, providing a basis for the subsequent generation of supply guarantee strategies and the output of scheduling instructions.

[0099] In some embodiments, based on supply and demand assessment results and risk identification, strategic recommendations or instructions for grain supply scheduling are generated, including: Based on the gap size and gap time window information in the supply and demand assessment results, and combined with risk indicators, the target supply guarantee strategy type is determined from the set of preset supply guarantee strategies; The target supply guarantee strategy type is parameterized according to the preset constraint rules to determine the execution time window and resource scheduling parameters corresponding to the target supply guarantee strategy type. The resource scheduling parameters include at least one of the following: import rhythm adjustment parameters, reserve mobilization parameters, transportation capacity allocation parameters, and processing scheduling parameters. The output includes strategy recommendations or instructions that include execution time windows and resource scheduling parameters.

[0100] Specifically, the strategy output module pre-maintains a set of preset supply guarantee strategies. This set includes at least one or more of the following: import schedule adjustment strategies, reserve mobilization strategies, transportation capacity allocation strategies, and processing production scheduling adjustment strategies. It may further include regional allocation strategies or combinations of multiple strategies. Based on the gap size and gap time window information from the supply and demand assessment results, and in conjunction with risk indicators, the strategy output module determines the target supply guarantee strategy type from the preset set of strategies.

[0101] In some implementations, the strategy output module selects the strategy type according to preset mapping rules, which at least associate the risk level with the location of the gap time window. For example, when the risk indicator is high and the gap time window is located within a continuous interval of one or more future time slices, a combination of import rhythm adjustment strategy and processing production scheduling adjustment strategy is preferred; when the gap time window is short and occurs within a time slice where port inventory is low, a reserve mobilization strategy and transportation capacity allocation strategy are preferred; when the gap size is small but there is a risk of harvest delay in the main producing areas, a transportation capacity allocation strategy and processing production scheduling adjustment strategy can be preferred to smooth short-term raw material supply fluctuations.

[0102] For example, if the system identifies a significant gap with a high risk level in the case of "Region Identifier = Port Service Area A, Time Slot Identifier = August 2025 to October 2025", the strategy output module determines the target supply guarantee strategy type as a combination of import rhythm adjustment strategy and reserve mobilization strategy, and locks the relevant execution time window. For short-term gaps in the case of "Region Identifier = Supply Radius Area of ​​a Certain Oil Plant, Time Slot Identifier = September 2025", the strategy output module can determine the target supply guarantee strategy type as a combination of processing production scheduling adjustment strategy and transportation capacity allocation strategy to ensure the oil plant's raw material supply.

[0103] Furthermore, after determining the target supply guarantee strategy type, the strategy output module parameterizes the target supply guarantee strategy type according to preset constraint rules to determine the execution time window and resource scheduling parameters corresponding to the target supply guarantee strategy type. The preset constraint rules are used to limit the execution of the strategy within the actual resources and operational boundaries. The constraint objects include at least port throughput capacity, arrival window constraints, inland transportation capacity, available reserve scale, oil mill crushing capacity and maintenance plan, etc.

[0104] The results of parameterized configuration include execution time windows and resource scheduling parameters, wherein the resource scheduling parameters include at least one or more of the following: import rhythm adjustment parameters, reserve mobilization parameters, transportation capacity allocation parameters, and processing scheduling parameters.

[0105] In some examples, import rhythm adjustment parameters can be used to characterize the amount and distribution of supplies that need to arrive at the port earlier or later within the gap time window; reserve mobilization parameters can be used to characterize the quantity, batches, and release areas of reserves mobilized in a specified time slice; transport capacity allocation parameters can be used to characterize the allocation of transport capacity from the port to the oil mill or consumption area, the occupation and allocation routes of vehicle and ship transport capacity; and processing production scheduling parameters can be used to characterize the adjustment range of oil mill crushing load, the adjustment of production sequence, or the rhythm of raw material substitution.

[0106] For example, when the gap time window falls between "August 2025 to October 2025" and there is an upper limit constraint on port throughput, the strategy output module decomposes the import rhythm adjustment parameter into multiple arrival batches during parameterization configuration, and sets the reserve mobilization parameter to be released first in "September 2025" to cover the most intense time slot. At the same time, it allocates the transport capacity allocation parameter in combination with the inland transport capacity constraint, so that the reserve release and arrival replenishment can spatially cover key oil mills and feed consumption areas. If an oil mill has a maintenance window in "September 2025", the processing production scheduling parameter can adjust the crushing load to be arranged in different time slots before and after the maintenance to avoid conflict with the maintenance constraint.

[0107] Furthermore, after completing the parameterization configuration, the strategy output module outputs strategy suggestions or instructions containing execution time windows and resource scheduling parameters. The strategy suggestions or instructions at least include a strategy type identifier, an applicable region identifier, an execution time window identifier, and resource scheduling parameter fields, enabling the scheduling execution object to perform actions such as adjusting shipping schedules, releasing reserves, allocating transport capacity, or adjusting crushing production schedules.

[0108] For example, the system can output strategy suggestions such as "Applicable area = port service area A, execution time window = August 2025 to October 2025, strategy type = import rhythm adjustment + reserve mobilization, resource scheduling parameters = arrival supply allocation and reserve release allocation"; for the oil mill side, it outputs instruction information such as "Applicable area = oil mill supply radius area B, execution time window = September 2025, strategy type = processing production adjustment + transportation capacity allocation, resource scheduling parameters = crushing load adjustment range and transportation capacity allocation", thereby realizing the implementation output from supply and demand gap and risk identification to executable scheduling actions.

[0109] Through the above embodiments, the system can determine the target supply guarantee strategy type from the preset supply guarantee strategy set based on the gap size and gap time window in the supply and demand assessment results and combined with risk indicators. Under the constraints of port throughput, reserve scale, transportation capacity and crushing capacity, the system can complete the parameterized configuration of the strategy and output strategy suggestions or instruction information containing execution time windows and resource scheduling parameters, thereby transforming the production forecast results into an execution basis that can be directly used for grain supply guarantee scheduling.

[0110] The above embodiments have described in detail the specific modules and functions of the crop yield prediction system based on automated modeling intelligent agents of this application. The implementation process of the crop yield prediction method based on automated modeling intelligent agents of this application will be described in detail below with reference to specific embodiments. Figure 2 This is a flowchart illustrating the crop yield prediction method based on automated modeling agents provided in this application embodiment, as shown below. Figure 2As shown, the crop yield prediction method based on automated modeling agents may specifically include the following steps: S201: Acquire multi-source data related to the growth of the target crop and the operation of the supply chain, perform standardization processing and spatiotemporal alignment on the multi-source data, and generate a unified dataset; S202, based on a unified dataset, constructs a feature set for prediction according to the stage rules and time window rules related to the growth process of the target crop; S203, the automated modeling agent performs candidate model selection and parameter optimization based on the feature set, generates regional adaptive yield prediction models corresponding to different regions, and iteratively updates the regional adaptive yield prediction models when new data arrives; S204, using the regional adaptive yield forecasting model to output regional-scale yield forecasting results, and performing scale aggregation on the regional-scale yield forecasting results to generate higher-scale yield forecasting results. S205 links production forecast results with supply chain operation data and generates supply and demand assessment results and risk labels according to preset supply and demand assessment rules; S206 generates strategic recommendations or instructions for grain supply scheduling based on supply and demand assessment results and risk identification.

[0111] Specifically, the data acquisition module first obtains data related to soybean growth and supply chain operation from multiple heterogeneous sources. The multi-source data includes at least two of the following: remote sensing observation data, meteorological observation data, area data, historical yield data, and supply chain operation data.

[0112] For remote sensing observation data, this embodiment can acquire remote sensing images covering the target area from a satellite remote sensing data interface, and calculate vegetation index and leaf area correlation indicators based on the remote sensing images to generate remote sensing index data. This remote sensing index data is used to characterize crop growth status, and may include one or more of NDVI, EVI, and LAI. For meteorological observation data, this embodiment acquires daily or hourly meteorological observation sequences such as precipitation and temperature, and performs spatial aggregation according to a preset regional granularity and temporal aggregation according to a preset time slice granularity to generate meteorological element data consistent with a unified primary key.

[0113] In addition, this embodiment acquires area data and historical yield data corresponding to the target crop to ensure consistency between yield per unit area forecast and total output forecast; at the same time, it acquires supply chain operation data such as inventory, crushing operation, import arrival schedule, consumption, and transportation capacity, and maps them to a record structure consistent with the regional identifier and time slice identifier.

[0114] During the data processing phase, the above-mentioned multi-source data undergoes standardization processes, including field caliber unification, unit conversion, missing data handling, and anomaly data handling. Subsequently, based on region identifiers and time slice identifiers, the standardized multi-source data is associated and aligned to form a data record set with a pre-defined structure, serving as a unified dataset.

[0115] Furthermore, the stage rule is used to divide the unified dataset into multiple growth stages along the regional dimension, such as sowing period, vegetative growth period, flowering and grain-filling period, and maturity and harvesting period, and to segment the unified dataset according to multiple growth stages. Since the sowing time and hydrothermal structure are different in different regions, this embodiment can determine the stage boundaries for each region separately, so that Mato Grosso State and Bahia State may correspond to different growth stages in the same time slice, thereby making the feature representation closer to the actual growth rhythm of the region.

[0116] In the stage aggregation processing, stage statistics and change characterization are performed on remote sensing index data and meteorological element data within each growth stage to generate stage characteristics that represent crop growth and meteorological conditions within the stage, such as stage cumulative precipitation, stage average temperature, and stage remote sensing index change trends.

[0117] The time window rule is used to perform rolling aggregation processing on a uniform dataset according to a preset window length to generate window features. Window features are used to characterize short-term or medium-term dynamic changes before the prediction time point, such as cumulative precipitation in the past 30 days, average temperature in the past 60 days, temperature fluctuations and abnormal deviations in the past 90 days, etc. Subsequently, the stage features and window features are combined to form a feature set indexed by "region identifier + prediction time slice identifier", providing a unified input for the automated modeling agent.

[0118] Furthermore, the automated modeling agent constructs historical sample sets for different regions. These historical sample sets are indexed by "region identifier + time slice identifier" and include feature sets and corresponding historical output supervision targets. To ensure temporal consistency and generalization reliability, the agent performs rolling partitioning of the historical sample sets according to chronological order, forming training and validation datasets.

[0119] Subsequently, the agent trains candidate models from a pre-defined candidate model set and uses a pre-defined parameter optimization strategy to determine the parameter configuration corresponding to the candidate models. The candidate model set can cover gradient boosting tree models, neural network models, and time series hybrid models, to adapt to scenarios with complex agricultural characteristics, significant nonlinear relationships, and strong regional heterogeneity; the parameter optimization strategy can employ Bayesian optimization or evolutionary search, achieving automatic parameter tuning through iterative comparison on a validation dataset.

[0120] Based on preset evaluation rules, the evaluation results of different candidate models and their parameter configurations are compared on the validation dataset. The candidate model and corresponding parameter configuration that meet the optimization criteria are selected as the regional adaptive yield prediction model associated with the corresponding region. The evaluation rules can use single indicators or combinations of multiple indicators such as mean absolute percentage error, root mean square error, and coefficient of determination.

[0121] Furthermore, the multi-scale yield prediction module, for each target region, inputs the feature set corresponding to that target region into the regional adaptive yield prediction model and outputs the yield prediction result for that target region. Subsequently, it reads area data matching the target region and prediction time slice from a unified dataset, and performs a fusion calculation based on the yield prediction result and the area data to obtain the total yield prediction result for the target region, which serves as the regional-scale yield prediction result. The area data can be either the sown area or the harvested area; the system can select the area type according to a preset caliber to ensure consistency between the total yield prediction and the supply caliber.

[0122] Furthermore, the supply chain security assessment module uses regional identifiers and time slice identifiers as matching primary keys to match and correlate production forecast results with supply chain operation data. The available supply represented by the production forecast results, together with the inventory, arrival replenishment, and processing consumption represented by the supply chain operation data, constitute the core inputs for supply and demand assessment. Supply and demand balance calculations are performed according to preset supply and demand assessment rules to obtain supply and demand gap information, including gap size and gap time window information, as the supply and demand assessment result.

[0123] After obtaining information on the supply-demand gap, risk categories or risk levels are determined based on the gap size and time window, corresponding to regional and time-segment identifiers, and risk identifiers are generated. For example, a higher risk level is generated when the gap size exceeds a threshold and continuously covers key peak crushing season time segments; a lower risk level is generated when the gap size is small or only occurs in the short term, so that risk identifiers can be directly used as input constraints in the strategy output stage.

[0124] Furthermore, the strategy output module pre-maintains a set of preset supply guarantee strategies, which includes at least one or more of the following: import schedule adjustment strategy, reserve mobilization strategy, transportation capacity allocation strategy, and processing production scheduling adjustment strategy. Based on the gap size information and gap time window information in the supply and demand assessment results, and in conjunction with risk indicators, the strategy output module determines the target supply guarantee strategy type from the preset set of supply guarantee strategies.

[0125] After determining the target supply guarantee strategy type, the strategy type is parameterized according to preset constraint rules to determine the corresponding execution time window and resource scheduling parameters. These resource scheduling parameters include at least one of the following: import schedule adjustment parameters, reserve mobilization parameters, transport capacity allocation parameters, and processing production scheduling parameters. They are subject to constraints such as port throughput capacity, available reserve scale, inland transport capacity, and crushing capacity. The final output includes strategy suggestions or instructions containing the execution time window and resource scheduling parameters, which are then used by the target entity to adjust shipping schedules, release reserves, allocate transport capacity, or adjust crushing production schedules.

[0126] Through the above steps, this application utilizes an automated modeling agent to select candidate models and optimize parameters in multiple regions based on a unified dataset and feature set constructed with stage rules and time window rules. It also completes rolling model updates when new data arrives, further outputs multi-scale production forecast results, and links them with supply chain operation data to form supply and demand gaps and risk indicators. Finally, it generates parameterizable and executable supply guarantee scheduling strategy suggestions or instruction information, thus forming an integrated methodology of "production forecasting - supply and demand assessment - strategy output".

[0127] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although the technical solutions of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A crop yield prediction system based on an automated modeling agent, characterized in that, include: The data acquisition module is used to acquire multi-source data related to the growth of the target crop and the operation of the supply chain, and to perform standardization processing and spatiotemporal alignment on the multi-source data to generate a unified dataset. The feature construction module is used to construct a feature set for prediction based on the unified dataset, according to the stage rules and time window rules related to the growth process of the target crop. An automated modeling module is used by an automated modeling agent to perform candidate model selection and parameter optimization based on the feature set, generate regional adaptive yield prediction models corresponding to different regions, and iteratively update the regional adaptive yield prediction models when new data arrives. The multi-scale yield prediction module is used to output regional-scale yield prediction results using the regional adaptive yield prediction model, and to perform scale aggregation on the regional-scale yield prediction results to generate higher-scale yield prediction results. The supply chain security assessment module is used to associate the production forecast results with the supply chain operation data and generate supply and demand assessment results and risk labels according to preset supply and demand assessment rules. The strategy output module is used to generate strategy suggestions or instruction information for grain supply scheduling based on the supply and demand assessment results and the risk identification.

2. The system according to claim 1, characterized in that, The process of acquiring multi-source data related to the growth of the target crop and the operation of its supply chain, and performing standardization and spatiotemporal alignment on the multi-source data to generate a unified dataset includes: Acquire remote sensing observation data related to the growth of the target crop, and generate remote sensing index data characterizing the crop growth based on the remote sensing observation data; Acquire meteorological observation data related to the growth of the target crop, and perform spatial aggregation on the meteorological observation data according to a preset regional granularity, and perform temporal aggregation according to a preset time slice granularity to generate meteorological element data; Obtain area data and historical yield data corresponding to the target crop, and obtain supply chain operation data related to supply chain operation; The standardization process involves unifying field definitions, converting units, handling missing data, and processing abnormal data on the multi-source data. Based on region identifiers and time slice identifiers, the multi-source data that has undergone standardization is associated and aligned to form a data record set with a preset structure, which serves as the unified dataset.

3. The system according to claim 1, characterized in that, Based on the unified dataset, a feature set for prediction is constructed according to stage rules and time window rules related to the growth process of the target crop, including: Based on the stage rules, multiple growth stages corresponding to the growth process of the target crop are determined in the unified dataset, and the unified dataset is segmented according to the multiple growth stages. Remote sensing index data and meteorological element data for each growth stage are aggregated in stages to generate stage characteristics that characterize crop growth and meteorological conditions within the growth stage. The unified dataset is subjected to rolling aggregation processing according to the time window rules to generate window features that characterize the changes of the unified dataset within the rolling time window. The stage features are combined with the window features to obtain the feature set used for prediction.

4. The system according to claim 1, characterized in that, The step of the automated modeling agent performing candidate model selection and parameter optimization based on the feature set to generate regional adaptive yield prediction models corresponding to different regions includes: Based on the feature set, historical sample sets are constructed for different regions, and the historical sample sets are rolled into training datasets and validation datasets in chronological order. Train candidate models in a preset candidate model set, and use a preset parameter optimization strategy to determine the parameter configuration corresponding to the candidate models; The evaluation results of different candidate models and parameter configurations on the validation dataset are compared according to the preset evaluation rules. The candidate model and corresponding parameter configuration that meet the preferred conditions are selected as the regional adaptive yield prediction model associated with the corresponding region.

5. The system according to claim 4, characterized in that, The iterative update of the regional adaptive yield prediction model upon the arrival of new data includes: New data corresponding to the unified dataset is accessed, and the new data is updated with features based on the stage rules and time window rules to form incremental samples; Whether to initiate a model update is determined based on preset update trigger conditions, wherein the update trigger conditions include at least one of the following: the time span of newly added data coverage meets a threshold condition and the change in model prediction error meets a threshold condition. When determining to initiate a model update, the regional adaptive yield prediction model is retrained or incrementally updated based on the incremental samples to obtain the updated regional adaptive yield prediction model.

6. The system according to claim 5, characterized in that, The regional adaptive yield prediction model includes: A feature input substructure is used to receive the feature set and generate the corresponding feature representation; The temporal fusion substructure is used to fuse the feature representations corresponding to different time slices to obtain fused feature representations; The regression prediction substructure is used to output the unit yield prediction result associated with the corresponding region based on the fused feature representation.

7. The system according to claim 1, characterized in that, The step of using the regional adaptive yield prediction model to output regional-scale yield prediction results, and performing scale aggregation on the regional-scale yield prediction results to generate higher-scale yield prediction results, includes: The feature set corresponding to the target region is input into the regional adaptive yield prediction model, and the yield prediction result of the target region is output. The total output prediction result of the target area is obtained by fusing the unit yield prediction result with the area data corresponding to the unified dataset, and is used as the output prediction result at the regional scale. The total production forecast results for multiple regions are weighted and aggregated or cumulatively aggregated according to preset aggregation rules to generate the higher-scale production forecast results.

8. The system according to claim 1, characterized in that, The step of linking the production forecast results with supply chain operation data and generating supply and demand assessment results and risk indicators according to preset supply and demand assessment rules includes: The production forecast results are matched and associated with the supply chain operation data based on the region identifier and time slice identifier; According to the preset supply and demand assessment rules, the available supply represented by the production forecast results and the inventory, port replenishment and processing consumption represented by the supply chain operation data are used to calculate the supply and demand balance, and the supply and demand gap information including gap size information and gap time window information is obtained, which is used as the supply and demand assessment result. Based on the supply and demand gap information, determine the risk category or risk level corresponding to the region identifier and time slice identifier, and generate the risk identifier.

9. The system according to claim 1, characterized in that, The step of generating strategy recommendations or instructions for grain supply scheduling based on the supply and demand assessment results and the risk identification includes: Based on the gap size information and gap time window information in the supply and demand assessment results, and in conjunction with the risk identification, the target supply guarantee strategy type is determined from the preset supply guarantee strategy set; The target supply guarantee strategy type is parameterized according to preset constraint rules to determine the execution time window and resource scheduling parameters corresponding to the target supply guarantee strategy type. The resource scheduling parameters include at least one of import rhythm adjustment parameters, reserve mobilization parameters, transportation capacity allocation parameters, and processing production scheduling parameters. The output includes strategy suggestions or instruction information containing the execution time window and the resource scheduling parameters.

10. A method for predicting crop yield based on an automated modeling agent using a system as described in any one of claims 1 to 9, characterized in that, include: Acquire multi-source data related to the growth of the target crop and the operation of the supply chain, perform standardization and spatiotemporal alignment on the multi-source data, and generate a unified dataset; Based on the unified dataset, a feature set for prediction is constructed according to the stage rules and time window rules related to the growth process of the target crop. An automated modeling agent performs candidate model selection and parameter optimization based on the feature set to generate regional adaptive yield prediction models corresponding to different regions, and iteratively updates the regional adaptive yield prediction models when new data arrives. The regional adaptive yield prediction model is used to output regional-scale yield prediction results, and the regional-scale yield prediction results are scaled to generate higher-scale yield prediction results. The production forecast results are correlated with supply chain operation data, and supply and demand assessment results and risk labels are generated according to preset supply and demand assessment rules. Based on the supply and demand assessment results and the risk identification, strategy recommendations or instructions for grain supply scheduling are generated.