Grain yield prediction method, device and medium based on large model and time series retrieval

By performing time alignment, spatial mapping, and growth stage labeling on multimodal data, and combining large-scale models with time-series retrieval techniques, the problem of unified modeling of multimodal data in grain yield prediction was solved, resulting in more stable and interpretable yield prediction.

CN122019798BActive Publication Date: 2026-07-31北京衔远有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
北京衔远有限公司
Filing Date
2026-04-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing grain yield forecasting methods struggle to achieve unified modeling and consistent alignment of cross-modal information under conditions such as inconsistent spatiotemporal resolution of multimodal data, significant differences in phenological stages, and frequent extreme weather events. They also fail to adequately utilize evolutionary patterns from similar historical years and lack traceable forecasting interpretation mechanisms.

Method used

By acquiring multimodal data (meteorological time series, remote sensing images, and agricultural information text), time alignment, spatial mapping, and growth stage labeling are performed to construct a temporal segment representation of the current growth stage. Then, using large models and temporal retrieval techniques, combined with cross-modal attention mechanisms, feature fusion and reinforcement learning are performed to output yield predictions and explanatory information.

Benefits of technology

It enhances the ability of multimodal fusion modeling, improves the stability and interpretability of predictions for extreme years, and strengthens the accuracy of predictions under extreme weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019798B_ABST
    Figure CN122019798B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and medium for grain yield prediction based on a large model and time-series retrieval. The method includes: acquiring multimodal data corresponding to a target region, target crop, and prediction period; processing the multimodal data to obtain aligned data; fusing the encoding results based on a cross-modal attention mechanism to obtain multimodal fusion features; constructing a time-series representation of the current growth stage and obtaining yield label information associated with similar historical segments; inputting the prediction context into a supervised-trained large model and outputting explanatory information; triggering self-reflective verification of the initial yield prediction value and explanatory information and correcting it based on a reinforcement learning strategy; outputting the yield prediction value and uncertainty interval of the target region and target crop, and outputting the influencing factors associated with the yield prediction value and similar historical segment reference information. This application can improve multimodal fusion modeling capabilities, enhance the prediction stability in extreme years, and improve the interpretability of prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and medium for predicting grain yield based on large models and time-series retrieval. Background Technology

[0002] Grain yield forecasting is a crucial foundation for agricultural production organization, grain storage and regulation, and disaster response. It typically requires combining meteorological conditions, growth monitoring, and agricultural information during crop growth to predict the yield or yield per unit area of ​​a target region and crop within the forecast period. Current yield forecasting methods largely rely on statistical regression, mechanistic models, or machine learning models based on remote sensing and meteorological features. One type of method establishes empirical relationships between historical yields and meteorological indicators; another type uses remote sensing features such as vegetation indices to characterize growth and performs regression estimation; and some methods attempt to incorporate textual agricultural information or disaster records as auxiliary variables.

[0003] However, while the aforementioned methods can achieve a certain predictive capability in regular years, they are difficult to achieve unified modeling and consistent alignment of cross-modal information under conditions such as inconsistent spatiotemporal resolution of multi-source data, significant differences in phenological stages, and frequent extreme weather events. At the same time, regression inference based on fixed features cannot effectively utilize the evolution patterns and corresponding yield labels of similar historical years, resulting in insufficient basis for extrapolation in extreme years. In addition, the inference process of existing methods is mostly a black box output, lacking referable historical evidence and traceable influencing factor links, and lacking self-reflection and continuous error correction mechanisms when prediction deviations occur, making it difficult to meet the requirements of the business side for stability and interpretability. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus and medium for predicting grain yield based on large models and time-series retrieval, in order to solve the problems of difficulty in unifying and merging multimodal data, insufficient utilization of historical similar cases and lack of traceable basis for prediction interpretation in the prior art.

[0005] A first aspect of this application provides a method for predicting grain yield based on a large model and time-series retrieval, comprising: acquiring multimodal data corresponding to a target region, a target crop, and a prediction period, wherein the multimodal data includes meteorological time series, remote sensing images, and agricultural text; performing time alignment, spatial mapping, and growth stage labeling on the multimodal data to obtain aligned data associated with a unified spatial unit and a unified time axis; performing feature encoding on the meteorological time series, remote sensing images, and agricultural text in the aligned data respectively, and fusing the encoding results based on a cross-modal attention mechanism to obtain multimodal fusion features; and constructing a time-series segment representation of the current growth stage based on the multimodal fusion features and the growth stage labeling. Using the current growth stage time-series segment representation as the retrieval condition, similar historical segments are retrieved from the historical time-series segment knowledge base, and yield label information associated with similar historical segments is obtained. The multimodal fusion features, the current growth stage time-series segment representation, similar historical segments, and yield label information are combined into a prediction context, and the prediction context is input into a supervised training large model, outputting the initial yield prediction value and the corresponding explanatory information. The initial yield prediction value and explanatory information are subject to self-reflective verification and correction based on a reinforcement learning strategy, outputting the yield prediction value and uncertainty interval of the target region and target crop, and outputting the influencing factors associated with the yield prediction value and similar historical segment reference information.

[0006] A second aspect of this application provides a grain yield prediction device based on a large model and time-series retrieval, comprising: an acquisition module for acquiring multimodal data corresponding to a target region, a target crop, and a prediction period, wherein the multimodal data includes meteorological time series, remote sensing images, and agricultural text; an alignment module for performing time alignment, spatial mapping, and growth stage labeling on the multimodal data to obtain aligned data associated with a unified spatial unit and a unified time axis; a fusion module for performing feature encoding on the meteorological time series, remote sensing images, and agricultural text in the aligned data respectively, and fusing the encoding results based on a cross-modal attention mechanism to obtain multimodal fusion features; and a construction module for constructing a time-series slice of the current growth stage based on the multimodal fusion features and the growth stage labeling. The system comprises the following modules: a segment representation; a retrieval module, which uses the current growth stage time-series segment representation as the retrieval condition to search for similar historical segments in the historical time-series segment knowledge base and obtain yield label information associated with similar historical segments; a prediction module, which combines multimodal fusion features, the current growth stage time-series segment representation, similar historical segments, and yield label information into a prediction context, inputs the prediction context into a supervised training large model, and outputs the initial yield prediction value and the corresponding explanatory information; and an output module, which triggers self-reflective verification of the initial yield prediction value and explanatory information and makes corrections based on reinforcement learning strategies, outputs the yield prediction value and uncertainty interval of the target region and target crop, and outputs the influencing factors associated with the yield prediction value and the reference information of similar historical segments.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0009] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: By acquiring multimodal data corresponding to the target region, target crop, and prediction period, including meteorological time series, remote sensing images, and agricultural text, the multimodal data undergoes time alignment, spatial mapping, and growth stage labeling to obtain aligned data associated with a unified spatial unit and a unified time axis. Feature encoding is then performed on the meteorological time series, remote sensing images, and agricultural text within the aligned data, and the encoding results are fused based on a cross-modal attention mechanism to obtain multimodal fusion features. Based on the multimodal fusion features and growth stage labeling, a temporal segment representation of the current growth stage is constructed. This temporal segment representation of the current growth stage is then used as the retrieval condition. This method retrieves similar historical segments from a historical time-series knowledge base and obtains yield label information associated with these segments. It combines multimodal fusion features, the current growth stage time-series segment representation, similar historical segments, and yield label information into a prediction context. This prediction context is then input into a supervised-trained large model, outputting an initial yield prediction value and corresponding explanatory information. The initial yield prediction value and explanatory information undergo self-reflective verification and are corrected based on a reinforcement learning strategy. The method outputs the yield prediction values ​​and uncertainty intervals for the target region and target crop, along with the influencing factors associated with the yield prediction values ​​and similar historical segment reference information. This application enhances multimodal fusion modeling capabilities, improves prediction stability in extreme years, and strengthens the interpretability of prediction results. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating the grain yield prediction method involved in a real-world scenario in this application embodiment; Figure 2 This is a flowchart illustrating the grain yield prediction method based on large model and time series retrieval provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the grain yield prediction device based on large model and time series retrieval provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] Before providing a detailed description of the embodiments of the technical solution of this application, the implementation process of the grain yield prediction method based on large model and time series retrieval of this application will be summarized first with reference to the accompanying drawings and embodiments. Figure 1 The flowchart of the grain yield prediction method involved in the embodiments of this application in a real-world scenario is shown in the figure below. Figure 1 As shown, the method includes the following main operations: Acquire multimodal data corresponding to the target area, target crop, and forecast period. The multimodal data includes meteorological time series, remote sensing images, and agricultural information text. Preprocessing and alignment of multimodal data are performed, including temporal alignment, spatial mapping and growth stage labeling, to form aligned data associated with a unified spatial unit and a unified temporal axis; Temporal encoding is performed on the meteorological time series in the aligned data to obtain meteorological embedding sequences; image encoding is performed on the remote sensing images in the aligned data to obtain image embedding sequences; text encoding is performed on the agricultural information text in the aligned data to obtain text semantic embeddings and extract event tags. Multimodal fusion features are obtained by fusing meteorological embedding sequences, image embedding sequences, text semantic embeddings, and event labels based on a cross-modal attention mechanism. Based on multimodal fusion features and combined with growth stage annotations, a temporal segment representation of the current growth stage is constructed; Using the current growth stage time series segment as the search criteria, similar historical segments are retrieved from the historical time series segment knowledge base, and yield tag information associated with similar historical segments is obtained; The multimodal fusion features, the temporal segment representation of the current growth stage, and similar historical segments and yield label information are combined into a prediction context. The prediction context is then input into a supervised training large model to generate an initial yield prediction value and the corresponding explanatory information. The system triggers self-reflective verification of the initial yield forecast and interpretation information and makes corrections based on reinforcement learning strategies. It outputs the yield forecast and uncertainty range of the target region and target crop, as well as the influencing factors associated with the yield forecast and reference information of similar historical segments.

[0014] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0015] Figure 2 This is a flowchart illustrating the grain yield prediction method based on large model and time-series retrieval provided in an embodiment of this application. Figure 2 As shown, the method may specifically include: S201, acquire multimodal data corresponding to the target area, target crop and prediction period. The multimodal data includes meteorological time series, remote sensing images and agricultural information text. S202 performs temporal alignment, spatial mapping, and growth stage labeling on multimodal data to obtain aligned data associated with a unified spatial unit and a unified time axis; S203, feature encoding is performed on meteorological time series, remote sensing images and agricultural text in the aligned data respectively, and the encoding results are fused based on cross-modal attention mechanism to obtain multimodal fusion features; S204, based on multimodal fusion features and combined with growth stage annotations, constructs a temporal segment representation of the current growth stage; S205, using the current growth stage time series segment as the retrieval condition, retrieve similar historical segments in the historical time series segment knowledge base and obtain the yield tag information associated with the similar historical segments; S206 combines multimodal fusion features, temporal segment representation of the current growth stage, similar historical segments, and yield label information into a prediction context, and inputs the prediction context into a supervised training large model, outputting the initial yield prediction value and the explanation information corresponding to the initial yield prediction value; S207 triggers self-reflective verification of the initial yield forecast and interpretation information and makes corrections based on reinforcement learning strategies. It outputs the yield forecast and uncertainty range of the target region and target crop, and outputs the influencing factors associated with the yield forecast and similar historical fragment reference information.

[0016] In some embodiments, temporal alignment, spatial mapping, and growth stage labeling are performed on multimodal data to obtain aligned data associated with a unified spatial unit and a unified temporal axis, including: Identify unified spatial units associated with the target area, and map meteorological time series, remote sensing images, and agricultural information texts to the unified spatial units to form multimodal data entries associated with the unified spatial units; A unified time axis corresponding to the forecast period is determined, and time scale unification and time index alignment are performed on the multimodal data entries to form a multimodal time series associated with the unified time axis; Based on the phenological stage division rules of the target crop, growth stage markers are labeled at each time position in the multimodal time series, and the growth stage markers are written into the corresponding multimodal time series to obtain aligned data.

[0017] Specifically, firstly, this embodiment will acquire multimodal data. For example, according to the task configuration, the system will pull all currently available information from different data sources, including but not limited to: Meteorological time series data: historical + recent daily / hourly data: temperature, precipitation, radiation, wind speed, humidity, etc.; Remote sensing / image data: optical satellite (NDVI / EVI, vegetation cover), SAR data (cloud-rich areas for observation); Text data: weekly agricultural reports, pest and disease bulletins, meteorological disaster warnings, agricultural statistical reports, and survey records; Other structured data includes historical yields, sown area, soil type, topography, irrigation conditions, and other basic parameters.

[0018] Furthermore, before the grain yield prediction task enters the feature encoding and cross-modal fusion stage, meteorological time series, remote sensing images, and agricultural text are first aligned in a unified spatiotemporal and growth stage manner to form aligned data that can be referenced by subsequent time series coding, image coding, and text coding. This aligned data uses a unified spatial unit and a unified time axis as the main index, and writes the growth stage marker corresponding to the target crop at each time position on the time axis, enabling different modal data to be associated and combined under the same spatial unit, the same time position, and the same growth stage semantics.

[0019] The unified spatial unit is used to carry the spatial mapping results of various types of data within the target area. It can use plot boundaries, administrative grids, or regular grids as spatial carriers, and assign a spatial unit identifier to each spatial unit. The unified time axis is used to carry the time alignment results of various types of data within the prediction period. It can use daily or hourly scales as the unified step size, and assign a time index to each time position. The growth stage marker is used to describe the phenological stage of the target crop at each time position on the unified time axis. It can be generated based on the sowing date, accumulated temperature threshold, calendar threshold, or phenological observation records, and together with the spatial unit identifier and time index, it forms the constraints for subsequent time series segment construction and historical retrieval.

[0020] Furthermore, when determining the unified spatial units associated with the target area, the spatial division method can be selected at the plot level or grid level according to the task configuration. For example, in the scenario of predicting corn yield in a county, the target area is divided into multiple regular grid units, and a unique identifier is generated for each grid unit; at the same time, spatial mapping rules are established to convert the spatial positioning results of different data sources into corresponding grid identifiers.

[0021] For meteorological time series, meteorological data is usually provided in the form of grid points or stations. Each meteorological record can be associated with the nearest grid cell or distributed to multiple grid cells according to area weight based on the station's latitude and longitude or the coverage relationship of meteorological grid points. For remote sensing images, remote sensing data is usually provided in the form of pixels. The set of pixels falling into grid cells can be aggregated to obtain grid-level vegetation indices, backscattering intensity, or texture statistics. For agricultural texts, the text usually carries administrative divisions, township names, place names, or survey locations. Text events can be associated with corresponding grid cells through administrative division code matching, place name resolution, or location grid placement. When it is not possible to pinpoint to a single grid cell, the event can be mapped to a set of grid cells within the administrative scope.

[0022] After completing the above spatial mapping, a multimodal data entry is formed for each grid cell. The multimodal data entry includes a set of meteorological records, a set of remote sensing observations, and a set of text events, and uses the spatial cell identifier as a unified index.

[0023] Furthermore, when determining the unified time axis corresponding to the forecast period, a unified time step is selected based on the forecast business scope, and a continuous time index sequence covering the forecast period is constructed. For example, the unified time step is set to a daily scale, and the unified time axis covers a continuous date sequence from broadcast to the current time.

[0024] For meteorological time series, if the original data is hourly, it is aggregated daily to generate daily maximum temperature, daily minimum temperature, daily cumulative precipitation, daily average humidity, etc., and the aggregated results are written into a unified time axis by date. For remote sensing images, if the observation time is not continuous or there is cloud cover, valid observations on the same or adjacent dates can be synthesized according to priority, and the synthesized grid-level remote sensing features are written into a unified time axis by date. On missing dates, a missing mark can be written and the reason for the missing data can be retained so that it can be processed in the subsequent coding stage in conjunction with the masking mechanism.

[0025] For agricultural information texts, text events can be located on a unified timeline based on the text's publication time or the event's occurrence time. Multiple text events within the same date can be merged and recorded to form a date-associated event set or event statistics. By unifying the time scale and aligning the time index for each modality, a multimodal time series associated with a unified timeline can be obtained, ensuring that each spatial unit has queryable meteorological features, remote sensing features, and text event information at every time point on the unified timeline.

[0026] Furthermore, when generating growth stage markers, the stage determination is performed and markers are written for each time position in a unified time axis according to the phenological stage division rules of the target crop. For example, stage boundary rules such as emergence, jointing, tasseling and silking, grain filling, and maturity can be preset for maize: when the sowing date is known, the sowing date can be used as the zero point of time, and then the stage transition node can be determined by combining the accumulated temperature value or the historical threshold; when historical phenological observation records are available, the observed stage dates can be used to calibrate the rules; when there are differences in sowing dates in different grid units, stage markers can be calculated separately according to the grid dimensions, so that the growth stage markers are bound to the spatial unit identifiers.

[0027] In some implementations, after generating growth stage markers, the growth stage markers are written to each time position corresponding to the multimodal time series to form aligned data. This enables subsequent modules to select time series window parameters based on the growth stage markers and retrieve similar historical segments with consistent stages from the historical time series segment knowledge base.

[0028] Through the above alignment process, aligned data with spatial unit identifiers, unified time axes, and growth stage markers as the skeleton is formed, making meteorological time series, remote sensing images, and agricultural texts usable under unified spatiotemporal and phenological semantics. This improves the consistency of subsequent cross-modal fusion and time series segment construction, reduces the impact of inconsistent resolution and sampling of different data sources on the prediction link, and provides a comparable segment representation basis for historical similar segment retrieval.

[0029] In some embodiments, feature encoding is performed on meteorological time series, remote sensing images, and agricultural text in the aligned data, respectively, and the encoding results are fused based on a cross-modal attention mechanism to obtain multimodal fused features, including: Sequence modeling is performed on meteorological time series input time encoders associated with unified spatial units and unified time axes to output meteorological embedded sequences carrying time sequence information; The remote sensing images associated with a unified spatial unit and falling within the prediction time period are segmented, and the segmented image blocks are input into an image encoder to output an image block embedding sequence. The text encoder generates text semantic embeddings from agricultural information texts associated with unified spatial units and prediction time periods, and extracts event tags related to disasters or agricultural management from the agricultural information texts. Based on a cross-modal attention mechanism, meteorological embedding sequences, image patch embedding sequences, text semantic embeddings, and event labels are aligned and fused to output multimodal fusion features.

[0030] Specifically, after completing the temporal alignment, spatial mapping, and growth stage labeling of multimodal data, this embodiment uses a unified spatial unit and a unified time axis as common indices to perform feature encoding on meteorological time series, remote sensing images, and agricultural texts, respectively. Furthermore, a cross-modal attention mechanism is used to achieve alignment and fusion between different modalities, thereby forming multimodal fusion features that can be used in subsequent time series segment construction, historical similar segment retrieval, and large-scale model inference. The key to this process lies in two aspects: firstly, converting the original data of each modality into comparable embedded representations; and secondly, introducing cross-modal attention weights during the fusion stage, enabling meteorological anomalies, growth changes, and disaster management events to establish corresponding relationships within the same semantic space.

[0031] In this embodiment, the vector sequence obtained by encoding the meteorological time series is called the meteorological embedding sequence; the vector sequence obtained by encoding each image block after dividing the remote sensing image is called the image block embedding sequence; the semantic vector obtained by encoding the agricultural text is called the text semantic embedding; the disaster or agricultural management elements extracted from the agricultural text are called event labels, and the event labels include one or more of the event type and event intensity; the attention mechanism used to calculate the association weights and perform information interaction between different modal embeddings is called the cross-modal attention mechanism.

[0032] In terms of meteorological time series feature encoding, for each unified spatial unit, the multivariate meteorological time series corresponding to that spatial unit is read within the prediction period and organized into a fixed-length sequence input according to a unified time axis. For example, in the scenario of county-level corn yield prediction, the unified time axis adopts a daily scale, the prediction window is selected 60 days prior to the current date, and the meteorological variables include daily maximum temperature, daily minimum temperature, daily cumulative precipitation, daily average relative humidity, and daily radiation. The above multivariate sequence is input into a time-series encoder for sequence modeling. The time-series encoder can adopt a sequence modeling structure with time position representation to output a meteorological embedded sequence that corresponds one-to-one with the position on the unified time axis.

[0033] If some dates have missing data, a missing data marker is written to the input and a time mask is generated simultaneously. This allows the time encoder to distinguish between valid and missing observations during calculation, preventing missing data from interfering with the sequence pattern representation. In this way, meteorological embedded sequences can represent temporal evolution patterns such as sustained high temperatures, periods of low rainfall, and concentrated precipitation while maintaining temporal order information, providing a foundation for subsequent alignment with remote sensing crop growth changes and textual events.

[0034] Furthermore, in terms of remote sensing image feature coding, for each unified spatial unit, the remote sensing observation data associated with that spatial unit are read within the prediction period, and the image or index raster within the spatial unit is obtained based on the spatial mapping results. To establish an alignable relationship with meteorological time series on the time axis, multi-temporal remote sensing observations within the prediction period can be synthesized or interpolated and aligned according to a unified time axis to form a remote sensing image sequence or index map sequence associated with the time index.

[0035] In some implementations, during image encoding, the remote sensing image for each time phase is segmented. The segmentation granularity can be configured according to the resolution and spatial unit size. For example, the remote sensing image within a spatial unit range can be divided into several regular image blocks, and each image block is assigned a block location label. The segmented image blocks are then input into an image encoder, which can employ a feature extraction structure based on image blocks to output an image block embedding sequence. To enhance the representation of growth change trajectories, image block embeddings from adjacent time phases can be concatenated in chronological order, or a temporal location representation can be introduced at the input of the image encoder, so that the image block embedding sequence carries both spatial structure information and temporal evolution information. If invalid observations are encountered due to cloud obstruction, a validity label can be written for the corresponding image block, and this label can be used to constrain attention weights in subsequent fusion stages.

[0036] Furthermore, regarding the encoding of agricultural text features and the extraction of event tags, for each unified spatial unit, text entries associated with that spatial unit, such as weekly agricultural reports, disaster warnings, pest and disease reports, survey records, and descriptions of policy measures, are aggregated within the prediction period. These are then mapped to a unified timeline based on the text publication time or event occurrence time. The agricultural text is input into a text encoder to generate semantic embeddings, which are used to express semantic information in the text regarding crop growth, disaster impact, and management measures.

[0037] Simultaneously, based on preset event extraction rules or the structured extraction capabilities of a text encoder, event tags are extracted from agricultural data text. These event tags include at least disaster-related event tags and agricultural management event tags. For example, "continuous heavy rainfall caused waterlogging in some fields" is extracted as a disaster type of waterlogging with a moderate intensity; "the proportion of lodging-resistant varieties being promoted is increasing" is extracted as a management measure type of variety adjustment with a relatively high intensity. Event tags are associated with corresponding dates on the timeline and can be further encoded into tag vector representations that can be fused with other modal embeddings.

[0038] Furthermore, in terms of cross-modal attention mechanism fusion, meteorological embedding sequences, image patch embedding sequences, text semantic embeddings, and event labels are input into the cross-modal attention fusion module. Within the same unified spatial unit, the fusion module performs alignment and interaction on the embedding representations of different modalities: on the one hand, using the temporal position of the meteorological embedding sequence as a query, attention weights are assigned to the growth representations at the same or adjacent temporal positions in the image patch embedding sequence, thereby establishing a correlation between meteorological changes and remotely sensed growth changes; on the other hand, using event labels or text semantic embeddings as queries, attention weights are assigned to the meteorological embedding sequence and the image patch embedding sequence, enabling the semantic information of disasters or management measures to be aggregated with their corresponding meteorological and growth evidence.

[0039] In some implementations, when the event tag indicates high-temperature heat damage during the grouting period, the fusion module assigns higher weights to the high-temperature meteorological embedding corresponding to the grouting period and the image patch embedding of declining growth during the same period, in order to form a fusion representation with consistent directionality. The fusion module outputs multimodal fusion features, which can be a fusion sequence organized along a unified time axis or a fusion vector obtained by aggregating the fusion sequence. These multimodal fusion features will serve as the input basis for subsequently constructing the temporal segment representation of the current growth stage and will be further used for historical similar segment retrieval and large model inference context construction.

[0040] For example, in a corn yield prediction scenario during the grain-filling stage in a certain county, the system encodes the meteorological time series of the target grid cell over the past 60 days to obtain a meteorological embedding sequence reflecting continuous high temperatures and low precipitation; it also encodes the remote sensing index map of the same period in blocks to obtain an image patch embedding sequence representing the downward trend of the index in the later stage of grain filling; and it encodes agricultural information text and extracts event tags such as "high temperature heat damage" and "irrigation restriction". The cross-modal attention fusion module aligns the event tags with the meteorological embedding and image patch embedding at the time position corresponding to the grain filling stage, and outputs the multimodal fusion features with weighted concentrations during the grain filling period. This allows subsequent time series segment construction to directly extract the fusion trajectory of this stage and retrieve it from the historical time series segment knowledge base, thus providing consistent evidence input for large model prediction and interpretation generation.

[0041] Through the above encoding and fusion processing, this embodiment can form a structurally consistent multimodal fusion feature under the constraints of a unified spatial unit and a unified time axis, enabling meteorological changes, remote sensing growth status and agricultural events to achieve aligned interaction in the same semantic space, thereby improving the comparability of subsequent growth stage time series segment construction and historical similar segment retrieval, and enhancing the comprehensive utilization and consistency of multi-source evidence when the large model generates predicted values ​​and interpretation information.

[0042] In some embodiments, a temporal segment representation of the current growth stage is constructed based on multimodal fusion features and combined with growth stage annotations, including: The target growth stage corresponding to the current time is determined based on the growth stage annotation, and the temporal window parameters associated with the target growth stage are determined from the multimodal fusion features. The temporal window parameters include window length or time span. According to the temporal window parameters, multimodal temporal segment features aligned with the unified time axis are extracted from the multimodal fusion features, and the multimodal temporal segment features are associated and encapsulated with the unified spatial unit and the target growth stage; The encapsulated multimodal temporal segment features are subjected to aggregation or sequence representation generation processing to obtain the temporal segment representation of the current growth stage.

[0043] Specifically, after completing the feature encoding of meteorological time series, remote sensing images, and agricultural text and obtaining multimodal fusion features, this embodiment further constructs a temporal segment representation of the growth stage corresponding to the current moment on a unified time axis, using growth stage annotation as a constraint. This enables subsequent retrieval of similar historical segments to be conducted under the premise of "comparability within the same stage." The key to this process is that instead of directly using the fusion sequence covering the entire prediction period for retrieval, it selects a temporal window parameter matching the phenological stage of the target crop at the current moment, extracts the fusion trajectory of that stage on a unified time axis, and then encapsulates it into a searchable segment representation.

[0044] Furthermore, regarding the determination of the target growth stage, the system obtains the time index of the current moment on a unified time axis based on the growth stage labels written in the aforementioned aligned data, and reads the growth stage label corresponding to that time index to determine the target growth stage corresponding to the current moment. For example, in a county-level maize yield prediction scenario, the unified time axis uses a daily scale. When the current date is in the grain-filling stage, the target growth stage is determined to be the grain-filling stage; when the current date is in the tasseling and silking stage, the target growth stage is determined to be the tasseling and silking stage. For situations where there are differences in sowing dates within the same target area, the system reads the growth stage labels separately at the unified spatial unit granularity, thereby allowing different spatial units to correspond to different target growth stages at the same prediction time, avoiding stage mismatch caused by substituting the regional average stage.

[0045] Furthermore, regarding the determination of time-series window parameters, the system determines time-series window parameters associated with the target growth stage from multimodal fusion features. These parameters include window length or time span and can vary with the growth stage. For example, to improve the targeting of stage features, the system configures different time spans for different growth stages: during the tasseling and silking stage, which focuses on the impact of short-term extreme high temperatures and sudden changes in precipitation, a shorter window length can be used; during the grouting stage, which focuses on the cumulative effects of persistent heat damage or water stress, a window length covering a longer time span can be used. Furthermore, the time-series window parameters can also be associated with the start and end boundaries of the target growth stage; that is, the window boundary can be limited to the target growth stage or extended forward to the preceding stage interval, in order to capture conditional changes before the stage transition.

[0046] Furthermore, regarding the truncation and encapsulation of multimodal temporal segment features, the system determines the truncation interval on a unified time axis according to the temporal window parameters, and extracts the multimodal temporal segment features corresponding to that interval from the multimodal fusion features. If the multimodal fusion features are a fusion sequence organized along a unified time axis, the subsequence is directly truncated according to the time index; if the multimodal fusion features are a set of fusion vectors with time positions, the segment set is obtained by filtering according to the time positions. After truncation, the system associates and encapsulates the multimodal temporal segment features with a unified spatial unit and the target growth stage to form a segment encapsulation body, ensuring that the segment encapsulation body carries at least the spatial unit identifier, growth stage marker, segment start and end time index, and segment feature payload. This encapsulation body serves as the basic object for subsequent generation of retrieval representations and also facilitates backtracking the stage range and spatial range corresponding to the segment when outputting interpretation information.

[0047] Furthermore, in terms of fragment representation generation, the system performs aggregation or sequence representation generation processing on the encapsulated multimodal time-series fragment features to obtain the current growth stage time-series fragment representation. For example, when it is necessary to preserve the evolutionary morphology of the daily sequence within the stage, sequence representation generation processing can be performed on the fragment subsequences to obtain a fragment representation carrying time sequence information; when it is only necessary to express the overall state of the stage and reduce the computational burden of retrieval, aggregation processing can be performed on the fragment features to obtain a stage summary fragment representation. Regardless of whether aggregation processing or sequence representation generation processing is used, the obtained current growth stage time-series fragment representation is bound to the unified spatial unit and the target growth stage, thus serving as a retrieval condition in subsequent historical time-series fragment knowledge base retrieval, and prioritizing the retrieval scope within the historical fragment set with consistent growth stages.

[0048] For example, in a scenario predicting the grain-filling stage of maize in a certain county, the system determines the target growth stage as the grain-filling stage from the growth stage annotations and configures a 45-day time span as the time series window parameter for the grain-filling stage. Then, it extracts a multimodal fusion sequence segment corresponding to the past 45 days from a unified time axis. This segment includes the coupled changes of meteorological embeddings and remote sensing growth embeddings within the grain-filling stage, as well as the impact of event labels related to disasters and irrigation restrictions. The system encapsulates this segment, along with the corresponding grid cell identifier and grain-filling stage marker, into a fragment encapsulation body and performs sequence representation generation processing on this segment to obtain a time series fragment representation of the current growth stage that can be used for retrieval. This fragment representation is then used to retrieve historical fragments with similar patterns and belonging to the same grain-filling stage from the historical time series fragment knowledge base to obtain yield label information and background description information for similar years.

[0049] Through the above processing, this embodiment can construct comparable temporal segment representations on a unified time axis with growth stage labels as constraints, enabling multimodal fusion features to form structured segment expressions at the stage granularity, thereby improving the matching accuracy and interpretability of historical similar segment retrieval, and providing more focused stage evidence input for subsequent large-scale model reasoning based on retrieval cases.

[0050] In some embodiments, using the current growth stage time series segment representation as the retrieval condition, similar historical segments are retrieved from the historical time series segment knowledge base, and yield tag information associated with similar historical segments is obtained, including: A historical time series fragment knowledge base is constructed, which stores multiple historical time series fragment entries. Each historical time series fragment entry includes a historical growth stage time series fragment representation and output label information associated with the historical growth stage time series fragment representation. The output label information includes actual output or deviation from the average annual output. A retrieval representation is generated based on the temporal segment representation of the current growth stage, and the similarity between the retrieval representation and the temporal segment representation of each historical growth stage is calculated in the historical temporal segment knowledge base. Based on similarity, a preset number of similar historical segments are selected from the historical time series segment knowledge base, and the output of the production tag information corresponding to the similar historical segments and the segment identification information associated with the similar historical segments are output.

[0051] Specifically, after obtaining the temporal segment representation of the current growth stage, this embodiment uses a temporal retrieval enhancement mechanism to search for historical segments with similar patterns in the historical temporal segment knowledge base. It then outputs yield label information and segment identification information associated with these similar historical segments, enabling subsequent large-scale model inference to perform trend extrapolation based on referable historical evidence. The key to this process is that the historical knowledge base does not merely store annual yields or static statistical features, but rather stores the correspondence between "historical growth stage temporal segment representation + yield label information" at the growth stage granularity. This ensures that the retrieved object and the current segment are comparable in stage semantics, and allows the retrieved results to directly provide reference labels and background clues for yield prediction.

[0052] In terms of constructing a historical time-series fragment knowledge base, the system generates and stores multiple historical time-series fragment entries offline based on multimodal datasets of historical years. Each historical time-series fragment entry uses a unified spatial unit and a unified time axis as the indexing benchmark, generates a historical growth stage time-series fragment representation according to alignment rules and encoding fusion rules consistent with online prediction, and binds and stores it with historical yield labels.

[0053] For example, in the scenario of county-level maize yield prediction, the system uses grid cells as unified spatial units and performs time alignment, spatial mapping, and growth stage labeling on historical data from 2010 to 2024 year by year. Then, it performs meteorological coding, remote sensing block coding, and agricultural text coding respectively, and obtains historical multimodal fusion features through cross-modal attention fusion. Then, based on the window parameters corresponding to the grain filling period of each grid cell in that year, it extracts the fusion sequence segment and generates a time-series representation of the historical growth stage of the grain filling period in that year.

[0054] The corresponding yield label information can be the actual yield per unit of the grid cell in that year, or the deviation from the average yield per unit in previous years, or the yield deviation converted under the regional statistical caliber, and is associated one-to-one with the historical growth stage time series segment representation. To facilitate subsequent reference and tracing, historical time series segment entries can also include segment identification information, which includes one or more of the following: year identifier, spatial unit identifier, crop identifier, growth stage marker, and segment start and end time index.

[0055] Furthermore, regarding retrieval representation generation, the system generates a retrieval representation for searching based on the temporal segment representation of the current growth stage. If the temporal segment representation of the current growth stage is a sequence representation, the retrieval representation can be directly constructed from this sequence representation, or from the segment vector obtained by dimensionality reduction and aggregation of the sequence representation; if the temporal segment representation of the current growth stage is an aggregated vector, the retrieval representation directly adopts this aggregated vector. To maintain consistency between online retrieval and offline database entry, the generation method of the retrieval representation is consistent with the generation method of the temporal segment representation of the historical growth stages, so that the similarity calculation reflects the degree of similarity of segment patterns under the same representation space.

[0056] Furthermore, regarding similarity calculation, the system calculates the similarity between the retrieved representation and the representation of each historical growth stage in the historical time-series fragment knowledge base. Constraint filtering can be introduced during the calculation process to prevent incomparable historical entries from entering the candidate set. For example, the system prioritizes performing similarity calculations on historical fragment sets with consistent crop identifiers and growth stage labels. When the target area spans a large area, it can also prioritize performing similarity calculations on historical fragment sets within the same spatial unit or the same ecological zoning to reduce bias caused by spatial heterogeneity. Similarity calculation can employ vector distance or sequence matching metrics to reflect the similarity of meteorological embeddings, remote sensing growth embeddings, and event label influences within fragments in terms of temporal evolution.

[0057] Furthermore, regarding the selection and output of similar historical segments, the system selects a preset number of similar historical segments from the historical time-series segment knowledge base based on similarity, and outputs the corresponding yield label information and segment identification information for each similar historical segment. For example, the preset number can be 5 or 10. The system selects several historical segments with the highest similarity as the search result set, and outputs the relative average yield deviation, actual yield per unit area, and segment identification information such as year identifier and growth stage marker for each search result. While outputting the yield label information, the system can also output a background description index associated with the segment identification information, so as to generate a structured description of "common features of similar years" when combining prediction contexts later.

[0058] For example, in a scenario predicting the grain-filling stage of maize in a certain county, the system has generated a time-series representation of the current growth stage of the current grid cell during the grain-filling stage, and generates a retrieval representation based on this. The system limits the set of historical entries in the historical time-series segment knowledge base to the crop identified as maize and the growth stage marked as grain-filling stage, calculates the similarity between the retrieval representation and the representation of each historical grain-filling stage segment, and obtains the similarity ranking results. The system selects the 5 historical segments with the highest similarity and outputs their year identifier, segment start and end time index, corresponding relative average yield deviation, and actual yield value. Subsequently, during large-scale model inference, these retrieval results can be organized into a historical evidence description of "several similar years experiencing sustained high temperatures and low rainfall during the grain-filling stage, accompanied by a decline in growth, with corresponding yield deviations concentrated in a certain range," thus providing a reference basis for the generation of predicted values ​​and explanatory information.

[0059] Through the above processing, this embodiment can retrieve historical fragments similar to the current fragment pattern from the historical time series fragment knowledge base under the constraint of consistent growth stages, and simultaneously obtain the yield tag information and fragment identification information bound to similar historical fragments, thereby providing referenceable comparison samples for subsequent large model inference based on historical evidence, and improving the traceability of trend inference and prediction results for abnormal weather years.

[0060] In some embodiments, the prediction context is input into a supervised-trained large model, which outputs an initial output prediction and corresponding explanatory information, including: The current state description information is generated based on the multimodal fusion features and the temporal segment representation of the current growth stage, and the historical case description information is generated based on the yield label information and segment identification information of similar historical segments. The current state description information and historical case description information are combined according to a preset context structure to obtain the prediction context. The prediction context includes the target area identifier, target crop identifier, prediction time period identifier, current growth stage marker, and historical case description information. The prediction context is input into a large model trained under supervision. The large model outputs an initial yield prediction and simultaneously outputs explanatory information associated with the initial yield prediction. The explanatory information includes descriptions of the main influencing factors and references to similar historical segments.

[0061] Specifically, in generating current state description information, the system uses a unified spatial unit as the granularity, reads the multimodal fusion features of that spatial unit on a unified time axis, and extracts state elements related to the current growth stage by combining the temporal segment representation of the current growth stage, thus generating current state description information. The current state description information includes both structured fields that can be directly consumed by large models and stage summary content used to express stage evolution. For example, the current state description information includes, but is not limited to: target area identifier, target crop identifier, prediction period identifier, current growth stage marker, stage meteorological summary for characterizing meteorological models, stage remote sensing summary for characterizing growth changes, and stage event summary for characterizing disasters and management measures.

[0062] The system provides a range of data, including: a stage-based meteorological summary generated from the aggregation of meteorological embedding sequences within the window corresponding to the current growth stage; a stage-based remote sensing summary generated from the magnitude, trend direction, or key time point changes of image patch embedding sequences within the window; a stage-based summary generated from the type distribution and intensity statistics of event labels within the window; and a stage-based summary generated from the occurrence of floods, droughts, pests and diseases, or management measures such as irrigation, fertilization, and variety adjustment. To maintain consistency with subsequent interpretation information, the system also records evidence indexes used for summary generation when generating the current state description information, such as the time index range corresponding to the summary, the occurrence date of key event labels, and the corresponding fragment encapsulation identifier.

[0063] Furthermore, in terms of generating historical case description information, the system reads the yield tag information and segment identification information of similar historical segments and organizes them into historical case description information that can be referenced by large models. The historical case description information includes: segment identification information of similar historical segments, similarity ranking information or candidate set number information, and yield tag information associated with each similar historical segment. For example, the system describes each similar historical segment as an entry containing a year identifier, spatial unit identifier, growth stage marker, relative average yield deviation, or actual yield value, and forms a set description for a preset number of similar historical segments.

[0064] In some implementations, to facilitate inductive reasoning in large models, the system can also extract common features from sets of similar historical segments, generating historical common summary information. This includes, for example, statistically analyzing the proportion of yield reduction and increase, the range of yield deviations, and extracting common meteorological patterns, growth change patterns, and event label patterns that occur in multiple similar historical segments during their growth stages. This historical common summary information, along with segment identification information, is written into the historical case description information, enabling large models to reference both individual similar years and the overall patterns of similar sets.

[0065] Furthermore, regarding the combination of prediction contexts, the system combines current state description information with historical case description information according to a preset context structure to obtain the prediction context. The preset context structure is used to limit the organization order and field positions of each information segment, enabling large models to stably parse and reuse the same structure during inference. For example, the prediction context may include: a task identifier segment for loading target area identifiers, target crop identifiers, and prediction period identifiers; a stage identifier segment for loading the current growth stage marker and stage window range; a current evidence segment for loading stage meteorological summaries, stage remote sensing summaries, and stage event summaries; a historical evidence segment for loading a set of similar historical fragment entries, yield label information, and historical commonality summary information; and an inference instruction segment for limiting the output content to include yield prediction values ​​and explanatory information, requiring the explanatory information to reference historical fragment identifier information and descriptions of major influencing factors. Through this structured combination, the prediction context places current multimodal evidence and historical similar evidence in the same input carrier, and maintains a sufficient basis for citation at the field level, avoiding the occurrence of untraceable sources when the large model generates explanations.

[0066] Furthermore, regarding the large model inference output, the system inputs the prediction context into a supervised-trained large model, which outputs an initial yield prediction value, along with explanatory information associated with the initial yield prediction value. The initial yield prediction value can be the yield or yield per unit area of ​​the target region and target crop during the prediction period, and can be bound to spatial unit identifiers to support grid-level or plot-level output.

[0067] In some implementations, the explanatory information includes descriptions of key influencing factors and references to similar historical segments. The descriptions of key influencing factors explain the meteorological patterns, growth changes, and event label elements that contribute significantly to the predicted value during the current growth stage. The references to similar historical segments reference segment identifiers within the prediction context, enabling the explanatory information to identify the historical year or segment number for comparison. For example, in a maize grain-filling stage prediction scenario, the large model can output the initial yield prediction for a specific grid cell and reference the year identifiers of similar historical segments in the explanatory information, indicating that "continuous high temperatures during the grain-filling stage, low rainfall, and agricultural events indicating limited irrigation" are the key influencing factors. It also points out that the yield deviation range is similar to that of several similar years, thus forming a traceable inference chain.

[0068] Through the above processing, this embodiment can organize multimodal fusion features, the temporal segment representation of the current growth stage, and the output label information and segment identification information of similar historical segments into a structured prediction context, and drive the supervised training large model to output the initial output prediction value and corresponding explanation information, so that the prediction results have a referenceable historical evidence basis and traceable influence factor description, providing a consistent input and verification basis for subsequent self-reflection verification and reinforcement learning strategy correction.

[0069] In some embodiments, the initial yield forecast and explanatory information are subject to self-reflective verification and correction based on a reinforcement learning strategy. The output includes the yield forecast values ​​and uncertainty intervals for the target region and target crop, as well as the output of influencing factors associated with the yield forecast values ​​and reference information for similar historical segments, including: Reflective inputs are generated based on the initial output forecast, explanatory information, and prediction context. These reflective inputs are then used as inputs to the large model to obtain consistency verification results. The consistency verification results are used to indicate the matching status between the influence factors and similar historical fragment reference information in the explanatory information and the prediction context. If the consistency verification results indicate a mismatch or omission, correction guidance information is generated based on preset reward rules. The preset reward rules include rules related to the deviation of actual output, the consistency of the direction of output change, and the consistency of interpretation. Based on the corrected guidance information, a reinforcement learning strategy is invoked to correct the initial output prediction and explanation information, resulting in corrected output prediction and corrected explanation information. An uncertainty interval is generated based on the revised production forecast, and the influence factors associated with the production forecast and similar historical data references are output based on the revised explanatory information.

[0070] Specifically, in terms of reflective input generation and consistency verification, the system encapsulates the initial output forecast, explanatory information, and prediction context at a unified spatial unit granularity to generate reflective input. For example, reflective input includes: the initial output forecast, key paragraphs or structured elements of the explanatory information, current state description information and historical case description information in the prediction context, and consistency verification instructions.

[0071] In some implementations, this consistency check instruction is used to constrain large models to check the following two types of consistency: one is reference consistency, that is, whether the reference information of similar historical fragments appearing in the explanation information can be found in the historical case description information of the prediction context; the other is evidence consistency, that is, whether the main influencing factors listed in the explanation information can be found in the current state description information of the prediction context with corresponding meteorological summaries, remote sensing summaries or event labels, and whether the time range pointed to by the influencing factors is consistent with the current growth stage label and its window range.

[0072] Furthermore, the system uses the reflection input as input to the large model to obtain the consistency verification result. The consistency verification result may include a matching status flag and problem location information, wherein the matching status flag is used to indicate at least one of the following: no mismatch, citation mismatch, evidence mismatch, or omission; the problem location information is used to indicate the location of the mismatched field, the category of the missing influence factor, or the identifier of the missing citation fragment.

[0073] For example, in a corn grain-filling stage forecast scenario, the initial interpretation information references a similar historical fragment year identifier and states that "sustained high temperatures during the grain-filling stage lead to reduced yield." The consistency check will examine whether the year identifier actually exists in the set of similar historical fragment entries in the forecast context, and whether "sustained high temperatures" can be mapped to the high temperature pattern representation within the grain-filling stage in the current state description information's stage meteorological summary. When the interpretation information mentions "flood damage impact" but the forecast context does not contain a corresponding event label or text summary, the consistency check result will indicate that there is a mismatch of evidence or an omission.

[0074] Furthermore, regarding the generation of corrective guidance information, when the consistency verification result indicates a mismatch or omission, the system generates corrective guidance information based on preset reward rules. The preset reward rules include at least three categories of rules: rules related to deviations from actual output, rules related to consistency in the direction of output changes, and rules related to interpretive consistency.

[0075] Rules related to deviation from actual output are used to compare actual output during the training phase, determine the magnitude of the error, and formulate correction directions. Rules related to consistency of output change direction are used to determine whether the predicted value's judgment on the direction of increased or decreased output is consistent with the output label trend of similar historical segments. Rules related to interpretive consistency are used to constrain interpretive information to only reference segment identification information and evidence elements existing in the prediction context, and to provide requirements for the completion of missing key elements. Correction guidance information can be organized into structured correction constraints, including: references that need to be deleted or replaced, categories of influencing factors that need to be supplemented, stage evidence that needs to be re-evaluated, and candidate ranges for upward or downward adjustment of predicted values.

[0076] For example, when the consistency check finds that the explanation information omits the "irrigation-restricted" event label that already exists in the prediction context and that the label is highly correlated with the yield reduction interval in similar historical segments, the correction guidance information will include the constraint of "supplementing irrigation restriction as an influencing factor and referencing the entry number of similar historical segments in the explanation"; when the training phase compares the actual yield and finds that the predicted value is too high and the yield labels of similar historical segments are concentrated in the yield reduction interval, the correction guidance information will include the constraint of "adjusting the predicted value according to the direction of yield reduction and simultaneously adjusting the weight of meteorological sensitivity in the explanation".

[0077] Furthermore, regarding reinforcement learning policy correction, the system invokes reinforcement learning policies based on correction guidance information to correct the initial output predictions and explanations, resulting in corrected output predictions and explanations. During the inference phase, the reinforcement learning policy can select candidate outputs through policy-based decoding or a reward model-based reordering method, ensuring the outputs satisfy the consistency constraints in the correction guidance information. During the training phase, the system compares the actual output with the predictions before and after correction, calculates the reward signal according to preset reward rules, and updates the policy parameters, making the large model more inclined to output results with smaller errors, consistent direction, and correct explanations under subsequent similar inputs. Simultaneously, when there are samples with significant bias, the system can record the consistency verification problem location information and correction guidance information together as reflective samples, allowing these reflective samples to be fed back into the training dataset to support subsequent supervised training and policy optimization.

[0078] Furthermore, regarding the generation and interpretation of uncertainty intervals, the system generates uncertainty intervals based on the revised yield forecast and outputs influencing factors associated with the yield forecast and reference information of similar historical segments based on the revised interpretation information. Uncertainty intervals can be generated based on the yield label distribution of similar historical segments, the yield deviation range weighted by similarity, or the predicted fluctuation amplitude before and after revision. For example, the system can calculate upper and lower bounds based on the yield label information of a preset number of similar historical segments and assign higher weights to segments with higher similarity, thereby making the uncertainty interval reflect the dispersion of historical comparison evidence. Simultaneously, the revised interpretation information outputs influencing factors sorted according to consistency constraints and only references segment identification information existing in the historical case description information, making the interpretation chain verifiable and traceable.

[0079] For example, in a corn grain-filling stage prediction scenario in a certain county, the initial prediction value showed a higher yield per unit area, and the explanation cited a year identifier not found in the search results. The consistency check result indicated a mismatch in the citation and pointed out the missing trade-off of "susceptibility to sustained high temperatures". The system generates correction guidance information based on preset reward rules, requiring the replacement with similar year identifiers from the search results, supplementing "high temperature during grain filling stage" and "lower rainfall" as the dominant influencing factors, and adjusting the prediction value according to the direction of yield reduction. The reinforcement learning strategy generates the corrected prediction value and explanation information accordingly, and generates upper and lower bounds based on the yield deviation range of similar historical segments. At the same time, the explanation cites the corresponding similar historical segment identifiers and provides an influencing factor ranking and risk warning items.

[0080] Through the above self-reflection verification and reinforcement learning strategy correction process, this embodiment can verify the consistency of evidence and citation of the initial output prediction value and interpretation information. When a mismatch or omission is found, correction guidance information is generated according to the reward rule to drive the strategic correction, thereby outputting an output prediction value and uncertainty range that is more in line with the contextual evidence constraints. At the same time, verifiable impact factors and similar historical fragment citation information are output, thereby enhancing the stability and interpretation consistency of the prediction results.

[0081] The above embodiments have provided a detailed description of the implementation process of the grain yield prediction method based on large model and time series retrieval in this application. The following is based on... Figure 1 The process shown, combined with examples, illustrates the implementation flow of the grain yield prediction method of this application. This method may specifically include the following operations: 1. Multimodal data acquisition: Based on the task configuration, the system pulls all currently available information from different data sources: Meteorological time series data: historical + recent daily / hourly data: temperature, precipitation, radiation, wind speed, humidity, etc.; Remote sensing / image data: optical satellite (NDVI / EVI, vegetation cover), SAR data (cloud-rich areas for observation); Text data: weekly agricultural reports, pest and disease bulletins, meteorological disaster warnings, agricultural statistical reports, and survey records; Other structured data includes historical yields, sown area, soil type, topography, irrigation conditions, and other basic parameters.

[0082] 2. Data preprocessing and alignment: To facilitate understanding of the larger model later on, all data needs to be mapped to the same coordinate system, including: Spatial alignment: Unified to plots / grids; remote sensing pixels, meteorological grids, and statistical data are all mapped to the same spatial unit; Time alignment: unifying the time step; Crop and phenological stage labeling: Label each spatial unit with what crop is planted and what growth stage it is in; Standardization and feature engineering: extreme value truncation, normalization, and simple derived features.

[0083] 3. Meteorological Time Series Encoder: Input: Meteorological time series for this region and this crop within the forecast window (e.g., daily temperature / precipitation over the past 60 days); Handling method: Use a time encoder (such as Transformer / LSTM+time position encoding) to convert a multidimensional time series into a string of "time token embeddings"; Encode "time patterns": such as sustained high temperatures, periods of drought, and concentrated bursts of precipitation; Output: A sequence of meteorological time-series vectors, containing information in chronological order.

[0084] 4. Remote sensing image coding (image features) Input: Satellite images of the same area within the prediction window, vegetation index grid map, etc.

[0085] Handling method: The remote sensing image is divided into patches and fed into CNN / ViT to obtain the embedding of each patch; It can be superimposed over time to encode the trajectory of growth changes; Output: A sequence of visual feature vector patches.

[0086] 5. Text encoding (agricultural information / policy / disaster information, etc.) Input: Text data within the corresponding region and time period, such as: "Continuous heavy rainfall this week caused waterlogging in some fields" and "The proportion of lodging-resistant varieties promoted this year reached 70%"; Handling method: Large Language Model (LLM) is used to convert text into vectors and extract structured information such as event types: drought / flood / pests / policy benefits, etc. Severity: Mild / Moderate / Severe; Impact: Positive / Negative to production; Output: Text semantic vector + optional structured tags (event type, intensity).

[0087] 6. Multimodal feature fusion (cross-modal attention) Input: the above meteorological time series embedding, the above remote sensing visual embedding, and the above text semantic embedding; Handling method: Use a unified Transformer / multimodal Encoder: Feature alignment is achieved through self-attention and cross-modal attention, allowing different modalities to "see" each other. For example, the model aligns the correspondence between "excessive precipitation" and "increased remote sensing NDVI / excessive growth". Output: Multimodal fusion feature vector (used for subsequent time-series slices, RAG, and prediction heads).

[0088] 7. Construct the time sequence segment of the current growth stage. Input: Multimodal features; Handling method: Select a window length, such as the past 30 / 45 / 60 days, and construct a time segment: It includes the evolution trajectory of all key variables (meteorology, NDVI, disaster labels, etc.) during this period; Align with the current crop phenological stage (heading stage, grain-filling stage, etc.), and use different length / feature combinations for different stages; Output: Current time-series slice features, used to find similar patterns in the historical database.

[0089] 8. Time-series RAG retrieval of similar historical segments and corresponding output labels. Input: The current time slice generated in the previous step; Handling method: Search the historical "Time-Sequence Fragment Knowledge Base": The knowledge base stores a large number of samples (including year, region, crop, and disaster background information) of [the deviation of actual yield derived from time series segments]. (Calculate the similarity between the current slice and fragments in the library). Return the K most similar segments, along with the following information: subsequent weather evolution in that year, actual yield and its increase or decrease relative to the average year, whether a disaster occurred, and whether there was any policy intervention. Output: A set of similar historical cases + corresponding output performance / background description.

[0090] 9. Combine input context: current time series segment + RAG retrieval case Input: Current multimodal fusion features, current temporal slice, and retrieved historical similar segments and corresponding information; Handling method: Transform historical cases into structured text descriptions, such as: "In the past 5 years similar to the present, there were 3 years with a 5-10% reduction in yield, and the common characteristics were sustained high temperatures during the grain-filling period and late sowing." Use this content as a prompt, and input it along with the current state into the large model for inference.

[0091] Output: Current information + historical experience prompt, used for prediction models.

[0092] 10. Generate initial predictions and related explanations from a large model trained with SFT. Input: The prompt prepared in the previous step; Handling method: Using a large model trained on specialized grain product forecasting data, the reasoning process is automatically generated according to the sequence of "information collection—analysis—conclusion": Assess the current phenology and growth status; Analyze key meteorological deviations, disasters, and management measures; Based on historical cases, estimate the possible magnitude of production increases or decreases; Output the initial production forecast results; Output explanatory text to accompany the prediction: reasons, risk points, and reference cases; Output: Initial prediction result y0 and initial interpretation E0.

[0093] 11. RL optimization strategy invocation, self-reflection and correction Input: Initial prediction y0 and explanation E0, and real data already available during training; Handling method: Reasoning stage: Run the "self-reflection prompt": let the model check its own inference chain for obvious contradictions or omissions; The prediction should be revised based on the output of the empirical rules / reward model, such as "Insufficient sensitivity to extreme high temperature periods, the prediction value needs to be lowered"; Training phase: Compare with the actual output y_true to calculate the quantitative rule rewards (such as error, directionality, interpretation consistency, etc.); Update policy parameters using algorithms such as PPO; For samples with large biases, "reflective generation" is triggered, and the failure analysis is used as new training data for re-injection. Output: Corrected prediction y* and explanation E*, and newly generated empirical samples from the training phase.

[0094] 12. Output the final results: production forecast and uncertainty interval. Input: Predicted y* after RL / self-reflection correction; The output includes: Predicted yield / yield per unit area for a target region and target crop at a specified time scale; The confidence interval or upper and lower bounds of the prediction.

[0095] 13. Output explanatory information: main influencing factors, historically similar years, and risk warnings. Input: Explanation text E*, yield forecasts, and uncertainty information; Processing method: Organize the explanation into structured content: The main influencing factors are ranked as follows: "High temperatures during the heading stage contribute to a 4-6% reduction in yield, while excessive rainfall leads to a slight increase in yield of 1-2%." Citing similar historical cases, such as "similar to the situation in this region in 2013, production decreased by 5% that year"; Risk and scenario descriptions, such as "If high temperatures continue for the next 10 days, the risk of production reduction will increase further; if rainfall returns to normal, production may rebound." Output format: Report text.

[0096] Through the above-described technical solution, this application has the following advantages: 1. Integrate a multimodal large model, a time-series RAG, and a closed-loop RL model into a grain yield prediction system.

[0097] 2. More robust predictions: Historical similarity fragment retrieval greatly improves robustness to extreme weather years.

[0098] 3. Richer explanations: It can output a full-chain explanation of the prediction results, including driving factors, historical case comparisons, and risk warnings.

[0099] 4. The model can self-correct: the reflection mechanism enables the model to evolve continuously over the years.

[0100] 5. High adaptability: It can be extended to wheat, corn, rice, soybeans, and multiple regions and scenarios. It can be used in agricultural forecasting platforms, government grain regulation systems, and intelligent decision-making systems for agricultural enterprises, and has industrial value.

[0101] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0102] Figure 3 This is a schematic diagram of the structure of a grain yield prediction device based on a large model and time-series retrieval provided in an embodiment of this application. Figure 3 As shown, the device includes: The acquisition module 301 is used to acquire multimodal data corresponding to the target area, target crop and prediction period. The multimodal data includes meteorological time series, remote sensing images and agricultural information text. Alignment module 302 is used to perform temporal alignment, spatial mapping and growth stage labeling on multimodal data to obtain aligned data associated with a unified spatial unit and a unified time axis; The fusion module 303 is used to perform feature encoding on the meteorological time series, remote sensing images and agricultural text in the aligned data respectively, and to fuse the encoding results based on the cross-modal attention mechanism to obtain multimodal fusion features; Construction module 304 is used to construct a temporal segment representation of the current growth stage based on multimodal fusion features and combined with growth stage annotations; The retrieval module 305 is used to retrieve similar historical segments in the historical time series segment knowledge base and obtain the yield tag information associated with the similar historical segments, using the current growth stage time series segment representation as the retrieval condition. Prediction module 306 is used to combine multimodal fusion features, temporal segment representation of the current growth stage, similar historical segments and yield label information into prediction context, and input the prediction context into a supervised training large model, and output the initial yield prediction value and the explanation information corresponding to the initial yield prediction value; The output module 307 is used to trigger self-reflective verification of the initial yield forecast and interpretation information and make corrections based on the reinforcement learning strategy. It outputs the yield forecast and uncertainty range of the target region and target crop, and outputs the influencing factors and similar historical fragment reference information associated with the yield forecast.

[0103] In some embodiments, Figure 3 The alignment module 302 determines the unified spatial unit associated with the target area and maps the meteorological time series, remote sensing images, and agricultural text to the unified spatial unit to form multimodal data entries associated with the unified spatial unit; it determines the unified time axis corresponding to the prediction period and performs time scale unification and time index alignment on the multimodal data entries to form multimodal time series associated with the unified time axis; based on the phenological stage division rules of the target crop, it marks the growth stage at each time position in the multimodal time series and writes the growth stage mark into the corresponding multimodal time series to obtain aligned data.

[0104] In some embodiments, Figure 3 The fusion module 303 performs sequence modeling on the meteorological time series input to the time encoder associated with the unified spatial unit and the unified time axis to output a meteorological embedding sequence carrying time sequence information; it performs block processing on the remote sensing images associated with the unified spatial unit and falling within the prediction period, and inputs the block image blocks into the image encoder to output an image block embedding sequence; it inputs the agricultural information text associated with the unified spatial unit and the prediction period into the text encoder to generate text semantic embedding, and extracts event labels related to disasters or agricultural management from the agricultural information text; it aligns and fuses the meteorological embedding sequence, image block embedding sequence, text semantic embedding, and event labels based on a cross-modal attention mechanism to output multimodal fusion features.

[0105] In some embodiments, Figure 3 The construction module 304 determines the target growth stage corresponding to the current time based on the growth stage annotation, and determines the temporal window parameters associated with the target growth stage from the multimodal fusion features. The temporal window parameters include window length or time span. According to the temporal window parameters, multimodal temporal segment features aligned with a unified time axis are extracted from the multimodal fusion features, and the multimodal temporal segment features are associated and encapsulated with a unified spatial unit and the target growth stage. Aggregation or sequence representation generation processing is performed on the encapsulated multimodal temporal segment features to obtain the temporal segment representation of the current growth stage.

[0106] In some embodiments, Figure 3 The retrieval module 305 constructs a historical time-series fragment knowledge base, which stores multiple historical time-series fragment entries. Each historical time-series fragment entry includes a historical growth stage time-series fragment representation and yield tag information associated with the historical growth stage time-series fragment representation. The yield tag information includes actual yield or deviation from the average annual yield. A retrieval representation is generated based on the current growth stage time-series fragment representation, and the similarity between the retrieval representation and the historical growth stage time-series fragment representations is calculated in the historical time-series fragment knowledge base. A preset number of similar historical fragments are selected from the historical time-series fragment knowledge base according to the similarity, and the yield tag information corresponding to the similar historical fragments and the fragment identification information associated with the similar historical fragments are output.

[0107] In some embodiments, Figure 3 The prediction module 306 generates current state description information based on multimodal fusion features and time-series segment representations of the current growth stage, and generates historical case description information based on yield label information and segment identification information of similar historical segments. The current state description information and historical case description information are combined according to a preset context structure to obtain the prediction context, which includes target area identifier, target crop identifier, prediction time period identifier, current growth stage marker, and historical case description information. The prediction context is input into a supervised training large model, which outputs an initial yield prediction value and simultaneously outputs explanatory information associated with the initial yield prediction value. The explanatory information includes descriptions of major influencing factors and reference information of similar historical segments.

[0108] In some embodiments, Figure 3 The output module 307 generates a reflective input based on the initial output forecast, explanatory information, and prediction context. This reflective input is then used as input to the large model to obtain a consistency check result. The consistency check result indicates the matching status between the influencing factors and similar historical fragment references in the explanatory information and the prediction context. If the consistency check result indicates a mismatch or omission, correction guidance information is generated based on preset reward rules. These preset reward rules include rules related to deviations from actual output, consistency of output change direction, and explanatory consistency. Based on the correction guidance information, a reinforcement learning strategy is invoked to correct the initial output forecast and explanatory information, resulting in corrected output forecasts and corrected explanatory information. An uncertainty interval is generated based on the corrected output forecast, and the influencing factors and similar historical fragment references associated with the output forecast are output based on the corrected explanatory information.

[0109] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0110] Figure 4 This is a schematic diagram of the electronic device 4 provided in an embodiment of this application. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, it implements the steps in the various method embodiments described above. Alternatively, when the processor 401 executes the computer program 403, it implements the functions of each module / unit in the various device embodiments described above.

[0111] Electronic device 4 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 4 may include, but is not limited to, processor 401 and memory 402. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 4 and does not constitute a limitation on electronic device 4. It may include more or fewer components than shown, or different components.

[0112] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0113] The memory 402 can be an internal storage unit of the electronic device 4, such as a hard disk or RAM of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 4. The memory 402 can also include both internal and external storage units of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0115] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0116] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for grain yield prediction based on large model and time series retrieval, characterized in that, include: Acquire multimodal data corresponding to the target area, target crop, and prediction period. The multimodal data includes meteorological time series, remote sensing images, and agricultural information text. Perform temporal alignment, spatial mapping, and growth stage labeling on multimodal data to obtain aligned data associated with a unified spatial unit and a unified temporal axis; Feature encoding is performed on the meteorological time series, remote sensing images, and agricultural text in the aligned data respectively, and the encoding results are fused based on a cross-modal attention mechanism to obtain multimodal fusion features; Based on multimodal fusion features and combined with growth stage annotations, a temporal segment representation of the current growth stage is constructed; Using the current growth stage time series segment as the retrieval condition, similar historical segments are retrieved from the historical time series segment knowledge base, and yield tag information associated with similar historical segments is obtained; The multimodal fusion features, the temporal segment representation of the current growth stage, similar historical segments, and yield label information are combined into a prediction context. The prediction context is then input into a supervised training large model, which outputs the initial yield prediction value and the corresponding explanatory information. The system triggers self-reflective verification of the initial yield forecast and interpretation information and makes corrections based on reinforcement learning strategies. It outputs the yield forecast and uncertainty range of the target region and target crop, as well as the influencing factors and similar historical fragment reference information associated with the yield forecast. The step of using the current growth stage time series segment as the retrieval condition to search for similar historical segments in the historical time series segment knowledge base and obtaining yield tag information associated with similar historical segments includes: A historical time series fragment knowledge base is constructed, which stores multiple historical time series fragment entries. Each historical time series fragment entry includes a historical growth stage time series fragment representation and output label information associated with the historical growth stage time series fragment representation. The output label information includes actual output or deviation from the average annual output. A retrieval representation is generated based on the current growth stage time series segment representation, and the similarity between the retrieval representation and the time series segment representation of each historical growth stage is calculated in the historical time series segment knowledge base; Based on similarity, a preset number of similar historical segments are selected from the historical time series segment knowledge base, and the production tag information corresponding to the similar historical segments and the segment identification information associated with the similar historical segments are output. In terms of building a historical time series fragment knowledge base, the system generates and stores multiple historical time series fragment entries offline based on multimodal datasets of historical years. Each historical time series fragment entry uses a unified spatial unit and a unified time axis as the index benchmark, generates a historical growth stage time series fragment representation according to alignment rules and encoding fusion rules consistent with online prediction, and binds and stores it with historical yield labels. To maintain consistency between online retrieval and offline data entry, the generation method of the retrieval representation is consistent with the generation method of the temporal segment representation in the historical growth stage, so that the similarity calculation reflects the degree of similarity of segment patterns under the same representation space.

2. The method of claim 1, wherein, The process of performing temporal alignment, spatial mapping, and growth stage labeling on multimodal data to obtain aligned data associated with a unified spatial unit and a unified temporal axis includes: A unified spatial unit associated with the target area is identified, and the meteorological time series, remote sensing images, and agricultural information texts are mapped to the unified spatial unit to form multimodal data entries associated with the unified spatial unit. A unified time axis corresponding to the prediction period is determined, and time scale unification and time index alignment are performed on the multimodal data entries to form a multimodal time series associated with the unified time axis; Based on the phenological stage division rules of the target crop, growth stage markers are marked at each time position in the multimodal time series, and the growth stage markers are written into the corresponding multimodal time series to obtain the aligned data.

3. The method of claim 1, wherein, The process involves performing feature encoding on the meteorological time series, remote sensing images, and agricultural text in the aligned data, and then fusing the encoding results based on a cross-modal attention mechanism to obtain multimodal fusion features, including: Sequence modeling is performed on meteorological time series input time encoders associated with unified spatial units and unified time axes to output meteorological embedded sequences carrying time sequence information; The remote sensing images associated with a unified spatial unit and falling within the prediction time period are segmented, and the segmented image blocks are input into an image encoder to output an image block embedding sequence. The text encoder generates text semantic embeddings from agricultural information text associated with a unified spatial unit and a prediction period, and extracts event tags related to disasters or agricultural management from the agricultural information text. The meteorological embedding sequence, image patch embedding sequence, text semantic embedding, and event label are aligned and fused based on a cross-modal attention mechanism to output the multimodal fusion feature.

4. The method according to claim 1, characterized in that, The construction of the temporal segment representation of the current growth stage based on multimodal fusion features and combined with growth stage annotations includes: Based on the growth stage annotation, the target growth stage corresponding to the current time is determined, and the temporal window parameters associated with the target growth stage are determined from the multimodal fusion features. The temporal window parameters include window length or time span. According to the time window parameters, multimodal time segment features aligned with a unified time axis are extracted from the multimodal fusion features, and the multimodal time segment features are associated and encapsulated with a unified spatial unit and target growth stage; The encapsulated multimodal temporal segment features are subjected to aggregation or sequence representation generation processing to obtain the temporal segment representation of the current growth stage.

5. The method according to claim 1, characterized in that, The step of inputting the prediction context into a supervised-trained large model and outputting initial output predictions and corresponding explanatory information includes: Based on the multimodal fusion features and the temporal segment representation of the current growth stage, current state description information is generated, and historical case description information is generated based on the yield label information and segment identification information of similar historical segments; The current state description information and historical case description information are combined according to a preset context structure to obtain the prediction context. The prediction context includes target area identifier, target crop identifier, prediction time period identifier, current growth stage marker, and historical case description information. The prediction context is input into the supervised training large model, which outputs an initial output prediction value and simultaneously outputs explanatory information associated with the initial output prediction value. The explanatory information includes descriptions of major influencing factors and references to similar historical segments.

6. The method according to claim 1, characterized in that, The process triggers self-reflective verification of the initial yield forecast and interpretation information, and makes corrections based on a reinforcement learning strategy. It outputs the yield forecast values ​​and uncertainty intervals for the target region and target crop, and also outputs the influencing factors associated with the yield forecast values ​​and reference information for similar historical segments, including: Based on the initial output forecast, explanatory information, and the prediction context, a reflective input is generated, and the reflective input is used as the input of the large model to obtain a consistency verification result. The consistency verification result is used to indicate the matching status between the influence factors and similar historical fragment reference information in the explanatory information and the prediction context. If the consistency verification result indicates a mismatch or omission, correction guidance information is generated based on preset reward rules. The preset reward rules include rules related to actual output deviation, consistency of output change direction, and interpretation consistency. Based on the corrected guidance information, a reinforcement learning strategy is invoked to correct the initial output prediction value and the explanation information, resulting in the corrected output prediction value and the corrected explanation information. The uncertainty interval is generated based on the revised production forecast, and the influence factors and similar historical fragment reference information associated with the production forecast are output based on the revised explanatory information.

7. A grain yield prediction device based on large model and time series retrieval, characterized in that, include: The acquisition module is used to acquire multimodal data corresponding to the target area, target crop, and prediction period. The multimodal data includes meteorological time series, remote sensing images, and agricultural information text. The alignment module is used to perform temporal alignment, spatial mapping, and growth stage labeling on multimodal data to obtain aligned data associated with a unified spatial unit and a unified time axis. The fusion module is used to perform feature encoding on the meteorological time series, remote sensing images and agricultural text in the aligned data respectively, and to fuse the encoding results based on the cross-modal attention mechanism to obtain multimodal fusion features; The construction module is used to construct a temporal segment representation of the current growth stage based on multimodal fusion features and growth stage annotations; The retrieval module is used to retrieve similar historical segments from the historical time series segment knowledge base and obtain the yield tag information associated with the similar historical segments, using the current growth stage time series segment representation as the retrieval condition. The prediction module is used to combine the multimodal fusion features, the temporal segment representation of the current growth stage, similar historical segments and yield label information into a prediction context, and input the prediction context into a supervised training large model, and output the initial yield prediction value and the explanation information corresponding to the initial yield prediction value; The output module is used to trigger self-reflective verification of the initial yield forecast and interpretation information and make corrections based on reinforcement learning strategies. It outputs the yield forecast and uncertainty range of the target region and target crop, as well as the influencing factors and similar historical fragment reference information associated with the yield forecast. The retrieval module is used to construct a historical time-series fragment knowledge base, which stores multiple historical time-series fragment entries. Each historical time-series fragment entry includes a historical growth stage time-series fragment representation and yield tag information associated with the historical growth stage time-series fragment representation. The yield tag information includes actual yield or deviation from the average annual yield. A retrieval representation is generated based on the current growth stage time-series fragment representation, and the similarity between the retrieval representation and the historical growth stage time-series fragment representations is calculated in the historical time-series fragment knowledge base. A preset number of similar historical fragments are selected from the historical time-series fragment knowledge base based on the similarity, and the yield tag information corresponding to the similar historical fragments and the fragment identification information associated with the similar historical fragments are output. In terms of constructing a historical time-series fragment knowledge base, the system generates and stores multiple historical time-series fragment entries offline based on multimodal datasets of historical years. Each historical time-series fragment entry uses a unified spatial unit and a unified time axis as the index benchmark, generates a historical growth stage time-series fragment representation according to alignment rules and encoding fusion rules consistent with online prediction, and binds and stores it with historical output labels. To maintain consistency between online retrieval and offline storage, the generation method of the retrieval representation is consistent with the generation method of the historical growth stage time-series fragment representation, so that the similarity calculation reflects the degree of similarity of fragment patterns under the same representation space.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.