Crop yield prediction method based on multi-modal fusion technology
By integrating multimodal fusion technology and deep learning methods, multi-source remote sensing data is combined to construct dynamic adaptive and cross-token attention modules, which solves the problem of insufficient modality selection and utilization in multimodal temporal fusion. This enables accurate prediction of crop yield and understanding of dynamic growth processes, supporting precision agricultural management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-17
AI Technical Summary
Existing multimodal time series fusion studies suffer from insufficient modality selection and utilization, crude fusion methods, and difficulty in fully reflecting the complex processes of farmland ecosystems. This results in inaccurate characterization of crop growth processes and affects the accuracy of crop yield prediction.
By employing multimodal fusion technology, integrating various remote sensing data and environmental information, and utilizing deep learning and multimodal temporal fusion techniques, a dynamic adaptive fusion module and a cross-token attention module are constructed to achieve accurate fusion of multimodal temporal data, extract key features, and perform pixel-level prediction.
It enables more accurate crop yield forecasting, enhances the dynamic capture of crop growth processes and insights into environmental responses, supports precision agriculture management, reduces resource waste, adapts to geographical and climatic differences, and provides pixel-level forecast results.
Smart Images

Figure CN121884145A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and remote sensing agriculture, and in particular to a crop yield prediction method based on multimodal fusion technology. Background Technology
[0002] With global climate change and food security becoming increasingly severe issues, accurately and timely acquiring information on crop growth, productivity, and yield has become a crucial research topic in agricultural monitoring and management. Remote sensing technology, with its advantages of wide coverage, high frequency of repeated observations, and non-contact observation, provides a key tool for agricultural monitoring at regional and even global scales. In particular, with the continuous development and operationalization of multi-source satellite sensors, information on surface vegetation across different spectral dimensions (optical, thermal infrared, radar, etc.) can be acquired continuously and over a long period, forming multimodal time-series remote sensing data with rich spectral and temporal dimensions. This provides unprecedented opportunities for deeply characterizing crop growth processes and their responses to environmental changes.
[0003] In the field of remote sensing agriculture, multimodal time-series data fusion has been explored and applied in several typical scenarios. Taking crop growth monitoring as an example, researchers typically combine time series data of vegetation indices acquired through optical remote sensing with meteorological reanalysis data and soil moisture products to characterize the comprehensive response of crops at different growth stages to light, temperature, and water conditions. In crop classification and phenological identification tasks, the spectral and textural information of optical images is often combined with the structural information of radar data to improve the ability to identify different crop types and their phenological differences. In applications such as drought monitoring, pest and disease assessment, and yield early warning, comprehensive analysis of multiple time-series variables such as surface temperature, vegetation indices, and precipitation / evapotranspiration is increasingly becoming an important means to improve monitoring accuracy and timeliness. These applications collectively demonstrate that a single modality is insufficient to fully reflect the complex processes of farmland ecosystems, and multimodal time-series data fusion has become an important direction for the development of remote sensing agriculture towards precision and intelligence.
[0004] However, existing multimodal time-series fusion research still faces some prominent problems: First, the selection and utilization of modes are insufficient. Many studies only fuse limited modes, such as fusion of only optical and SAR images, or simply combining optical vegetation indices with meteorological variables, failing to systematically integrate multi-source information such as vegetation status and biophysical parameters, resulting in an incomplete characterization of crop growth processes; Second, the fusion methods are relatively crude. Common practices include channel stitching of images from different modes before fusion, or simple feature stitching, static weighting, or result combination at the decision level during fusion. These methods are difficult to show the correlation and complementarity between different modes over time, and are also difficult to specifically distinguish between "key information modes" and "auxiliary supplementary modes".
[0005] Based on the above problems, there is an urgent need to provide a multimodal time series fusion method to achieve accurate crop yield prediction. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a crop yield prediction method based on multimodal fusion technology. By integrating multiple remote sensing data and environmental information, and employing advanced deep learning and multimodal temporal fusion technology, multimodal temporal fusion is achieved, thereby enabling accurate crop yield prediction.
[0007] To achieve the above objectives, the present invention provides a crop yield prediction method based on multimodal fusion technology, comprising: The acquired multi-source remote sensing data is preprocessed to obtain multi-source remote sensing data with consistent temporal and spatial dimensions. The preprocessed remote sensing images of each modality are split into fine-grained segments, and a fixed-size temporal three-dimensional cube set is constructed according to the spatial location. The three-dimensional cube sets in the same spatial location of each modality are stitched together to obtain multiple sets of multi-source temporal cube remote sensing data, which are then input into the model. The model divides multi-source time-series cube remote sensing data into high-correlation mode branches and low-correlation mode branches based on prior knowledge, and extracts key features from each branch through a dynamic adaptive fusion module. The key features of both the high-relevance modality branch and the low-relevance modality branch are input into the cross-token attention module. The learnable token of the high-relevance modality branch is used as a proxy and combined with the attention mechanism to learn the complementary information of the low-relevance modality branch, so as to obtain the deep features after feature fusion. The deep features are input into the spatial encoder, and global features are extracted by adding the position information of the deep features and applying a stacked Transformer architecture. The pixel-level prediction results are output through the segmentation head. The prediction results of each group of multi-source time-series cube remote sensing data are stitched together in the original spatial order to obtain a pixel-level predicted image of crop yield for the entire region. The predicted images were compared with the actual yield data, and the model accuracy was verified using evaluation metrics, and the model parameters were optimized.
[0008] As a further improvement of the present invention, the three-dimensional cube sets in the same spatial location in each modality are spliced together, and after model processing, a crop yield estimation map for that spatial location is generated. By stitching together the estimated crop yield maps of all spatial locations obtained from the model processing in the original spatial order, a pixel-level predicted image of the overall regional crop yield is obtained.
[0009] As a further improvement of the present invention, the dynamic adaptive fusion module evaluates the importance of each modality through a 3D spatiotemporal attention mechanism and dynamically allocates the weights of each modality using a self-learning weighting method.
[0010] As a further improvement of the present invention, each branch extracts key features through a dynamic adaptive fusion module, including: The modalities of each branch first pass through a dynamic adaptive fusion module. This module uses a hybrid feature extractor, a 3D channel attention mechanism, a temporal feature encoder, and a learnable weight mechanism to automatically adjust the modal weights at different times and locations, thereby extracting the key features of that branch.
[0011] As a further improvement of the present invention, the cross-token attention module uses the high-dimensional features of the highly correlated modality branch as the main branch to extract learnable tokens. As a proxy, the system learns complementary information of the low-relevance modal branches through an attention mechanism, projects the proxy token back to the feature label, and concatenates it with the features of the high-relevance modal branches to obtain the deep features after feature fusion.
[0012] As a further improvement of the present invention, the spatial encoder retains class labels for the deep features after feature fusion, adds spatial location encoding and performs spatial attention calculation, and extracts global features through a multi-layer Transformer.
[0013] As a further improvement of the present invention, the segmentation head processes the deep feature vectors of each group of multi-source time-series cube remote sensing data through MLP processing including LayerNorm and linear layers to recover the pixel values of each cube block, and stitches multiple cube blocks together to restore the original-size image, thereby obtaining pixel-level prediction results, with each pixel point corresponding to a regression prediction value.
[0014] As a further improvement of the present invention, the acquired multi-source remote sensing data is preprocessed, including: The spatial and temporal resolutions of multi-source remote sensing data are uniformly processed, and feature normalization or standardization is performed.
[0015] As a further improvement of the present invention, the prior knowledge quantifies the correlation coefficient between multi-source remote sensing data and crop yield, and a preset correlation threshold is set. If the correlation coefficient between multi-source remote sensing data and crop yield is higher than the correlation threshold, the mode is divided into a high correlation mode branch; if it is lower than the correlation threshold, the mode is divided into a low correlation mode branch.
[0016] As a further improvement of the present invention, the evaluation indicators include correlation coefficient R², root mean square error RMSE, and mean absolute error MAE. The evaluation indicators are calculated by comparing the predicted image with the actual yield data to comprehensively evaluate the model accuracy. The model parameters are adjusted by backpropagation algorithm and loss function minimization.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention, by integrating multiple remote sensing data and advanced deep learning technologies, significantly outperforms existing technologies in terms of accuracy, efficiency, and application scope. Specifically, this invention achieves more accurate crop yield prediction through multimodal temporal fusion and dynamic adaptive mechanisms. Compared to existing methods that often suffer from large errors due to insufficient modality utilization or coarse fusion methods (such as simple splicing or static weighting), this invention introduces a spatiotemporal attention mechanism and a cross-token attention module, which can deeply mine complementary information between modalities. This helps agricultural managers make decisions in advance, reduces resource waste caused by prediction bias, and supports food security.
[0018] This invention effectively integrates multi-source information (such as meteorological data, vegetation indices, and radar data), and enhances the dynamic capture of the entire crop growth process through time-series modeling. Compared with existing technologies that are mostly limited to single-modal or static analysis and are difficult to reflect environmental interaction effects, the layer-by-layer fusion strategy of this invention can analyze the mutual influences in the time dimension, improve the insight into scenarios such as climate change response and pest and disease early warning, and promote precision agricultural management.
[0019] This invention utilizes a dynamic adaptive fusion module to automatically adjust modal weights across different times and locations, effectively addressing geographical and climatic variations. Compared to existing methods that often employ fixed weights and cannot adapt to regional heterogeneity, this invention's learnable mechanism ensures model stability in diverse environments, maintaining high accuracy even in resource-constrained areas (such as remote farmland).
[0020] This invention provides pixel-level prediction results that can be directly used to generate crop yield distribution maps, contributing to precision agriculture and sustainable development. Compared to existing methods that often output regional means and lack detail, this invention's segmentation head outputs pixel-level regression values, and combined with evaluation metrics, ensures reliability. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating a crop yield prediction method based on multimodal fusion technology disclosed in one embodiment of the present invention. Figure 2 This is a schematic diagram of the processing flow of a multimodal adaptive fusion module disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the processing flow of the cross-token attention module disclosed in one embodiment of the present invention; Figure 4 This is a schematic diagram of a pixel-level regression prediction process disclosed in one embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] The present invention will now be described in further detail with reference to the accompanying drawings: like Figure 1 As shown, the crop yield prediction method based on multimodal fusion technology provided by this invention utilizes the advantages of various remote sensing data and provides accurate crop yield predictions through systematic processing and fusion. This method effectively integrates information from different sources, enhances the understanding of crop growth dynamics, and has broad application prospects. For the collected different modal data, a specialized feature extractor is used to extract key features; and by introducing a spatiotemporal attention mechanism, the temporal features of different modalities are fused, strengthening the model's understanding of the dynamic interactions between various modalities in the time dimension; a layer-by-layer fusion strategy is adopted to promote feature sharing and information enhancement among modalities; after temporal fusion processing, the final features are input into the segmentation head for crop yield prediction; by comparing with actual yield data, the prediction accuracy of the model is evaluated, model parameters are optimized, and its effectiveness and reliability in practical applications are ensured. Specifically, it includes: S1. Preprocess the acquired multi-source remote sensing data to obtain multi-source remote sensing data with consistent time and spatial dimensions; This involves unifying the spatial and temporal resolution of multi-source remote sensing data and performing feature normalization or standardization.
[0024] Specifically, the spatial resolution of various remote sensing data was unified to 500 meters, and the temporal resolution was unified to 8 days to ensure consistency in the dimensions of all data. In addition, features were normalized or standardized to obtain multi-source remote sensing data that were consistent in both temporal and spatial dimensions for subsequent analysis and processing.
[0025] S2. The preprocessed remote sensing images of each modality are split into fine-grained segments, and a fixed-size temporal three-dimensional cube set is constructed according to the spatial location. The three-dimensional cube sets in the same spatial location of each modality are stitched together to obtain multiple sets of multi-source temporal cube remote sensing data, which are then input into the model. Specifically, the set of three-dimensional cubes in the same spatial location in each modality is stitched together and input into the model; Furthermore, according to a predetermined order, sequential sets of three-dimensional cubes of fixed size located in the same spatial position are used as model input, resulting in N sets of data. All sets of three-dimensional cubes in the same spatial position are stitched together and input into the model for processing.
[0026] Specifically, for multimodal time series data A set of temporally ordered three-dimensional cubes located in the same spatial position, arranged in a certain order. As input to the model, a set of three-dimensional cubes representing all modalities in the same spatial location. After model processing, the final crop yield estimation map for the corresponding location is generated. To obtain the overall partitioning results, the prediction results of the model are obtained by processing all 3D cube sets. The images are then reassembled to obtain the final crop prediction pixel map for the region. .
[0027] S3. Based on prior knowledge, the model divides multi-source time-series cube remote sensing data into high-correlation mode branches and low-correlation mode branches. Key features are extracted from each branch through a dynamic adaptive fusion module. Among them, prior knowledge quantifies the correlation coefficient between multi-source remote sensing data and crop yield, and presets a correlation threshold. If the correlation coefficient between multi-source remote sensing data and crop yield is higher than the correlation threshold, the mode is classified as a high correlation mode branch, and if it is lower than the correlation threshold, the mode is classified as a low correlation mode branch.
[0028] The dynamic adaptive fusion module evaluates the importance of each modality through a 3D spatiotemporal attention mechanism and dynamically allocates the weights of each modality using a self-learning weighting method to adapt to the geographical and climatic differences at different times and locations, thereby improving prediction accuracy.
[0029] Specifically, such as Figure 2As shown, the modal features of the two branches are input into a dual-branch cross-token attention module. In this module, the token agent of the high-relevance modality branch is used in conjunction with the attention mechanism, enabling it to fully learn the key features and complementary information of the low-relevance modality branch. In this way, the high-relevance modality branch with potential synergistic relationships will enhance the proportion of key features in the fused representation; at the same time, the low-relevance modality branch will convey differential information through interaction with the high-relevance modality branch, thus forming a more complete semantic representation. Because the importance of various modalities varies depending on the geographical location of the corresponding image and the climatic characteristics of the corresponding time when the model processes massive input data from different times, locations, and climatic conditions, this means that ecosystems in different regions respond differently to specific climatic factors. By introducing learnable weights, our model can dynamically adjust the current optimal weight of each modality based on the importance of these different features, ensuring that geographical and climatic differences are fully considered in crop yield estimation. This flexible feature fusion method facilitates accurate prediction of the dynamic response of large-scale ecosystems and improves the overall performance of the model.
[0030] Each branch extracts key features through a dynamic adaptive fusion module, including: The modalities of each branch first pass through a dynamic adaptive fusion module. This module utilizes a hybrid feature extractor, a 3D channel attention mechanism, a temporal feature encoder, and a learnable weighting mechanism to automatically adjust the modal weights at different times and locations, thereby extracting the key features of that branch. S4. Input the key features of both the high-relevance modal branch and the low-relevance modal branch into the cross-token attention module. Use the learnable token of the high-relevance modal branch as a proxy and combine it with the attention mechanism to learn the complementary information of the low-relevance modal branch to obtain the deep features after feature fusion. The cross-token attention module uses high-dimensional features from highly correlated modal branches as the main branch to extract learnable tokens. As a proxy, it learns complementary information of low-relevance modal branches through an attention mechanism, projects the proxy token back to the feature label, and concatenates it with the features of high-relevance modal branches to obtain deep features after feature fusion.
[0031] Specifically, such as Figure 3 As shown, cross-token attention aims to extract complementary information by dynamically weighting modal features. This effectively captures the correlations between these modalities and optimizes the feature representation of each modality using existing information. The module's proxy and attention mechanisms allow highly correlated modal branches to fully learn the complementary information of less correlated branches, while residual connections ensure smooth information flow and gradient stability.
[0032] S5. Input the deep features into the spatial encoder, extract global features by adding the position information of the deep features and applying a stacked Transformer architecture, and output pixel-level prediction results through the segmentation head. Among them, the spatial encoder retains class labels for deep features after feature fusion, adds spatial location encoding and performs spatial attention calculation, extracts global features through multi-layer Transformer, and combines the dynamic and spatial features of time series data to enhance the model's time series analysis capabilities.
[0033] The segmentation head processes the deep feature vectors of each group of multi-source time-series cube remote sensing data through an MLP containing LayerNorm and a linear layer to recover the pixel values of each cube block. It then stitches together multiple cube blocks to restore the original-size image, obtaining pixel-level prediction results. Each pixel corresponds to a regression prediction value.
[0034] Specifically, deep features are further extracted by adding location information and applying a stacked Transformer architecture. Finally, the corresponding pixel-level prediction results are output through the segmentation head, and the loss is calculated by comparing the obtained results with the true values to train the model.
[0035] S6. The prediction results of each group of multi-source time-series cube remote sensing data are stitched together in the original spatial order to obtain a pixel-level prediction image of the overall regional crop yield. In this process, a set of three-dimensional cubes at the same spatial location in each modality is stitched together and input into the model. After model processing, a crop yield estimation map for that spatial location is generated. The crop yield estimation maps for all spatial locations obtained from the model processing are stitched together in their original spatial order to obtain a pixel-level predicted image of the overall regional crop yield, such as... Figure 4 As shown, this invention divides an image into multiple temporal cube modes and feeds each temporal cube into a network model for prediction. Therefore, to obtain the result of the original whole image, we reassemble the prediction results of all the three-dimensional cube sets to finally obtain a pixel-level predicted image of the overall crop yield of the region.
[0036] Specifically, following the multimodal temporal cube partitioning order, all temporal cubes are placed into the model for prediction. The prediction results of the N sets of three-dimensional cube sets are then reassembled to finally obtain a pixel-level predicted image of the overall crop yield of the region.
[0037] S7. Compare the predicted images with the actual yield data, use evaluation indicators to verify the model accuracy, and optimize the model parameters.
[0038] The evaluation metrics include correlation coefficient R², root mean square error (RMSE), and mean absolute error (MAE). These metrics are calculated using predicted images and actual yield data to comprehensively assess model accuracy. Model parameters are adjusted using backpropagation algorithm and loss function minimization. Example
[0039] Step 1: Using the MODIS dataset, which contains various time-series remote sensing data from January 1, 2010 to the present, covering a global area. We selected five time-series remote sensing data from parts of China's terrestrial ecosystem and preprocessed them, aligning the spatial resolution to 500m and the temporal resolution to 8 days. We then normalized or standardized these modal data. Furthermore, the images of all modalities in this region were divided into multiple time-series cubes, which served as input to the model.
[0040] Step 2: The multi-source temporal cubic remote sensing data processed in Step 1 is divided according to prior knowledge to determine high-correlation mode branches and low-correlation mode branches. The modes of each branch first pass through a dynamic adaptive fusion module. This module utilizes a hybrid feature extractor, a 3D channel attention mechanism, a temporal feature encoder, and a learnable weighting mechanism to automatically adjust the mode weights at different time locations, extracting key features within each branch. Proceed to Step 3; Step 3: Input the modal features of the two branches obtained in Step 2 into the dual-branch cross-token attention module. In this module, the token proxy of the highly relevant modal branch is used in conjunction with the attention mechanism, enabling it to fully learn the key features and complementary information of the low-relevance modal branch, thus obtaining the fused features of the two branches. Proceed to Step 4; Step 4: Input the deep features obtained in Step 3 into the spatial encoder to extract global features. This step further extracts deep features by adding positional information and applying a stacked Transformer architecture. Finally, the corresponding pixel-level prediction results are output through the segmentation head. Proceed to Step 5.
[0041] Step 5: Following the multimodal temporal cube partitioning order in Step 5, input all temporal cubes into the network model for prediction. The prediction results of the N sets of 3D cubes are then reassembled to obtain a pixel-level predicted image of the overall crop yield for the region. Proceed to Step 6.
[0042] Step 6: Compare the overall pixel-level predicted image obtained by the model with the real image, and conduct experiments using multiple evaluation metrics to verify the prediction accuracy of the model.
[0043] The model uses correlation coefficients on the dataset. The method was evaluated using three metrics: root mean square error (RMSE), mean absolute error (MAE), and average mean square error (RMSE). It achieved the best results compared to other advanced spatiotemporal deep learning models. Furthermore, the prediction results of this invention compared to other methods are visually illustrated using charts.
[0044] Advantages of this invention: This invention combines spatiotemporal deep learning models with multimodal time series fusion technology to study crop yield prediction. By employing advanced deep learning techniques and innovative fusion strategies, this patent provides an effective method for crop yield prediction. The implementation of this invention will promote the further development of precision agriculture and improve agricultural productivity and the achievement of sustainable development goals.
[0045] In the overall architecture of this invention, all modal images are divided into high-relevance branches and low-relevance modal branches based on prior knowledge, and the image is further split into a set of N temporal cubes as input. Each set of inputs passes through a dynamic adaptive fusion module, a temporal encoder, a cross-token attention module, a spatial encoder, and a segmentation head according to its respective branch, finally obtaining the pixel-level prediction result of that set of temporal cubes.
[0046] In the dynamic adaptive fusion module of the present invention, all modalities in each branch are first evaluated for importance by a hybrid feature extractor and 3D channel attention. Then, the importance of each modality is dynamically assigned by a self-learning weight method. Finally, a learnable token mechanism is used for weighted fusion to obtain the fused features within each branch.
[0047] In the cross-token attention module of this invention, the high-dimensional features of highly correlated branches serve as the main branch, and learnable token classes are extracted from the main branch. As a proxy, the learnable classes of the geocorrelation branch are directly discarded. An attention mechanism is used to learn complementary and differential information in the low-correlation branch, and this proxy class token is projected back onto the feature labels. Finally, this proxy class token is concatenated with the original high-dimensional features to obtain the key features after the two branches are fully integrated.
[0048] In the pixel-level regression prediction of this invention, each set of modal cubes is... The final prediction result obtained through the model is then restored to its original size image according to the original row and column stitching order, resulting in the final pixel-level prediction result of crop yield. Its size is the same as that of each set of modal cubes in the input, and each pixel corresponds to a predicted regression value.
[0049] This invention maintains high performance while controlling computational complexity, resulting in lower memory usage, fewer parameters, and fewer FLOPS, making it suitable for resource-constrained devices. Compared to existing deep learning models that are often impractical due to their high computational demands, this invention achieves efficient inference through optimized modules (such as token brokers and spatial encoders).
[0050] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A crop yield prediction method based on a multi-modal fusion technology, characterized in that, include: The acquired multi-source remote sensing data is preprocessed to obtain multi-source remote sensing data with consistent temporal and spatial dimensions. The preprocessed remote sensing images of each modality are split into fine-grained segments, and a fixed-size temporal three-dimensional cube set is constructed according to the spatial location. The three-dimensional cube sets in the same spatial location of each modality are stitched together to obtain multiple sets of multi-source temporal cube remote sensing data, which are then input into the model. The model divides multi-source time-series cube remote sensing data into high-correlation mode branches and low-correlation mode branches based on prior knowledge, and extracts key features for each branch through a dynamic adaptive fusion module. The key features of both the high-relevance modality branch and the low-relevance modality branch are input into the cross-token attention module. The learnable token of the high-relevance modality branch is used as a proxy and combined with the attention mechanism to learn the complementary information of the low-relevance modality branch, so as to obtain the deep features after feature fusion. The deep features are input into the spatial encoder, and global features are extracted by adding the position information of the deep features and applying a stacked Transformer architecture. The pixel-level prediction results are output through the segmentation head. The prediction results of each group of multi-source time-series cube remote sensing data are stitched together in the original spatial order to obtain a pixel-level predicted image of crop yield for the entire region. The predicted images were compared with the actual yield data, and the model accuracy was verified using evaluation metrics, and the model parameters were optimized.
2. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: By stitching together the sets of three-dimensional cubes at the same spatial location in each modality and processing the model, a crop yield estimation map for that spatial location is generated. By stitching together the estimated crop yield maps of all spatial locations obtained from the model processing in the original spatial order, a pixel-level predicted image of the overall regional crop yield is obtained.
3. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: The dynamic adaptive fusion module evaluates the importance of each modality through a 3D spatiotemporal attention mechanism and dynamically allocates the weights of each modality using a self-learning weighting method.
4. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: Each branch extracts key features through a dynamic adaptive fusion module, including: The modalities of each branch first pass through a dynamic adaptive fusion module. This module uses a hybrid feature extractor, a 3D channel attention mechanism, a temporal feature encoder, and a learnable weight mechanism to automatically adjust the modal weights at different times and locations, thereby extracting the key features of that branch.
5. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: The cross-token attention module uses the high-dimensional features of the highly relevant modal branch as the main branch, extracts the learnable token t_HCM^cls as a proxy, learns the complementary information of the low-relevance modal branch through the attention mechanism, projects the proxy token back to the feature label, and concatenates it with the features of the highly relevant modal branch to obtain the deep features after feature fusion. 6.The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: The spatial encoder retains class labels for deep features after feature fusion, adds spatial location encoding and performs spatial attention calculation, and extracts global features through a multi-layer Transformer.
7. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: The segmentation head processes the deep feature vectors of each group of multi-source time-series cube remote sensing data through an MLP containing LayerNorm and a linear layer to recover the pixel value of each cube block, and stitches multiple cube blocks together to restore the original-size image, obtaining pixel-level prediction results, with each pixel corresponding to a regression prediction value.
8. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: Preprocessing of the acquired multi-source remote sensing data includes: The spatial and temporal resolutions of multi-source remote sensing data are uniformly processed, and feature normalization or standardization is performed.
9. The crop yield prediction method based on multi-modal fusion technology according to claim 1, characterized in that: The prior knowledge quantifies the correlation coefficient between multi-source remote sensing data and crop yield. A preset correlation threshold is set. If the correlation coefficient between multi-source remote sensing data and crop yield is higher than the correlation threshold, the mode is classified as a high-correlation mode branch. If it is lower than the correlation threshold, the mode is classified as a low-correlation mode branch.
10. The crop yield prediction method based on multimodal fusion technology according to claim 1, characterized in that: The evaluation metrics include correlation coefficient R², root mean square error (RMSE), and mean absolute error (MAE). These metrics are calculated using predicted images and actual yield data to comprehensively assess model accuracy. Model parameters are adjusted using backpropagation algorithm and loss function minimization.