Machine learning based method for predicting high temperature yield loss in corn
By constructing a high-temperature-drought coupled feature tensor and a deep learning model, the model drift problem in predicting high-temperature loss of maize yield under unsteady climate conditions was solved, achieving high-precision and robust prediction results, supporting agricultural decision-making and disaster early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 湖南省作物研究所
- Filing Date
- 2025-11-04
- Publication Date
- 2026-05-05
AI Technical Summary
Existing machine learning-based models for predicting maize yield loss due to high temperatures are unable to effectively distinguish the interaction between high temperature and drought stress under non-steady-state climate conditions, leading to model drift and yield prediction bias, especially under combined stress conditions, resulting in misjudgment and deviation.
By constructing a high-temperature-drought composite coupled feature tensor, performing nonlinear orthogonal decomposition using covariance and mutual information matrices, and combining a deep learning model constrained by a composite stress confusion distortion index, the network parameters are optimized to generate a stable and robust high-temperature loss prediction model. Feature contribution analysis and spatial grid mapping are then performed.
It enables refined, dynamic, and spatial prediction of high-temperature losses in maize yield, significantly improving prediction accuracy and generalization ability. It can cope with climate instability and combined heat and drought stress, providing a scientific basis for agricultural decision-making.
Smart Images

Figure CN121436286B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-temperature loss prediction technology, and more specifically, to a method for predicting high-temperature loss in corn yield based on machine learning. Background Technology
[0002] In machine learning-based studies predicting maize yield loss due to high temperatures, models typically rely on historical meteorological sequences, remote sensing features, and yield samples to construct a mapping relationship between high-temperature stress and yield response. However, as the global climate system enters a highly unstable phase, the spatiotemporal superposition of abnormal temperature rise events and regional drought events is becoming increasingly frequent, forming the so-called "heat-drought combined stress" phenomenon. Against this backdrop, climate instability not only alters the statistical distribution characteristics of high-temperature events but also causes a drift in the input-output relationship upon which the model relies. The "temperature rise—yield decline" pattern learned during model training often fails to maintain stable generalization in complex situations involving drought superimposed on top of this, resulting in a significant model drift effect.
[0003] Meanwhile, high temperature and drought stress exhibit a high degree of coupling at both physiological and spectral levels: drought leads to stomatal closure and increased leaf temperature, while high temperature further exacerbates transpiration imbalance and water deficit, both affecting canopy heat balance and photosynthetic efficiency. Because these combined effects manifest as a synchronous signal of decreasing NDVI (Density Spectrum Indicator) and increasing LST (Land Surface Temperature) in the remote sensing feature space, models struggle to distinguish the independent contributions of the two stresses at the feature layer. Parameter drift under unstable climate conditions further amplifies this confusion, causing models to misinterpret drought-induced spectral changes as high-temperature responses, thus forming a closed-loop amplification chain of "climate drift—feature aliasing—parameter instability—re-drift" during gradient learning.
[0004] This interaction not only undermines the model's ability to accurately represent the intensity of high-temperature stress but also leads to systematic distortion of yield sensitivity parameters to high-temperature stress. As a result, high-temperature loss predictions exhibit spatial shifts and temporal lags, and in severe drought years, they may even incorrectly underestimate or overestimate the risk of yield loss, forming a typical "compound stress confounding distortion" problem. The root cause lies in the mutually reinforcing coupling between climate non-steady-state-driven distribution drift and the characteristic interaction of heat-drought stress. This causes machine learning models, in the absence of explicit causal constraints, to mistakenly mix cross-variables into a single response path, resulting in high-temperature loss predictions deviating from true physiological patterns. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a machine learning-based method for predicting high-temperature losses in corn yield, thereby addressing the problems mentioned in the background section.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A machine learning-based method for predicting high-temperature losses in maize yield includes the following steps:
[0008] Collect ground meteorological, remote sensing imagery and measured yield data, perform temporal alignment and spatial resampling, and generate a multi-source fusion dataset with a unified spatiotemporal scale;
[0009] Based on the multi-source fusion dataset, identify temperature anomaly segments during the critical growth period of crops, calculate the duration and intensity of high temperatures, form a high-temperature stress feature sequence, and obtain a continuous high-temperature stress function through time smoothing.
[0010] Using the high-temperature stress function as a time reference, the soil moisture, evapotranspiration and vegetation index sequences were time-aligned, drought intensity and recovery principal components were extracted, and a drought stress response matrix was constructed.
[0011] Based on the high temperature stress function and drought response matrix, the covariance and mutual information matrix are calculated, and nonlinear orthogonal decomposition is performed to generate the high temperature-drought composite coupled feature tensor.
[0012] Based on the high temperature-drought composite coupling characteristic tensor, calculate the climate drift KL divergence and coupling covariance coefficient, and fuse them to generate the composite stress confusion distortion index;
[0013] The composite stress representation vector and historical yield loss samples are input into the deep learning model. The composite stress confusion distortion index is used as the regularization term to constrain the loss function. The network parameters are optimized and the prediction bias is dynamically adjusted to generate a stable and robust high-temperature loss prediction model.
[0014] Feature contribution analysis is performed on the prediction results of the high temperature loss prediction model, and the corrected production loss is mapped to a spatial grid to generate a high temperature loss distribution map, thereby realizing the risk visualization of the compound stress area.
[0015] In a preferred embodiment, abnormal temperature segments during key crop growth periods are identified based on the multi-source fusion dataset, the duration and intensity of high temperatures are calculated to form a high-temperature stress feature sequence, and a continuous high-temperature stress function is obtained through time smoothing, as follows:
[0016] Based on crop type and regional planting history, the corresponding growth period time range is extracted from the multi-source fusion dataset; the start and end times of the growth stage are calculated using the growth period division table. ,in This represents the cumulative effective accumulated temperature of the crop at time t. This represents the daily average temperature on day i. This represents the physiological zero degree of crop growth. This indicates the start date of sowing and germination, and t represents the current detection date number.
[0017] When the cumulative effective accumulated temperature reaches the stage threshold, the transition point of the reproductive period is determined;
[0018] Daily maximum temperature data were extracted during the critical reproductive period to establish temperature thresholds. If the number of consecutive days meets the requirement If the temperature is abnormal, it is marked as a high temperature abnormality section, and the start and end positions of the temperature abnormality section are detected by a sliding time window.
[0019] For each high-temperature anomaly zone, calculate the duration of the high temperature: ,in This represents the number of days the k-th high-temperature anomaly segment lasted. This indicates the starting time of the k-th high-temperature anomaly segment. This indicates the end time of the k-th high-temperature anomaly segment;
[0020] And calculate the high-temperature strength index: ,in This represents the average high-temperature intensity of the k-th high-temperature anomaly zone;
[0021] All high-temperature anomaly zones were arranged in chronological order to form a high-temperature stress characteristic sequence;
[0022] High temperature stress characteristic sequence Smoothing was performed using Gaussian weighted smoothing to obtain a continuous high-temperature stress function. .
[0023] In a preferred embodiment, the high-temperature stress function is used as a time reference to perform time-series alignment of soil moisture, evapotranspiration, and vegetation index sequences, extract drought intensity and recovery principal components, and construct a drought stress response matrix, as detailed below:
[0024] Based on the obtained high temperature stress function Using time as a reference, time mapping is performed on time series from different sources, and time scale unification is achieved using a linear interpolation function;
[0025] In high temperature stress function Extract the corresponding soil moisture during the period when the rate of change reaches a local extreme. , evapotranspiration With vegetation index The subsequence, defined as the drought response window by calculating the high temperature-soil moisture cross-correlation coefficient;
[0026] Drought intensity and restorative principal components are extracted within the drought response window to construct a drought stress response matrix.
[0027] In a preferred embodiment, the calculation of the high temperature-soil moisture cross-correlation coefficient, defined as the drought response window, is as follows:
[0028] The specific formula for calculating the high temperature-soil moisture cross-correlation coefficient is as follows: ,in , They represent , Time average, Indicates the time lag. This represents the correlation coefficient between high temperature and soil moisture.
[0029] Will The time lag at which the maximum value is obtained is marked as the optimal lag, and the time window corresponding to the optimal lag is marked as the drought response window.
[0030] In a preferred embodiment, the step of extracting drought intensity and restorative principal components within the drought response window to construct a drought stress response matrix is as follows:
[0031] Calculation of normalized soil moisture deficit index ;
[0032] Calculate the normalized evapotranspiration index ;
[0033] Constructing the overall drought intensity: ,in These represent the normalized soil moisture deficit index. Normalized Evapotranspiration Index Weighting coefficients;
[0034] vegetation index Perform principal component analysis within the drought response window to obtain the first principal component. The formula for representing vegetation restoration potential is as follows: ,in In principal component analysis (PCA), the eigenvector corresponding to the largest eigenvalue is... Let the mean vector of vegetation indices be used; the first principal component is... Marked as a restorative principal component, It is the transpose symbol;
[0035] The drought intensity and recovery principal components were aligned by time to form a drought stress response matrix. : .
[0036] In a preferred embodiment, based on the high-temperature stress function and the drought response matrix, the covariance and mutual information matrix are calculated, and a nonlinear orthogonal decomposition is performed to generate a high-temperature-drought composite coupled feature tensor, as follows:
[0037] High temperature stress function With drought stress response matrix Forming a feature matrix by time alignment ; Calculate the characteristic matrix covariance matrix ;
[0038] Calculate the characteristic matrix Mutual information between each pair of features ;
[0039] Obtain the mutual information matrix ;
[0040] The covariance matrix With mutual information matrix Fusion to form a weighted composite matrix : ,in For a weighted composite matrix, These are linear weighting coefficients;
[0041] right Perform nonlinear orthogonal decomposition to solve for orthogonal eigenvectors. With eigenvalue matrix ;
[0042] The eigenvectors after orthogonal decomposition Stacking them over time forms a three-dimensional tensor: ,in The high-temperature-drought composite coupled feature tensor is a three-dimensional matrix, consisting of time × feature × orthogonal components, where t is the time index. For feature index, For time t, the first The feature in the first Projected values on each orthogonal component.
[0043] In a preferred embodiment, the climate drift KL divergence and coupling covariance coefficient are calculated based on the high temperature-drought composite coupling characteristic tensor, and then fused to generate a composite stress confusion distortion index, as follows:
[0044] High temperature-drought coupled characteristic tensor By statistically analyzing the combined distribution of historical data from multiple years and current year data, we obtain the historical probability distribution. and the current probability distribution ;
[0045] Based on historical probability distribution For reference, calculate the current probability distribution. Climate drift KL divergence : ,in This represents the probability distribution of historical high-temperature-drought characteristics. This represents the probability distribution of the high-temperature-drought characteristics for the current year. The state vector of the feature space;
[0046] High temperature-drought coupled characteristic tensor For each orthogonal component, calculate the covariance between pairwise features. ; Calculate the coupling covariance coefficient ;
[0047] Climate drift KL divergence and mean coupled covariance coefficient Perform weighted fusion calculation of the composite stress confusion distortion index : ,in The weighting coefficients for climate drift KL divergence and average coupling covariance coefficient are determined by the minimum error criterion.
[0048] In a preferred embodiment, the composite stress representation vector and historical yield loss samples are input into a deep learning model. The composite stress confusion distortion index is used as a regularization term to constrain the loss function. The network parameters are optimized and the prediction bias is dynamically adjusted to generate a stable and robust high-temperature loss prediction model, as detailed below:
[0049] The composite stress characterization vector Corresponding historical production loss samples Align the training samples according to time and space to form a training sample set. ,in Let be the composite stress representation vector of the q-th sample. Let q be the measured yield loss value for the q-th sample. The number of samples;
[0050] The deep learning model is constructed from a multi-branch neural network structure, with the main branch capturing high-temperature features and auxiliary branches modeling drought and recovery features. ,in This represents the corn yield loss value predicted by the deep learning model. Indicates network parameters Nonlinear mapping model;
[0051] Constructing a composite loss function with regularization terms : ,in For composite loss function, These are the regularization weight coefficients. For the compound stress confusion distortion constraint function, ;
[0052] An adaptive learning rate optimization algorithm is used to iteratively update network parameters, train a high-temperature loss prediction model, and finally generate a stable and robust high-temperature loss prediction model.
[0053] The technical effects and advantages of this invention are as follows:
[0054] 1. This invention constructs a machine learning prediction system based on the combined characteristics of high temperature and drought, achieving refined, dynamic, and spatial prediction of high temperature loss in maize yield. It can effectively capture the interaction between high temperature and drought stress under unsteady climate conditions. Through joint analysis of the high temperature stress function and drought response matrix, the combined stress characteristics are quantified. Nonlinear orthogonal decomposition and mutual information analysis are used to reduce misjudgments caused by feature aliasing, thereby significantly reducing the impact of model drift on yield prediction. The combined stress confusion distortion index is used to constrain the deep learning model training, optimize network parameters, and dynamically correct prediction biases, achieving stable and robust prediction of yield loss under extreme climate conditions. At the same time, through feature contribution analysis and spatial grid mapping, the spatiotemporal distribution of yield loss under combined stress conditions can be intuitively displayed, achieving accurate visualization of high temperature risk areas.
[0055] 2. On the one hand, this invention significantly improves the accuracy and generalization ability of high-temperature loss prediction, enabling it to cope with complex environments under climate instability and combined heat and drought stress; on the other hand, through quantitative feature contribution and spatial correction, it realizes the interpretability and visualization of yield loss prediction results, providing a scientific basis for agricultural decision-making, planting regulation and disaster early warning. Attached Figure Description
[0056] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings;
[0057] Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation
[0058] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0059] Example: Figure 1This invention presents a machine learning-based method for predicting high-temperature losses in maize yield, comprising the following steps:
[0060] Collect ground meteorological, remote sensing imagery and measured yield data, perform temporal alignment and spatial resampling, and generate a multi-source fusion dataset with a unified spatiotemporal scale;
[0061] Based on the multi-source fusion dataset, identify temperature anomaly segments during the critical growth period of crops, calculate the duration and intensity of high temperatures, form a high-temperature stress feature sequence, and obtain a continuous high-temperature stress function through time smoothing.
[0062] Using the high-temperature stress function as a time reference, the soil moisture, evapotranspiration and vegetation index sequences were time-aligned, drought intensity and recovery principal components were extracted, and a drought stress response matrix was constructed.
[0063] Based on the high temperature stress function and drought response matrix, the covariance and mutual information matrix are calculated, and nonlinear orthogonal decomposition is performed to generate the high temperature-drought composite coupled feature tensor.
[0064] Based on the high temperature-drought composite coupling characteristic tensor, calculate the climate drift KL divergence and coupling covariance coefficient, and fuse them to generate the composite stress confusion distortion index;
[0065] The composite stress representation vector and historical yield loss samples are input into the deep learning model. The composite stress confusion distortion index is used as the regularization term to constrain the loss function. The network parameters are optimized and the prediction bias is dynamically adjusted to generate a stable and robust high-temperature loss prediction model.
[0066] Feature contribution analysis is performed on the prediction results of the high temperature loss prediction model, and the corrected production loss is mapped to a spatial grid to generate a high temperature loss distribution map, thereby realizing the risk visualization of the compound stress area.
[0067] Collect ground meteorological, remote sensing imagery and measured yield data, perform temporal alignment and spatial resampling, and generate a multi-source fusion dataset with a unified spatiotemporal scale;
[0068] In this embodiment of the invention, raw data is obtained from ground meteorological monitoring stations, satellite remote sensing image databases, and farmland measured yield records. The collected ground meteorological, remote sensing, and measured yield data are identified by source and formatted uniformly, establishing a data coding system including timestamps, geographic coordinates, sensor numbers, and crop type codes to ensure traceability and consistency in subsequent fusion. A unified time reference table is established to address the temporal resolution differences between different data sources (e.g., meteorological data is hourly, remote sensing data is daily, and yield data is quarterly). Time interpolation (e.g., linear interpolation, spline interpolation, PCHIP, etc.) and sliding window weighted smoothing algorithms (e.g., simple moving average, weighted moving average, Gaussian smoothing, etc.) are used to upsample low-frequency data to the target temporal resolution, while high-frequency data is time-aggregated to ensure that data from the same time slice is consistent across different time segments. The data from various sources are comparable; a unified spatial grid is established based on the spatial resolution of remote sensing images (e.g., 10 meters or 30 meters); ground meteorological observation point data are extended to grid cells using inverse distance weighted interpolation or kriging interpolation; spatial matching and local weighted regression are performed on measured yield point data to form corresponding spatial rasterized data; cloud cover, noise, or missing areas in remote sensing images are filled using neighborhood averaging and principal component reconstruction; bilinear or cubic convolution resampling is performed on data with inconsistent spatial resolutions to unify them to the target resolution; outliers with significant temporal or spatial deviations are removed and corrected using median filtering and standard deviation constraints; the time- and spatially aligned meteorological, remote sensing, and yield data are overlaid on a unified spatiotemporal grid to finally generate a multi-source fusion dataset with a unified spatiotemporal scale.
[0069] Based on the multi-source fusion dataset, identify temperature anomaly segments during the critical growth period of crops, calculate the duration and intensity of high temperatures, form a high-temperature stress feature sequence, and obtain a continuous high-temperature stress function through time smoothing.
[0070] In this embodiment of the invention, based on crop type and regional planting history, the corresponding growth period time range (such as jointing stage, heading stage, grain-filling stage, etc.) is extracted from a multi-source fusion dataset; the start and end times of the growth stage are calculated using a growth period division table. ,in This represents the cumulative effective accumulated temperature of the crop at time t, used to determine the crop's growth stage. This represents the daily average temperature on day i. This represents the physiological zero degree for crop growth, which is the lowest temperature threshold at which crop growth begins. This indicates the start date of sowing and germination, and t represents the current detection date number.
[0071] When the cumulative effective accumulated temperature reaches the stage threshold, the transition point of the reproductive period is determined;
[0072] Daily maximum temperature data were extracted during the critical reproductive period to establish temperature thresholds. (Determined based on the crop's critical heat tolerance temperature, such as 35℃); if the number of consecutive days meets the requirements... If the temperature is abnormal, it is marked as a high temperature abnormality section, and the start and end positions of the temperature abnormality section are detected by a sliding time window.
[0073] It should be noted that, This indicates the highest temperature on a given day. The number of consecutive days can be set to a minimum threshold, such as 3 days, to exclude short-term fluctuations.
[0074] For each high-temperature anomaly zone, calculate the duration of the high temperature: ,in This represents the number of days the k-th high-temperature anomaly segment lasted. This indicates the starting time of the k-th high-temperature anomaly segment. This indicates the end time of the k-th high-temperature anomaly segment;
[0075] And calculate the high-temperature strength index: ,in This represents the average high-temperature intensity of the k-th high-temperature anomaly zone;
[0076] Arrange all high-temperature anomaly zones in chronological order to form a high-temperature stress characteristic sequence: ,in The characteristic value of high-temperature stress (dimensionless or normalized thermal intensity unit) represents the degree of thermal stress at time t. The mapping function can be in linear weighted form. ( It can be used as a weighting coefficient for the duration and intensity of high temperature, or a nonlinear combination (such as an exponential or sigmoid function) to comprehensively reflect the interaction effect between the duration and intensity of high temperature.
[0077] High temperature stress characteristic sequence Smoothing is performed using Gaussian weighted smoothing: ,in This represents the high-temperature stress characteristic value at step j within the time window, and n represents the radius of the time smoothing window (usually taken as 3–7 days). The standard deviation of the Gaussian weights is used to control the smoothness; the exponential term is also present. This represents a time-distance weighting factor, ensuring that nearby times have a greater impact; thus, a continuous high-temperature stress function is obtained. It is used to reflect the temporal evolution trend of heat stress during the reproductive period.
[0078] Using the high-temperature stress function as a time reference, the soil moisture, evapotranspiration and vegetation index sequences were time-aligned, drought intensity and recovery principal components were extracted, and a drought stress response matrix was constructed.
[0079] In this embodiment of the invention, the already obtained high-temperature stress function is used. Using soil moisture as a time reference, time series from different sources (including soil moisture) , evapotranspiration With vegetation index Perform time mapping and use a linear interpolation function to achieve time scale unification: ,in This represents the sequence of variables synchronized with the high-temperature stress function after time remapping. This refers to the original sequence (referring to the original time series variables, including soil moisture). , evapotranspiration With vegetation index The start time of the sequence. Indicates the time step. , Represents the high temperature stress function The start and end times;
[0080] It should be noted that soil moisture refers to the amount of water contained in a unit volume or mass of soil. It is an important indicator for measuring the water available to crop roots. Its dynamic changes reflect the degree of drought and the process of soil moisture deficit, and it is a key mediating variable for the impact of high temperature stress on crops. Evapotranspiration includes both plant transpiration and soil evaporation. It characterizes the intensity of crop water consumption and atmospheric evaporation demand, and is an important variable for judging crop water deficit and energy balance. The vegetation index is a dimensionless index formed by the combination of multi-band reflectance. It reflects crop growth, leaf area, photosynthetic capacity and physiological state, and is often used to characterize the response and recovery of vegetation to stress.
[0081] In high temperature stress function Extract the corresponding soil moisture during the period when the rate of change reaches a local extreme. , evapotranspiration With vegetation index The subsequence, defined as the drought response window by calculating the high temperature-soil moisture cross-correlation coefficient;
[0082] The drought response window is defined by calculating the high temperature-soil moisture cross-correlation coefficient, as follows:
[0083] The specific formula for calculating the high temperature-soil moisture cross-correlation coefficient is as follows: ,in , They represent , Time average, This represents the time lag, used to measure the time delay in the response of soil moisture to changes in high temperature. This represents the correlation coefficient between high temperature and soil moisture.
[0084] Will The time lag at which the maximum value is obtained is marked as the optimal lag, and the time window corresponding to the optimal lag is marked as the drought response window;
[0085] Drought intensity and restorative principal components are extracted within the drought response window to construct a drought stress response matrix;
[0086] The process of extracting drought intensity and restorative principal components within the drought response window to construct a drought stress response matrix is as follows:
[0087] Calculation of normalized soil moisture deficit index : ,in These represent the maximum and minimum soil moisture values within the entire drought response window, respectively.
[0088] Calculate the normalized evapotranspiration index : ,in These represent the maximum and minimum evapotranspiration values within the entire drought response window, respectively.
[0089] By combining the evapotranspiration rate with the evaporative stress effect, a comprehensive drought intensity is constructed: ,in These represent the normalized soil moisture deficit index. Normalized Evapotranspiration Index The weighting coefficients are determined using the minimum error criterion;
[0090] vegetation index Perform principal component analysis (PCA) within the drought response window to obtain the first principal component. The formula for representing vegetation restoration potential is as follows: ,in In principal component analysis (PCA), the eigenvector corresponding to the largest eigenvalue is... Let the mean vector of vegetation indices be used; the first principal component is... Marked as a restorative principal component, It is the transpose symbol;
[0091] The drought intensity and recovery principal components were aligned by time to form a drought stress response matrix. : The drought stress response matrix characterizes the intensity, persistence, and vegetation response features of drought stress driven by high temperature, providing an input basis for subsequent calculation of the "high temperature-drought composite coupling feature tensor".
[0092] Based on the high temperature stress function and drought response matrix, the covariance and mutual information matrix are calculated, and nonlinear orthogonal decomposition is performed to generate the high temperature-drought composite coupled feature tensor.
[0093] In this embodiment of the invention, the high temperature stress function is... With drought stress response matrix Forming a feature matrix by time alignment : , characteristic matrix Includes high temperature stress intensity, drought intensity, and restorative principal components; calculates the characteristic matrix. covariance matrix Quantify the linear dependencies between features: ,in For time steps, As the characteristic mean vector, the covariance matrix can reveal the covariance of high temperature and drought stress at different time periods;
[0094] Calculate the characteristic matrix Mutual information between each pair of features : ,in , Representation of the characteristic matrix A single feature column in the data. for and The joint probability distribution, The marginal probability distribution of each feature;
[0095] Obtain the mutual information matrix Capture non-linear dependencies;
[0096] The covariance matrix With mutual information matrix Fusion to form a weighted composite matrix : ,in It is a weighted composite matrix that integrates linear covariance and nonlinear mutual information. These are linear weighting coefficients that control the contribution ratio of the covariance matrix to the composite matrix;
[0097] right Perform nonlinear orthogonal decomposition (NOD) to solve for orthogonal eigenvectors. With eigenvalue matrix : ;
[0098] The eigenvectors after orthogonal decomposition Stacking them over time forms a three-dimensional tensor: ,in The high-temperature-drought composite coupled feature tensor is a three-dimensional matrix, consisting of time × feature × orthogonal components, where t is the time index. For feature index, corresponding , These are the indices of the orthogonal components, corresponding to the orthogonal directions obtained from the decomposition. For time t, the first The feature in the first Projected values on each orthogonal component.
[0099] Based on the high temperature-drought composite coupling characteristic tensor, calculate the climate drift KL divergence and coupling covariance coefficient, and fuse them to generate the composite stress confusion distortion index;
[0100] In this embodiment of the invention, the high-temperature-drought composite coupling feature tensor is used. By statistically analyzing the combined distribution of historical data from multiple years and current year data, we obtain the historical probability distribution. and the current probability distribution The kernel density estimation method is used to probabilistically process the tensor data to ensure a smooth distribution in the high-dimensional feature space.
[0101] Based on historical probability distribution For reference, calculate the current probability distribution. Climate drift KL divergence : ,in This represents the probability distribution of historical high-temperature-drought characteristics. This represents the probability distribution of the high-temperature-drought characteristics for the current year. Let be the state vector of the feature space, representing the value of each orthogonal feature at a certain time point;
[0102] Climate drift KL divergence measures the degree of deviation of the current high temperature-drought composite characteristics relative to historical models under unsteady climate conditions, reflecting the potential drift risk of the model;
[0103] High temperature-drought coupled characteristic tensor For each orthogonal component, calculate the covariance between pairwise features. : ,in The time step is the length of the tensor along the time dimension. Let g be the mean of the feature in the time dimension. The mean of feature h over the time dimension; calculate the coupling covariance coefficient. : ,in , These are the standard deviations of features g and h, respectively;
[0104] Climate drift KL divergence and mean coupled covariance coefficient ( The composite stress confusion distortion index is calculated by weighting and fusing the features based on the number of feature dimensions. : ,in The weighting coefficients for climate drift KL divergence and average coupling covariance coefficient are determined by the minimum error criterion.
[0105] The composite stress confusion index is used to comprehensively characterize the prediction bias risk caused by the interaction of high temperature and drought under unsteady climate conditions.
[0106] The composite stress representation vector and historical yield loss samples are input into the deep learning model. The composite stress confusion distortion index is used as the regularization term to constrain the loss function. The network parameters are optimized and the prediction bias is dynamically adjusted to generate a stable and robust high-temperature loss prediction model.
[0107] In this embodiment of the invention, the composite stress characterization vector is... Corresponding historical production loss samples Align the training samples according to time and space to form a training sample set. ,in Let be the composite stress representation vector of the q-th sample. Let q be the measured yield loss value for the q-th sample. The number of samples;
[0108] The deep learning model is constructed from a multi-branch neural network structure, with the main branch capturing high-temperature features and auxiliary branches modeling drought and recovery features. ,in This represents the corn yield loss value predicted by the deep learning model. Indicates network parameters Nonlinear mapping models can employ multi-layer convolutional networks (CNNs) or long short-term memory networks (LSTMs) to handle mixed temporal and spatial features;
[0109] Constructing a composite loss function with regularization terms : ,in This is a composite loss function, used as the objective function to guide model training. These are regularization weights, used to control the intensity of the influence of the composite stress constraint terms on the model optimization process. This is a composite stress-induced confusion distortion constraint function, used to penalize samples with large prediction bias in regions with high distortion exponents. This indicates the sensitivity of the predicted value to the distortion index under conditions of compound distortion;
[0110] An adaptive learning rate optimization algorithm (such as Adam or RMSProp) is used to iteratively update the network parameters to train a high-temperature loss prediction model. ,in Let be the model parameter vector for the t-th iteration. For the updated model parameters, The learning rate for the current iteration step controls the magnitude of parameter updates. This represents the gradient of the composite loss function with respect to the model parameters.
[0111] Cross-Climate Validation is performed on the trained high-temperature loss prediction model. This involves independently evaluating the stability of the model's output under different climate phases (such as historical cold and wet periods and current hot and dry periods) to ultimately generate a stable and robust high-temperature loss prediction model. .
[0112] Feature contribution analysis is performed on the prediction results of the high temperature loss prediction model, and the corrected production loss is mapped to a spatial grid to generate a high temperature loss distribution map, thereby realizing the risk visualization of the compound stress area.
[0113] In this embodiment of the invention, an optimized high-temperature loss prediction model is used to characterize the combined stress during the verification period and the current period. Inference is performed to obtain the predicted yield loss sequence. The marginal contribution strength is obtained by calculating the sensitivity of the high-temperature loss prediction model output to the input features using gradient-weighted feature attribution (GWFA). ,in Let y be the marginal contribution strength of the y-th feature to the prediction result; normalization is performed on all features: Where z is the feature index; the standardized feature contribution weight vector is obtained. This is used to distinguish the relative impact of each stress factor on yield loss;
[0114] Forecasted production loss Compared with actual production loss Perform residual analysis: ,in For production loss residuals;
[0115] Using feature contribution weights Factorize and weight the residuals: ,in The predicted yield loss is after factor weighting adjustment. The weights for the y-th feature in the feature contribution weight vector. For the local residuals of the characteristic component y;
[0116] Establish a spatial coordinate index matrix based on remote sensing geographic grids (e.g., 1km×1km). The corrected predicted yield loss is mapped to the corresponding grid points. Spatial interpolation (such as kriging or radial basis function interpolation) can be used to form a continuous surface to obtain a high-temperature loss distribution map.
[0117] This invention constructs a machine learning prediction system based on the combined characteristics of high temperature and drought, achieving refined, dynamic, and spatial prediction of high temperature loss in maize yield. It can effectively capture the interaction between high temperature and drought stress under unsteady climate conditions. By jointly analyzing the high temperature stress function and the drought response matrix, the combined stress characteristics are quantified. Nonlinear orthogonal decomposition and mutual information analysis are used to reduce misjudgments caused by feature aliasing, thereby significantly reducing the impact of model drift on yield prediction. The combined stress confusion distortion index is used to constrain the deep learning model training, optimize network parameters, and dynamically correct prediction biases, achieving stable and robust prediction of yield loss under extreme climate conditions. At the same time, through feature contribution analysis and spatial grid mapping, the spatiotemporal distribution of yield loss under combined stress conditions can be intuitively displayed, achieving accurate visualization of high temperature risk areas.
[0118] This invention significantly improves the accuracy and generalization ability of high-temperature loss prediction, enabling it to cope with complex environments under climate instability and combined heat and drought stress. On the other hand, by quantifying feature contributions and spatial correction, it achieves interpretability and visualization of yield loss prediction results, providing a scientific basis for agricultural decision-making, planting regulation, and disaster early warning.
[0119] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0120] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0121] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A machine learning-based method for predicting high-temperature loss in maize yield, characterized by: Includes the following steps: Collect ground meteorological, remote sensing imagery and measured yield data, perform temporal alignment and spatial resampling, and generate a multi-source fusion dataset with a unified spatiotemporal scale; Based on the multi-source fusion dataset, identify temperature anomaly segments during the critical growth period of crops, calculate the duration and intensity of high temperatures, form a high-temperature stress feature sequence, and obtain a continuous high-temperature stress function through time smoothing. Using the high-temperature stress function as a time reference, the soil moisture, evapotranspiration and vegetation index sequences were time-aligned, drought intensity and recovery principal components were extracted, and a drought stress response matrix was constructed. Based on the high temperature stress function and drought response matrix, the covariance and mutual information matrix are calculated, and nonlinear orthogonal decomposition is performed to generate the high temperature-drought composite coupled feature tensor. Based on the high temperature-drought composite coupling characteristic tensor, calculate the climate drift KL divergence and coupling covariance coefficient, and fuse them to generate the composite stress confusion distortion index; The composite stress representation vector and historical yield loss samples are input into a deep learning model. The loss function is constrained by the composite stress confusion distortion index as a regularization term. The network parameters are optimized and the prediction bias is dynamically adjusted to generate a stable and robust high-temperature loss prediction model. The composite stress representation vector includes the high-temperature stress function, drought comprehensive intensity, restorative principal component, and composite stress confusion distortion index. Feature contribution analysis is performed on the prediction results of the high temperature loss prediction model, and the corrected production loss is mapped to a spatial grid to generate a high temperature loss distribution map, thereby realizing risk visualization of the compound stress area. Based on the high-temperature-drought composite coupling characteristic tensor, the climate drift KL divergence and coupling covariance coefficient are calculated, and a composite stress confusion distortion index is generated by fusing them, as follows: High temperature-drought coupled characteristic tensor By statistically analyzing the combined distribution of historical data from multiple years and current year data, we obtain the historical probability distribution. and the current probability distribution ; Based on historical probability distribution For reference, calculate the current probability distribution. Climate drift KL divergence : ,in This represents the probability distribution of historical high-temperature-drought characteristics. This represents the probability distribution of the high-temperature-drought characteristics for the current year. The state vector of the feature space; High temperature-drought coupled characteristic tensor For each orthogonal component, calculate the covariance between pairwise features. ; Calculate the coupling covariance coefficient ; Climate drift KL divergence and mean coupled covariance coefficient Weighted fusion is performed, where For feature index, As orthogonal components, calculate the composite stress confusion distortion index. : ,in The weighting coefficients for climate drift KL divergence and average coupling covariance coefficient are determined by the minimum error criterion.
2. The method for predicting high-temperature loss of corn yield based on machine learning according to claim 1, characterized in that: Based on the multi-source fusion dataset, abnormal temperature segments during key crop growth periods are identified, and the duration and intensity of high temperatures are calculated to form a high-temperature stress feature sequence. A continuous high-temperature stress function is then obtained through time smoothing, as detailed below: Based on crop type and regional planting history, the corresponding growth period time range is extracted from the multi-source fusion dataset; the start and end times of the growth stage are calculated using the growth period division table. ,in This represents the cumulative effective accumulated temperature of the crop at time t. This represents the daily average temperature on day i. This represents the physiological zero degree of crop growth. This indicates the start date of sowing and germination, and t represents the current detection date number. When the cumulative effective accumulated temperature reaches the stage threshold, the transition point of the reproductive period is determined; Daily maximum temperature data were extracted during the critical reproductive period to establish temperature thresholds. If the number of consecutive days meets the requirement If the temperature is abnormal, it is marked as a high temperature abnormality section, and the start and end positions of the temperature abnormality section are detected by a sliding time window. For each high-temperature anomaly zone, calculate the duration of the high temperature: ,in This represents the number of days the k-th high-temperature anomaly segment lasted. This indicates the starting time of the k-th high-temperature anomaly segment. This indicates the end time of the k-th high-temperature anomaly segment; And calculate the high-temperature strength index: ,in This represents the average high-temperature intensity of the k-th high-temperature anomaly zone; All high-temperature anomaly zones were arranged in chronological order to form a high-temperature stress characteristic sequence; High temperature stress characteristic sequence Smoothing was performed using Gaussian weighted smoothing to obtain a continuous high-temperature stress function. .
3. The method for predicting high-temperature loss of corn yield based on machine learning according to claim 2, characterized in that: Using the high-temperature stress function as a time reference, the soil moisture, evapotranspiration, and vegetation index sequences were time-aligned, and the drought intensity and recovery principal components were extracted to construct a drought stress response matrix, as detailed below: Based on the obtained high temperature stress function Using time as a reference, time mapping is performed on time series from different sources, and time scale unification is achieved using a linear interpolation function; In high temperature stress function Extract the corresponding soil moisture during the period when the rate of change reaches a local extreme. , evapotranspiration With vegetation index The subsequence, defined as the drought response window by calculating the high temperature-soil moisture cross-correlation coefficient; Drought intensity and restorative principal components are extracted within the drought response window to construct a drought stress response matrix.
4. The method for predicting high-temperature loss of corn yield based on machine learning according to claim 3, characterized in that: The drought response window is defined by calculating the high temperature-soil moisture cross-correlation coefficient, as follows: The specific formula for calculating the high temperature-soil moisture cross-correlation coefficient is as follows: ,in , They represent , Time average, Indicates the time lag. This represents the correlation coefficient between high temperature and soil moisture. Will The time lag at which the maximum value is obtained is marked as the optimal lag, and the time window corresponding to the optimal lag is marked as the drought response window.
5. The method for predicting high-temperature loss of corn yield based on machine learning according to claim 3, characterized in that: The process of extracting drought intensity and restorative principal components within the drought response window to construct a drought stress response matrix is as follows: Calculation of normalized soil moisture deficit index ; Calculate the normalized evapotranspiration index ; Constructing the overall drought intensity: ,in These represent the normalized soil moisture deficit index. Normalized Evapotranspiration Index Weighting coefficients; vegetation index Perform principal component analysis within the drought response window to obtain the first principal component. The formula for representing vegetation restoration potential is as follows: ,in The eigenvector corresponding to the largest eigenvalue in principal component analysis (PCA) Let the mean vector of vegetation indices be used; the first principal component is... Marked as a restorative principal component, It is the transpose symbol; The drought intensity and recovery principal components were aligned by time to form a drought stress response matrix. : .
6. The method for predicting high-temperature loss of maize yield based on machine learning according to claim 3, characterized in that: Based on the high-temperature stress function and drought response matrix, the covariance and mutual information matrices are calculated, and nonlinear orthogonal decomposition is performed to generate the high-temperature-drought composite coupled feature tensor, as follows: High temperature stress function With drought stress response matrix Forming a feature matrix by time alignment ; Calculate the characteristic matrix covariance matrix ; Calculate the characteristic matrix Mutual information between each feature ; Obtain the mutual information matrix ; The covariance matrix With mutual information matrix Fusion, forming a weighted composite matrix : ,in For a weighted composite matrix, These are linear weighting coefficients; right Perform nonlinear orthogonal decomposition to solve for orthogonal eigenvectors. With eigenvalue matrix ; The eigenvectors after orthogonal decomposition Stacking them over time forms a three-dimensional tensor: ,in The high-temperature-drought composite coupled feature tensor is a three-dimensional matrix, consisting of time × feature × orthogonal components, where t is the time index. For feature index, For time t, the first The feature in the first Projected values on each orthogonal component.
7. The method for predicting high-temperature loss of maize yield based on machine learning according to claim 1, characterized in that: The composite stress representation vector and historical yield loss samples are input into a deep learning model. The composite stress confusion distortion index is used as a regularization term to constrain the loss function. The network parameters are optimized and the prediction bias is dynamically adjusted to generate a stable and robust high-temperature loss prediction model, as detailed below: The composite stress characterization vector Corresponding historical production loss samples Align the training samples according to time and space to form a training sample set. ,in Let be the composite stress representation vector of the q-th sample. Let q be the measured yield loss value for the q-th sample. The number of samples; The deep learning model is constructed from a multi-branch neural network structure, with the main branch capturing high-temperature features and auxiliary branches modeling drought and recovery features. ,in This represents the corn yield loss value predicted by the deep learning model. Indicates network parameters Nonlinear mapping model; Constructing a composite loss function with regularization terms : ,in For composite loss function, These are the regularization weight coefficients. For the compound stress confusion distortion constraint function, ; An adaptive learning rate optimization algorithm is used to iteratively update network parameters, train a high-temperature loss prediction model, and finally generate a stable and robust high-temperature loss prediction model.