A method for predicting pasture yield based on multi-modal feature fusion explainability
By using multispectral drones to collect images and combining them with multimodal feature fusion methods, the problem of large errors in traditional field measurement methods has been solved, enabling efficient and accurate prediction of forage yield, supporting precision agricultural management, and improving agricultural production efficiency and sustainability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN AGRI UNIV
- Filing Date
- 2026-06-23
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional methods for measuring ryegrass yield in the field rely on manual measurement, which is greatly affected by growth height, weather and terrain, resulting in large errors and poor timeliness. They are difficult to adapt to the needs of precision agriculture in large-scale planting areas. Furthermore, existing methods lack multi-dimensional data fusion, and the prediction results are inaccurate and difficult to interpret.
Multispectral images were acquired using the Mavic 3M multispectral UAV. Combined with multimodal feature fusion methods, vegetation index and texture feature extraction were used, and an XGBoost model was constructed using multimodal convolutional neural network and causal inference model to predict forage yield, thereby achieving the fusion of spectral information, trait structure features and climate information.
It enables efficient and accurate forecasting of forage yield, reduces labor input, improves timeliness, provides highly interpretable forecast results, adapts to large-scale planting scenarios, supports precision fertilization and management decisions, and promotes sustainable agricultural development.
Smart Images

Figure CN122434005A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ryegrass production technology, specifically a method for predicting forage yield based on the interpretability of multimodal feature fusion. Background Technology
[0002] Ryegrass is a grass belonging to the Poaceae family, widely adapted to temperate and subtropical cool regions. It holds a dominant position in global livestock farming, and in southern and central my country, it is one of the most widely planted and highest-value forage grasses during the winter and spring seasons. Ryegrass is a clump-forming grass with rapid growth, strong tillering ability, and a typical plant height of 80-120 cm. It has a dense root system, tender leaves, good regeneration ability, strong resistance to adverse conditions, simple cultivation and management, stable yield, and excellent palatability. Ryegrass has significant application value in green fodder, haymaking, grassland rotation, green manure utilization, ecological restoration, and landscaping.
[0003] To breed high-quality, high-yield ryegrass varieties (lines), breeders typically need to efficiently, promptly, and accurately monitor the growth status of ryegrass and have a thorough understanding of its yield information. However, traditional field yield measurement methods mainly rely on manual measurement, which is difficult due to the influence of forage height and is susceptible to subjective factors from random sampling, leading to significant errors in the results. Furthermore, these methods require substantial manpower, are easily affected by weather and terrain, and lack timeliness for large-scale planting areas, failing to meet the needs of modern agricultural development and making it difficult to develop targeted solutions based on measurement results in actual production.
[0004] To select high-quality, high-yield, and multiflora ryegrass varieties and to facilitate growers' timely monitoring of ryegrass growth, timely yield measurement is necessary. Traditional field yield measurement methods primarily rely on manual methods. However, given the large planting area of ryegrass, these methods are unsuitable. The measurement process is difficult due to the limited area of forage growth, and random sampling is susceptible to subjective factors, leading to significant errors in the results. Furthermore, these methods require substantial manpower and are easily affected by weather and terrain factors. For large-scale planting areas, they lack timeliness, and the data recording methods in traditional methods are entirely manual, which no longer meets the needs of modern agricultural development. Additionally, it is difficult to develop targeted solutions based on the measurement results during actual production.
[0005] A multispectral sensor is a detection device equipped with multiple discrete spectral band channels, typically covering key bands such as blue, green, red, red-edge, and near-infrared. Its core capability lies in overcoming the limitations of visible light sensors that rely solely on the red, green, and blue channels. By extending to near-infrared and red-edge bands, it enables deep inversion of vegetation physiological parameters. This technology significantly improves the accuracy of crop nitrogen content assessment, chlorophyll dynamic monitoring, and biomass estimation, and is widely used in precision agriculture for growth diagnosis, stress early warning, and yield prediction, especially suitable for remote sensing monitoring tasks at regional scales.
[0006] As unmanned aerial vehicles controlled by programs or remotely guided, drones are gradually becoming an important carrier of aerial remote sensing due to their advantages such as adjustable flight altitude and speed, low-altitude operational flexibility, takeoff and landing adaptability, and low operating costs. Compared with manned aircraft, drones have lower operational risks and higher safety in complex environments. Based on their functional scope, they can be divided into two main categories: military and civilian. Civilian drones have deeply penetrated fields such as farmland management, geological exploration, disaster emergency response, environmental assessment, and ecological protection. For example, they are used to perform crop phenotypic analysis by carrying multispectral cameras or to perform nighttime search and rescue missions by combining thermal infrared sensors. Currently, the integration of multispectral sensors and drone platforms is driving the rapid development of the "low-altitude remote sensing+" model. As sensors further evolve towards miniaturization, low power consumption, and multimodal fusion, and as drones' intelligent networking and AI computing capabilities are strengthened, the potential of this joint technology system in areas such as comprehensive ecological monitoring, smart city governance, and high-precision digital agriculture will continue to be released.
[0007] Vegetation indices quantify photosynthetic activity and physiological status of vegetation by utilizing spectral bands collected by multispectral sensors, such as a combination of near-infrared and red light. A typical example is the NDVI (Normalized Differential Vegetation Index), which, according to previous studies, can effectively assess leaf area index, chlorophyll content, and vegetation cover, and is currently widely used in crop growth monitoring, biomass estimation, and stress response analysis. Furthermore, indices such as EVI2 (Enhanced Vegetation Index) can reduce atmospheric and soil background interference and are suitable for high biomass areas; while the NDRE (Normalized Differential Red Edge Index) is sensitive to canopy nitrogen in the mid-to-late stages of crop growth and is often used for precision fertilization decisions.
[0008] Texture features describe the structural characteristics of vegetation canopy or leaf arrangement by analyzing the spatial grayscale variation patterns of pixels. Common indicators include mean, variance, contrast, homogeneity, entropy, second moment of angle, correlation, dissimilarity, and energy. These features can effectively distinguish different crop types, identify canopy closure, detect disease and pest patches, and help identify structural anomalies such as missing plants and lodging in the field. For example, high homogeneity usually indicates a uniform canopy surface, while high entropy reflects a complex spatial distribution of branches and leaves.
[0009] Currently, breeders face many challenges in timely understanding of ryegrass yield, particularly in accurately integrating data from diverse sources such as spectral data, textural features, structural traits, and environmental and management data. Most existing crop yield prediction methods rely on single-modal data or simple data fusion methods, lacking sufficient multi-dimensional information and interpretability. Especially when dealing with the high dimensionality and complexity of different data types, existing methods often fail to fully consider the relationships and mechanisms of interaction between different modalities, leading to inaccurate and difficult-to-interpret prediction results.
[0010] To address this issue, we propose a method for predicting forage yield based on the interpretability of multimodal feature fusion. Summary of the Invention
[0011] The purpose of this invention is to provide a method for predicting forage yield based on multimodal feature fusion interpretability, which solves the problems in the background art mentioned above.
[0012] To achieve the above objectives, the present invention provides the following technical solution: a method for predicting forage yield based on multimodal feature fusion interpretability, comprising the following steps: (1) The Mavic 3M multispectral UAV was used to plan the flight route and set the flight parameters, including a flight altitude of 12m, a heading overlap of 80%, and a lateral overlap of 75%. Visible light images and multispectral images of the study area in four bands (green, red, red edge, and near-infrared) were collected by the visible light sensor and multispectral sensor carried by the UAV. (2) Obtain field trait data, including selecting 12 individual plants in the plot, measuring plant height, chlorophyll content, leaf length, leaf width and number of tillers, and harvesting and yield measurement in accordance with national standard GB / T30395-2013 to calculate the actual yield; (3) The multispectral image is preprocessed using DJITerra software. The preprocessing includes radiometric correction, noise reduction, smoothing and sharpening to obtain a positive image. (4) Using ENVI software, ArcGISPro software and self-developed Python script, nine vegetation indices and nine texture features were extracted from the preprocessed image. The vegetation indices include Green Band Ratio Vegetation Index (GRVI), Normalized Difference Vegetation Index (NDVI), etc. The texture features were extracted based on the Gray Co-occurrence Matrix (GLCM) and included mean, variance, contrast, etc. (5) Principal component analysis (PCA) is used to reduce the dimensionality and remove redundancy of texture features. High-level features of different modalities are extracted by multimodal convolutional neural networks (CNN). The weights of each modality are dynamically adjusted by the self-attention mechanism of multimodal fusion to achieve feature fusion of spectral information, morphological structure features and climate information. (6) Construct causal inference and structural equation model (SEM) to clarify the causal relationship between spectral information, crop trait characteristics, climate data and yield. Calculate path coefficients through maximum likelihood estimation (MLE), and use CFI and RMSEA indicators to test and optimize the model fit to obtain causal prior information. (7) Input the fused features and causal prior information into the XGBoost model for training, and use the SHAP method to quantify the contribution of each feature to the prediction results, so as to achieve interpretable prediction of forage yield.
[0013] Preferably, the calculation formulas for the nine vegetation indices in step (4) are as follows: Green band ratio vegetation index (GRVI): GRVI = (NIR - G) / (NIR + G) Normalized Difference Vegetation Index (NDVI): NDVI = (NIR - R) / (NIR + R) Green normalized vegetation index (GNDVI): GNDVI = (NIR - G) / (NIR + G) Normalized Difference Red Edge Index (NDRE): NDRE = (NIR - RE) / (NIR + RE) Soil-modified vegetation index (SAVI): SAVI = (1 + L) × (NIR - R) / (NIR + R + L), where L is the soil modifier, with a value of 0.5. Optimized Soil Adjusted Vegetation Index (OSAVI): OSAVI = (NIR - R) / (NIR + R + 0.16) Improved chlorophyll uptake index (MCARI): MCARI = [(R700 - R670) - 0.2 × (R700 - R550)] × (R700 / R670) Converted Chlorophyll Uptake Index (TCARI): TCARI = 3 × [(R700 - R670) - 0.2 × (R700 - R550) × (R700 / R670)] Enhanced Vegetation Index (EVI2): EVI2 = (NIR - R) / (NIR + R + 2.4) Wherein, R represents the red light band, G the green light band, B the blue light band, NIR the near-infrared band, and RE the red edge band.
[0014] Preferably, the nine texture features mentioned in step (4) are calculated based on the gray-level co-occurrence matrix (GLCM), and the calculation formula is as follows: Mean: Variance: Homogeneity: Contrast: Dissimilarity: Energy: Angular second moment (ASM): Entropy Where i and j are the pixel values represented by the GLCM index, and p(i,j) are the element values of the i-th row and j-th column in the GLCM.
[0015] Preferably, the formula for the multimodal convolutional neural network in step (5) is: Where W is the weight matrix of the fully connected layer, and b is the bias term. f is the activation function. CNNA (X A ), f CNNB (X B The images show the convolutional feature extraction results for different modalities of data.
[0016] Preferably, the specific process of principal component analysis (PCA) in step (5) is as follows: the original texture feature data matrix X is standardized to obtain Z, and the covariance matrix is calculated. Eigenvalue decomposition is performed on S to obtain eigenvalues λk and eigenvectors vk. The first m eigenvectors are selected to form the principal component loading matrix Vm. The principal component scores are obtained by projecting Z onto this subspace. The value of m is determined based on the cumulative variance contribution rate Rm.
[0017] Preferably, the calculation formula for the self-attention mechanism of multimodal fusion in step (5) is as follows: Where Qi is the query matrix of the i-th modality, Ki is the key matrix of the i-th modality, Vi is the value matrix of the i-th modality, and dk is the dimension of the query and key.
[0018] Preferably, the structural equation model (SEM) in step (6) includes a measurement model and a structural model, and the formula for the measurement model is: The structural model formula is η=Bξ+Γζ+η, where y is the vector of observed variables, Λ_y is the factor loading matrix, η is the vector of latent variables, ε is the error term, B is the path coefficient matrix of latent independent variables to dependent variables, ξ is the vector of latent variables of independent variables, Γ is the path coefficient matrix of latent exogenous variables to dependent variables, and ζ is the vector of latent variables of exogenous variables.
[0019] Preferably, the formula for calculating CFI in step (6) is as follows: The formula for calculating RMSEA is: , where χ² model χ² represents the chi-square value of the fitted model, df_model represents the degrees of freedom of the fitted model, and χ² represents the chi-square value of the model. null For the chi-square value of the independent model, df null denoted by , where represents the degrees of freedom of the independent model, and N represents the sample size.
[0020] Preferably, the loss function of the XGBoost model in step (7) is: ,in The loss function Ω(f) is used to measure the error between the predicted and actual values. k ) is a regularization term that controls the complexity of the model.
[0021] Preferably, the formula for calculating the SHAP value in step (7) is: Where N is the total set of features, S is the subset of features that does not contain feature x_i, f(S) is the model prediction on the feature subset S, and f(S∪{i}) is the prediction value on the feature subset S and feature x_i. i The model prediction values are given by |S|, where |S| is the number of features in set S, and n is the total number of features.
[0022] This invention provides a method for predicting forage yield based on multimodal feature fusion and interpretability. This method for predicting forage yield based on multimodal feature fusion and interpretability has the following advantages:
[0023] 1. Improved efficiency: Low-altitude remote sensing by drones replaces manual measurement, covering large planting areas, significantly reducing manpower input and yield measurement time, and greatly enhancing timeliness; 2. Predictive accuracy: By fusing multimodal data such as spectral, texture, trait, and climate, and combining causal inference and self-attention fusion mechanisms, the bias of a single data source is reduced, and the prediction accuracy is better than that of traditional methods. 3. High interpretability: SHAP analysis quantifies the contribution of each feature, and SEM model clarifies the causal relationship between variables, making the prediction results traceable and assisting agricultural decision-making; 4. Wide adaptability: It is not limited by terrain, weather or other factors, and is suitable for large-scale forage planting scenarios, providing a scientific basis for precision fertilization, irrigation and harvesting, and promoting sustainable agricultural development. Attached Figure Description
[0024] Figure 1 This is a model prediction structure diagram of a method for predicting forage yield based on multimodal feature fusion interpretability according to the present invention; Figure 2 This is a scatter plot of the contribution of SHAP to the interpretability prediction of forage yield based on multimodal feature fusion, a method of the present invention. Detailed Implementation
[0025] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described with reference to the accompanying drawings.
[0026] A preferred embodiment of the method for predicting forage yield based on multimodal feature fusion interpretability provided by the present invention is as follows: Figures 1 to 2 As shown: A method for predicting forage yield based on multimodal feature fusion interpretability includes the following steps: (1) The Mavic 3M multispectral UAV was used to plan the flight route and set the flight parameters, including a flight altitude of 12m, a heading overlap of 80%, and a lateral overlap of 75%. Visible light images and multispectral images of the study area in four bands (green, red, red edge, and near-infrared) were collected by the visible light sensor and multispectral sensor carried by the UAV. (2) Obtain field trait data, including selecting 12 individual plants in the plot, measuring plant height, chlorophyll content, leaf length, leaf width and number of tillers, and harvesting and yield measurement in accordance with national standard GB / T30395-2013 to calculate the actual yield; (3) The multispectral image is preprocessed using DJITerra software. The preprocessing includes radiometric correction, noise reduction, smoothing and sharpening to obtain a positive image. (4) Using ENVI software, ArcGISPro software and self-developed Python script, nine vegetation indices and nine texture features were extracted from the preprocessed image. The vegetation indices include Green Band Ratio Vegetation Index (GRVI), Normalized Difference Vegetation Index (NDVI), etc. The texture features were extracted based on the Gray Co-occurrence Matrix (GLCM) and included mean, variance, contrast, etc. Here are the formulas for calculating the nine vegetation indices: Green band ratio vegetation index (GRVI): GRVI = (NIR - G) / (NIR + G) Normalized Difference Vegetation Index (NDVI): NDVI = (NIR - R) / (NIR + R) Green normalized vegetation index (GNDVI): GNDVI = (NIR - G) / (NIR + G) Normalized Difference Red Edge Index (NDRE): NDRE = (NIR - RE) / (NIR + RE) Soil-modified vegetation index (SAVI): SAVI = (1 + L) × (NIR - R) / (NIR + R + L), where L is the soil modifier, with a value of 0.5. Optimized Soil Adjusted Vegetation Index (OSAVI): OSAVI = (NIR - R) / (NIR + R + 0.16) Improved chlorophyll uptake index (MCARI): MCARI = [(R700 - R670) - 0.2 × (R700 - R550)] × (R700 / R670) Converted Chlorophyll Uptake Index (TCARI): TCARI = 3 × [(R700 - R670) - 0.2 × (R700 - R550) × (R700 / R670)] Enhanced Vegetation Index (EVI2): EVI2 = (NIR - R) / (NIR + R + 2.4) Wherein, R represents the red light band, G the green light band, B the blue light band, NIR the near-infrared band, and RE the red edge band; Furthermore, the nine texture features are calculated based on the Gray-Level Co-occurrence Matrix (GLCM), using the following formula: Mean: Variance: Homogeneity: Contrast: Dissimilarity: Energy: Angular second moment (ASM): Entropy Where i and j are the pixel values represented by the GLCM index, and p(i,j) is the element value of the i-th row and j-th column in the GLCM; (5) Principal component analysis (PCA) is used to reduce the dimensionality and remove redundancy of texture features. High-level features of different modalities are extracted by multimodal convolutional neural networks (CNN). The weights of each modality are dynamically adjusted by the self-attention mechanism of multimodal fusion to achieve feature fusion of spectral information, morphological structure features and climate information. Here, the formula for a multimodal convolutional neural network is: Where W is the weight matrix of the fully connected layer, and b is the bias term. fCNNA(XA) and fCNNB(XB) are the activation functions, respectively, and the convolutional feature extraction results of data from different modalities. Furthermore, the specific process of Principal Component Analysis (PCA) is as follows: The original texture feature data matrix X is standardized to obtain Z, and the covariance matrix is calculated. Eigenvalue decomposition is performed on S to obtain eigenvalues λk and eigenvectors vk. The first m eigenvectors are selected to form the principal component loading matrix Vm. The principal component scores are obtained by projecting Z onto this subspace. The value of m is determined based on the cumulative variance contribution rate Rm; Furthermore, the calculation formula for the self-attention mechanism of multimodal fusion is as follows: Where Qi is the query matrix of the i-th modality, Ki is the key matrix of the i-th modality, Vi is the value matrix of the i-th modality, and dk is the dimension of the query and key; (6) Construct causal inference and structural equation model (SEM) to clarify the causal relationship between spectral information, crop trait characteristics, climate data and yield. Calculate path coefficients through maximum likelihood estimation (MLE), and use CFI and RMSEA indicators to test and optimize the model fit to obtain causal prior information. Here, structural equation modeling (SEM) includes a measurement model and a structural model. The measurement model formula is as follows: The structural model formula is η=Bξ+Γζ+η, where y is the observed variable vector, Λ_y is the factor loading matrix, η is the latent variable vector, ε is the error term, B is the path coefficient matrix of the latent independent variable to the dependent variable, ξ is the latent variable vector of the independent variable, Γ is the path coefficient matrix of the latent exogenous variable to the dependent variable, and ζ is the latent variable vector of the exogenous variable. Furthermore, the formula for calculating CFI is as follows: The formula for calculating RMSEA is: , where χ²model is the chi-square value of the fitted model, df_model is the degree of freedom of the fitted model, χ²null is the chi-square value of the independent model, dfnull is the degree of freedom of the independent model, and N is the sample size; (7) Input the fused features and causal prior information into the XGBoost model for training, and use the SHAP method to quantify the contribution of each feature to the prediction results, so as to achieve interpretable prediction of forage yield. Here, the loss function of the XGBoost model is ,in The loss function is used to measure the error between the predicted and the true values, and Ω(fk) is a regularization term that controls the complexity of the model. Furthermore, the formula for calculating the SHAP value is as follows: , where N is the total set of features, S is the feature subset that does not contain feature x_i, f(S) is the model prediction value on feature subset S, f(S∪{i}) is the model prediction value on feature subset S and feature xi, |S| is the number of features in set S, and n is the total number of features.
[0027] This invention combines a drone's multispectral sensor with a high-precision yield prediction model to achieve rapid, real-time monitoring of ryegrass growth, effectively overcoming the lag and inefficiency of traditional yield measurement methods. The system can acquire crop spectral and growth information over large areas in a timely manner and rapidly convert it into yield prediction results, providing agricultural managers with accurate and reliable decision-making support. The real-time monitoring mechanism not only supports refined management of different regions and growth stages, making fertilization, irrigation, and harvesting more scientific and personalized, but also significantly improves production efficiency and avoids the blind decision-making common in traditional methods. Simultaneously, this invention provides key technological support for the development of precision agriculture. With automated monitoring and intelligent analysis, agricultural managers can continuously monitor crop growth dynamics without relying on extensive manual labor, thereby optimizing production scheduling, improving resource utilization efficiency, and reducing fertilizer and pesticide inputs, promoting agricultural production towards high efficiency, intelligence, and sustainability. The real-time monitoring and intelligent decision-making system constructed by this invention helps to comprehensively improve the scientific nature of agricultural management and is an important component of modern precision agriculture.
[0028] The above description is merely an illustrative embodiment of the present invention and is not intended to limit the scope of the invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention should fall within the scope of protection of the present invention. Furthermore, it should be noted that the components of the present invention are not limited to the overall application described above. Each technical feature described in the specification can be used individually or in combination as needed. Therefore, the present invention naturally covers other combinations and specific applications related to the inventive points of this case.
Claims
1. A method for predicting forage yield based on multimodal feature fusion interpretability, characterized in that, Includes the following steps: (1) The Mavic 3M multispectral UAV was used to plan the flight route and set the flight parameters, including a flight altitude of 12m, a heading overlap of 80%, and a lateral overlap of 75%. Visible light images and multispectral images of the study area in four bands (green, red, red edge, and near-infrared) were collected by the visible light sensor and multispectral sensor carried by the UAV. (2) Obtain field trait data, including selecting 12 individual plants in the plot, measuring plant height, chlorophyll content, leaf length, leaf width and number of tillers, and harvesting and yield measurement in accordance with national standard GB / T30395-2013 to calculate the actual yield; (3) The multispectral image is preprocessed using DJITerra software. The preprocessing includes radiometric correction, noise reduction, smoothing and sharpening to obtain a positive image. (4) Using ENVI software, ArcGISPro software and self-developed Python script, nine vegetation indices and nine texture features were extracted from the preprocessed image. The vegetation indices include Green Band Ratio Vegetation Index (GRVI), Normalized Difference Vegetation Index (NDVI), etc. The texture features were extracted based on the Gray Co-occurrence Matrix (GLCM) and included mean, variance, contrast, etc. (5) Principal component analysis (PCA) is used to reduce the dimensionality and remove redundancy of texture features. High-level features of different modalities are extracted by multimodal convolutional neural networks (CNN). The weights of each modality are dynamically adjusted by the self-attention mechanism of multimodal fusion to achieve feature fusion of spectral information, morphological structure features and climate information. (6) Construct causal inference and structural equation model (SEM) to clarify the causal relationship between spectral information, crop trait characteristics, climate data and yield. Calculate path coefficients through maximum likelihood estimation (MLE), and use CFI and RMSEA indicators to test and optimize the model fit to obtain causal prior information. (7) Input the fused features and causal prior information into the XGBoost model for training, and use the SHAP method to quantify the contribution of each feature to the prediction results, so as to achieve interpretable prediction of forage yield.
2. The method according to claim 1, characterized in that, The calculation formulas for the nine vegetation indices mentioned in step (4) are as follows: Green band ratio vegetation index (GRVI): GRVI = (NIR - G) / (NIR + G) Normalized Difference Vegetation Index (NDVI): NDVI = (NIR - R) / (NIR + R) Green normalized vegetation index (GNDVI): GNDVI = (NIR - G) / (NIR + G) Normalized Difference Red Edge Index (NDRE): NDRE = (NIR - RE) / (NIR + RE) Soil-modified vegetation index (SAVI): SAVI = (1 + L) × (NIR - R) / (NIR + R + L), where L is the soil modifier, with a value of 0.
5. Optimized Soil Adjusted Vegetation Index (OSAVI): OSAVI = (NIR - R) / (NIR + R + 0.16) Improved chlorophyll uptake index (MCARI): MCARI = [(R700 - R670) - 0.2 × (R700 - R550)] × (R700 / R670) Converted Chlorophyll Uptake Index (TCARI): TCARI = 3 × [(R700 - R670) - 0.2 × (R700 - R550) × (R700 / R670)] Enhanced Vegetation Index (EVI2): EVI2 = (NIR - R) / (NIR + R + 2.4) Wherein, R represents the red light band, G the green light band, B the blue light band, NIR the near-infrared band, and RE the red edge band.
3. The method according to claim 1, characterized in that, The nine texture features mentioned in step (4) are calculated based on the Gray-Level Co-occurrence Matrix (GLCM), and the calculation formula is as follows: Mean: Variance: Homogeneity: Contrast: Dissimilarity: Energy: Angular second moment (ASM): Entropy Where i and j are the pixel values represented by the GLCM index, and p(i,j) are the element values of the i-th row and j-th column in the GLCM.
4. The method according to claim 1, characterized in that, The formula for the multimodal convolutional neural network in step (5) is: Where W is the weight matrix of the fully connected layer, and b is the bias term. f is the activation function. CNNA (X A ), f CNNB (X B The images show the convolutional feature extraction results for different modalities of data.
5. The method according to claim 1, characterized in that, The specific process of principal component analysis (PCA) in step (5) is as follows: the original texture feature data matrix X is standardized to obtain Z, and the covariance matrix is calculated. Eigenvalue decomposition is performed on S to obtain eigenvalues λk and eigenvectors vk. The first m eigenvectors are selected to form the principal component loading matrix Vm. The principal component scores are obtained by projecting Z onto this subspace. The value of m is determined based on the cumulative variance contribution rate Rm.
6. The method according to claim 1, characterized in that, The formula for calculating the self-attention mechanism of multimodal fusion in step (5) is as follows: Where Qi is the query matrix of the i-th modality, Ki is the key matrix of the i-th modality, Vi is the value matrix of the i-th modality, and dk is the dimension of the query and key.
7. The method according to claim 1, characterized in that, In step (6), the structural equation model (SEM) includes a measurement model and a structural model. The formula for the measurement model is: The structural model formula is η=Bξ+Γζ+η, where y is the vector of observed variables, Λ_y is the factor loading matrix, η is the vector of latent variables, ε is the error term, B is the path coefficient matrix of latent independent variables to dependent variables, ξ is the vector of latent variables of independent variables, Γ is the path coefficient matrix of latent exogenous variables to dependent variables, and ζ is the vector of latent variables of exogenous variables.
8. The method according to claim 1, characterized in that, The formula for calculating CFI in step (6) is as follows: The formula for calculating RMSEA is: , where χ² model χ² represents the chi-square value of the fitted model, df_model represents the degrees of freedom of the fitted model, and χ² represents the chi-square value of the model. null For the chi-square value of the independent model, df null denoted by , where represents the degrees of freedom of the independent model, and N represents the sample size.
9. The method according to claim 1, characterized in that, The loss function of the XGBoost model in step (7) is: ,in The loss function Ω(f) is used to measure the error between the predicted and actual values. k ) is a regularization term that controls the complexity of the model.
10. The method according to claim 1, characterized in that, The formula for calculating the SHAP value in step (7) is: Where N is the total set of features, S is the subset of features that does not contain feature x_i, f(S) is the model prediction on the feature subset S, and f(S∪{i}) is the prediction value on the feature subset S and feature x_i. i The model prediction values are given by |S|, where |S| is the number of features in set S, and n is the total number of features.