Construction and application of wheat biomass estimation model fusing phenological information and hyperspectral data

By fusion of phenological information and hyperspectral data, using the ATDNN model of Boruta-SHAP and attention mechanism, the efficiency and accuracy of biomass estimation in hyperspectral remote sensing are solved, and robust estimation of wheat biomass and modern agricultural decision support are achieved.

CN120493679APending Publication Date: 2025-08-15HENAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510368404.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture crop growth and development dynamics in hyperspectral remote sensing, and traditional biomass estimation methods are labor-intensive and time-efficient, making it difficult to meet the real-time decision-making needs of modern agriculture.

Method used

To integrate phenological information and hyperspectral data, the spectral characteristics were selected using the Boruta-SHAP algorithm, and combined with the attention mechanism's ATDNN model, dynamic weighting of the feature importance, and a wheat biomass estimation model was established.

Benefits of technology

Improves prediction accuracy and robustness of biomass estimation, reduces training time, and provides scalable agricultural decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493679A_ABST
    Figure CN120493679A_ABST
Patent Text Reader

Abstract

The invention relates to construction and application of a wheat biomass estimation model fusing phenological information and hyperspectral data, and aims to solve the technical problem that wheat biomass estimation prediction precision and interpretability are difficult to balance. According to the method, Boruta-SHAP feature selection is combined with an attention-based neural network to estimate the wheat biomass, and the wheat biomass comprises phenological indexes (GDD, GD and ZS) with physiological significance and spectral features which are optimally selected by using a hierarchical attention mechanism. The model can effectively capture the crop development time and spectral dynamics. SHAP analysis shows that the contribution of phenological information to the model performance is 31.5%, key spectral characteristics comprise NIR (22.9%) and SWIT1 (12.3%), and the estimation precision is further improved. The robust and explainable framework model has a potential wide application prospect in the aspects of wheat yield prediction, fertilization optimization and irrigation management by providing an extensible biomass estimation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crop biomass estimation, and in particular to the construction and application of a wheat biomass estimation model integrating phenological information and hyperspectral data. Background Art

[0002] Accurately estimating crop biomass is a key challenge facing modern agricultural science, with important implications for food security, precision agriculture, and sustainable resource management. Reliable biomass estimation is crucial for guiding agricultural production decisions, such as fertilization, irrigation scheduling, and yield forecasting. However, while relatively accurate, traditional biomass estimation methods are inherently labor-intensive, time-inefficient, and lack scalability, making them increasingly inadequate for modern agricultural monitoring and real-time decision-making.

[0003] Research has shown that the reflectance spectrum of plants (350-2500nm) contains rich information about vegetation characteristics. Based on this, hyperspectral remote sensing technology is increasingly being used in agriculture, revolutionizing crop monitoring. Hyperspectral remote sensing, which enables repeated and non-destructive monitoring and measurement of crop growth and development dynamics, holds great promise. However, the high dimensionality of hyperspectral data also presents significant challenges, including computational inefficiency, model instability, and multicollinearity. While hyperspectral remote sensing technology offers unique advantages, traditional dimensionality reduction methods often struggle to capture and exploit the complex, nonlinear relationships inherent in spectral data.

[0004] At present, the deep integration of machine learning technology with specific knowledge in the field of agricultural remote sensing to achieve the prediction of relevant biological indicators has also made great progress. However, it still faces some difficult problems, such as the mechanistic explanation of feature importance, the integration of temporal spectral dynamics, and the management of environmental heterogeneity. These problems are particularly evident when it comes to the processing of long-term data sets spanning multiple growing seasons and different management practices. In addition, the relationship between spectral reflectance and its biomass is inherently non-stationary and evolves in time and space. To address this complexity, complex modeling methods are needed to balance its interpretability and prediction accuracy.

[0005] The information disclosed in this background technology section is only used to deepen the understanding of the background technology of the present disclosure and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0006] The inventors have discovered that phenological indices with key temporal context, such as growing days (GD), accumulated temperature (GDD), and growth period indices (e.g., Zadoks and BBCH), can provide a standardized framework for wheat development modeling. Integrating these indices can more accurately simulate the relationship between spectral data and biomass accumulation at different growth stages. Based on this, the inventors employed the Boruta-SHAP algorithm to select spectral features to identify important spectral features under different temporal and environmental conditions. They also developed an integrated temporal-spectral modeling approach that combines phenological indices (GDD, GD, ZS) with the selected spectral features using an attention mechanism. This dynamic architecture assigns different weights to features based on crop growth stage, capturing the non-stationary relationship between spectral reflectance and biomass accumulation. The framework was validated using a long-term dataset (spanning ten years) of wheat varieties under different nitrogen fertilizer and irrigation treatments, demonstrating its robustness across diverse environmental conditions and timescales.

[0007] According to one aspect of the present disclosure, a method for constructing a wheat biomass estimation model that integrates phenological information and hyperspectral data is provided, comprising the following steps: (1) Select phenological indicators that can characterize different wheat growth and development stages: growing days (GD), accumulated temperature (GDD), and growth period index (ZS); (2) Identify and screen the spectral features related to wheat biomass at the corresponding growth and development stages based on the Boruta-SHAP algorithm; (3) An ATDNN model was established to combine the phenological indicators GDD, GD, and ZS with the selected spectral features using an attention mechanism and dynamically weight the feature importance to capture the non-stationary relationship between spectral reflectance and biomass accumulation; (4) The established ATDNN model is trained iteratively based on the historical spectral feature dataset and the corresponding phenological indicator dataset until convergence.

[0008] In some embodiments of the present disclosure, the ATDNN model includes an input layer, a first attention module, a feature extractor, a feature processor, and a regressor. The input layer is used to implement the input of the spectral features and phenological indicators; the first attention module is arranged before the feature extractor, and adopts a four-head attention mechanism to dynamically weight the importance of the input features; the feature extractor performs adaptive dimensionality conversion while retaining physiological related information through batch normalization and ReLU activation function; the feature processor is enhanced by the corresponding second attention module, and promotes the integration of processed features through nonlinear transformation; the regressor combines the processed features with linear transformation and ReLU activation to generate biomass prediction.

[0009] In some embodiments of the present disclosure, the four-head attention mechanism is expressed as follows: ; in, Q, K, V represent the query matrix, key matrix and value matrix respectively, d k Indicates the dimension of the key vector.

[0010] In some embodiments of the present disclosure, in step (4), the model training uses the mean square error (MSE) loss function to quantify the prediction accuracy, and the learning rate of the Adam optimizer is initialized to 5×10 -4 ; The training process uses mixed precision calculation to optimize numerical stability and computational efficiency while maintaining prediction accuracy.

[0011] In some embodiments of the present disclosure, in the ATDNN model, a 0.3% dropout regularization is used, and a predefined threshold is used. τ Gradient norm clipping (with a value in the range of 0.5 to 1.0) is used to maintain optimization stability, and an early stopping mechanism with a patience window of 30 epochs is used.

[0012] In some embodiments of the present disclosure, all phenological indicators are standardized using the z-score normalization method to facilitate integration with hyperspectral data in subsequent modeling.

[0013] According to another aspect of the present disclosure, a method for estimating wheat biomass is provided, comprising the following steps: (1) Collect and obtain wheat hyperspectral data or / phenological index data of the area to be tested during the corresponding period; (2) The obtained hyperspectral data were subjected to wavelength resampling, splicing correction, and noise removal; the obtained phenological index data were subjected to standardization and Z-score conversion; (3) Inputting the hyperspectral data and / or phenological index data obtained by preprocessing in step (2) into the wheat biomass estimation model of claim 1 to generate and output a corresponding biomass prediction value.

[0014] In some embodiments of the present disclosure, in the step (1), wheat hyperspectral data within the wavelength range of 350-2500 nm is collected; the phenological index data includes growing days GD, growth accumulated temperature GDD and growth period index ZS.

[0015] In some embodiments of the present disclosure, in step (2), the obtained hyperspectral data is first subjected to stitching correction at wavelengths of 1000 nm and 1800 nm to eliminate detector knot artifacts, and then the spectral regions affected by atmospheric water absorption and low signal-to-noise ratio are systematically removed. Then, a smoothing algorithm is applied to eliminate or reduce random noise while retaining basic spectral features.

[0016] According to yet another aspect of the present disclosure, the wheat biomass estimation model is applied to wheat yield prediction, fertilization optimization and / or irrigation management.

[0017] One or more technical solutions provided in the embodiments of this application have at least any of the following technical effects or advantages: 1. A novel wheat biomass estimation model was proposed by integrating spectral feature selection, phenological indicators, and attention-based deep learning. Comprehensive validation using a dataset spanning ten years demonstrated significant improvements in prediction accuracy. Specifically, Boruta-SHAP achieved a 25.6% improvement in R² compared to XGB-FS, while the inclusion of phenological information further boosted model performance by 16.7%, to an R² of 0.847. The hierarchical attention architecture successfully balanced computational efficiency and predictive power, reducing training time by 49.3%.

[0018] SHAP analysis revealed that phenological information contributed 31.5% to model performance, with key spectral features including NIR (22.9%) and SWIR1 (12.3%) further improving estimation accuracy. This robust and interpretable framework, by providing a scalable biomass estimation method, has potential applications in wheat yield prediction, fertilization optimization, and irrigation management. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the experimental arrangement in one embodiment of the present invention, showing the layout and geographical location of the Xiaotangshan experimental area in Changping District, Beijing, China.

[0020] Figure 2 This is a flowchart for constructing a wheat biomass estimation model in one embodiment of the present invention, which clearly shows the complete process from data acquisition to model construction and evaluation.

[0021] Figure 3 This is a schematic diagram of the hierarchical attention neural network architecture in one embodiment of the present invention, showing the structural composition of the ATDNN model, including the input layer, the first attention module, the feature extractor, the feature processor and the regressor.

[0022] Figure 4 This is a schematic diagram of the workflow of the multi-head attention mechanism in one embodiment of the present invention, showing how the four-head attention mechanism processes input features and assigns dynamic weights.

[0023] Figure 5 This is a schematic diagram comparing the model training convergence in one embodiment of the present invention, which compares the convergence speed and performance differences between the ATDNN model and the ordinary DNN model during the training process.

[0024] Figure 6 This is a schematic diagram of the spatiotemporal feature contribution analysis in one embodiment of the present invention, which visualizes the contribution and importance of different types of features to model performance through SHAP values. DETAILED DESCRIPTION

[0025] In order to better understand the technical solution of the present application, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0026] Example 1: Construction of a wheat biomass estimation model integrating phenological information and hyperspectral data (1) Data acquisition and processing 1. Study area and experimental design This study was conducted at the Xiaotangshan Experimental Base in Changping District, Beijing, China (40.17°N, 116.42°E) over a 10-year period (2012-2022). The study site experiences a temperate monsoon climate with distinct seasonal variations, with hot and humid summers and cold and dry winters. The annual mean temperature is approximately 12°C, with seasonal extremes ranging from -10°C in winter to 30°C in summer. The monthly mean temperature ranges from -4°C in January to 26°C in July. Figure 1 The average annual precipitation in the experimental area is approximately 500 mm, with approximately 70% of the rainfall occurring in the summer (June to August). The experimental site is located at an altitude of approximately 40 meters and has a relatively flat terrain (slope <2%). The predominant soil type is silt loam, which provides optimal conditions for wheat growth due to its balanced water retention and nutrient availability.

[0027] Field experiments were systematically conducted over 10 consecutive growing seasons (2012-2022). Treatments varied nitrogen application rates (0-440 kg / ha), irrigation rates (0-384 mm), and wheat varieties (Table 1). This long-term experimental framework facilitated the collection of an extensive dataset integrating hyperspectral measurements (350-2500 nm), phenological information, and ground-truth biomass measurements across multiple growing seasons and management scenarios.

[0028] Table 1. Summary of experimental design and conditions from 2012 to 2022 .

[0029] Hyperspectral data acquisition During the pilot study period (2012-2022), hyperspectral reflectance measurement data were collected using an ASD FieldSpec High-Res spectrometer (ASD Spectral Acquisition Device) system. This instrument recorded spectral data from 350 to 2500 nm with a spectral resolution of 3 nm in the visible / near-infrared region (350-1000 nm) and 8 nm in the shortwave infrared region (1000-2500 nm). To ensure accuracy, system calibration intervals were performed using a Spectral® white reference panel (99% reflectance) to correct for variations in solar illumination.

[0030] Spectral data were collected under clear, cloudless conditions between 10:00 AM and 2:00 PM local time to minimize atmospheric influences and variations in sun angle. The fiber-optic probe was mounted on a tripod at a fixed height with a 25° field of view, corresponding to a ground sampling area of approximately 0.44 m². Within each experimental plot, spectral measurements were taken at four locations to account for spatial variability, with five replicate scans performed at each location. These rigorous operating procedures ensured high-quality hyperspectral data that accurately reflected canopy-level characteristics and minimized potential sources of error during data collection.

[0031] Biomass data collection (1) Biomass sampling plan Aboveground biomass sampling was performed immediately after hyperspectral data collection on the same day, ensuring temporal synchronization between spectral measurements and biomass data. This approach minimized potential physiological variations in plant state and the systematic errors that could be introduced into the relationship between spectral reflectance and biomass parameters. Sampling was performed between 2:00 PM and 4:00 PM local time to maintain consistent diurnal conditions across all measurement campaigns.

[0032] Sampling was carried out using a standardized destructive harvesting method. A 0.5 m² sampling area was demarcated within each experimental plot using a rigid quadrat frame and its location determined by stratified random sampling. Each plot was divided into a 3 × 2 grid, with cells randomly selected to ensure spatial representation while avoiding duplicate sampling of previously sampled areas. The sampling frame was carefully positioned to include three groups of intact wheat plants. All plants within the sampling frame were cut at ground level (approximately 1 cm from the soil surface) using sharp, clean shears to ensure precise cutting and complete biomass collection. Fresh samples were immediately processed to minimize water loss, placed in pre-labeled moisture-proof bags, and weighed on-site using a calibrated electronic balance (XS205, Mettler, Toledo; resolution: 0.01 g). A subset of each sample (~200 g) was retained for moisture content determination, while the remaining material was kept for further analysis.

[0033] The samples were dried in a forced air drying oven (DHG-9070A, Shanghai Kinoh) at 85 ± 1 °C for at least 48 h, with intermediate weighing every 12 h until a constant mass was achieved (defined as <0.1% variation between consecutive measurements). After drying, the samples were cooled to room temperature in a desiccator to prevent moisture resorption, and the final dry weight was recorded.

[0034] (2) Spatiotemporal patterns and growth stage-specific dynamics of wheat AGB from 2012 to 2021 The temporal dynamics of wheat aboveground biomass (AGB) across the growing seasons from 2012 to 2021 showed a distinct and consistent pattern. In the early growth stage (ZS ≈ 31), AGB values ranged from 1500 to 4000 kg / ha, gradually increased in the middle growth stage, and reached over 15000 kg / ha in the late growth stage (ZS ≈ 80).

[0035] The observed consistency of AGB dynamics across multiple seasons highlights the strong influence of heat accumulation (GDD) and growing date (GD) on biomass development. Furthermore, the predictable relationship between phenological progression and biomass accumulation emphasizes the value of incorporating phenological metrics into remote sensing models for biomass estimation.

[0036] Calculation of phenological metrics The developmental process of wheat was systematically characterized using three phenological indices: growing days (GD), accumulated temperature (GDD), and the Zadoks scale (ZS). This multi-metric approach comprehensively quantifies the temporal, thermal, and morphological aspects of crop development throughout the growing season.

[0037] Growing Days (GD) represent the cumulative calendar days from sowing to each sampling event, providing a straightforward timeframe for assessing developmental progress. While simple to calculate, this metric serves as a basic baseline for comparing growth rates across growing seasons and management treatments.

[0038] The accumulated temperature (GDD) quantifies the temporal accumulation of heat, capturing the biologically effective temperature during crop development. Temperature data were obtained from the ERA5-Land reanalysis dataset (European Centre for Medium-Range Weather Forecasts) using Google Earth Engine. GDD values were calculated using the sine wave method: (1); Where n represents the number of days from sowing to sampling, T (max,i) and T (min,i) represent the daily maximum and minimum temperatures (°C), T baseis the crop-specific base temperature, which is 0℃ for winter wheat; a maximum function is used to ensure non-negative daily heat accumulation, avoiding the occurrence of T base negative contribution to GDD.

[0039] The Zadoks Scale (ZS) provides detailed morphological characteristics of phenological stages using a standardized decimal coding system, ranging from 00 (dry seeds) to 99 (secondary dormancy). Field observations by trained researchers assessed key morphological indicators of multiple plants within each plot to ensure a representative assessment of major developmental stages. Particular emphasis was placed on major growth stages, including tillering, stem elongation, heading, and grain filling, with careful attention paid to uniformity within the sampling area.

[0040] Table 2 Summary of field experiments, sampling dates, and dataset distribution (2012-2022) .

[0041] All phenological indicators were normalized using the z-score normalization method to facilitate integration with hyperspectral data in subsequent modeling. The normalization procedure is expressed as: (2); in, x is the original measurement value, μ is the overall mean, σ This standardization ensures computational stability during neural network training while preserving the relative relationships between developmental stages.

[0042] Data preprocessing and partitioning strategies (1) Data preprocessing and partitioning strategy: Hyperspectral measurement data were initially processed using ViewSpec Pro software (Analytical Spectral Instruments, Boulder, CO, USA) with a wavelength resampling interval of 1 nm. The main quality control included averaging multiple scans of each plot to obtain representative spectra and excluding obvious instrument artifacts. Subsequent data preprocessing was implemented using the "hsdar" package in R (version 4.2.0). First, a stitching correction was performed at wavelengths of 1000 nm and 1800 nm to eliminate detector artifacts; spectral regions affected by atmospheric water absorption and low signal-to-noise ratio (1355-1410 nm, 1820-1942 nm, and 2400-2500 nm) were systematically removed. The Savitzky-Golay smoothing algorithm (window size: 11 points, polynomial order: 2) was then applied to reduce random noise while preserving basic spectral features.

[0043] (2) Dataset partitioning strategy: The dataset spanning ten years (2012-2022) consists of 1,336 samples. The dataset partitioning strategy (Table 2) is as follows: The training set (n = 854) includes samples from 2012–2020, representing 80% of the data for that period. The validation set (n = 214) consists of the remaining 20% of samples from 2012–2020, selected through stratified sampling to ensure independence from the training set. The test set (n = 268) includes all samples from 2021–2022, providing an independent evaluation dataset representative of operational forecasting scenarios. This partitioning strategy maintains temporal consistency and ensures robust evaluation under diverse environmental and management conditions.

[0044] (2) Model construction 1. Feature Selection Framework: Due to the inherent high dimensionality and multicollinearity of hyperspectral data, selecting optimal spectral features is a key challenge. To address this issue, this example designs and compares two spectral feature selection methods: Boruta-SHAP and XGBoost-based Feature Selection (XGB-FS).

[0045] The Boruta-SHAP algorithm integrates Boruta's iterative feature selection with Shapley's additive explanations (SHAP) value to achieve robust feature importance estimation in different machine learning architectures. This example uses BorutaShap (version 1.0.17) and XGBoost (version 2.1.3) as the base learner, with optimized hyperparameters (learning rate = 0.1, maximum depth = 6, number of trees = 100).

[0046] For each original feature in the dataset X i , the algorithm generates the corresponding shadow features by value transformation S i , establish a baseline for significance evaluation. The importance of features is quantified using Shap’s tree interpreter: (2); in, N represents the sample count, Represents the SHAP value of the feature in the sample.

[0047] Features are classified as “important” when their importance score exceeds the maximum importance of the shadow features: (3).

[0048] For tentative features, the significance test is two-sided. tTest, the importance score is compared with the shadow feature threshold.,The algorithm iteratively refines the classification until,convergence or an iteration limit is reached.

[0049] The implementation of XGB-FS follows a systematic optimization framework with three key components. First, the method uses GPU-accelerated XGBoost for initial feature importance evaluation, configured with optimized hyperparameters (n_estimators=100, learning_rate=0.1, max_depth=6). Second, an iterative evaluation strategy systematically evaluates feature subsets via five-fold cross-validation, examining different subset sizes from 5 to a predefined maximum. Finally, the optimal feature subset is determined via an elbow point detection mechanism, analyzing the relationship between feature count and model performance. This optimization process incorporates a 5% mean squared error (MSE) tolerance threshold to balance model complexity and performance, ensuring the selection of a parsimonious yet effective feature set.

[0050] Hierarchical Attention Neural Network This example implements a two-layer attention mechanism based on the attention-based deep neural network (ATDNN) ( Figure 2 ) to process spectral features and integrate phenological indicators by dynamically weighting feature importance and capturing complex interactions.

[0051] The ATDNN architecture consists of three main functional modules: feature extractor, feature processor and regressor ( Figure 3 The feature extractor performs adaptive dimensionality transformation while preserving physiologically relevant information through batch normalization and ReLU activation function. The first attention module is located before the feature extractor and adopts a four-head attention mechanism to dynamically weight the importance of input features ( Figure 4 ). The mechanism is expressed as: (4); in, Q, K, V represent the query matrix, key matrix and value matrix respectively, d _k Represents the dimension of the key vector. Each attention head independently projects the input features into a high-dimensional space, which can capture different feature interactions.

[0052] The feature processor is augmented by a second attention module, which facilitates the integration of processed features through nonlinear transformations. This hierarchical attention design enables the model to capture feature importance at multiple levels of abstraction, distinguishing it from baseline DNNs that rely solely on standard feed-forward layers. The processor incorporates dropout regularization (rate = 0.3) and residual connections to maintain robust feature propagation. The final regression module combines the processed features with linear transformations and ReLU activations to generate biomass predictions.

[0053] The key difference between ATDNN and DNN in architecture lies in their feature processing mechanism. DNN uses traditional feed-forward layers for feature transformation, while ATDNN introduces a dual attention module to achieve dynamic feature weighting ( Figure 3 This design enhances the model's ability to capture complex relationships among input features, whether they are purely spectral or combined with phenological indicators.

[0054] Model training and evaluation strategies The model training framework incorporates comprehensive optimization strategies to ensure robustness and reproducibility. All experiments and data processing in this example were performed on a workstation equipped with an NVIDIA GeForce RTX 6000 GPU, an Intel Xeon 3423 CPU, and 256 GB of RAM using Python 3.9, PyTorch (version 2.5.1), and the scikit-learn (version 1.5.2) libraries. Model training used the mean squared error (MSE) loss function to quantify prediction accuracy, and the Adam optimizer was initialized with a learning rate of 5×10. -4 ; The training process uses mixed precision calculation.

[0055] In order to prevent overfitting and ensure the generalization of the model, a systematic regularization strategy is adopted: a dropout regularization of 0.3 rate is used in the entire network architecture, and a predefined threshold is used. τ Gradient norm clipping is used to maintain optimization stability, and an early stopping mechanism with a patience window of 30 epochs is used. The training process uses a mini-batch size of 16 samples to optimize the balance between computational efficiency and gradient update stability. When the verification performance stabilizes, the detection mechanism dynamically adjusts the learning rate through multiplicative decay, as follows: (5); Where, t represents the learning rate during training iterations, is the learning rate decay factor, is the loss on the validation set.

[0056] Model evaluation uses two metrics to ensure a comprehensive performance assessment: the coefficient of determination (R²) quantifies the proportion of variance explained by the model, while the root mean square error (RMSE) measures the absolute prediction accuracy: (6); (7); To ensure the robustness of the evaluation, the above metrics are calculated consistently across training, validation, and test sets.

[0057] For model interpretability analysis, the SHAP (SHapley Additive Extensions) method was used to quantify feature contributions. SHAP analysis was implemented using the TreeInterpreter framework, which efficiently computes Shapley values for tree-based models. Feature importance was assessed at the individual feature level using SHAP value distribution and at the feature group level using aggregate contribution analysis. This two-level interpretability framework provided insights into local feature effects and global importance patterns, enhancing the practical applicability of the model in agricultural settings.

[0058] Model evaluation analysis (1) Comparative analysis of feature selection methods for wheat AGB estimation The Boruta-SHAP and XGB-FS methods demonstrated unique capabilities in selecting spectral features for wheat biomass estimation (Table 3). Boruta-SHAP identified 27 wavelengths across the visible (350–585 nm), red edge (708–749 nm), near-infrared (1063–1350 nm), and SWIR (1453–2276 nm) regions, reflecting diverse physiological parameters such as chlorophyll content, photosynthetic activity, and canopy structure. XGB-FS, on the other hand, selected 15 wavelengths, primarily in the SWIR (1456–2278 nm) region, which are often associated with vegetation moisture content and structural characteristics.

[0059] Table 3. Selected spectral features for wheat biomass estimation using Boruta-SHAP and XGB-FS methods method Selected spectral features (nm) Boruta-SHAP 350, 387, 500, 511, 585, 708, 709, 718, 733, 740, 745, 749, 1063, 1069, 1086,1118, 1335, 1349, 1350, 1453, 1714, 1716,1722, 2002, 2039, 2263, 2276 XGB-FS 721, 1069, 1187, 1456, 1505, 1655, 1710, 1712, 1715, 1717, 1720, 1721, 2077,2124, 2278

[0060] In summary, the integration of the Boruta and SHAP algorithms demonstrated significant advantages in identifying relevant spectral features for wheat biomass estimation. The hybrid approach achieved superior prediction accuracy (R² = 0.726), a 25.6% improvement over XGB-FS (R² = 0.578). This performance gain stems from Boruta-SHAP's comprehensive spectral coverage and robust feature selection process, which effectively identified statistically significant wavelengths across different spectral regions. Key spectral regions identified included the NIR (733–1118 nm), associated with canopy structure and biomass density, and the SWIR (1714–2276 nm), associated with water content and biochemical changes. Furthermore, the red edge band (708–749 nm), reflecting variations in chlorophyll content and leaf area, demonstrates its key role in characterizing biomass accumulation. These findings demonstrate Boruta-SHAP's ability to integrate physiological relevance into feature selection.

[0061] (2) Integration effect and performance analysis of multi-source phenological indicators in wheat AGB estimation Combining phenological metrics with spectral features selected by Boruta-SHAP significantly improved model performance. Among the individual metrics, the Zadoks Scale (ZS) achieved the highest accuracy (R² = 0.829; RMSE = 1662.65 kg / ha), followed by GDD (R² = 0.824; RMSE = 1690.18 kg / ha) and GD (R² = 0.791; RMSE = 1840.03 kg / ha). Combining the three metrics (ZS, GDD, and GD) with spectral features achieved the highest accuracy, reaching R² = 0.847 and RMSE = 1575.11 kg / ha. This represents a 16.7% improvement in R² and a 25.2% reduction in RMSE compared to the individual spectral features (R² = 0.726; RMSE = 2106.87 kg / ha).

[0062] These results demonstrate the complementary roles of these phenological indices: ZS provides precise stage-specific context, GDD reflects heat accumulation, and GD explains temporal progression. Together, they construct a robust temporal framework that improves the reliability and accuracy of model biomass estimates.

[0063] (3) Comparative analysis of the hierarchical attention mechanism in the phenologically enhanced wheat AGB estimation model In the integrated phenology enhancement model (combining GD, GDD, and ZS), the application of the hierarchical attention mechanism showed significant improvements in computational efficiency and prediction accuracy. The attention-enhanced architecture (ATDNN) significantly improved training efficiency, achieving optimal convergence in 73 epochs compared to the standard model (DNN) without attention mechanism, which reduced training iterations by 49.3%. Convergence trajectory analysis, such as Figure 5 As shown, this shows a faster initial learning rate and consistent performance improvement throughout the training process.

[0064] A detailed computational analysis, as shown in Table 4, shows that the ATDNN increases the number of model parameters from 10,753 to 31,617 (a 193.9% increase), while the actual training time increases slightly from 36.68 seconds to 42.09 seconds (a 14.6% increase). In terms of predictive performance, the ATDNN achieves higher accuracy, with a test set R² of 0.847, an 8.7% improvement over the baseline DNN (R² = 0.779), while the RMSE decreases by 16.8%, from 1892.68 to 1575.11 kg / ha. These quantitative improvements validate the effectiveness of the hierarchical attention architecture in balancing computational efficiency and predictive accuracy in agricultural biomass estimation applications.

[0065] Table 4. Performance evaluation indicators of ATDNN and DNN models in wheat biomass estimation parameter ATDNN DNN Model complexity Total parameters 31,617 10,753 Training efficiency Training time(s) 42.09 36.68 Best Times 73 144 Inference speed (ms) 1.56 0.55 Prediction performance Test set R² 0.847 0.779 Test set RMSE (kg / ha) 1,575.11 1,892.68

[0066] (4) Characteristic contribution analysis of wheat AGB estimation based on SHAP values SHAP-based feature influence analysis showed different contribution patterns of temporal spectral features in biomass estimation ( Figure 6 ). Phenological indicators are the main contributors, and ZS and GDD show consistent and substantial SHAP value distributions. In the spectral domain, NIR bands (such as 1118 nm, 1086 nm, 1069 nm) show strong contributions, reflecting their sensitivity to vegetation structure; red edge bands (such as 733 nm, 745 nm) show dynamic responses, while SWIR bands (such as 2276 nm, 1714 nm) have moderate but stable effects. Comprehensive analysis showed that phenological information was the main feature category (31.5%), followed by NIR (22.9%), SWIR1 (12.3%) and red edge bands (8.6%). These findings validate the feature selection strategy and highlight the synergistic effect of temporal spectral integration in achieving robust biomass estimation.

[0067] Although some preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0068] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of the inventive concept. Thus, if such changes and modifications fall within the scope of the claims of this application and their equivalents, this application is intended to include such changes and modifications.

Claims

1. A method for constructing a wheat biomass estimation model that integrates phenological information and hyperspectral data, comprising the following steps: (1) Select phenological indicators that can characterize different wheat growth and development stages: growing days (GD), accumulated temperature (GDD), and growth period index (ZS); (2) Identify and screen the spectral features related to wheat biomass at the corresponding growth and development stages based on the Boruta-SHAP algorithm; (3) An attention temporal deep neural network (ATDNN) model was established to combine the phenological indicators GDD, GD, and ZS with the selected spectral features using an attention mechanism and dynamically weight the feature importance to capture the non-stationary relationship between spectral reflectance and biomass accumulation. (4) The established ATDNN model is trained iteratively based on the historical spectral feature dataset and the corresponding phenological indicator dataset until convergence.

2. The wheat biomass estimation model construction method according to claim 1, characterized in that: The ATDNN model includes an input layer, a first attention module, a feature extractor, a feature processor, and a regressor. The input layer is used to input the spectral features and phenological indicators. The first attention module is set before the feature extractor and adopts a four-head attention mechanism to dynamically weight the importance of the input features. The feature extractor performs adaptive dimensionality conversion while retaining physiological relevant information through batch normalization and ReLU activation function; The feature processor is enhanced by a corresponding second attention module to facilitate the integration of processed features through nonlinear transformation; the regressor combines the processed features with linear transformation and ReLU activation to generate biomass predictions.

3. The wheat biomass estimation model construction method according to claim 2, characterized in that: The four-head attention mechanism is expressed as follows: ; in, Q, K, V represent the query matrix, key matrix and value matrix respectively, d k Indicates the dimension of the key vector.

4. The wheat biomass estimation model construction method according to claim 1, characterized in that: In step (4), the model training uses the mean square error (MSE) loss function to quantify the prediction accuracy, combined with the Adam optimizer, and the initial learning rate is 5×10^ -4 ; The training process uses mixed precision calculations to optimize numerical stability and computational efficiency while maintaining prediction accuracy.

5. The wheat biomass estimation model construction method according to claim 1, characterized in that: In the ATDNN model, a dropout regularization of 0.3 rate is adopted, using a predefined threshold τ Gradient norm clipping is used to maintain the stability of the optimization, and an early stopping mechanism with a patience window of 30 epochs is used.

6. The wheat biomass estimation model construction method according to claim 1, characterized in that: All phenological indices were standardized using the z-score normalization method.

7. A wheat biomass estimation method, characterized in that: The steps include: (1) Collect and obtain wheat hyperspectral data and / or phenological index data of the target area during the corresponding period; (2) The obtained hyperspectral data were subjected to wavelength resampling, splicing correction, and noise removal; the obtained phenological index data were subjected to standardization and Z-score conversion; (3) Inputting the hyperspectral data and / or phenological index data obtained by preprocessing in step (2) into the wheat biomass estimation model of claim 1 to generate and output a corresponding biomass prediction value.

8. The method for predicting biomass production according to claim 7, wherein: In the step (1), wheat hyperspectral data within the wavelength range of 350-2500 nm is collected; the phenological index data includes growing days GD, growing accumulated temperature GDD and growth period index ZS.

9. The biomass estimation method according to claim 7, characterized in that: In step (2), the obtained hyperspectral data are first subjected to splicing correction at wavelengths of 1000 nm and 1800 nm to eliminate detector knot artifacts, and then the spectral regions affected by atmospheric water absorption and low signal-to-noise ratio are systematically removed. Then, a smoothing algorithm is applied to eliminate or reduce random noise while retaining the basic spectral features.

10. Application of the wheat biomass estimation model according to claim 1 in wheat yield prediction, fertilization optimization and / or irrigation management.

Citation Information

Cited By

  • Dam multi-point deformation prediction method and system

    CN120781631A