Vegetation canopy nitrogen content prediction method based on remote sensing image spectrum
Through the improved Black-winged Kire Optimization Algorithm (IBKA) optimization of the random forest model, the limitations of the random forest model in the inversion of nitrogen content in vegetation leaves in complex data scenarios are solved, and more efficient and reliable nitrogen content prediction is achieved.
Patent Information
- Application Number
- CN202510089262.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-20
AI Technical Summary
In the current technology, in high-dimensional complex data scenarios, random forest models are prone to overfitting and parameter sensitivity in the nitrogen content in vegetation leaves inversion, and the performance of existing metaheuristic optimization algorithms in complex data scenarios still needs to be improved.
An improved black-winged kite optimization algorithm (IBKA) is proposed, and population initialization is performed through Sobol sequences, position update is used to adopt somersault foraging strategy, and a multi-scale collaborative variation mechanism is introduced to optimize the random forest model to improve the prediction accuracy of the nitrogen content in vegetation canopy.
It effectively improves the algorithm's global search ability, local optimization ability, convergence efficiency and stability, overcomes problems such as uneven group initialization, blindness of initial search, easy to fall into local optimization and unbalanced search and development capabilities, and significantly improves the accuracy of vegetation nitrogen content prediction.
Smart Images

Figure CN120182804A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for predicting leaf nitrogen content, belonging to the field of agricultural remote sensing, and particularly to a method for predicting nitrogen content in vegetation canopy based on remote sensing image spectra. Background Art
[0002] Leaf nitrogen content (LNC) is a key indicator reflecting crop nutrient status and photosynthesis efficiency. Traditional LNC measurement methods rely on field sampling and laboratory analysis. Although they have high accuracy, the process is cumbersome and costly, limiting their popularization in large-scale applications.
[0003] In recent years, remote sensing technology, with its characteristics of being fast, efficient, and low-cost, has become the core tool for crop growth monitoring in precision agriculture. Especially low-altitude unmanned aerial vehicle (UAV) remote sensing technology, due to its high spatial resolution, flexibility in data acquisition, and simplicity in operation, has shown significant potential in estimating indicators such as nitrogen, phosphorus, potassium content, SPAD value, and LAI of crops. Research shows that remote sensing technology combined with machine learning algorithms provides a reliable solution for solving non-linear and high-dimensional data problems in agricultural monitoring.
[0004] In machine learning models, the random forest (RF) has been widely used in the inversion research of crop nutrient components due to its excellent non-linear fitting ability and advantages in dealing with high-dimensional data. However, the RF model may still have limitations when facing high-dimensional complex data, such as being prone to overfitting, sensitive to parameters, etc. To improve its performance, in recent years, meta-heuristic optimization algorithms based on swarm intelligence have been gradually adopted. By simulating collaborative behaviors and dynamic search processes in nature, these algorithms have significant advantages in global search ability and avoiding local optima. For example, ant colony algorithm, particle swarm algorithm, etc. have achieved preliminary applications in the field of agricultural remote sensing, but their performance in complex data scenarios still needs to be further improved. Summary of the Invention
[0005] The present application provides a method for predicting nitrogen content in vegetation canopy based on remote sensing image spectra. The method proposes a new optimization algorithm, IBKA, which can be used to optimize the random forest model to solve the complex problem of inverting nitrogen content in vegetation leaves from high-dimensional spectral data, providing an efficient and reliable method for dynamic monitoring of LNC in precision agriculture.
[0006] The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectra includes:
[0007] Select and calculate spectral indices and texture indices according to the spectral images of vegetation;
[0008] The machine learning model is optimized by an improved black-winged kite optimization algorithm, and the optimized model is trained using the nitrogen content of the samples, spectral indices, and texture indices to obtain a vegetation canopy nitrogen content prediction model. The input of this model is spectral indices and texture indices, and the output is the predicted value of vegetation nitrogen content. Among them, the improved black-winged kite optimization algorithm includes: in the swarm intelligence optimization algorithm, the Sobol sequence is used for population initialization, the somersault foraging strategy is used for position update, and the multi-scale collaborative mutation mechanism is used to determine the escape position for each iteration.
[0009] The spectral indices and texture indices of the vegetation to be measured are input into the vegetation canopy nitrogen content prediction model to obtain the predicted value of the canopy nitrogen content of the vegetation to be measured.
[0010] Optionally, the spectral index uses a vegetation index.
[0011] Optionally, the vegetation index includes but is not limited to: vegetation index based on soil line, ratio-based vegetation index, atmospheric correction index, derivative vegetation index.
[0012] Optionally, the acquisition of the texture index includes: extracting several texture features for each band from the multi-spectral remote sensing image of the leaves of the vegetation canopy based on the gray-level co-occurrence matrix.
[0013] Optionally, the texture features include but are not limited to: mean, variance, co-occurrence, contrast, dissimilarity, information entropy, second-order moment, and correlation.
[0014] Optionally, this method further includes: performing feature screening on the spectral indices and texture indices to eliminate non-sensitive features.
[0015] Optionally, the use of the Sobol sequence for population initialization includes:
[0016] Set the optimal value range to [ω min ,ω max ;
[0017] The Sobol sequence is used to map the initialized population, as shown in the following formula:
[0018] ω n =ω min +L n ×(ω max -ω min )
[0019] where ω n is the initial position of the population, L n ∈[0,1] is the random number generated by the Sobol sequence, n = 1,..., N, and N is the size of the initial population.
[0020] Optionally, the position update using the somersault foraging strategy is as shown in the following formula:
[0021] K e (t + 1)=K e (t)+e×(λ1×K bs -λ2×K e (t))
[0022] Where K e (t + 1) is the new position of the particle in the population, and K e (t) is the current position of the particle in the population, e is the somersault factor, and K bs is the probability of the current global optimum, and λ1 and λ2 are random numbers in [0, 1] respectively.
[0023] Optionally, the escape position for each iteration determined by the multi-scale cooperative mutation mechanism is specifically to introduce a Gaussian mutation operator in the mutation mechanism, which has J mutation scales;
[0024] Define the variance of the high-speed mutation operator, and its initialization is as shown in the following formula:
[0025]
[0026] Where represents the variance of the i-th mutation scale in the t0-th iteration, and its initial value is defined as the entire solution space;
[0027] During the iteration process, the variance is adjusted as follows:
[0028] Sort the particles in the population according to the fitness value;
[0029] Divide the sorted particles into J subgroups, each subgroup contains Q = N / J particles, and the fitness value of the j-th subgroup is calculated according to the following formula:
[0030]
[0031] Where k represents selecting k subgroups from J subgroups; is the fitness value of subgroup j in k subgroups; is the fitness value of the q-th particle in subgroup j, and the position update of particle q uses the somersault foraging strategy;
[0032] Dynamically adjust the standard deviation of each subgroup according to the fitness value of each subgroup;
[0033] Calculate the position of the population according to the standard deviation of each subgroup.
[0034] Dynamically adjust the standard deviation of each subgroup according to the fitness value of the subgroup;
[0035] Calculate the position of the population according to the standard deviation of each subgroup.
[0036] Optionally, the dynamic adjustment includes:
[0037] Obtain the maximum and minimum fitness values of all subgroups, denoted as Fit max and Fit min respectively, as shown in the following formula:
[0038]
[0039] where is the fitness value of the j-th subgroup among k subgroups;
[0040] Then, at the t-th iteration, the standard deviation of each subgroup is as shown in the following formula:
[0041]
[0042] where is the standard deviation of subgroup j at the t-th iteration.
[0043] Optionally, the calculating the position of the population according to the standard deviation of each subgroup includes:
[0044] Normalize the standard deviation of each subgroup to obtain a high-speed mutation operator, as shown in the following formula:
[0045]
[0046] where D is the domain of the given optimization problem, t is the number of iterations, and the dynamic adjustment process is repeated until is the standard deviation of subgroup j at the t-th iteration; is the high-speed mutation operator at the t-th iteration;
[0047] Define the position of the population, as shown in the following formula:
[0048]
[0049] The above conditions are:
[0050]
[0051] where t is the number of iterations, x q (t) is the position of the q-th particle at the t-th iteration, f(·) is the fitness function, and rand(0,1) is a random value between 0 and 1.
[0052] Optionally, the machine learning model selects a random forest or a deep extreme learning machine.
[0053] Preferably, the machine learning model is a random forest model.
[0054] Optionally, the method further includes testing and validating the vegetation canopy nitrogen content prediction model: comparing the training and test results of the vegetation canopy nitrogen content LNC at different growth stages, where the growth stages include: leaf expansion stage, initial flowering stage, and full flowering stage.
[0055] The beneficial effects that this application can produce include:
[0056] 1) The method for predicting the nitrogen content of the vegetation canopy based on the remote sensing image spectrum provided by this application proposes a new optimization algorithm, namely the improved black-winged kite optimization algorithm (IBKA). Through the synergistic effect of the Sobol sequence, somersault foraging strategy, and multi-scale cooperative mutation mechanism, an unprecedented optimization mechanism is formed, improving the global search ability, local development ability, convergence efficiency, and stability of the algorithm, effectively overcoming problems such as uneven swarm initialization, blind initial search, easy entrapment in local optima, and imbalance between search and development capabilities existing in the traditional BKA algorithm.
[0057] 2) The improved black-winged kite optimization algorithm (IBKA) provided by this application can be applied to the optimization of the random forest model (IBKA-RF), that is, the vegetation canopy nitrogen content prediction model described in this application is obtained. Experiments prove that compared with applying IBKA to other machine learning models, such as the BP neural network (BP Neural Network, BP) and the deep extreme learning machine (Deep Extreme Learning Machine, DELM), the IBKA-RF algorithm of this application still shows excellent performance in solving high-dimensional complex optimization problems, and high fitting accuracies have been achieved for the inversion of the nitrogen content of vegetation leaves at different growth stages, which provides an efficient and reliable tool for the modeling of precise agricultural remote sensing data. Description of the Drawings
[0058] Figure 1 It is a comparison graph of the initial population distribution generated after running 100 times each using the Sobol sequence and the pseudo-random sequence respectively in an embodiment of this application;
[0059] Figure 2.1 - 2.23 They are respectively comparison graphs of the convergence performance of using IBKA and other optimization algorithms on benchmark functions in an embodiment of this application, and the drawings are sorted in sequence according to the benchmark functions F1 - F23;
[0060] Figure 3 It is a comparison graph of the nitrogen content of leaves at different growth stages in an embodiment of this application, where A represents the leaf expansion stage, B represents the initial flowering stage, and C represents the full flowering stage;
[0061] Figure 4 This is a correlation heatmap of vegetation indices and texture features with nitrogen content at different growth stages (leaf expansion stage, early flowering stage, and flowering stage) in an embodiment of this application;
[0062] Figures 5(a), 5(b), and 5(c) are comparison charts of the true values and predicted nitrogen content values of the training set at different growth stages (leaf expansion stage, early flowering stage, full flowering stage) predicted by the IBKA-RF model in an embodiment of this application; Figures 5(d), 5(e), and 5(f) are comparison charts of the true values and predicted nitrogen content values of the test set when the vegetation is at the leaf expansion stage, early flowering stage, and full flowering stage;
[0063] Figures 6(a), 6(b), and 6(c) are comparison charts of the true values and predicted nitrogen content values of the training set at different growth stages (leaf expansion stage, early flowering stage, full flowering stage) predicted by the IBKA-BP model in an embodiment of this application; Figures 6(d), 6(e), and 6(f) are comparison charts of the true values and predicted nitrogen content values of the test set when the vegetation is at the leaf expansion stage, early flowering stage, and full flowering stage;
[0064] Figures 7(a), 7(b), and 7(c) are comparison charts of the true values and predicted nitrogen content values of the training set at different growth stages (leaf expansion stage, early flowering stage, full flowering stage) predicted by the IBKA-DELM model in an embodiment of this application; Figures 7(d), 7(e), and 7(f) are comparison charts of the true values and predicted nitrogen content values of the test set when the vegetation is at the leaf expansion stage, early flowering stage, and full flowering stage. Detailed implementation manners
[0065] The following describes this application in detail with reference to embodiments, but this application is not limited to these embodiments.
[0066] This application provides a method for predicting nitrogen content in a vegetation canopy based on remote sensing image spectra, including:
[0067] Select and calculate spectral indices and texture indices according to the spectral images of the vegetation;
[0068] Optimize a machine learning model using an improved black-winged kite optimization algorithm, and train the optimized model with the nitrogen content of the samples as well as the spectral indices and texture indices to obtain a prediction model for nitrogen content in the vegetation canopy. The input of this model is the spectral indices and texture indices, and the output is the predicted value of vegetation nitrogen content; wherein, the improved black-winged kite optimization algorithm includes: in the swarm intelligence optimization algorithm, use the Sobol sequence for population initialization, use the somersault foraging strategy for position update, and use the multi-scale collaborative mutation mechanism to determine the escape position for each iteration;
[0069] Input the spectral index and texture index of the vegetation to be measured into the vegetation canopy nitrogen content prediction model to obtain the predicted value of the canopy nitrogen content of the vegetation to be measured.
[0070] Example 1
[0071] A method for predicting the nitrogen content of vegetation canopy based on remote sensing image spectrum, comprising the following steps:
[0072] (1) Vegetation selection and determination of sample nitrogen content
[0073] In this example, the leaf nitrogen content of jujube trees was studied. The study area was selected in the Alar Reclamation Area in southern Xinjiang (40°30'39"N, 81°13'14"E), located at the confluence of the Taklimakan Desert and the Tarim River. It belongs to a typical temperate continental arid desert climate with a significant diurnal temperature difference. The average annual precipitation is 78 mm, the average annual temperature is 11.7 °C, and the total annual solar radiation is between 539.12 and 592.73 kJ / cm 2 Between, the annual sunshine duration is 2,843 to 2,963 hours, the frost-free period lasts for 201 to 224 days, and the average annual effective accumulated temperature reaches 3,936 °C. The experimental area is flat, the soil is mainly sandy loam, the field water holding rate is 26% (volume water content), the wilting water content is 5.6% (volume water content), and the groundwater depth exceeds 10 meters. The organic matter content in the 0-100 cm soil layer is between 5.31 and 8.94 g / kg, which belongs to the medium level.
[0074] In this example, a DJI Phantom 4 unmanned aerial vehicle (UAV) equipped with a multispectral camera was used as the remote sensing image acquisition platform. At the same time, the Kjeldahl nitrogen analyzer was used to measure the LNC value of jujube trees, and the sampling point positions were measured by Huace S8. The DJI Phantom 4 UAV is equipped with a 1 / 2.3-inch CMOS image sensor, with 12 million effective pixels, an f / 2.8 aperture, and a 94° viewing angle, including five bands: Blue (450 nm), Green (560 nm), Red (650 nm), Red edge (730 nm), Near infrared (840 nm).
[0075] To reduce the influence of the solar radiation angle on the quality of multispectral images, this study selected to collect images at noon on sunny days. The specific test process includes the following steps: multispectral image acquisition. The dates when the UAV collected images were between 12:00 and 14:00 on May 9, 2023 (leaf expansion stage), June 9, 2023 (early flowering stage), and June 25, 2023 (full flowering stage). A total of 270 leaf samples were collected synchronously, 90 samples in each period. The Huace S8 was used to record the GPS positions of the jujube trees of the leaves to be measured, so as to ensure that according to the longitude and latitude coordinates of the ground sampling points, the same-name feature points could be found on the images.
[0076] For the determination of nitrogen content in jujube tree leaves, the collected leaf samples are placed in an oven at 105 °C for 30 min for deactivation, continuously dried to a constant weight at 80 °C, pulverized and passed through a 0.5 mm sieve, digested with H2O2-H2SO4 solution. For every 50 ml of digestion solution, 10 ml of the sample is used, and this sample is added to a GK-700 automatic Kjeldahl nitrogen analyzer for distillation. 1 / 2H2SO4 is used to titrate the distillation solution to determine the nitrogen concentration of the sample.
[0077] (2) Select and calculate the vegetation index.
[0078] In this embodiment, the spectral index selects the vegetation index, which can reflect the growth status of crops and provide detailed and valuable agricultural information for monitoring vegetation health. According to the calculation methods and application scenarios of vegetation indices, they can generally be divided into four categories:
[0079] 1. Vegetation indices based on soil lines
[0080] Such as SAVI, SARE, etc. This index aims to reduce the influence of soil background on the vegetation index by introducing a soil correction factor, and it performs well especially in the case of sparse vegetation cover.
[0081] 2. Ratio-based vegetation indices
[0082] Such as RVI, NDVI, GNDVI, GRNDVI, NDRE, NGRDI, RERVI, etc. This index differentiates vegetation from other surface substances such as soil and water through the ratio between different bands. Among them, NDVI is one of the most widely used indices in the field of remote sensing.
[0083] 3. Atmospheric correction indices
[0084] Such as EVI, MEVI, SIPI, etc. This index is designed specifically to reduce atmospheric interference. These indices are not only more sensitive to vegetation changes but also have higher stability in the face of atmospheric changes.
[0085] 4. Derivative vegetation indices
[0086] Such as NRI, CIGreen, DVI, etc. This index can more accurately capture the changes in specific bands by calculating the difference or derivative between bands, and is usually used to analyze the health status, chlorophyll content or nitrogen content of vegetation.
[0087] In this embodiment, a total of 22 vegetation indices are selected, and the specific information is shown in Table 1 below.
[0088] Table 1 List of vegetation indices and their formulas used in Example 1
[0089]
[0090] In the above table, b1, b2, b3, b4, and b5 represent the blue band, green band, red band, red edge band, and near-infrared band, respectively.
[0091] The application of these vegetation indices provides strong technical support for remote sensing data analysis and crop health monitoring.
[0092] (3) Select texture features
[0093] Use ENVI5.3 software to process the multi-spectral remote sensing images of jujube trees. Based on the gray-level co-occurrence matrix, 8 texture features are extracted from the multi-spectral remote sensing images of the jujube tree canopy leaves for each band. These features include: mean, variance, homogeneity, contrast, dissimilarity, entropy, second moment, and correlation. In this embodiment, multi-spectral images are collected in the bands of 450nm, 540nm, 650nm, 730nm, and 840nm, and 8 texture features are extracted for each band. Therefore, a total of 40 texture features are obtained from 5 spectral bands.
[0094] (4) Improved Black Kite Optimization Algorithm (IBKA)
[0095] This application improves the original Black Kite Algorithm (BKA) by combining the following three strategies to obtain the Improved Black Kite Optimization Algorithm (IBKA).
[0096] (4.1) Initialize the population position based on the Sobol sequence
[0097] In metaheuristic algorithms, the distribution of the initialized population largely affects the convergence speed and accuracy of the algorithm. In the original Black Kite Optimization Algorithm, the search space generates the initialized population in a random number manner. This method has low convenience, and the particle distribution is uneven and unpredictable, which affects the performance of the algorithm to a certain extent. To improve the global search ability, some scholars use chaotic search to optimize the initialization sequence. For example, using the Tent chaos method to initialize the population, adopting the Circle chaos method to initialize the population, and introducing the Cubic mapping algorithm to initialize the population. Although these chaotic algorithms have a certain ability to jump out of the local optimum, they have deficiencies: strong randomness, and the algorithm faces great uncertainty during operation; adjacent points are closely related, and if iterated to an unstable point, it cannot continue to run, such as the small cycle points (0.2, 0.4, 0.6, 0.8) and unstable points (0, 0.25, 0.5, 0.75) of the Tent mapping.
[0098] This application uses a deterministic low-discrepancy sequence to replace the pseudo-random sequence. By selecting a reasonable sampling direction, points are filled into the multi-dimensional hypercube cells as evenly as possible to obtain higher efficiency and uniformity when dealing with probability problems. Further, this application uses the Sobol sequence, which has a shorter calculation cycle, a faster sampling speed, and higher efficiency in dealing with high-dimensional series, to map the initial population. The specific steps are as follows.
[0099] Let the optimal value range be [ω min , ω max , and the random number L n generated by the Sobol sequence ∈ [0, 1]. Then the initial position of the population can be defined as:
[0100] ω n = ω min + L n × (ω max - ω min )
[0101] where ω n is the initial position of the population, L n ∈ [0, 1] is the random number generated by the Sobol sequence, n = 1,..., N, and N is the size of the initial population.
[0102] Assume that the upper and lower bounds of the population are 0 and 1 respectively. The distribution of the initial population generated by running the Sobol sequence and the pseudo-random sequence 100 times respectively is as Figure 1 shown. Among them, the horizontal axis represents the value size, and the vertical axis represents the number of times the particles fall into a certain range. It can be seen that compared with the pseudo-random numbers, the number of particles in the initial population generated by the Sobol sequence falling into each range is roughly the same, and they are more uniform and have a wider ergodicity.
[0103] (4.2) Use the somersault foraging strategy for position update
[0104] In the later stage of the algorithm, all black-winged stilt particles move closer to the area where the optimal particle in the current group is located, resulting in a loss of group diversity. If the current optimal particle is not the global optimal solution, the algorithm will fall into a local optimum, which is an inherent shortcoming of the swarm intelligence optimization algorithm.
[0105] To overcome this shortcoming, this application introduces the somersault foraging strategy into the IBKA to obtain the following mathematical expression:
[0106] P s (t + 1) = P s (t) + S · (r1 · P bs - r2 · X(t))
[0107] K e(t + 1) = K e (t) + e×(λ1×K bs -λ2×K e (t))
[0108] Among them, K e (t + 1) is the new position of the particles in the population, K e (t) is the current position of the particles in the population, e is the somersault factor, K bs is the probability of the current global optimum, and λ1 and λ2 are random numbers in [0, 1] respectively.
[0109] (4.3) Introduce the multi-scale collaborative mutation mechanism
[0110] In each iteration of BKA, the escape position is determined by the uniform mutation scale. The mutation scale enables BKA to escape from the population positions of the previous generation. However, the distance between the local extrema of a given benchmark function is unpredictable, so an appropriate mutation scale cannot be determined in advance. To enable BKA to efficiently find the optimal solution, the following problems must be solved when selecting an appropriate mutation scale: If the mutation scale is too large, some extrema may be skipped, including the global optimum; if the mutation scale is too small, a large number of iterations may be required to traverse the entire search space, which will significantly reduce the convergence speed of the algorithm. To solve these problems, the IBKA of this application adopts a Gaussian mutation operator with different scaling variances to escape from the local optimum.
[0111] (4.3.1) Assume that the Gaussian mutation operator has J mutation scales, then its variance is initialized as:
[0112]
[0113] Among them, represents the variance of the i-th mutation scale in the t0-th iteration, and its initial value is defined as the entire solution space.
[0114] (4.3.2) Sort the particles in the population according to the fitness value;
[0115] Divide the sorted particles into J subgroups, each subgroup contains Q = N / J particles, and calculate the fitness value of each subgroup as shown in the following formula:
[0116]
[0117] Among them, k represents selecting k subgroups from J subgroups; is the fitness value of subgroup j in k subgroups; is the fitness value of the q-th particle in subgroup j, and the position update of particle q adopts the somersault foraging strategy.
[0118] (4.3.3) Dynamically adjust the standard deviation of each subgroup according to the fitness value of the subgroup. The adjustment process is as follows:
[0119] Obtain the maximum and minimum values of all subgroups, denoted as Fit max and Fit min , as shown in the following formula:
[0120]
[0121] where is the fitness value of the j-th subgroup among the k subgroups;
[0122] Then, at the t-th iteration, the standard deviation of each subgroup is shown in the following formula:
[0123]
[0124] where is the standard deviation of subgroup j at the t-th iteration.
[0125] (4.3.4) Calculate the position of the population according to the standard deviation of each subgroup.
[0126] First, as the algorithm iterates, the mutation operator may become very large. Therefore, it is necessary to first perform normalization to obtain the standard deviation of the high-speed mutation operator, as shown in the following formula:
[0127]
[0128] where D is the domain of the given optimization problem, t is the number of iterations, and the dynamic adjustment process is repeated until is the standard deviation of subgroup j at the t-th iteration; is the high-speed mutation operator at the t-th iteration;
[0129] Then, according to the standard deviation of the high-speed mutation operator, define the position of the population as follows:
[0130]
[0131] The above conditions are:
[0132]
[0133] where t is the number of iterations, x q (t) is the position of the q-th particle at the t-th iteration, f(·) is the fitness function, and rand(0,1) is a random value between 0 and 1.
[0134] In this application, the Sobol sequence initialization, somersault foraging strategy, and multi-scale collaborative mutation mechanism play a synergistic role in the BKA algorithm, thus forming a dynamic and efficient search mechanism. The coupling of the above three strategies produces new technical effects, rather than simply adding up the effects of each strategy. Specifically, it is reflected in:
[0135] 1. Improve search efficiency and accuracy: Due to the uniform initialization of the Sobol sequence and the global jump of somersault foraging, particles can detect valuable solution space regions more quickly without being trapped in local optima, providing more potential high-quality solutions for subsequent optimization, thus ensuring that the IBKA has strong exploration ability and accuracy during the convergence process.
[0136] 2. Enhance global optimization ability and diversity: The introduction of the somersault foraging strategy breaks the premature convergence problem of the original BKA algorithm, enabling the algorithm to continuously explore potential optimal solutions while maintaining global search ability.
[0137] 3. Optimize the convergence process: Adopting the technical route of the synergistic effect of global search and local optimization enables the IBKA to converge to a position close to the global optimal solution in a shorter time, while avoiding the dilemma of local optimal solutions.
[0138] This synergistic effect significantly improves the performance of the IBKA. Especially when dealing with high-dimensional complex problems, it can achieve a good balance between diversity and search efficiency. Compared with the existing technologies, the optimization effect of the IBKA is difficult to be predicted by simple strategy combinations because it forms an optimization ability and solution that exceed traditional algorithms through new combination methods.
[0139] (4.4) IBKA Performance Test and Results
[0140] To verify the performance of the IBKA, this embodiment uses the CEC2005 test set to evaluate and compare the ability of the optimization algorithm IBKA to handle various objective functions. The test environment is: carried out on MATLAB R2021a, using a 2.0GHz 64-bit Core i7 processor and 16GB of main memory.
[0141] The test set includes unimodal problems (f1 - f5), basic multimodal problems (f6 - f12), extended multimodal problems (f13 - f14), and hybrid composite problems (f15 - f23). IBKA is used to solve these function optimization problems, and the results are compared with those of well-known algorithms such as the original Black-winged Stilt Optimization Algorithm (BKA), Artificial Protozoa Optimizer (APO), Artificial Rabbit Optimization Algorithm (ARO), Pigeon In swarm Optimization Algorithm (PIO), and Differential Evolution Algorithm (DE). For the fairness of comparison, the parameter settings of the six algorithms are the same, that is, the population size is set to 30, and the maximum number of iterations is 500. In the experiment, the number of bits for all optimization problems is set to 30. Each algorithm runs independently 30 times for the 23 function optimization problems to reduce errors. In addition, four evaluation functions are used to comprehensively analyze the performance of the algorithms: optimal value, standard deviation, mean, median, and worst value, as shown in Table 2 below.
[0142] Table 2 Statistical Comparison of IBKA and Other Optimization Algorithms on Benchmark Functions (F1 - F23)
[0143]
[0144]
[0145]
[0146]
[0147] It can be seen from the comparison results in the above table that when IBKA is dealing with unimodal problems (such as F1 and F3) and basic multimodal problems (such as F9 and F11), all its statistical indicators - optimal value, standard deviation, mean, median, and worst value are 0. This performance indicates that IBKA shows extremely high solution efficiency and consistency in these problems, demonstrating excellent robustness.
[0148] Furthermore, in this embodiment, the convergence performance of IBKA is also tested. Figure 2.1 - Figure 2.23The convergence curves of IBKA and other algorithms in 10 dimensions are shown. The vertical axis in the figure is the fitness value, and the horizontal axis is the number of iterations. On unimodal functions, IBKA demonstrates amazing convergence speed. Especially in F1 and F4, its ability to quickly approach the optimal solution far exceeds other algorithms, which is extremely important in practical applications that require efficient optimization. IBKA can not only converge quickly but also maintain a low fitness value for a long time. Even in the more challenging F3 and F5, facing potential premature convergence problems, IBKA can still explore solutions close to the global optimum, fully demonstrating its excellent performance in the optimization field. Although the advantages in F2 and F5 are not as significant as those in F1 and F4, the performance of IBKA is still stable, comparable to other top algorithms, showing its wide adaptability to different types of unimodal problems.
[0149] When dealing with multimodal problems (such as F6 to F12), IBKA also performs excellently. Its powerful global search ability enables the algorithm to quickly locate the global optimum value and effectively avoid local optimum traps. In these tests, IBKA quickly reduces the fitness value on multiple multimodal functions (such as F6, F7, F8, F9, F10, F11, and F12), showing a very fast convergence trend. In some scenarios, IBKA can approach the theoretical optimum value in the early iteration stage. Especially in the F10 and F11 functions, its fitness value is almost 0, demonstrating its excellent global search performance. When facing complex multimodal functions (such as F12 and F14), although other algorithms often fall into local optima, IBKA can continuously optimize and find better global solutions. This fully proves the significant advantages of IBKA in expanding the search space and avoiding local optima. In all multimodal problems, IBKA demonstrates excellent adaptability and efficient optimization ability. Especially in F9 and F11, when other algorithms encounter difficulties in the optimization process, IBKA can still continuously improve, showing extremely strong robustness and reliability.
[0150] In the processing of extended functions (such as F13 and F14), IBKA also demonstrates remarkable optimization ability. Its convergence curve is significantly faster than other algorithms, quickly approaching the theoretical optimum value, and maintaining a low fitness value throughout the iteration process, demonstrating its efficiency and reliability in dealing with extended problems. In F14, IBKA converges quickly in the initial iteration, showing its excellent performance in dealing with multimodality and complex fitness landscapes.
[0151] For mixed multimodal functions, the advantages of IBKA are also significant. On some functions (such as F15, F17, F19, F21, and F23), IBKA almost reaches the theoretical optimal solution, demonstrating its excellent global search ability and optimization efficiency in complex multimodal problems. However, in some functions (such as F18, F20, and F22), the performance of IBKA is similar to that of other algorithms. Although it fails to obtain the optimal solution, the gap between its result and the optimal solution is extremely small, further proving the strong potential and competitiveness of IBKA in the optimization process.
[0152] In summary, IBKA has demonstrated excellent global search ability and robustness in various optimization problems. Whether it is unimodal, multimodal, or extended and mixed multimodal problems, IBKA shows a very fast convergence rate and a strong ability to approximate the optimal solution. Even in a complex search space, IBKA can effectively avoid local optima and maintain stable optimization performance. Although its performance is close to that of other algorithms on some functions, IBKA's overall performance is more robust and suitable for a wide range of practical application scenarios.
[0153] (5) Construction and optimization of machine models
[0154] (5.1) Theoretical support for predicting nitrogen content in vegetation canopy
[0155] Nitrogen is a key nutrient element for plant growth and development. In this embodiment, jujube trees are taken as an example. When they are in different growth periods, the change in nitrogen content has a significant impact on the physiological activities of plants. Such as Figure 3The box plot shown below demonstrates the nitrogen content in jujube tree leaves at different growth stages: During the leaf expansion stage (A), the nitrogen content in the leaves is relatively low, which may be due to the early growth stage of the jujube tree and its relatively low demand for nitrogen. However, despite the low nitrogen content, nitrogen is still crucial for the structural construction and functional maintenance of the leaves at this time. After entering the early flowering stage (B), the nitrogen content increases significantly. As a key element in chlorophyll synthesis and photosynthesis, nitrogen directly affects energy supply, flower formation, and the maintenance of plant metabolism. At the full flowering stage (C), the nitrogen content reaches its maximum value, which may be related to the gradual maturation of the plant's reproductive organs. At this time, the demand for nitrogen increases significantly to support flower and fruit formation and meet the requirements of the peak photosynthesis period. The results of analysis of variance (ANOVA) show that there are significant statistical differences in nitrogen content among the three growth stages (F = 212.7, P < 0.0001), indicating that these differences are not randomly generated but are significantly driven by the different growth stages. In addition, the results of Tukey's multiple comparison tests further confirm the differences in nitrogen content among different growth stages. All pairwise comparisons (A vs. B, A vs. C, B vs. C) show significant differences (P < 0.0001), and the confidence intervals also confirm the existence of these differences. This result supports the trend of the gradually increasing nitrogen demand of jujube trees at different growth stages, which is highly correlated with the physiological requirements of plants at each growth stage. The results of the homogeneity of variance test show that both the Brown-Forsythe and Bartlett tests indicate significant differences in the standard deviations among different groups (P < 0.0001), suggesting that the fluctuations in nitrogen content at each growth stage are heterogeneous. This may be related to the different physiological requirements of plants at different growth stages and the environmental pressures they are subjected to, resulting in uneven changes in nitrogen distribution. Therefore, in practical applications, the application of nitrogen fertilizer should be precisely regulated according to the growth stage requirements of the plants.
[0156] However, the overall correlation between existing spectral indices and texture features and the nitrogen content (LNC) of jujube tree leaves at different growth stages is weak, and it is difficult to achieve accurate LNC prediction solely based on these features. As Figure 4 shown, by using the random forest algorithm to analyze the correlation between the nitrogen content (LNC) of jujube tree leaves at different growth stages and spectral indices and texture features, it can be found that the overall correlation is weak and cannot achieve an ideal prediction effect. Among them, Figure 4 the outermost circle represents the leaf expansion stage. Although PSRI and SAVI show relatively high correlations, with correlation coefficients reaching 0.4 and 0.35 respectively, they still cannot fully reflect the complex changes in LNC. The correlations of the remaining spectral indices such as RVI and NDVI are even lower, with correlation coefficients between 0.2 and 0.3, making it difficult to effectively use them for accurate prediction. Compared with other indices, the correlations of B5-Correlation and B5-Second Moment in texture features are relatively high, but the correlation coefficients are only about 0.3, indicating that there is still room for improvement in their prediction ability.Figure 4 The middle layer represents the early flowering stage. The correlation between the spectral indices and LNC further weakens, especially for RVI and SAVI, with the correlation coefficients dropping to 0.25 and 0.33 respectively. Among the texture features, B4-Homogeneity and B4-Correlation show a slight improvement, but the correlation still fails to reach the ideal prediction accuracy, with the correlation coefficient around 0.3. Figure 4 The innermost circle represents the full flowering stage, which also does not show a significant correlation. Although the correlation between PSRI and SAVI remains around 0.4, the correlations of the remaining spectral indices and texture features are still relatively low, especially for RVI and NDVI, showing obvious deficiencies. Among the texture features, B3-Homogeneity and B3-Entropy slightly improve during the full flowering stage, but their correlation coefficients still hover around 0.3, failing to fully reflect the dynamic changes of LNC.
[0157] Therefore, in this embodiment, it is expected to introduce the improved Black-winged Kite Algorithm (IBKA) into the machine learning model to improve the prediction accuracy and overcome the problem of insufficient correlation in the current nitrogen content prediction.
[0158] (5.2) Construction of the vegetation canopy nitrogen content prediction model
[0159] Based on the original spectra, spectral indices, and texture indices, an estimation model for the nitrogen content of jujube tree leaves was established. The true values of the training set and test set of the nitrogen content in jujube tree leaves during the leaf expansion stage, early flowering stage, and full flowering stage were compared with the nitrogen content values predicted by the algorithms. In this application, the IBKA-RF, IBKA-BP, and IBKA-DELM algorithms were used respectively. As shown in Figures 5 - 7, there are significant differences in their performances.
[0160] During the leaf expansion stage, as shown in Figures 5(a) and 5(d), IBKA-RF demonstrates a powerful ability to capture non-linear data, with the Rsquare values of training and testing being 0.92746 and 0.6628 respectively, and the RMSE values corresponding to 0.44224 and 1.5108. This excellent performance benefits from the ensemble learning characteristics of the RF algorithm, which effectively reduces overfitting and enhances generalization ability through the combination of multiple decision trees. In contrast, as shown in Figures 6(a) and 6(d), IBKA-BP is prone to falling into local optima when dealing with extreme values. Its training Rsquare value is 0.81196, and the RMSE value is 0.85708, while the test results are significantly reduced, with the Rsquare value dropping to 0.50254 and the RMSE value being 1.3873. As shown in Figures 7(a) and 7(d), the training results of IBKA-DELM at this stage reach 0.81627, with the RMSE value being 0.88854, and the Rsquare value during testing is 2 respectively 0.92746 and 0.6628, and the RMSE values are 0.44224 and 1.5108 respectively. This excellent performance benefits from the ensemble learning characteristics of the RF algorithm, which effectively reduces overfitting and enhances generalization ability through the combination of multiple decision trees. In contrast, as shown in Figures 6(a) and 6(d), IBKA-BP is prone to falling into local optima when dealing with extreme values. Its training Rsquare value is 0.81196, and the RMSE value is 0.85708, while the test results are significantly reduced, with the Rsquare value dropping to 0.50254 and the RMSE value being 1.3873. As shown in Figures 7(a) and 7(d), the training results of IBKA-DELM at this stage reach 0.81627, with the RMSE value being 0.88854, and the Rsquare value during testing is 2 0.81196, and the RMSE value is 0.85708, while the test results are significantly reduced, with the Rsquare value dropping to 2 0.50254, and the RMSE value is 1.3873. As shown in Figures 7(a) and 7(d), the training results of IBKA-DELM at this stage reach 2 0.81627, and the RMSE value is 0.88854, and the Rsquare value during testing is 2is 0.7592, and the corresponding RMSE is 0.84587, indicating that it performs well in processing simple non-linear data, but the accuracy still needs to be improved in complex data scenarios. At the early flowering stage, as shown in Figures 5(b) and 5(e), IBKA-RF continues to show stability, and the training R 2 reaches 0.94495, with an RMSE of 0.62596, while the test result R 2 is 0.60209, and the RMSE is 1.5451. As shown in Figures 6(b) and 6(e), the training performance of IBKA-BP is slightly inferior, and its R 2 is 0.81543, with an RMSE of 1.1744. The test result R 2 drops to 0.42716, and the corresponding RMSE is 1.7227. As shown in Figures 7(b) and 7(e), the training and test results R 2 of IBKA-DELM are 0.75218 and 0.69091 respectively, and the corresponding RMSEs are 1.2613 and 1.5089. Although it is better than the IBKA-BP algorithm, it still shows deficiencies in processing complex data. At the full flowering stage, as shown in Figures 5(c) and 5(f), the stability of IBKA-RF is still prominent, and the training R 2 is 0.95926, with an RMSE of 0.82389. When tested, R 2 is 0.73644, and the corresponding RMSE is 1.853. IBKA-BP performs weakly at this stage. As shown in Figures 6(c) and 6(f), the training R 2 is 0.84694, with an RMSE of 1.5405. The test result R 2 drops to 0.55362, and the corresponding RMSE is 2.5634. As shown in Figures 7(c) and 7(f), the training result of IBKA-DELM is R 2 0.82627, with an RMSE of 1.6693. When tested, R 2 is 0.7678, and the corresponding RMSE is 1.7106. Although it is slightly better than IBKA-BP on the test set, it still faces accuracy challenges in complex data scenarios. In summary, IBKA-RF performs excellently in all stages and is suitable for processing complex non-linear data; IBKA-BP performs poorly due to the local optimum problem, while IBKA-DELM performs excellently in some stages, but still needs to improve its accuracy in complex data situations.
[0161] Therefore, as a preferred embodiment, the present application combines IBKA with the random forest (RF) algorithm to obtain the IBKA-RF model, which is used as the vegetation canopy nitrogen content prediction model and applied to the monitoring of the leaf nitrogen content (LNC) of jujube trees. The experimental results prove that this model can significantly improve the LNC prediction accuracy in different growth periods: in the leaf expansion period, the early flowering period and the full flowering period, the R2 of the training set of IBKA-RF all exceed 0.92, while the R 2 reach 0.6628, 0.6021 and 0.7364 respectively. The introduction of IBKA not only effectively optimizes the variable combination of high-dimensional spectral and texture features, but also significantly improves the fitting ability of the random forest in high-dimensional non-linear problems, solving the technical problem of insufficient correlation in nitrogen content prediction when directly using the RF (random forest) algorithm in section (5.1) above.
[0162] This embodiment only takes the jujube trees planted in Xinjiang region as an example. In fact, the vegetation canopy nitrogen content prediction model of the present application can be extended to other varieties and growth stages.
[0163] The above are only several embodiments of the present application, and do not impose any form of limitation on the present application. Although the present application is disclosed as above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art, without departing from the scope of the technical solution of the present application, makes some changes or modifications using the technical content disclosed above, which are equivalent to equivalent implementation cases and all fall within the scope of the technical solution.
Claims
1. A method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum, characterized in that: The method includes: According to the spectral image of vegetation, the spectral index and texture index are selected and calculated; The improved black-winged kite optimization algorithm is used to optimize the machine learning model, and the optimized model is trained using the nitrogen content, spectral index and texture index of the sample to obtain a vegetation canopy nitrogen content prediction model, the input of the model is the spectral index and the texture index, and the output is the vegetation nitrogen content prediction value; wherein the improved black-winged kite optimization algorithm includes: in the swarm intelligence optimization algorithm, the Sobol sequence is used for population initialization, the somersault foraging strategy is used for position update, and the multi-scale collaborative mutation mechanism is used to determine the escape position of each iteration; The spectral index and texture index of the vegetation to be tested are input into the vegetation canopy nitrogen content prediction model to obtain the predicted value of the canopy nitrogen content of the vegetation to be tested.
2. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 1, characterized in that: The spectral index adopts a vegetation index; the vegetation index includes at least one of the following: a soil line-based vegetation index, a ratio-based vegetation index, an atmospheric correction index, and a derivative vegetation index.
3. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 1, characterized in that: The acquisition of the texture index includes: extracting a number of texture features for each band from a multispectral remote sensing image of a vegetation crown leaf based on a gray level co-occurrence matrix; The texture features include at least one of the following: mean, variance, synergy, contrast, dissimilarity, information entropy, second-order moment and correlation.
4. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 1, characterized in that: The use of the Sobol sequence to initialize the population includes: The Sobol sequence is used to map the initialized population, as shown in the following formula: oh n =ω min +L n ×(ω max -oh min ) Among them, ω n is the initial position of the population, [ω min ,ω max ] is the value range of the optimal solution, L n ∈[0,1] is a random number generated by the Sobol sequence, n=1,...,N, and N is the initial population size.
5. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 1, characterized in that: The position update using the somersault foraging strategy is shown in the following formula: K e (t+1)=K e (t)+e×(λ1×K bs -λ2×K e (t)) Among them, K e (t+1) is the new position of the particle in the population, K e (t) is the current position of the particle in the population, e is the flip factor, K bs is the current global optimal probability, λ1 and λ2 are random numbers in [0,1] respectively.
6. According to the method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum in claim 1, the method of using a multi-scale collaborative variation mechanism to determine the escape position of each iteration specifically comprises introducing a Gaussian variation operator with J variation scales into the variation mechanism; The variance of the high-speed mutation operator is defined and initialized as shown in the following formula: in, represents the variance of the i-th variation scale at the t0-th iteration, and its initial value is defined as the entire solution space; During the iteration process, the variance is adjusted as follows: Sort particles in the population according to their fitness values; The sorted particles are divided into J subgroups, each of which contains Q = N / J particles. The fitness value of the j-th subgroup is calculated according to the following formula: Among them, k means selecting k subgroups from J subgroups; is the fitness value of subgroup j among k subgroups; is the fitness value of the qth particle in subgroup j, and the position update of particle q adopts the somersault foraging strategy; Dynamically adjust the standard deviation of each subgroup according to the fitness value of each subgroup; The position of the population is calculated based on the standard deviation of each subpopulation.
7. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 6, characterized in that: The dynamic adjustment includes: Get the maximum and minimum values of the fitness of all subgroups, denoted as Fit max and Fit min , as shown in the following formula: in, is the fitness value of the jth subgroup among the k subgroups; Then at the tth iteration, the standard deviation of each subgroup is as follows: in, is the standard deviation of subgroup j at the tth iteration.
8. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 7, characterized in that: Calculating the position of the population according to the standard deviation of each subgroup includes: The standard deviation of each subgroup is normalized to obtain the high-speed mutation operator, as shown in the following formula: Where D is the domain of the given optimization problem, t is the number of iterations, and the dynamic adjustment process is repeated until is the standard deviation of subgroup j at the tth iteration; is the high-speed mutation operator at the tth iteration; Define the location of the population as shown in the following formula: The above conditions are: conditional: others: Among them, t is the number of iterations, x q (t) is the position of the qth particle at the tth iteration, f(·) is the fitness function, and rand(0,1) is a random value between 0 and 1.
9. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 1, characterized in that: The machine learning model uses random forest or deep extreme learning machine.
10. The method for predicting nitrogen content in vegetation canopy based on remote sensing image spectrum according to claim 1, characterized in that: The method also includes testing and verifying the vegetation canopy nitrogen content prediction model: comparing the training and test results of the vegetation canopy nitrogen content LNC in different growth periods, wherein the growth periods include: leaf expansion period, initial flowering period, and peak flowering period.
Citation Information
Cited By
Urban cinnamomum camphora health monitoring method based on hyperspectrum of unmanned aerial vehicle
CN120997667A