Multi-modal network leaf area index inversion method based on thin cloud pollution treatment
By employing a multi-head attention mechanism and a multimodal network structure, the problem of data accuracy in leaf area index measurement under thin cloud pollution was solved, achieving high-precision leaf area index inversion, reducing computational costs, and improving the precision of agricultural production.
Patent Information
- Application Number
- CN202610365716.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-24
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies for measuring leaf area index under thin cloud pollution conditions suffer from data accuracy issues, especially in complex canopy crops. Traditional methods are low in accuracy and expensive in equipment, making it difficult to achieve high-precision leaf area index inversion.
A fully connected neural network based on a multi-head attention mechanism was used for thin cloud pollution data repair. Combined with a multimodal network structure, the optimal band combination was selected, and the leaf area index was inverted through the BLIP2-DMoE network, including data acquisition, preprocessing and inversion processes.
It effectively mitigates the impact of thin cloud noise, improves the accuracy of leaf area index retrieval, reduces computational costs, enhances adaptability in complex environments, and promotes the development of precision agriculture.
Smart Images

Figure CN122049412A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of agricultural information technology and remote sensing technology, and in particular to a multimodal network leaf area index inversion method based on thin cloud pollution treatment. Background Technology
[0002] Leaf area index (LAI) is a key indicator used in agricultural production and breeding to assess crop growth and photosynthetic efficiency. LAI is defined as the total area of all green leaves per unit area of ground. Traditional methods of LAI measurement fall into two main categories: direct measurement and indirect measurement. Direct measurement involves removing all leaves from the crop within a given area and measuring the area of each leaf. This is time-consuming, labor-intensive, and highly destructive, making it unsuitable for breeding and other research. Indirect measurement, on the other hand, mostly utilizes non-destructive equipment. For example, the LAI-2200c device captures multi-angle scattering of radiation from the canopy and measures LAI within a block based on the light interception rate above and below the canopy. However, this method often requires a considerable amount of time to photograph the area above and below the canopy. If cloud movement during this time causes inconsistent light intensity, the accuracy of the final LAI result will be low. Furthermore, the high cost of the equipment deters many agricultural breeding researchers.
[0003] Using instantaneous photography can solve the data error caused by excessive measurement time. However, when conducting field trials, researchers need to pay extra attention to the appearance of thin clouds. In most cases, thin clouds that appear suddenly cannot be detected in time, and it is too late to remedy the data problems when they are discovered during data processing.
[0004] In recent years, the development of deep learning technology has enabled intelligent solutions to many critical problems. For example, in the calculation of vegetation index, this patent utilizes a multispectral camera that captures images across 27 bands. The selection of key bands becomes particularly important. Assuming the green band ranges approximately from 500nm to 580nm, the device simultaneously has five bands: 500nm, 520nm, 540nm, 560nm, and 580nm. Each band selection results in numerous possible combinations. Therefore, utilizing deep learning technology, combined with the leaf area index to be retrieved, to automatically select the most suitable spectral band representing the desired result becomes crucial.
[0005] Traditional PROSAIL transport models often result in low accuracy for leaf area index (LAI) retrieval, particularly for crops with complex canopies. While machine learning can address these issues, its versatility across various crops is limited. When applying deep learning to LAI retrieval, especially with diverse crop types and growth stages, constructing a single-mode network model is optimal. In conclusion, eliminating the interference of thin cloud pollution on spectral data and achieving high-precision LAI retrieval in complex scenarios has become a critical technical challenge in precision agriculture. Therefore, proposing a multimodal network-based LAI retrieval method based on thin cloud pollution treatment has significant application value. Summary of the Invention
[0006] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a multimodal network leaf area index inversion method based on thin cloud pollution treatment.
[0007] To achieve the above objectives, the present invention provides the following solution: This invention provides a multimodal network leaf area index inversion method based on thin cloud pollution treatment, comprising: Step 1: Acquire multispectral data and RGB image data; Step 2: Repair the collected multispectral data based on the thin cloud pollution data repair model; Step 3: Based on band selection, choose the optimal band combination; Step 4: Perform preprocessing operations on the RGB image data and the repaired multispectral data; Step 5: Based on the leaf area index inversion model, perform leaf area index inversion using the preprocessed data and the optimal band combination.
[0008] Preferably, in step 1, the acquisition of multispectral data and RGB image data specifically includes: Multispectral data and RGB image data are collected using a handheld device. If the handheld device is not used for a long time or the sky environment changes significantly, the handheld device is calibrated using a standard reflectance correction white plate.
[0009] Preferably, in step 2, the collected multispectral data is repaired based on the thin cloud pollution data repair model, specifically as follows: A thin cloud pollution data restoration model was constructed using a fully connected neural network based on a multi-head attention mechanism. Construct a training dataset for a thin cloud pollution data restoration model; A data restoration model for thin cloud pollution was trained based on the training dataset; The collected multispectral data is input into the thin cloud pollution data restoration model to obtain the restored multispectral data.
[0010] Preferably, the training dataset for constructing the thin cloud pollution data repair model is as follows: Obtain a dataset polluted by thin clouds and a normal dataset whose shooting time differs from the dataset by less than five minutes. Construct a training dataset based on the two datasets.
[0011] Preferably, in step 3, the optimal band combination is selected based on band filtering, specifically as follows: Multiple vegetation indices were calculated from the restored multispectral data, and the correlation between each vegetation index and LAI was assessed. Representative band combinations are selected through a weighted score ranking mechanism.
[0012] Preferably, the representative band combination includes blue light band, green light band, red light band, red edge band, and near-infrared band.
[0013] Preferably, step 4: preprocessing the RGB image data and the repaired multispectral data, specifically as follows: Normalize the RGB image data and the repaired multispectral data; The outlier detection method is used to detect the normalized RGB image data and the repaired multispectral data. The outlier data is cleaned to obtain the preprocessed data. The vegetation index is calculated based on the preprocessed data and the optimal band combination.
[0014] Preferably, in step 5, the leaf area index inversion model is constructed based on the BLIP2-DMoE network structure.
[0015] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention provides a multimodal network leaf area index (LAI) inversion method based on thin cloud pollution treatment. The method includes acquiring multispectral data and RGB image data; repairing the acquired multispectral data based on a thin cloud pollution data repair model; selecting the optimal band combination based on band filtering; preprocessing the RGB image data and the repaired multispectral data; and performing LAI inversion based on the preprocessed data and the optimal band combination using the LAI inversion model. This invention effectively captures the complex relationships between different bands, better corrects data affected by thin cloud noise, and the proposed band selection function filters characteristic bands before large-scale data computation, reducing large-scale computation on edge devices and effectively saving computational costs. The invention aims to optimize algorithms, improve spectral data quality, enhance the adaptability of crops to multiple stages and complex environments, and promote the development of precision agriculture. It fills a gap in existing technologies and provides an innovative and substantial solution for agricultural modernization. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 An overall structure diagram provided for multimodal network inversion of leaf area index based on thin cloud pollution treatment; Figure 3 This describes the treatment method used in this experiment; Figure 4 The diagram shows the structure of a fully connected neural network based on a multi-head attention mechanism. Figure 5 The network structure of BLIP2-DMoE, a multimodal leaf area index inversion model for thin cloud pollution remediation. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0018] First, the invention will be introduced in general, which has two objectives: the first is to improve the quality of spectral data through a series of means, and the second is to propose an inversion model for leaf area index (LAI) of multiple crops based on thin cloud pollution remediation.
[0019] First, this invention designs a fully connected network with a multi-head attention mechanism to repair spectral data captured under thin cloud conditions that were not detected in time during field experiments. Leveraging the superiority of deep learning neural networks in feature learning and their advantages in nonlinear mapping, the missing spectral portions affected by thin cloud pollution are efficiently filled in, improving data quality and thus effectively enhancing the accuracy of leaf area index (LAI) retrieval. Subsequently, vegetation indices are calculated using multiple bands, and a linear regression model is used to evaluate the correlation between different combinations of these bands and the target LAI value, outputting a set of weighted recommended band combination values to optimize the computational efficiency of handheld devices. Since spectral equipment is often affected by factors such as equipment noise or outliers, a data cleaning function is added to the data preprocessing stage to remove some interfering data from the final results.
[0020] This invention also proposes a multimodal network for leaf area index (LAI) inversion. This network extracts features through spectral and RGB branches and fuses them using a self-attention mechanism, dynamically adjusting feature weights. This enables feature extraction from multiple vegetation index data, thereby achieving high-precision inversion of the LAI of the measured region.
[0021] Leaf area index (LAI) is a key indicator driving smart agriculture science, providing constructive guidance for agricultural management, such as nitrogen fertilizer application and irrigation area ratios. Traditional measurement methods suffer from destructive sampling and high costs in terms of financial and human resources. This invention utilizes a portable handheld device to simply and accurately retrieve the LAI of the measured area, lowering the technical barrier. This invention provides the experimental location and initial method, such as... Figure 3 As shown.
[0022] like Figure 1 and Figure 2 As shown, this invention provides a method for inverting the leaf area index of a multimodal network based on thin cloud contamination treatment, comprising: Step 1: Acquire multispectral data and RGB image data; Step 2: Repair the collected multispectral data based on the thin cloud pollution data repair model; Step 3: Based on band selection, choose the optimal band combination; Step 4: Perform preprocessing operations on the RGB image data and the repaired multispectral data; Step 5: Based on the leaf area index inversion model, perform leaf area index inversion using the preprocessed data and the optimal band combination.
[0023] In step 1, multispectral data and RGB image data are acquired using a handheld device, specifically as follows: If the device has not been used for an extended period or if there are significant changes in the ambient sky conditions, it is necessary to calibrate the device using a standard reflectance correction white plate.
[0024] In step 2, the collected multispectral data is repaired based on the thin cloud pollution data repair model, specifically as follows: like Figure 4 As shown, this invention proposes a fully connected neural network based on the multi-head self-attention (MHSA) mechanism, and constructs a thin cloud pollution data repair model based on it to repair data affected by thin cloud noise. The training dataset includes a dataset of thin cloud pollution and a normal dataset taken around the same time. By learning from the abnormal data of clear weather and thin cloud weather, the model is trained using 67 datasets of thin cloud pollution and 357 sets of normal data taken around the same time that are not affected by thin cloud pollution. The dataset is divided into an 80% training set and a 20% validation set. The cloud layer changes very quickly during the collection, so the time difference between the thin cloud pollution dataset and the normal dataset is basically within five minutes.
[0025] The following section provides a detailed introduction to the thin cloud pollution data restoration model: This network model processes data stepwise through fully connected layers and optimizes model parameters using backpropagation. During the data input phase, after inputting 27 bands, a linear transformation using the FC1 layer is employed for dimensionality mapping, providing the network with more mapping space and helping it learn more complex mapping relationships. This linear layer projects a hidden layer containing 64 neurons and utilizes the ReLU activation function. Subsequently, a multi-head attention mechanism is introduced to focus on the complex relationships between bands, enhancing the model's attention to bands affected by thin cloud contamination. This allows the model to learn the error range caused by thin cloud contamination and the inter-band dependencies within the spectrum, better repairing data affected by thin cloud noise. In the MHSA module of the multi-head attention mechanism, after linear transformation, each band is mapped using three independent weight matrices, yielding Q (query), K (key), and V (value). The dot product of Q and K is then calculated, as shown in the formula below. ; In the formula, The matrix representing the transpose of K. The dimension representing the key vector is 64 in this embodiment, and the multi-head attention mechanism includes 4 attention heads. The value is 16.
[0026] The calculation results of all attention heads are concatenated and output through a linear mapping. In the subsequent FC2 module, the network performs a nonlinear transformation on each band to improve the model's expressive power. This stage uses two fully connected layers: the first fully connected layer maps the features from 64 dimensions to 64 dimensions, and the second fully connected layer uses ReLU activation to enhance the network's nonlinearity. In the final output layer FC3, the 64-dimensional features are mapped back to the reflectance of 27 bands.
[0027] The above FC1, FC2 and FC3 are integrated into a three-layer MLP structure, as shown in the following formula: ; In this formula, This indicates the input spectral data contaminated by thin clouds; , , These represent the weight matrices of the first (FC1), second (FC2), and third (FC3) fully connected layers, respectively. , , These represent the bias vectors of the first, second, and third fully connected layers, respectively. This represents the set of all learnable parameters in the model.
[0028] During model training, the mean squared error (MSE) loss function is used to measure the error between the predicted and target values. Specifically, this embodiment constructs a sample-by-sample, band-by-band MSE loss function, the calculation formula of which is as follows: ; In the formula, N Indicates the total number of samples (or batch size) during the training process; i Indicates the index of the current sample; k The index represents the spectral band, with a value ranging from 1 to 27; This represents the true target value corresponding to the k-th band of the i-th sample.
[0029] Finally, after model training, the input thin cloud noise data is repaired through model inference. The repaired result output by the model is then de-standardized to restore the reflectivity to physical units. The calculation formula is as follows: ; In the formula, The first output of the model represents the... i The first sample k Standardized repair prediction values for each band; Indicates the first training set k The standard deviation of reflectance data for each band. Indicates the first training set k The mean of reflectance data for each band.
[0030] Its network structure diagram is as follows Figure 4 As shown.
[0031] This invention provides evaluation indicators before and after remediation, as shown in Table 1. In Table 1 and below, "0407" and "0420" refer to data collected on April 7, 2025, and April 20, 2025, respectively, at two different rice growth stages. By comparing data from different time dimensions, the generalization ability and robustness of the thin cloud pollution data remediation model proposed in this invention under different time spans and environmental changes are verified.
[0032] Table 1 Evaluation indicators before and after restoration
[0033] Table 1 shows that the deep learning model based on the multi-head self-attention mechanism performs excellently in correcting cloud-sky spectral data. Analysis of the MSE values with the current day's data and the 0420 data indicates that the corrected spectrum is basically consistent with the clear-sky data. The Pearson correlation coefficient also proves that the corrected data has a strong correlation with the clear-sky data. Furthermore, comparing the spectral angle mapping SAM value with the 0407 and 0420 data demonstrates that the corrected spectrum has a high angular correlation with the clear-sky spectrum. The calculation methods for MSE, Pearson correlation coefficient r, and SAM values are shown below: ; ; ; In the formula, Indicates the first i The first sample k The true values of normal spectral reflectance for each band under clear weather conditions. n Indicates the total number of samples; t The index variable representing the sample; and Let represent the predicted value and the actual observed value of the t-th data to be evaluated, respectively; and These represent the average of the predicted values and the average of the actual observed values in all the evaluation data, respectively. This represents the spectral angle between the predicted spectral eigenvector and the true spectral eigenvector; and They represent the first i The repair prediction spectral feature vector and the true spectral feature vector of a sample on a sunny day. Indicates the firsti The first sample k The restoration prediction spectral feature vectors corresponding to each band.
[0034] Table 2 Evaluation of LAI inversion experimental results before and after restoration
[0035] Finally, experiments were conducted comparing the corrected and uncorrected data, and the results were analyzed using several machine learning models. As shown in Table 2, most models showed a significant improvement in R² after correction. For example, the R² of the Ridge Regression model increased from 0.17 to 0.56, and the R² of the SVR (RBF Kernel) model increased from 0.23 to 0.56. This indicates that the corrected data can be effectively fitted by these models, demonstrating the effectiveness of spectral correction. Future research can further enhance the model's performance and application scope by increasing the sample size and modeling time dynamics.
[0036] In step 3, based on band selection, the optimal band combination is chosen, specifically as follows: Since this device can collect data from 27 wavelengths, in practical applications, when deriving reflectance formulas, most vegetation index formulas only require two or three wavelengths. These wavelengths are often referred to as the blue light band, green light band, red light band, red-edge band, and near-infrared band. These wavelengths are usually a range; for example, the red light band ranges from approximately 620–700 nm. The device used in this invention can collect data from 400 nm to 920 nm, averaging once every 20 nm. The specific wavelength range chosen as the representative wavelength for calculating the vegetation index, and the combination of wavelengths used, also need to be considered.
[0037] This step assigns weights to the R² values of each vegetation index and band combination. Combinations with higher weights are considered to have a positive correlation with the accuracy of LAI inversion. The frequency of occurrence of these high-performing bands is then counted, and their importance is statistically analyzed based on their weights. Finally, the most relevant bands for each type are obtained. These bands can be effectively used in the derivation of vegetation indices for LAI inversion. The bands selected in this experiment are shown in Table 3.
[0038] Table 3 Band Selection Results
[0039] In step 4, preprocessing operations are performed on the RGB image data and the repaired multispectral data, specifically as follows: Acquire RGB image data and restored multispectral data; Normalize the RGB image data and the repaired multispectral data; The acquired RGB images were first preprocessed, including normalization, noise suppression, texture and local feature enhancement, and size standardization. Specifically, the gray values of each channel were linearly mapped to a unified range to reduce brightness differences caused by different imaging conditions, exposure parameters, and sensor gains, thereby improving the comparability between images. Based on the normalization results, noise suppression was performed using mean filtering and median filtering to suppress random and spike noise. Subsequently, texture and local feature enhancement were applied to the noise-suppressed images. Contrast enhancement methods were used to moderately enhance the edges, texture structure, and subtle gradient changes of crop leaves to improve the sensitivity of subsequent anomaly detection and feature extraction.
[0040] For spectral data, an outlier detection method based on physical thresholds is used to process reflectance data in 27 bands. The preset physical upper limit threshold is 1.0. Overexposure or specular reflection usually causes reflectance values to exceed this upper limit. These outlier data are marked as outlier data and deleted to avoid affecting subsequent data analysis and modeling.
[0041] Subsequently, vegetation indices were calculated using the specific wavelengths selected in step 3. A total of 28 vegetation indices were selected. In the table, NIR represents the near-infrared band, RE represents the red-edge band, R represents the red light band, G represents the green light band, and B represents the blue light band. Some formulas with subscripts, such as R700, indicate that the current specified wavelength band of 700nm is used for calculation, and the representative wavelength band calculated in step 3 is not used. For the Soil Adjusted Vegetation Index (SAVI). L This represents the soil brightness adjustment coefficient, used to reduce the influence of soil background. In this embodiment... L The value is 0.5; for the Wide Dynamic Range Vegetation Index (WDRVI). α This represents the weighted parameter used to adjust the reflectivity in the near-infrared band, and its value ranges from 0 to... α <1. The specific vegetation indices and calculation formulas are shown in Table 4.
[0042] Table 4. Vegetation Indices and Calculation Formulas
[0043]
[0044] In step 5, the leaf area index (LAI) is inverted based on the preprocessed data using the leaf AAI inversion model. Specifically: A leaf area index inversion model was constructed based on the BLIP2-DMoE network structure. The BLIP-2 network model was selected as the baseline model. The BLIP-2 network model is a multimodal framework, which is used in this invention to jointly use spectral data and image data. Subsequently, the MoE mechanism was used for feature weighting and fusion, and the DMPL (Dynamic Prototype Layer) mechanism was used to learn the dynamic prototype to enhance the multimodal representation capability of the model. The model is divided into two branches based on the input. In the Image Token Encoder branch, the input image is a 3D image. ResNet-101 is used to extract image feature information. First, a high-dimensional feature is obtained through a convolutional layer, and then it is projected through a fully connected layer to extract the feature. In the other branch, the Spectral Token Encoder, the spectral data of each band is mapped to a total of 55-dimensional vectors, which include reflectance data for 27 spectral bands and vegetation indices for 28 spectral bands. These spectral data include spectral reflectance data and vegetation index data.
[0045] In the MoE Block, image and spectral tokens are concatenated and fused using the MoE mechanism. MoE comprises four experts (linear layers) and a gating network, allowing the model to dynamically select among multiple experts (sub-networks). The gating network selects the appropriate expert for each input. Softmax is used to assign weights to select the expert output, followed by multi-head self-attention to enhance feature interactions, ultimately outputting fused tokens. A MHA self-attention module is then inserted, fusing the self-attention-computed features with the expert outputs to obtain the final MoE module output.
[0046] In the DPML module, the principle is to enhance the model's representational ability by learning prototypes. Prototypes are fixed patterns learned during training. DMPL calculates the matching degree between input features and prototypes using cosine similarity, generates soft-assigned weights, and then weights and fuses prototype features to obtain enhanced feature representations.
[0047] Calculate the cosine similarity between the input features and each prototype: ; The method for calculating weights based on similarity: ; All prototypes are weighted and fused using calculated weights: ; Finally, during each training session, the prototype is updated in real time, calculated using the following formula: ; In the formula, F This represents the feature vector fused from the image and spectrum, output by the front-end module (MoE module) and currently input. and They represent the first k The and the first i Feature vectors of each prototype; K This represents the total number of prototypes learned. Representing input features F With the k A prototype Cosine similarity between them; This indicates that the value allocated to the first element after calculation by the Softmax function is... k Soft weighting of each prototype; This represents the enhanced multimodal feature representation obtained after weighted fusion of prototypes; This represents the momentum update coefficient during the prototype dynamic update process.
[0048] BLIP2-DMoE is a multimodal deep learning framework that combines the BLIP-2 framework, the Hybrid Expert (MoE) mechanism, and Dynamic Multimodal Prototype Learning (DMPL). By combining spectral and image data, the model can simultaneously utilize these two different types of information to predict LAI, thereby improving prediction accuracy. The network structure diagram is shown below. Figure 5 As shown, this experiment compared the LAI inversion results using the same data source and under the same experimental environment. Traditional machine learning models were used to combine spectral data with vegetation indices: Linear Regression, RidgeRegression, Random Forest, Gradient Boosting, SVR (RBF Kernel), K-Nearest Neighbors (KNN), CatBoost, LightGBM, and XGBoost. Simultaneously, a scheme combining RGB images and spectral data with vegetation indices was implemented using deep learning models: a multimodal network based on DenseNet, EfficientNet, and VGG16 backbone. The comparison results are shown in Table 4. Additionally, more mainstream multimodal networks were used for comparison, such as MMBT, CLIP, CoCa, ViLT, VisualBERT, UNITER, OSCAR, ALBEF, BEiT-3, SigLIP, and LLaVA-NeXT. The comparison results are shown in Table 5.
[0049] Table 5 Comparison of BLIP2-DMoE multimodal network with other methods
[0050] As shown in Table 5, in the comparison of machine learning models, CatBoost (R² = 0.69) and K-NearestNeighbors (R² = 0.68) performed best, outperforming some tree-structured models such as Random Forest, GradientBoosting, LightGBM, and XGBoost. In the comparison, we found that Ridge Regression performed the worst, with an R² of 0.37, possibly because the linear model failed to capture the linear relationship between the ground truth LAI and the spectral bands. In contrast, the BLIP2-DMoE method significantly improved performance, increasing its R² from CatBoost's 0.69 to 0.81, and achieving the best performance across all four evaluation metrics. Its superior performance (R² = 0.81, RMSE = 0.34, MAE = 0.29, MAPE = 8.56%) surpassed both machine learning models and multimodal models based on DenseNet, EfficientNet, and VGG16. This experiment demonstrates that combining spectral reflectance with vegetation indices and RGB images can effectively improve the accuracy of LAI inversion, validating the effectiveness of deep learning multimodal networks. Future work could focus on improving the model's robustness and generalization capabilities to adapt it to various crops. Table 6 Comparison of BLIP2-DMoE multimodal network with other methods
[0051] As can be seen from Table 6, BLIP2-DMoE, through the combination of BLIP-2 with MoE dynamic expert selection and DMPL prototype aggregation, can more fully capture the complementary relationship between spectrum, texture and canopy structure, thus leading in the overall fitting (R2) and squared error (RMSE) dimensions; while MAPE ranks second, which means that there is still room for improvement in extreme or low-value samples in terms of relative error. Overall, BLIP2-DMoE has achieved a robust and comprehensive lead in key indicators.
[0052] Embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0053] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0054] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0055] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0056] Contents not described in detail in this specification are prior art known to those skilled in the art. It is hereby indicated that the above description is intended to help those skilled in the art understand this invention, but does not limit the scope of protection of this invention. Any equivalent substitutions, modifications, improvements, or simplifications of the above descriptions that do not depart from the essential content of this invention fall within the scope of protection of this invention.
Claims
1. A method for inverting the leaf area index of a multimodal network based on thin cloud contamination treatment, characterized in that, include: Step 1: Acquire multispectral data and RGB image data; Step 2: Repair the collected multispectral data based on the thin cloud pollution data repair model, which is constructed based on a fully connected neural network with a multi-head attention mechanism; Step 3: Based on band selection, choose the optimal band combination; Step 4: Perform preprocessing operations on the RGB image data and the repaired multispectral data; Step 5: Based on the leaf area index inversion model, the leaf area index is inverted according to the preprocessed data and the optimal band combination. The leaf area index inversion model is constructed based on the BLIP2-DMoE network structure.
2. The method according to claim 1, characterized in that, In step 1, multispectral data and RGB image data are acquired, specifically as follows: Multispectral data and RGB image data are collected using a handheld device. If the handheld device is not used for a long time or the sky environment changes significantly, the handheld device is calibrated using a standard reflectance correction white plate.
3. The method according to claim 2, characterized in that, In step 2, the collected multispectral data is repaired based on the thin cloud pollution data repair model, specifically as follows: A thin cloud pollution data restoration model was constructed using a fully connected neural network based on a multi-head attention mechanism. Construct a training dataset for a thin cloud pollution data restoration model; A data restoration model for thin cloud pollution was trained based on the training dataset; The collected multispectral data is input into the thin cloud pollution data restoration model to obtain the restored multispectral data.
4. The method according to claim 3, characterized in that, The training dataset for constructing the thin cloud pollution data restoration model is as follows: Obtain a dataset polluted by thin clouds and a normal dataset whose shooting time differs from the dataset by less than five minutes. Construct a training dataset based on the two datasets.
5. The method according to claim 4, characterized in that, In step 3, based on band selection, the optimal band combination is chosen, specifically as follows: Multiple vegetation indices were calculated from the restored multispectral data, and the correlation between each vegetation index and LAI was assessed. Representative band combinations are selected through a weighted score ranking mechanism.
6. The method according to claim 5, characterized in that, The representative band combination includes blue light band, green light band, red light band, red edge band, and near-infrared band.
7. The method according to claim 6, characterized in that, Step 4: Perform preprocessing operations on the RGB image data and the repaired multispectral data, specifically: Normalize the RGB image data and the repaired multispectral data; The outlier detection method is used to detect the normalized RGB image data and the repaired multispectral data. The outlier data is cleaned to obtain the preprocessed data. The vegetation index is calculated based on the preprocessed data and the optimal band combination.