A Method for Inverting Physicochemical Parameters of Plant Leaves Based on Two-Dimensional Spectral Representation and Multi-Task Deep Learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-08-14
AI Technical Summary
[0009]尽管此类技术在其他领域中已有初步应用,现阶段仍未有研究系统探讨将一维光谱数据重构为二维图像的方法在植物叶片理化参数反演中的适配性与有效性,亦未形成与植物光谱特性相契合的转换策略与建模方法
[0126]本发明将一维光谱数据转化为二维图像形式,有效挖掘光谱曲线的结构特征与波段间非线性关联,解决同谱异物现象导致的参数反演干扰问题;采用CNN与Transformer级联的共享编码器架构,结合旋转位置编码技术,实现局部精细特征提取与全局上下文关系的精确建模,增强模型对光谱-理化参数复杂映射关系的解析能力;通过任务专属注意力机制与混合加权损失函数,在硬参数共享基础上实现多理化参数任务的协同优化,解决参数间光谱响应重叠导致的反演精度瓶颈,提升模型对叶绿素、类胡萝卜素等易混淆参数的区分能力,最终实现植物叶片理化参数的高精度、鲁棒性同步反演,为森林生态监测与资源管理提供高效无损的技术支撑。
Smart Images

Figure CN121659710B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for inverting physicochemical parameters of plant leaves based on two-dimensional spectral characterization and multi-task deep learning, belonging to the field of plant physiological detection technology. Background Technology
[0002] Forest ecosystems, as a vital component of Earth's terrestrial ecosystems, play a crucial role in global carbon cycling, water regulation, and biodiversity maintenance. Physicochemical parameters of plant leaves (such as chlorophyll, carotenoids, water content, and nitrogen content) not only reflect the physiological state of individual plants but are also important indicators for assessing stand health, productivity, and environmental adaptability. These parameters are closely related to key ecological processes such as photosynthesis, energy conversion, and nutrient cycling. Therefore, accurately obtaining leaf physicochemical parameters can provide fundamental support for forest carbon sequestration capacity assessment, ecosystem modeling, species functional diversity analysis, and dynamic monitoring of forest resources.
[0003] Traditional methods for obtaining physicochemical parameters of plant leaves mainly rely on manual sampling and laboratory chemical analysis. Although these methods have high precision, they suffer from problems such as high labor costs, poor timeliness, and high destructiveness. They are difficult to achieve efficient, continuous, and non-destructive monitoring of large areas and are not conducive to long-term dynamic tracking research.
[0004] Hyperspectral imaging technology provides a new technical approach for the non-destructive, high-throughput, and quantitative inversion of physicochemical parameters of plant leaves. In recent years, parameter inversion methods integrating hyperspectral data and machine learning models have gradually become a research hotspot. Among the existing technologies, Chinese invention patent CN202410175222.0, entitled "A Quantitative Inversion Method for Photosynthetic Pigments in Catalpa Tree Leaves Based on Hyperspectral Reflectance," discloses the acquisition of hyperspectral reflectance data of Catalpa tree leaves, as well as spectral preprocessing and feature analysis. It identifies specific bands and characteristic spectral parameters that are most sensitive to the content of chlorophyll a, chlorophyll b, carotenoids, and total chlorophyll. Based on the photosynthetic pigment content and optimized spectral vegetation index parameters, a machine learning training set is constructed. Furthermore, a high-precision and high-stability quantitative inversion model is established through a random forest regression algorithm, realizing the synchronous and rapid inversion of the content of multiple photosynthetic pigments in Catalpa tree leaves.
[0005] With the development of deep learning and multi-task learning, it has become possible to simultaneously invert multiple leaf physicochemical parameters based on a shared feature space. Multi-task learning can effectively explore the potential correlations between various parameters, improve the synergy of modeling and overall prediction performance, and significantly enhance the efficiency and accuracy of hyperspectral inversion, showing promising development prospects. Chinese invention patent CN202410630586.3, entitled "A Method for Inverting Soil Physicochemical Parameters Based on a Multi-Task Deep Convolutional Neural Network," discloses a method for preprocessing and feature enhancement of soil hyperspectral image data. It utilizes a multi-task deep convolutional neural network to automatically extract deep nonlinear features from soil spectral data. Based on the complex mapping relationship between multiple soil physicochemical parameters and deep spectral features, it constructs an inversion model that shares underlying features and outputs multiple parameters in parallel. Furthermore, it synchronously trains multiple inversion tasks through joint loss function optimization and weight sharing mechanisms. This network model enables high-precision and high-efficiency simultaneous inversion and prediction of various key soil physicochemical parameters. Reference: CHERIF H, MEZGHANII, FOHLER B, et al. From spectra to plant functional traits: Transferable multi-trait models from heterogeneous and sparse data[J]. Methods in Ecology and Evolution[J]. 2023, 14(5): 1143-1159. This paper proposes a multi-trait joint modeling method based on convolutional neural networks (CNN), integrating heterogeneous and sparse spectral-trait pairing data from 42 different ecosystems and sensors. The study utilizes a weakly supervised learning strategy to address the issues of missing and incomplete labels, and reveals the key bands on which different traits depend through SHAP interpretive analysis. The results show that this method significantly outperforms the traditional PLSR model on most traits and exhibits good transferability on external datasets.
[0006] Currently, most mainstream hyperspectral modeling methods still directly model using one-dimensional spectral data as input, which has significant shortcomings in expressing inter-spectral structural features and nonlinear interactions between bands. One-dimensional spectral data ignores the structural features inherent in the spectrum as a continuous curve, such as the shape of absorption valleys and the position of reflection peaks, which are intrinsic relationships between bands. This lack of representation of nonlinear relationships between spectra makes it difficult for models to understand spectral information from a higher dimension, resulting in poor robustness and generalization ability in complex ecological environments. In addition, there is overlap in the spectral responses of different physicochemical parameters. Taking pigments in healthy plants as an example, the absorption peaks of chlorophyll and carotenoids overlap in the blue light band. In addition, chlorophyll content is dominant, and the proportion of carotenoid content is very small. This causes the weak signal of carotenoids to be masked in the one-dimensional feature space, making remote sensing inversion of carotenoid content a challenge.
[0007] Research on converting one-dimensional time series data into two-dimensional images to enhance feature representation has made progress in fields such as medical signal processing, soil property estimation, and precision agriculture. These methods map the original one-dimensional sequence into two-dimensional structural images, such as Gram angle fields or Markov transition fields, reconstructing its spatial structure while preserving the original information. This effectively improves the model's ability to recognize local patterns, sequence correlations, and structural features.
[0008] Reference: JIN X, ZHOU J, RAO Y, et al. An innovative approach for integrating two-dimensional conversion of Vis-NIR spectra with the SwinTransformer model to leverage deep learning for predicting soil properties[J]. 2023, 436: 116555. This paper proposes a modeling method that combines two-dimensional conversion of visible-near-infrared spectra with SwinTransformer for predicting various soil properties. The study utilizes methods such as Cutting Reshape, GADF, GASF, and MTF to map one-dimensional spectra into two-dimensional images, and captures the deep structural features of the spectral images through a pre-trained Swin Transformer. The results show that the GADF-based and Swin Transformer-based approach outperforms traditional PLSR and one-dimensional CNN models in predicting soil organic carbon, nitrogen content, cation exchange capacity, pH, and particle composition, validating the potential and superiority of integrating two-dimensional spectral conversion with Swin Transformer in soil spectral modeling.
[0009] Although such technologies have been preliminarily applied in other fields, no research has yet systematically explored the adaptability and effectiveness of methods for reconstructing one-dimensional spectral data into two-dimensional images in the inversion of physicochemical parameters of plant leaves. Furthermore, no conversion strategy or modeling method has been developed that aligns with the spectral characteristics of plants. Therefore, there is an urgent need to propose a technical solution that combines the characteristics of hyperspectral data with the advantages of image modeling to address the bottlenecks in feature representation and task coupling of existing one-dimensional modeling methods, thereby improving the accuracy of inversion of physicochemical parameters of plant leaves. Summary of the Invention
[0010] To address the problems existing in the background technology, this invention provides a method for inverting the physicochemical parameters of plant leaves based on two-dimensional spectral characterization and multi-task deep learning.
[0011] To achieve the above objectives, the present invention adopts the following technical solution: a method for inverting physicochemical parameters of plant leaves based on two-dimensional spectral characterization and multi-task deep learning, the method comprising the following steps:
[0012] S1: Construction of a hyperspectral plant leaf dataset;
[0013] S2: Preprocessing of hyperspectral plant leaf dataset;
[0014] S3: Reconstruct the preprocessed one-dimensional spectral vector into a two-dimensional image;
[0015] S4: Using the two-dimensional image generated in S3 as input, construct a multi-task learning model and train and evaluate it;
[0016] S5: Apply the trained model to the test set to obtain the inversion results of the blade physicochemical parameters.
[0017] Furthermore, step S1 includes the following steps:
[0018] S101: Acquire leaf spectral information and calculate wavelength using a hyperspectral analyzer. reflectivity at :
[0019] (1)
[0020] In formula (1):
[0021] The leaf sample at wavelength The original reading at the location;
[0022] Is it a standard whiteboard at wavelength? The reading at the location;
[0023] It is dark current at wavelength The reading at the location;
[0024] S102: Physicochemical parameters of leaf samples were determined by laboratory chemical analysis.
[0025] S103: Match the spectral data obtained in S101 with the physicochemical parameters obtained in S102 to complete the construction of the hyperspectral plant leaf dataset.
[0026] Furthermore, step S2 includes the following steps:
[0027] S201: Divide the dataset in an 8:2 ratio, take 20% of the data as the test set, and the remaining data as the training set and validation set;
[0028] S202: Use SNV to independently transform the spectrum of each sample in the dataset:
[0029] (7)
[0030] In equation (7):
[0031] It is the transformed spectral data;
[0032] It is the raw spectral data;
[0033] It is the average value of the original spectrum;
[0034] It is the standard deviation of the original spectrum;
[0035] S203: The transformed spectral data is smoothed using an SG filter.
[0036] (8)
[0037] In equation (8):
[0038] It is the radius of the window, and the size of the sliding window is... ;
[0039] It is smoothed in position New data points;
[0040] In position The original data points within the surrounding window;
[0041] It is the first The polynomial fitting coefficients for each point.
[0042] Furthermore, step S3 includes the following steps:
[0043] S301: Scale the preprocessed one-dimensional spectral data obtained in S2 to the interval [-1, 1] to obtain a new spectral sequence. Calculate each data point Corresponding scaled data points :
[0044] (9)
[0045] In equation (9):
[0046] and These are spectral sequences The maximum and minimum values in;
[0047] S302: Scale the data points Considered as an angle The cosine value is obtained, and the angle is calculated using the inverse cosine function. :
[0048] , , (10)
[0049] S303: Angle conversion using spectral-image conversion rules By constructing a two-dimensional Gram matrix using the converted angle information, the relationship between adjacent bands is transformed into a local spatial structure relationship in the image, resulting in a two-dimensional image that corresponds one-to-one with the blade sample and contains structured information.
[0050] Furthermore, step S4 includes the following steps:
[0051] S401: Extracting depth spatial features from a two-dimensional image obtained after converting one-dimensional spectral data using a two-dimensional convolutional neural network EfficientNetV2-S;
[0052] S402: Input the feature sequence extracted by the convolutional neural network into the Transformer encoder integrated with RoPE for shared feature extraction across all tasks;
[0053] Assuming the wavelength sequence is in the 1st... The feature vector at each position is Define the base velocity of rotation for each eigenvector. :
[0054] (twenty one)
[0055] In equation (21):
[0056] It is the total dimension of the feature vector;
[0057] It is an index of feature pairs;
[0058] eigenvectors The dimensions are grouped pairwise, that is... , , ..., for the first Group feature pairs According to its location and base speed Perform two-dimensional rotation:
[0059] (twenty two)
[0060] In equation (22):
[0061] These are new feature pairs obtained after rotation;
[0062] S403: Design TSA;
[0063] The output sequence of the Transformer encoder and learnable task query matrix As input, perform dot product attention operations to generate... A matrix composed of task-specific feature vectors :
[0064] (twenty three)
[0065] In equation (23):
[0066] This is used to indicate normalization processing;
[0067] It's a query about attention mechanisms;
[0068] It is the key to the attention mechanism;
[0069] It is the value of the attention mechanism;
[0070] S404: Design a hybrid weighted loss function;
[0071] Calculate the base loss for each individual task:
[0072] (twenty four)
[0073] In equation (24):
[0074] It represents the number of samples for the task;
[0075] It is a task The true value;
[0076] It is a task The predicted value;
[0077] The model assigns learnable parameters to each task. Then calculate the total loss for all tasks. :
[0078] (25)
[0079] In equation (25):
[0080] B is the batch size;
[0081] The weights are fixed based on prior knowledge.
[0082] The model is for the task The automatically learned log-variance;
[0083] It is a dynamic weight;
[0084] S405: Introduces an early stopping mechanism to prevent overfitting.
[0085] Furthermore, when the input for the depth space feature extraction described in S401 is multiple images, any of the following fusion strategies shall be adopted:
[0086] Blending Strategy 1: Pixel Weighting;
[0087] The single-channel two-dimensional image matrices of GASF, GADF, and MTF are weighted according to a preset formula. Perform linear weighted summation to generate a new single-channel grayscale image. As the final model input:
[0088] (13)
[0089] In equation (13):
[0090] , , It is a single-channel two-dimensional image matrix of GASF, GADF, and MTF;
[0091] The prediction process chain of the model is as follows:
[0092] (14)
[0093] In equation (14):
[0094] This is the basic feature extraction performed by EfficientNetV2-S on the input image;
[0095] It is the global relational modeling performed by the Transformer encoder on the features extracted by the CNN;
[0096] It is a task-specific attention head that performs task-level filtering on the features output by the Transformer;
[0097] Each of the multiple task prediction branches corresponds to a prediction task for a physicochemical parameter;
[0098] These are the predicted results of physicochemical parameters;
[0099] Fusion Strategy 2: Pixel Stitching;
[0100] The three single-channel images, GASF, GADF, and MTF, are stacked along the new channel dimension to form a three-channel pseudo-color image. , For the height of the image, Use the width of the image as the final model input:
[0101] (15)
[0102] The prediction process chain of the model is as follows:
[0103] (16)
[0104] Fusion Strategy 3: Feature Fusion;
[0105] The model has three independent feature extractors Each is applied to the corresponding input image:
[0106] (17)
[0107] In equation (17):
[0108] It is the extracted high-level feature vector;
[0109] The fusion is completed at the feature layer through a splicing operation:
[0110] (18)
[0111] The prediction process chain of the model is as follows:
[0112] (19)
[0113] When fully unfolded, it looks like this:
[0114] (20)
[0115] Furthermore, step S5 includes the following steps:
[0116] S501: The training and validation sets are trained using a five-fold cross-validation method, and the optimal model for each fold is saved.
[0117] S502: Use the test set as input to the optimal model saved in each fold, and employ the coefficient of determination. Root mean square error Relative root mean square error based on 95% of the data range and relative prediction bias As an evaluation result of the model's effectiveness:
[0118] (26)
[0119] (27)
[0120] (28)
[0121] (29)
[0122] In equations (28)-(29):
[0123] It represents the range of data remaining after removing the minimum and maximum values (2.5%) from the dataset.
[0124] It is the standard deviation of the true value.
[0125] Compared with the prior art, the beneficial effects of the present invention are:
[0126] This invention transforms one-dimensional spectral data into two-dimensional image form, effectively mining the structural features of spectral curves and nonlinear correlations between bands, and solving the parameter inversion interference problem caused by heterogeneous objects in the same spectrum. It employs a shared encoder architecture cascaded with CNN and Transformer, combined with rotational position encoding technology, to achieve precise local feature extraction and accurate modeling of global contextual relationships, enhancing the model's ability to analyze complex spectral-physicochemical parameter mapping relationships. Through a task-specific attention mechanism and a hybrid weighted loss function, it achieves collaborative optimization of multiple physicochemical parameter tasks based on hard parameter sharing, solving the inversion accuracy bottleneck caused by overlapping spectral responses between parameters, and improving the model's ability to distinguish easily confused parameters such as chlorophyll and carotenoids. Ultimately, it achieves high-precision and robust synchronous inversion of plant leaf physicochemical parameters, providing efficient and lossless technical support for forest ecological monitoring and resource management. Attached Figure Description
[0127] Figure 1 This is a flowchart of the present invention;
[0128] Figure 2 This is a schematic diagram illustrating the conversion of one-dimensional spectral data into a two-dimensional spectral image;
[0129] Figure 3 This is a structural diagram of a multi-task physicochemical parameter inversion model for plant leaves based on spectral-image conversion. Detailed Implementation
[0130] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0131] A method for inverting physicochemical parameters of plant leaves based on two-dimensional spectral characterization and multi-task deep learning, the method comprising the following steps:
[0132] S1: Construction of a hyperspectral plant leaf dataset;
[0133] S1 includes the following steps:
[0134] S101: Acquire spectral information of the leaf blade in the 400-2500nm band using a hyperspectral analyzer and calculate the wavelength. reflectivity at :
[0135] (1)
[0136] In formula (1):
[0137] The leaf sample at wavelength The original reading at the location;
[0138] Is it a standard whiteboard at wavelength? The reading at the location;
[0139] It is dark current at wavelength The reading at the location;
[0140] S102: The main physicochemical parameters of leaf samples were determined by laboratory chemical analysis.
[0141] Preferably, chlorophyll a, chlorophyll b, and carotenoids can be measured using a spectrophotometer, extracted with 95% ethanol, and measured. =663nm =645nm and absorbance and leaf area at a wavelength of 470 nm The specific calculation formula is as follows:
[0142] Chlorophyll a content ( The formula for calculating ) is:
[0143] (2)
[0144] Chlorophyll b content ( The formula for calculating ) is:
[0145] (3)
[0146] Carotenoid content ( The formula for calculating ) is:
[0147] (4)
[0148] In equations (2)-(4):
[0149] It is the leaf area;
[0150] It is the absorbance at the corresponding wavelength;
[0151] It is the volume of the extract;
[0152] It is the concentration of chlorophyll a;
[0153] It is the concentration of chlorophyll b;
[0154] Equivalent water thickness The formula for calculating (mm) is:
[0155] (5)
[0156] In equation (5):
[0157] It is the fresh weight of the leaves;
[0158] It is the dry weight of the leaf;
[0159] The Kjeldahl method was used to measure the nitrogen content per unit area of leaves. ):
[0160] (6)
[0161] In formula (6):
[0162] It is the standard acid concentration;
[0163] This refers to the amount of standard acid consumed in the sample titration.
[0164] This is the amount of standard acid consumed in the titration of the blank sample;
[0165] It refers to sample quality;
[0166] S103: Strictly match the spectral data obtained in S101 with the physicochemical parameters obtained in S102 to complete the construction of the hyperspectral plant leaf dataset.
[0167] S2: Preprocessing of hyperspectral plant leaf dataset;
[0168] S2 includes the following steps:
[0169] S201: Divide the dataset in an 8:2 ratio, take 20% of the data as the test set, and the remaining data as the training set and validation set;
[0170] S202: SNV (Standard Normal Variance) is used to independently transform the spectrum of each sample in the dataset to eliminate or reduce baseline shift and scale scaling effects caused by irrelevant interference factors in spectral measurements, thereby improving the accuracy and stability of modeling. SNV processes each spectrum independently by subtracting its own mean and dividing by its own standard deviation to transform each spectrum to a relatively uniform scale.
[0171] (7)
[0172] In equation (7):
[0173] It is the transformed spectral data;
[0174] It is the raw spectral data;
[0175] It is the average value of the original spectrum;
[0176] It is the standard deviation of the original spectrum;
[0177] S203: The transformed spectral data is smoothed using an SG (Savtzky-Golay) filter, with a window size (window_length) of 11 and a polynomial order (polyorder) of 3. SG smoothing is a classic method in spectral preprocessing used for noise reduction and feature preservation. It smooths the spectral curve by fitting local data points with a polynomial, while preserving key features such as peaks and valleys as much as possible, avoiding feature loss due to over-smoothing.
[0178] (8)
[0179] In equation (8):
[0180] It is the radius of the window, and the size of the sliding window is... ;
[0181] It is smoothed in position New data points;
[0182] In position The original data points within the surrounding window;
[0183] It is the first The polynomial fitting coefficients for each point.
[0184] S3: Reconstruct the preprocessed one-dimensional spectral vector into a two-dimensional image. The aim is to explicitly encode the dynamic relationships and correlations implicit in the one-dimensional sequence data into the spatial structure of the two-dimensional image, so as to facilitate feature extraction by the subsequent deep learning network.
[0185] S3 includes the following steps:
[0186] S301: Scale the preprocessed one-dimensional spectral data obtained in S2 to the interval [-1, 1] to obtain a new spectral sequence. Calculate each data point Corresponding scaled data points :
[0187] (9)
[0188] In equation (9):
[0189] and These are spectral sequences The maximum and minimum values in;
[0190] S302: Scale the data points Considered as an angle The cosine value is obtained, and the angle is calculated using the inverse cosine function. :
[0191] , , (10)
[0192] S303: Angle conversion using spectral-image conversion rules By constructing a two-dimensional Gram matrix using the converted angle information, the relationship between adjacent bands is transformed into a local spatial structure relationship in the image, resulting in a two-dimensional image that corresponds one-to-one with the blade sample and contains structured information.
[0193] Existing spectral-image conversion rules include direct reshaping, continuous wavelet transform (CWT), short-time Fourier transform (STFT), Gram angle field (GAF), and Markov transfer field (MTF). These methods aim to explicitly encode the implicit and complex dependencies in one-dimensional spectral sequences into the texture and structure of two-dimensional images.
[0194] Combination Figure 2 The three conversion methods used in this invention will be described.
[0195] GAF (Gramian Angular Fields) includes GASF (Gramian Angular Summation Field) and GADF (Gramian Angular Difference Field). Its core principle is to encode a one-dimensional time series into a polar coordinate system, and then use trigonometric relationships of angles to generate a two-dimensional matrix, thereby preserving the time dependencies of the series in two-dimensional space.
[0196] GASF constructs images by calculating the cosine sum of angles corresponding to different wavelength positions, aiming to capture the superposition effects and correlations between points in the sequence. The GASF matrix is 128×128 in size, with the values on the main diagonal retaining the main information of the original spectral data, and the image texture representing the correlation structure between different bands of the spectral sequence. The GASF transformation formula is:
[0197] (11).
[0198] GADF constructs an image by calculating the sinusoidal difference of angles corresponding to different wavelength positions, focusing on capturing the relative changes and differences between points in the sequence. The GADF matrix is also 128×128 in size, with the values on the main diagonal always being 0. The texture of the image represents the dynamic changes of the spectral sequence between different bands. The GADF conversion formula is:
[0199] (12)
[0200] MTF (Markov Transition Field) constructs an image by building a probability transition matrix. It treats spectral data as a Markov process and uses its state transition probability matrix as the image. The MTF transformation includes the following steps:
[0201] S30301: State discretization;
[0202] The input spectrum has 2001 feature points and Q=12 state intervals. The quantile strategy ensures that each interval contains approximately the same number of spectral points (2001 / 12≈167 spectral points per interval). Therefore, a continuous spectral vector of length 2001 is transformed into a discrete state sequence of equal length consisting of integers from 1 to 12.
[0203] S30302: Construct the Markov transition matrix;
[0204] Based on Q=12 states, construct a Markov transition matrix of size 12×12. Spectral intensity from state space To state space The global probability is The global probability can be obtained by statistically analyzing the frequency of all adjacent state pairs in the entire 2001-state sequence and performing normalization calculations. The resulting 12×12 matrix quantifies the dynamic variation pattern of the spectral signal.
[0205] S30303: Generate Markov transition field;
[0206] Based on 2001 feature points from the original spectral data, this process generates a first image of size 2001×2001, namely the Markov transfer field. This image is the most complete and lossless two-dimensional representation of the intrinsic structure of the spectrum.
[0207] S30304: Image size normalization.
[0208] The 2001×2001 feature image obtained from S30303 is downsampled to 128×128 to adapt to the input size in the deep learning model. This process is achieved through bilinear interpolation, which can prevent information distortion during downsampling and preserve the structural texture and patterns in the original MTF image to the greatest extent.
[0209] S4: Using one or more 2D images of GASF, GADF and MTF generated by S3 as input, construct a multi-task learning model and train and evaluate it. Utilizing the complementarity of multi-source information, design a strategy to perform pixel weighting, pixel stitching and feature fusion on multiple images. This architecture can extract local and global features in a hierarchical manner and achieve accurate decoding and adaptive optimization of task-specific features through attention query mechanism and hybrid weighted loss function.
[0210] S4 includes the following steps:
[0211] S401: Extracting depth spatial features from a two-dimensional image obtained after converting one-dimensional spectral data using a two-dimensional convolutional neural network EfficientNetV2-S;
[0212] When a single image is used as input, the overall structure of the model is as follows: Figure 3 (a) is shown in the upper frame diagram.
[0213] When multiple images are used as input, the overall structure diagram of the model is as follows: Figure 3 (a) is shown in the lower frame diagram.
[0214] When the input for the depth space feature extraction described in S401 is multiple images, any of the following fusion strategies shall be adopted:
[0215] Two fusion methods, pixel-level fusion and feature-level fusion, are employed to explore information complementarity. Taking GASF, GADF, and MTF transformation methods as examples, the effectiveness of different fusion strategies is compared. Pixel-level fusion utilizes two strategies: pixel weighting and pixel stitching.
[0216] Fusion Strategy 1: Pixel Weighting; This strategy is implemented in the data preprocessing stage and belongs to data-level fusion.
[0217] The single-channel two-dimensional image matrices of GASF, GADF, and MTF are weighted according to a preset formula. Perform linear weighted summation to generate a new single-channel grayscale image. As the final model input, the model's first convolutional layer is now configured to receive one channel of input:
[0218] (13)
[0219] In equation (13):
[0220] , , It is a single-channel two-dimensional image matrix of GASF, GADF, and MTF, all of which are of size ;
[0221] In this invention, all weights are set to 1 / 3.
[0222] The prediction process chain of the model is as follows:
[0223] (14)
[0224] In equation (14):
[0225] This is the basic feature extraction performed by EfficientNetV2-S on the input image;
[0226] It is the global relational modeling performed by the Transformer encoder on the features extracted by the CNN;
[0227] It is a task-specific attention head that performs task-level filtering on the features output by the Transformer;
[0228] Each of the multiple task prediction branches corresponds to a prediction task for a physicochemical parameter;
[0229] These are the predicted results of physicochemical parameters;
[0230] Fusion Strategy 2: Pixel stitching; this strategy is implemented in the data preprocessing stage.
[0231] The three single-channel images, GASF, GADF, and MTF, are stacked along the new channel dimension to form a three-channel pseudo-color image. , For the height of the image, The image width, similar to processing a standard RGB image, is used as the final model input. The first convolutional layer of the model is configured with a three-channel input:
[0232] (15)
[0233] The prediction process chain of the model is as follows:
[0234] (16)
[0235] Fusion Strategy 3: Feature Fusion; this strategy is implemented within the model through parallel feature extraction. The model structure is as follows: Figure 3 As shown in (b).
[0236] During preprocessing, the data is also concatenated into a three-channel tensor, but the model has multiple parallel and independent branches. During the model's forward propagation, the input three-channel tensor is split, and the image from each channel is fed into its own dedicated, independent backbone network to extract high-level features. After obtaining their respective feature vectors, these feature vectors are concatenated to form a longer fused feature vector, which is then processed by a unified prediction head.
[0237] The model has three independent feature extractors Each is applied to the corresponding input image:
[0238] (17)
[0239] In equation (17):
[0240] It is the extracted high-level feature vector;
[0241] The fusion is completed at the feature layer through a splicing operation:
[0242] (18)
[0243] The prediction process chain of the model is as follows:
[0244] (19)
[0245] When fully unfolded, it looks like this:
[0246] (20)
[0247] S402: Input the feature sequence extracted by the convolutional neural network into the Transformer encoder that integrates RoPE (Rotation Position Encoding) for shared feature extraction for all tasks;
[0248] Traditional Transformer encoders use absolute position encoding and cannot perceive the order of the input sequence. This invention uses a Transformer encoder integrated with RoPE. The core idea of RoPE is to inject absolute position information into the feature vector through rotation, thereby allowing the self-attention mechanism to capture relative position information. After rotation, the dot product of the vectors at two positions depends only on their relative positions.
[0249] Assuming the wavelength sequence is in the 1st... The feature vector at each position is Define the base velocity of rotation for each eigenvector. :
[0250] (twenty one)
[0251] In equation (21):
[0252] It is the total dimension of the feature vector;
[0253] It is an index of feature pairs;
[0254] eigenvectors The dimensions are grouped pairwise, that is... , , ..., for the first Group feature pairs According to its location and base speed Perform two-dimensional rotation:
[0255] (twenty two)
[0256] In equation (22):
[0257] These are new feature pairs obtained after rotation;
[0258] The query vector after RoPE transformation and key vector When performing the dot product, The value will include a term that is only related to the relative position pz, thus enabling the attention score to perceive relative position information.
[0259] S403: Design TSA (Task-Specific Attention).
[0260] This module is based on an attention-based query mechanism, enabling each independent task to dynamically learn and extract the most relevant feature information from a shared feature sequence. The core idea of the TSA module is to set a learnable task query vector for each of the N different tasks. each of the rows Representing the A dedicated information probe for each task.
[0261] The output sequence of the Transformer encoder and learnable task query matrix As input, perform dot product attention operations to generate... A matrix composed of task-specific feature vectors :
[0262] (twenty three)
[0263] In equation (23):
[0264] This is used to indicate normalization processing; in the output sequence It is performed on the sequence length dimension.
[0265] It's a query about attention mechanisms;
[0266] It is the key to the attention mechanism;
[0267] It is the value of the attention mechanism;
[0268] The The line is the first Task-specific feature vectors The vector is then fed into the corresponding task prediction branch.
[0269] For the query vector of each task It will be related to the output sequence In Perform a dot product attention operation on each feature vector to make the query vector It can evaluate the relevance of features at each position in a sequence to its own task. The output sequence... Simultaneously serving as a key to attention mechanisms Sum , Query the task matrix As a query, the final generated A matrix composed of task-specific feature vectors .
[0270] S404: Design a hybrid weighted loss function;
[0271] A hybrid weighted loss function is adopted, which constructs a two-layer weighting system that integrates static weights based on prior knowledge and dynamic weights based on model homoscedasticity uncertainty, in order to achieve stable guidance and adaptive adjustment of the multi-task optimization process.
[0272] Calculate the base loss for each individual task:
[0273] (twenty four)
[0274] In equation (24):
[0275] It represents the number of samples for the task;
[0276] It is a task The true value;
[0277] It is a task The predicted value;
[0278] The model assigns learnable parameters to each task. Then calculate the total loss for all tasks. :
[0279] This parameter can be viewed as a measure of the model's uncertainty in predicting the task. The base loss for each task is determined based on its corresponding learnable parameters. Weighting is applied.
[0280] (25)
[0281] In equation (25):
[0282] B is the batch size;
[0283] The weights are fixed based on prior knowledge, with 2 for difficult tasks and 1 for easy tasks;
[0284] The model is for the task The automatically learned log-variance, reflecting the uncertainty of the model's predictions for the task, acts as a regularization term in the second part of the formula, penalizing excessively high uncertainty in predictions and preventing the model from over-regulating the log-variance. Set it to the maximum value to make the loss approach 0;
[0285] It is a dynamic weight. When the uncertainty is high, the weight is reduced, thereby reducing the proportion of the corresponding task error in the total loss.
[0286] S405: An early stopping mechanism is introduced to prevent overfitting. The model is trained for a maximum of 400 epochs, with a patience value of 50 and a batch size of 32. Hyperparameter optimization is performed using Optune, resulting in the following settings: learning rate of 6.32e-05, weight decay of 1.06e-06, dropout rate of 0.323, number of Transformer layers of 2, and number of Transformer multi-head attention heads of 4.
[0287] S5: Apply the trained model to the test set to obtain the inversion results of the blade physicochemical parameters.
[0288] S5 includes the following steps:
[0289] S501: The training and validation sets are trained using a five-fold cross-validation method, and the optimal model for each fold is saved to ensure the generalization ability and robustness of the model. After the dataset is divided, the training and validation sets are preprocessed.
[0290] S502: Use the test set as input to the optimal model saved in each fold, and employ the coefficient of determination. Root mean square error Relative root mean square error based on 95% of the data range and relative prediction bias As an evaluation result of the model's effectiveness:
[0291] (26)
[0292] (27)
[0293] (28)
[0294] (29)
[0295] In equations (28)-(29):
[0296] It is the range of data remaining after removing the minimum and maximum values of 2.5% from the dataset (i.e., the difference between the 97.5th percentile and the 2.5th percentile after sorting).
[0297] It is the standard deviation of the true value;
[0298] The larger the coefficient of determination, the smaller the root mean square error, the smaller the relative root mean square error based on 95% of the data range, and the larger the relative prediction bias. In other words, the model has better prediction performance and stability, which means that the model has high goodness of fit, small error, high relative accuracy, strong discrimination ability, and excellent performance.
[0299] In this embodiment, the trained model is applied to the test set samples. The test set data undergoes the same spectral preprocessing and 2D image conversion operations as during the training phase to ensure consistency in the data processing flow. The average inversion result of the optimal model saved in each round of cross-validation on the independent test set is used as the evaluation metric for the model's final generalization ability. Subsequently, the model's prediction results are analyzed based on the performance evaluation metric to obtain the final performance evaluation result of the model on the test set.
[0300] Example 1:
[0301] This invention uses three methods to convert one-dimensional spectral data into two-dimensional images. The comparison results of the three methods are shown in Table 1.
[0302] The results show that, in the leaf spectral test set, using GASF transformation as the model input achieved the best inversion results. EWT determination coefficients. The relative root mean square error (RMSE) is 0.951, and the RMSE for 95% of the data range is 5.16%. Compared to using GADF and MTF transformations as model inputs, the coefficient of determination is... These improvements were 0.002 and 0.048 respectively, while the relative root mean square error (%RMSE) for 95% of the data range decreased by 0.07% and 2.09% respectively.
[0303] Determinant coefficient of leaf nitrogen content The relative root mean square error (RMSE) is 0.872, and the 95% data range has a relative root mean square error (RMSE) of 10.02%. Compared to using GADF and MTF transformations as model inputs, the coefficient of determination is... These improvements were 0.02 and 0.106 respectively, while the relative root mean square error (%RMSE) for 95% of the data range decreased by 0.74% and 3.51% respectively.
[0304] The coefficient of determination of chlorophyll a content The coefficient of determination is 0.829, and the relative root mean square error (RMSE) for 95% of the data range is 13.72%. Compared to using GADF and MTF transformations as model inputs, the coefficient of determination is... The relative root mean square error (RMSE) for 95% of the data range decreased by 0.03% and 1.29%, respectively, with improvements of 0.001 and 0.034, respectively.
[0305] The coefficient of determination of chlorophyll b content The relative root mean square error (RMSE) is 0.812, and the RMSE for 95% of the data range is 14.1%. Compared to using GADF and MTF transformations as model inputs, the coefficient of determination is... These improvements were 0.007 and 0.038 respectively, while the relative root mean square error (%RMSE) for 95% of the data range decreased by 0.25% and 1.34% respectively.
[0306] The coefficient of determination of carotenoid content The relative root mean square error (RMSE) is 0.802, and the 95% data range has a relative root mean square error (RMSE) of 14.22%. Compared to using GADF and MTF transformations as model inputs, the coefficient of determination is... The results were improved by 0.012 and 0.051 respectively, and the relative root mean square error (RMSE) for 95% of the data range decreased by 0.43% and 1.73% respectively.
[0307] In this invention, regardless of the method used for spectral-graphic conversion, the relative prediction deviation (RPD) values for all five indicators are greater than 2, indicating that the model has good effectiveness in predicting each physicochemical parameter. Among the three methods, when GASF is used as the model input, it provides the best prediction results for all five leaf parameters (…). , , , as well as The inversion accuracy of both methods reached the highest level. Compared with GADF conversion, GASF brought a small but consistent performance improvement; while compared with MTF conversion, its advantages were more significant.
[0308] The above results confirm the effectiveness of the multi-task learning strategy. According to the results analysis, the model achieved the best performance on the EWT task, but also achieved good optimization for highly correlated pigment parameters, with an overall determination coefficient... The score remained above 0.80. This suggests that the task-specific attention heads and uncertainty-based weighted loss incorporated into the model may have played a crucial role, allowing the model to allocate sufficient resources and attention to each task while leveraging shared knowledge, thereby optimizing overall performance.
[0309] Table 1 Comparison of the predictive performance of different transformation methods on the model
[0310]
[0311] In this invention, after fusing three different map transformation methods, feature fusion is found to be the optimal fusion strategy. Pixel stitching and pixel weighting, two data-level fusion methods, show comparable results, as shown in Table 2. However, none of the three fusion strategies outperforms the optimal single method. This is because GADF and GASF are both Gram-squared methods, and the features they extract may be highly similar or correlated. Fusing them with the poorly performing MTF introduces noise or useless information, thus diluting the effectiveness of the optimal features.
[0312] Table 2 Comparison of Predictive Performance of Different Fusion Methods on the Model
[0313]
[0314] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0315] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for inverting physicochemical parameters of plant leaves based on two-dimensional spectral characterization and multi-task deep learning, characterized in that: The method includes the following steps: S1: Construction of a hyperspectral plant leaf dataset; S2: Preprocessing of hyperspectral plant leaf dataset; S3: Reconstruct the preprocessed one-dimensional spectral vector into a two-dimensional image; S4: Using the two-dimensional image generated in S3 as input, construct a multi-task learning model and train and evaluate it; S4 includes the following steps: S401: Extracting depth spatial features from a two-dimensional image obtained after converting one-dimensional spectral data using a two-dimensional convolutional neural network EfficientNetV2-S; S402: Input the feature sequence extracted by the convolutional neural network into the RoPE-integrated Transformer encoder for shared feature extraction across all tasks; Assuming the wavelength sequence in the 1st... The feature vector at each position is Define the base velocity of rotation for each eigenvector. : (21) In equation (21): It is the total dimension of the feature vector; It is an index of feature pairs; eigenvectors The dimensions are grouped pairwise, that is... , , ..., For the Group feature pairs According to its location and base speed Perform two-dimensional rotation: (22) In equation (22): These are new feature pairs obtained after rotation; S403: Design TSA; The output sequence of the Transformer encoder and learnable task query matrix As input, perform dot product attention operations to generate... A matrix composed of task-specific feature vectors : (23) In equation (23): This is used to indicate normalization processing; It's a query about attention mechanisms; It is the key to the attention mechanism; It is the value of the attention mechanism; S404: Design a hybrid weighted loss function; Calculate the base loss for each individual task: (24) In equation (24): It represents the number of samples for the task; It is a task The true value; It is a task The predicted value; The model assigns learnable parameters to each task. Then calculate the total loss for all tasks. : (25) In equation (25): B is the batch size; The weights are fixed based on prior knowledge. The model is for the task The automatically learned log-variance; It is a dynamic weight; S405: Introduces an early stopping mechanism to prevent overfitting; S5: Apply the trained model to the test set to obtain the inversion results of the blade physicochemical parameters.
2. The method for inverting plant leaf physicochemical parameters based on two-dimensional spectral characterization and multi-task deep learning according to claim 1, characterized in that: S1 includes the following steps: S101: Acquire leaf spectral information and calculate wavelength using a hyperspectral analyzer. reflectivity at : (1) In formula (1): The leaf sample at wavelength The original reading at the location; Is the standard whiteboard at wavelength The reading at the location; It is dark current at wavelength The reading at the location; S102: Physicochemical parameters of leaf samples were determined by laboratory chemical analysis. S103: Match the spectral data obtained in S101 with the physicochemical parameters obtained in S102 to complete the construction of the hyperspectral plant leaf dataset.
3. The method for inverting plant leaf physicochemical parameters based on two-dimensional spectral characterization and multi-task deep learning according to claim 2, characterized in that: S2 includes the following steps: S201: Divide the dataset in an 8:2 ratio, take 20% of the data as the test set, and the remaining data as the training set and validation set; S202: Use SNV to independently transform the spectrum of each sample in the dataset: (7) In equation (7): It is the transformed spectral data; It is the raw spectral data; It is the average value of the original spectrum; It is the standard deviation of the original spectrum; S203: The transformed spectral data is smoothed using an SG filter. (8) In equation (8): It is the radius of the window, and the size of the sliding window is... ; It is smoothed in position New data points; In position The original data points within the surrounding window; It is the first The polynomial fitting coefficients for each point.
4. The method for inverting plant leaf physicochemical parameters based on two-dimensional spectral characterization and multi-task deep learning according to claim 3, characterized in that: S3 includes the following steps: S301: Scale the preprocessed one-dimensional spectral data obtained in S2 to the interval [-1, 1] to obtain a new spectral sequence. Calculate each data point Corresponding scaled data points : (9) In equation (9): and These are spectral sequences The maximum and minimum values in; S302: Scale the data points Considered as an angle The cosine value is obtained, and the angle is calculated using the inverse cosine function. : (10) S303: Angle conversion using spectral-image conversion rules By constructing a two-dimensional Gram matrix using the converted angle information, the relationship between adjacent bands is transformed into a local spatial structure relationship in the image, resulting in a two-dimensional image that corresponds one-to-one with the blade sample and contains structured information.
5. The method for inverting plant leaf physicochemical parameters based on two-dimensional spectral characterization and multi-task deep learning according to claim 4, characterized in that: When the input for the depth space feature extraction described in S401 is multiple images, any of the following fusion strategies shall be adopted: Blending Strategy 1: Pixel Weighting; The single-channel two-dimensional image matrices of GASF, GADF, and MTF are weighted according to a preset formula. Perform linear weighted summation to generate a new single-channel grayscale image. As the final model input: (13) In equation (13): , , It is a single-channel two-dimensional image matrix of GASF, GADF, and MTF; The prediction process chain of the model is as follows: (14) In equation (14): This is the basic feature extraction performed by EfficientNetV2-S on the input image; It is the global relational modeling performed by the Transformer encoder on the features extracted by the CNN; It is a task-specific attention head that performs task-level filtering on the features output by the Transformer; Each of the multiple task prediction branches corresponds to a prediction task for a physicochemical parameter; These are the predicted results of physicochemical parameters; Fusion Strategy 2: Pixel Stitching; The three single-channel images, GASF, GADF, and MTF, are stacked along the new channel dimension to form a three-channel pseudo-color image. , For the height of the image, Use the width of the image as the final model input: (15) The prediction process chain of the model is as follows: (16) Fusion Strategy 3: Feature Fusion; The model has three independent feature extractors Each is applied to the corresponding input image: (17) In equation (17): It is the extracted high-level feature vector; The fusion is completed at the feature layer through a splicing operation: (18) The prediction process chain of the model is as follows: (19) When fully unfolded, it looks like this: (20)。 6. The method for inverting plant leaf physicochemical parameters based on two-dimensional spectral characterization and multi-task deep learning according to claim 5, characterized in that: S5 includes the following steps: S501: The training and validation sets are trained using a five-fold cross-validation method, and the optimal model for each fold is saved. S502: Use the test set as input to the optimal model saved in each fold, and employ the coefficient of determination. Root mean square error Relative root mean square error based on 95% of the data range and relative prediction bias As an evaluation result of the model's effectiveness: (26) (27) (28) (29) In equations (26)-(29): It is a task The average value; It represents the range of data remaining after removing the minimum and maximum values (2.5%) from the dataset. It is the standard deviation of the true value.
Citation Information
Patent Citations
Catalpa bungei leaf photosynthetic pigment quantitative inversion method based on hyperspectral reflectivity
CN118209496A
Soil physical and chemical parameter inversion method based on multi-task deep convolutional neural network
CN118569069A
One-dimensional spectrum classification method and system
CN113313059A
Multi-task plant biochemical parameter inversion method and system under unbalanced data
CN115270630A