Limited sample spectral data enhancement and physiological and biochemical component inversion method based on de-noising diffusion probability model
By generating high-fidelity spectral data based on the denoised diffusion probability model (DDPM), the problem of insufficient generalization ability of spectral modeling under small sample conditions was solved, and the ability to accurately measure and manage nitrogen nutritional status was improved.
Patent Information
- Application Number
- CN202510769703.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing spectral modeling methods have insufficient generalization capabilities under small sample conditions, making it difficult to generate high-quality spectral data, which affects the accurate measurement and management of nitrogen nutritional status.
The spectral data is enhanced by using a denoising diffusion probability model (DDPM). Noise is added through forward diffusion and reverse denoising is performed using a learnable Markov chain to generate synthetic samples that are highly similar to the real data. An extended training set is constructed to improve model performance.
The prediction performance of the spectral model under small sample conditions was significantly improved, the generalization ability and robustness of the model were improved, and the prediction accuracy of nitrogen content was enhanced.
Smart Images

Figure CN120673887A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and modern agricultural technology, and particularly relates to a finite sample spectral data enhancement and physiological and biochemical component inversion method based on a denoising diffusion probability model. Background Art
[0002] Nitrogen plays a vital role in plant growth and development. However, excessive nitrogen fertilizer application can lead to soil contamination. Failure to monitor crop nutritional status can result in insufficient nitrogen supply during critical growth periods, compromising yield and quality. Therefore, precise measurement and scientific management of crop nitrogen nutritional status are crucial for improving nitrogen use efficiency and reducing environmental pollution. Existing research indicates that accurate assessment of crop nitrogen nutritional status often relies on laboratory chemical analysis. While this method offers high precision, it is time-consuming and hinders dynamic feedback control of crop nutritional information. Furthermore, chemical analysis typically requires destructive sampling and relies on toxic reagents and complex sample preparation procedures. Vis / NIR spectroscopy has been widely used to monitor crop nitrogen content due to its non-destructive, rapid response, and environmentally friendly nature. However, its application across species and phenological periods is often affected by biological differences, environmental factors, and variations in sample distribution. This results in insufficient generalization of models based on traditional modeling approaches across diverse application scenarios. Therefore, in recent years, a growing number of studies have begun exploring data-driven machine learning (ML) and deep learning (DL) methods to enhance the modeling capabilities of Vis / NIR spectroscopy for accurate LNC prediction. These methods can fully exploit the complex nonlinear relationships in high-dimensional spectral data and possess greater modeling potential across species and growth stages. However, it is important to note that in the actual construction of these classification or regression models, dataset size is a key factor influencing model performance. Generally, training an accurate and robust model requires a large number of samples, especially for deep learning models with higher data requirements. However, in practical applications, spectral data collection is often limited by human and material resources, resulting in a limited number of available samples, which increases the difficulty of building high-performance prediction models under limited data conditions. Data augmentation, as an effective method to address the small sample size issue, can generate additional training samples that are highly similar to the original data based on the existing dataset. This not only helps alleviate the problem of model overfitting but also improves the model's generalization and robustness.
[0003] Currently, in the field of spectral analysis, a variety of data enhancement methods have been proposed to increase the number of samples, improve data diversity, and meet the data scale requirements of high-precision prediction models. Based on different implementation methods, these data enhancement technologies can be mainly divided into two categories:
[0004] 1. Transformation-based spectral data enhancement method:
[0005] In early studies, Conlin et al. performed data enhancement by adding Gaussian noise to infrared spectral datasets and analyzed them using PLSR. Bjerrum et al. further explored the effects of baseline shift and intensity change on spectral data, using a method of randomly adding offsets, multiplications, and slope changes to simulate the perturbations that may occur during spectral acquisition, and combined extended multiplicative signal correction (EMSC) to enhance the near-infrared spectral dataset of tablets. On this basis, Blazhko et al. optimized this method and proposed an extended multiplicative signal enhancement (EMSA) framework. Experimental results show that the EMSA method can not only replace the preprocessing step of convolutional neural networks (CNNs), but also improve the prediction performance of CNNs. Although the above methods improve data diversity to a certain extent, they mainly rely on simple noise perturbations or interpolation transformations and may not fully retain the key features of spectral data in classification or regression tasks.
[0006] 2. Spectral data enhancement method based on generative model:
[0007] Generative models, by learning the distribution of raw data, can automatically synthesize samples that are consistent with the style of real data. In recent years, they have been widely used in data augmentation tasks with limited sample sizes. Recently, generative adversarial networks (GANs) and their variants have garnered significant attention for spectral data augmentation, demonstrating particularly promising performance in classification tasks. These methods significantly improve the predictive performance of classification models by generating spectra with features similar to the original samples. However, current research on GANs has primarily focused on classification tasks, with limited research on regression tasks. Some existing studies have explored the application of GANs in spectral regression tasks. For example, Zhang et al. first used a DCGAN to establish a regression model. By simultaneously expanding the sample's spectral data and the corresponding actual chemical values, they improved the performance of PLSR and SVR models on the prediction set for two maize varieties, validating the effectiveness of the proposed method in improving the predictive performance of regression models. Hu et al. combined hyperspectral imaging technology with WGAN-GP for data augmentation, generating high-quality synthetic spectra. They then evaluated the method using three regression models, further demonstrating the feasibility of augmenting spectral datasets for regression tasks. Despite the progress made in the aforementioned research, GAN-like models can still be affected by issues like vanishing gradients and mode collapse during training, limiting further improvements in the quality of generated samples. Furthermore, training GANs is often difficult due to the difficulty in maintaining a balance between the generator and the discriminator, an inherent limitation of GANs.
[0008] Diffusion models offer new insights into addressing these challenges. Sohl-Dickstein et al. first formalized the diffusion process as a Markov chain and proposed a basic framework for diffusion models. The core idea is to incrementally add noise to the data through an iterative forward diffusion process, destructing the data structure, and then reconstruct the data through a learned backward denoising process, thereby constructing a highly flexible and easily optimized data generation model. However, early research on diffusion models was limited, and it was not until the introduction of the denoised diffusion probability model (DDPM) that diffusion models became mainstream in image generation tasks. Recent studies have demonstrated that denoised diffusion probability models exhibit excellent performance in generating one-dimensional sequence data. This provides new insights into data generation tasks in the spectral domain. Meanwhile, existing research has primarily focused on data augmentation for classification tasks, while research on data augmentation for regression tasks based on denoised diffusion probability models (particularly in the context of quantitative prediction of crop nutritional parameters) remains an understudied area. Summary of the Invention
[0009] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a finite sample spectral data enhancement and physiological and biochemical component inversion method based on a denoising diffusion probability model, so as to effectively overcome the sample size limitation of spectral modeling, and simultaneously generate leaf spectral data and its corresponding nitrogen content. The synthetic spectrum not only maintains the characteristics of the original spectrum, but also significantly improves the generalization ability of the constructed prediction model, thereby improving the prediction performance of the spectral model under small sample conditions.
[0010] In order to achieve the above object, the technical solution adopted by the present invention is:
[0011] A first aspect of the present invention provides a finite sample spectral data enhancement method based on a denoising diffusion probability model, comprising the following steps:
[0012] Step 1: Collect leaf spectral data of several plant samples, measure the leaf nitrogen content, and concatenate the leaf spectral data with the corresponding nitrogen content to obtain spectrum-nitrogen content pairs, thereby constructing a leaf spectrum-nitrogen content dataset.
[0013] Step 2: Input the spectral-nitrogen content data of the leaves into the denoising diffusion probability model for data enhancement;
[0014] The denoising diffusion probability model, in which the forward diffusion process degenerates the original data distribution into a resolvable distribution by gradually adding Gaussian noise, and in which the reverse denoising process utilizes a learnable Markov chain to inversely transform the noise distribution into the target data distribution, thereby reconstructing synthetic data from random samples of known distribution.
[0015] In one embodiment, in step 1, each plant sample is cultured to the seedling stage using the same nitrogen concentration, and then nitrogen stress treatments of different gradients are applied, and the culture medium is regularly replaced until the end of the seedling stage; then, several leaves of the same leaf position are selected from the plants under the same nitrogen treatment gradient as a sample, and their Vis / NIR spectra and nitrogen content are measured to construct a leaf spectrum-nitrogen content dataset.
[0016] In one embodiment, the collected Vis / NIR spectral data is preprocessed by one or more of the following means:
[0017] Cut off the starting and ending bands of the spectrum and select the spectral data within the range of 400.717-914.63nm;
[0018] Smoothing preprocessing was applied to the spectral data;
[0019] The spectral data were normalized using the maximum-minimum method.
[0020] In one embodiment, the forward diffusion process uses the spectrum-nitrogen content pair as the original data x0, and gradually adds Gaussian noise to degenerate it into a standard Gaussian distribution. Noise x T , where I is a diagonal covariance matrix with the same dimension as the data;
[0021] The reverse denoising process first randomly samples a noise x that obeys the standard Gaussian distribution. t , and then the noise is predicted to be x through the deep neural network t Gradually denoise and restore new data with the same distribution as the original data x0 Finally, Separation into leaf spectra and nitrogen content completes the data generation process.
[0022] In one embodiment, the stepwise addition of Gaussian noise is performed as follows:
[0023]
[0024] Among them, β t ∈(0,1), is a hyperparameter used to control the noise intensity, x t Represents the sample after adding t-step Gaussian noise, corresponding to each time step t, a β is defined t ;
[0025] Derive the distribution at any time step t from x0:
[0026]
[0027] in According to the distribution q(xt |x0), x t The formula for fast sampling from x0 is:
[0028]
[0029] When the time step is T, the noise x is obtained T ;
[0030] where ∈ is standard Gaussian noise.
[0031] In one embodiment, the reverse denoising process starts from the sample x after adding t steps of Gaussian noise. t Gradually recover new data with the same distribution as the original data x0 Its conditional probability distribution is:
[0032]
[0033] Where θ is the trainable parameter of the neural network, and the mean μ θ (x t ,t) and variance ∑ θ (x t ,t) are parameterized by deep neural networks.
[0034] In one embodiment, the deep neural network predicts the noise ∈ injected during the forward diffusion process θ (x t ,t), and reconstruct the mean parameter as follows:
[0035]
[0036] where α t =1-β t represents the noise scheduling coefficient, variance ∑ θ (x t ,t) adopt a fixed strategy;
[0037] The training goal of the denoising diffusion probability model is to minimize the forward distribution q(x 1:T |x0) and the reverse distribution p θ (x 1:T ), using variational inference, we get the following optimization objective:
[0038]
[0039] Among them, ∈ θ (x t ,t) is the noise estimated by the model when taking the parameter θ.
[0040] In one embodiment, the reverse denoising process generates data by reversely executing the trained Markov chain, as follows:
[0041] Given a noise prediction network ∈ θ (x t ,t) and a fixed variance ∑ θ (x t ,t), the inference process is from the standard Gaussian distribution Starting from this, the target spectrum - nitrogen content sample is gradually reconstructed through iterative denoising, and finally
[0042] In one embodiment, a network architecture for spectral data denoising is constructed based on U-Net, and a cosine time scheduling strategy is adopted to replace the linear time scheduling in U-Net. At the same time, a self-conditioning strategy is adopted. The denoising process is as follows:
[0043] First, a one-dimensional convolution is applied to the noisy spectral data, and position embeddings associated with the noise level are calculated;
[0044] Subsequently, a series of downsampling and upsampling blocks are used to extract and restore data features. In the encoding phase, each downsampling block captures multi-scale features through two ResNet blocks, an attention mechanism, a residual connection, and a downsampling operation. In the decoding phase, a symmetrical module is used, and upsampling is used instead of downsampling. The encoder and decoder are connected with a bottleneck layer.
[0045] Finally, through a Resnet block and a one-dimensional convolutional layer, the noise estimate added to the current input data is output, thus completing the denoising process.
[0046] A second aspect of the present invention further provides a method for inverting physiological and biochemical components based on a denoised diffusion probability model, comprising the following steps:
[0047] Step 1, obtaining synthetic data using the finite sample spectral data enhancement method based on the denoising diffusion probability model described in the first aspect of the present invention;
[0048] Step 2: Combine the original data with the synthetic data to construct an extended training set;
[0049] Step 3: Using the extended training set to train a regression model for inverting plant physiological and biochemical components, the regression model takes spectral data as input and outputs corresponding component prediction values.
[0050] Compared with the existing technology, the data enhancement method based on the denoising diffusion probability model of the present invention can generate synthetic samples that are highly similar to the real sample distribution, which can provide richer data support for the regression model, thereby enhancing the prediction performance of the model.
[0051] Experimental results show that the introduction of synthetic data can significantly improve the model performance, especially the spectral model based on the one-dimensional convolutional neural network algorithm, which has an R 2 It increased by 0.129 and RMSE decreased by 2.229 mg / g. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This invention is a denoising diffusion probability model framework based on Vis-NIR spectrum.
[0053] Figure 2 This is the overall structure of the U-net used for spectral data denoising in the present invention.
[0054] Figure 3 This is the specific structure of the downsampling module and the upsampling module of the present invention.
[0055] Figure 4 It is the 1D-CNN used in the present invention to predict the nitrogen content of tomato leaves.
[0056] Figure 5 are the spectral curves of all real samples in the embodiments of the present invention.
[0057] Figure 6 1 and 2 are the average spectral curves and standard deviations of the samples at three different nitrogen levels in the embodiment of the present invention.
[0058] Figure 7 Spectra generated by the Diffusion model of the embodiment of the present invention at different training steps.
[0059] Figure 8 These are spectra generated by different generation frameworks under the optimal number of training steps in the embodiments of the present invention.
[0060] Figure 9 : This is the leaf nitrogen content distribution of the samples generated by DDPM under different training steps in the embodiment of the present invention.
[0061] Figure 10 This is a visualization of the distribution of the real samples and generated samples in the three-dimensional principal component space after using PCA to extract the first three principal components of the variance contribution rate of the embodiment of the present invention.
[0062] Figure 11 This embodiment of the present invention visualizes real samples and generated samples through t-SNE. DETAILED DESCRIPTION
[0063] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.
[0064] Rapid, non-destructive measurement of leaf nitrogen content is crucial for precise production management. Visible / near-infrared (Vis / NIR) spectroscopy enables rapid nitrogen content measurement by capturing leaf reflectance spectra. Building a robust spectral model typically requires a sufficient sample volume, and obtaining reference values for nitrogen content in a large number of leaves is time-consuming and costly.
[0065] To address the lack of robustness of spectral models under small sample sizes, this paper proposes a data augmentation method based on a denoised diffusion probability model (DDPM) to generate high-fidelity spectra and corresponding nitrogen contents. Through multiple iterations, this method can generate synthetic samples with a distribution that closely resembles the original data.
[0066] In this embodiment of the present invention, using tomato leaves as an example, a DDPM for one-dimensional data is proposed to enhance the spectral data and nitrogen content of tomato leaves. The quality of the spectral data generated by the denoised diffusion probability model is qualitatively and quantitatively evaluated to ensure its similarity in feature distribution to the real spectral data. To verify the effectiveness, the present invention uses PLSR, SVR, and 1D-CNN algorithms to construct a prediction model for nitrogen content in tomato leaves, comparing the prediction performance of each prediction model after data enhancement. The impact of different numbers of generated samples on the performance of the regression model and the dependence of the regression model complexity on the amount of data are also analyzed.
[0067] Specifically, the data enhancement method of the present invention includes:
[0068] Step 1: Collect leaf spectral data of several plant samples, measure the leaf nitrogen content, and construct a leaf spectrum-nitrogen content dataset; wherein, different gradients of nitrogen are applied to the plant samples during the cultivation process.
[0069] In this embodiment, the tomato variety "Kaide Yali" was selected as the research object, and a hydroponic experiment was carried out under a controlled environment at the Key Laboratory of Agricultural Internet of Things, Ministry of Agriculture and Rural Affairs, Northwest A&F University (36°16'N, 108°4'E). In order to accurately construct a nitrogen gradient and eliminate matrix interference, the present invention uses perlite to fix the root system, and sets 5 nitrogen treatment gradients (N1-N5) based on the Japanese garden test standard nutrient solution. Among them, based on the conventional concentration of 17.5mmol / L as the benchmark (N4), 0% (N1), 25% (N2), 50% (N3), and 150% (N5) concentration nutrient solutions were respectively configured, and each treatment group contained 10 biological replicates (Table 1). In the seedling stage, all plants were cultured with N4 concentration, and after entering the seedling stage, nitrogen stress treatments of different gradients were applied. Throughout the experiment, environmental parameters were strictly controlled: the temperature was maintained at 25±0.2°C (day) / 18±0.2°C (night), the relative humidity was 41% (day) / 65% (night), the CO2 concentration was maintained at a natural state, white LED lights were used for illumination, and the culture medium was replaced every three days until the tomato seedling stage was completed. Subsequently, five leaves from the same leaf position were selected from plants under the same nitrogen treatment gradient as a sample, and their Vis / NIR spectra and nitrogen content were measured to construct a leaf spectral-nitrogen content dataset. Ultimately, a total of 120 samples were collected for subsequent analysis.
[0070] Table 1 Changes in nutrient element concentrations in nutrient solution under different nitrogen stress conditions (mmol·L -1 )
[0071] Nutrient Element N1 N2 N3 N4 N5 <![CDATA[N(mmol·L -1 )]]> 0 4.375 8.75 17.5 26.25 <![CDATA[P(mmol·L -1 )]]> 1.3 1.3 1.3 1.3 1.3 <![CDATA[K(mmol·L -1 )]]> 8 8 8 8 8 <![CDATA[Ca(mmol·L -1 )]]> 4 4 4 4 4 <![CDATA[Mg(mmol·L -1 )]]> 1 1 1 1 1
[0072] Subsequently, the leaf reflectance spectrum was collected. The acquisition system consisted of a spectrometer (QE Pro-VIS-NIR, Ocean Optics, USA), a halogen tungsten lamp (HL-2000-HP, Ocean Optics, USA), a reflective fiber, a leaf clip (SpectroClip-TR, Ocean Optics, USA), and a computer. The measurement wavelength range covered 350–925 nm, with a spectral resolution of 1.2 nm and a pixel count of 1024. Spectral data were acquired using the spectrometer's built-in OceanView 2.0 spectral analysis software.
[0073] To reduce interference from ambient light, all experiments were conducted in a controlled darkroom environment. Before data collection, the halogen tungsten lamp needed to be preheated for 30 minutes to ensure the stability of the light source. During the experiment, the integration time was set to 100ms, 5 scans were averaged, and the sliding average width was set to 5. In addition, dark noise correction was used to reduce the impact of noise, and the trigger mode was set to continuous scanning. During spectral measurement, in order to reduce the influence of leaf vein structure on the spectral signal, a leaf clamp was used to fix the leaf, and 5 measurement points were randomly selected at positions avoiding the leaf veins to obtain Vis / NIR reflectance spectra. Finally, the average spectrum of the 5 measurements was used as the original reflectance spectrum data of the leaf.
[0074] After completing the Vis-NIR spectral data collection of tomato leaves, the samples were first placed in a 105°C oven for 30 minutes to terminate enzyme activity. Subsequently, the temperature was adjusted to 75°C and the samples were dried for 48 hours until the weight of the samples was constant. After the samples were completely dried, they were ground into powder and 0.1g was taken for subsequent total nitrogen content determination. The experiment used a fully automatic nitrogen analyzer (Kjeltec TM 8400, FOSS, Denmark) was used to determine the total nitrogen content of tomato leaves.
[0075] Since dark noise exists at both ends of the collected reflectance spectrum, in order to ensure the reliability of the Vis-NIR data and reduce the systematic error caused by the instrument, the present invention cuts off the starting band and the ending band of the spectrum. Finally, the spectrum within the range of 400.717-914.63nm (a total of 679 wavelengths) was selected for subsequent research. In order to reduce the influence of noise in the original spectrum, Savitzky-Golay (SG) smoothing preprocessing was adopted for the spectral data. In addition, in order to eliminate the difference in the numerical range between the spectral data and the nitrogen content of the leaves, and to improve the training stability of the model, the present invention adopts maximum-minimum normalization processing to promote the effective training and convergence of the model. It is easy to understand that the preprocessing means of the present invention can be used by only one or several items in combination.
[0076] Step 2: Input the spectral-nitrogen content data of the leaves into the denoising diffusion probability model for data enhancement.
[0077] The proposed spectral denoising diffusion probability model framework consists of two coupled stochastic processes: a forward diffusion process that progressively adds Gaussian noise, degrading the original data distribution into a resolvable, simpler distribution (e.g., a standard Gaussian distribution); and a backward denoising process that utilizes a learnable Markov chain to inversely transform the noise distribution into the target data distribution, reconstructing synthetic data from random samples of a known distribution. This process is guided by a neural network that progressively optimizes sample quality, enabling the model to generate high-quality spectral data.
[0078] Specifically, the forward diffusion process of the present invention is as follows: given the original spectrum-nitrogen content pair as the original data x0, wherein x0=[λ1,λ2,...,λ 679 ,N content ] contains 679 discrete wavelength spectra and leaf nitrogen content, x0~q(x0), and forward diffusion degrades it to a standard Gaussian distribution by gradually adding Gaussian noise Noise x T , the noise adding process is expressed as:
[0079]
[0080] Among them, β t ∈(0,1), is a hyperparameter used to control the noise intensity. For each time step t, a β is defined. t , x t represents the sample after t steps of Gaussian noise have been added, and I is a diagonal covariance matrix with the same dimensions as the data. By gradually introducing noise, the forward diffusion process gradually adds noise to the original data. If the number of diffusion steps is large enough, pure Gaussian noise is ultimately obtained.
[0081] By using the properties of the forward diffusion process, the distribution of any time step t can be derived directly from the original data x0:
[0082]
[0083] in From the distribution q(x t |x0), x t The formula for fast sampling from x0 can be written as follows:
[0084]
[0085] When the time step is T, the noise x is obtained T , where ∈ is standard Gaussian noise.
[0086] The reverse denoising process is the reverse process of the forward diffusion process, that is, reconstructing synthetic data from random samples of known distribution, and reconstructing the noise x from the random samples of known distribution. t The original spectrum-chemical value pair is gradually restored. The process can be described as follows: first, a noise x that obeys the standard Gaussian distribution is randomly sampled. t , and then the noise is predicted to be x through the deep neural network t Gradually denoise and restore new data with the same distribution as the original data x0 Finally, Separation into leaf spectra and nitrogen content completes the data generation process.
[0087] Specifically, it is derived from the noise xt Step by step, the original spectrum-chemical value pairs are restored, and the conditional probability distribution is:
[0088]
[0089] Where θ is the trainable parameter of the neural network, and the mean μ θ (x t ,t) and variance ∑ θ (x t ,t), can be parameterized by a deep neural network. In a typical implementation, the deep neural network parameterizes the distribution through a noise prediction mechanism. Specifically, the deep neural network directly predicts the noise ∈ injected by the forward diffusion process θ (x t ,t), and reconstruct the mean parameter as follows:
[0090]
[0091] where α t :=1-β t represents the noise scheduling coefficient, variance ∑ θ (x t ,t) usually adopts a fixed strategy and only reconstructs the mean μ θ (x t ,t). The training goal of DDPM is to minimize the forward distribution q(x 1:T |x0) and the reverse distribution p θ (x 1:T ), using variational inference, we can get the following optimization objectives:
[0092]
[0093] However, in practice, directly optimizing the variational lower bound may be very complex. Therefore, DDPM simplifies the optimization objective and converts it into the mean square error (MSE) form of noise estimation:
[0094]
[0095] Among them, the standard Gaussian noise ∈ is the real noise added in the forward diffusion process, ∈ θ (x t ,t) is the noise estimated by the model when taking the parameter θ.
[0096] The inference phase of the denoising diffusion probability model (i.e., the reverse denoising process) realizes data generation by inversely executing the trained Markov chain. Specifically, given a noise prediction network ∈ θ (x t ,t) and a fixed variance ∑ θ (xt ,t), the inference process is from the standard Gaussian distribution Starting from, the target spectrum - nitrogen content sample is gradually reconstructed through iterative denoising. The inference algorithm of the denoised diffusion probability model is as follows. Ideally, a synthetic sample that looks like it comes from the real data distribution will be obtained in the end, that is,
[0097] 1:x T ~N(0,I)
[0098] 2: for t=T,……,1do
[0099] 3:∈~N(0,I)if t>1,else∈=0
[0100] 4:
[0101] 5: end for
[0102] 6: return x0
[0103] The above algorithm is interpreted as:
[0104] 1. Initialization: Sample an initial noise vector x from a standard Gaussian distribution T ~N(0,I) is used as the sampling starting point;
[0105] 2. Stepwise reverse denoising:
[0106] From time step t = T down to 1, repeat the following steps:
[0107] If t>1, a random noise ∈~N(0,I) is sampled from a standard Gaussian distribution;
[0108] Otherwise (ie t = 1), let ∈ = 0
[0109] Calculate the sample x of the previous time step according to the following formula t-1 :
[0110]
[0111] 3. Return the final result: When the sampling process is completed, the denoised sample x0 is returned, which is the output data generated by the model.
[0112] In this paper, DDPM, which is suitable for one-dimensional data, is used as the model framework and is adjusted accordingly for spectral data. The goal of this paper is to use DDPM to simultaneously generate tomato leaf spectral data and nitrogen content data. Its structure is as follows Figure 1In the forward diffusion process, the tomato leaf spectral data and the corresponding nitrogen content are first spliced to obtain the spectrum-nitrogen content pair x0. Then, by gradually adding Gaussian noise, it is degraded to conform to the standard Gaussian distribution. Noise x T In the reverse denoising process, we first randomly sample a noise that obeys the standard Gaussian distribution. The noise is then predicted by the deep neural network as Gradually remove noise and restore data It is new data with the same distribution as the original data x0. Finally, Separation into spectra and leaf nitrogen content completes the data generation process.
[0113] In the embodiment of the present invention, a network architecture for spectral data denoising is constructed based on U-Net, and its overall structure is as follows: Figure 2 As shown in Figure 1. First, a one-dimensional convolution is applied to the spectral data containing noise, and the position embedding related to the noise level is calculated. Subsequently, a series of downsampling blocks and upsampling blocks are used to extract and restore data features. The specific structure is shown in Figure 1. Figure 3 As shown in the figure. During the encoding phase, each downsampling block effectively captures multi-scale features through two Resnet blocks, an attention mechanism, residual connections, and downsampling operations. During the decoding phase, a symmetrical module is used, the main difference being that upsampling operations are used instead of downsampling. The bottleneck layer is used to connect the encoder and decoder to ensure that the model can effectively extract the most critical information. At the same time, the residual connection between the encoder and decoder significantly improves gradient propagation and contributes to the stability of training. Finally, through a Resnet block and a one-dimensional convolutional layer, the model outputs an estimate of the noise added to the current input data, thus completing the denoising process.
[0114] Table 2 details the structure and training process of U-net. To further optimize the model performance, this paper adopts a cosine time scheduling strategy instead of the traditional linear time scheduling and a self-conditioning strategy to achieve better generation results. The denoising diffusion probability model was trained for a total of 35,000 steps, with a time step of 1,000 and a learning rate of 8×10 -5 , using exponential moving average (EMA) to update the model parameters, and its weight is set to 0.995. Finally, to evaluate the performance of DDPM, it is compared with several GAN networks (GANRAM, DCGAN, WGAN-GP) that are also used for spectral data augmentation tasks.
[0115] Table 2. The proposed U-net architecture adopted in DDPM-based data augmentation.
[0116]
[0117]
[0118] On this basis, the complete physiological and biochemical component inversion method of the present invention includes the following steps:
[0119] Step 1: collect visible-near infrared spectral data of the plant sample to be tested and the corresponding physiological and biochemical component measurement values, and perform preprocessing operations such as normalization and smoothing on the spectral data to form an original sample data set.
[0120] Step 2: Use the diffusion probability model to perform forward diffusion and reverse denoising on the original sample spectrum, generate synthetic spectrum samples similar to the real spectrum characteristics and corresponding physiological and biochemical indicators through sampling, and construct synthetic sample pairs.
[0121] Step 3: Combine the original sample data pairs with the synthesized sample pairs to construct an extended training set.
[0122] Step 4: training a regression model (such as a one-dimensional convolutional neural network, SVR, PLSR, etc.) for inverting plant physiological and biochemical components on the extended training set. The regression model is used to receive corresponding spectral data input and output corresponding component prediction values.
[0123] This method can solve the problem of insufficient component inversion accuracy under the condition of limited number of plant samples and improve the generalization ability and stability of the model.
[0124] Based on the visualization of synthetic samples, the present invention comprehensively evaluates the sample quality through qualitative and quantitative assessment.
[0125] In terms of qualitative analysis, the present invention uses boxplots and spectral data distribution feature analysis to intuitively assess the quality of the synthesized samples. Since each spectral curve corresponds to a specific LNC, a high degree of similarity between the generated LNC and the actual LNC indicates a higher degree of credibility in the generated spectral data. Therefore, the present invention uses boxplots to assess the rationality of the LNC distribution of the synthesized samples and analyzes the differences between the LNC distribution and the actual LNC distribution. Furthermore, the present invention uses principal component analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE) to perform feature analysis on the generated spectral samples, enabling a more comprehensive assessment of the similarity between the synthesized data and the actual data at different scales.
[0126] Although qualitative evaluation methods can directly observe the quality and diversity of generated spectral curves, relying solely on this method to evaluate the quality of synthetic samples will result in a lack of objectivity. Therefore, to further verify the effectiveness of the denoising diffusion probability model in the spectral data generation task, this paper uses the Pearson Correlation Coefficient (PCC), Mahalanobis Distance (MD), and Maximum Mean Discrepancy (MMD) to quantitatively analyze the synthetic data.
[0127] The application of PCC in spectral analysis is mainly based on its ability to quantify the linear correlation between two variables. Since spectral data are usually expressed as high-dimensional vectors, PCC can be used to measure the similarity between two spectra. and Its PCC is defined as follows:
[0128]
[0129] in, and The PCC value range is [-1, 1]: 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no linear correlation.
[0130] In order to more accurately measure the potential similarity between synthetic samples and real samples, the present invention uses the following method to calculate the PCC between two spectral data sets: First, the PCC between each generated spectrum and all real spectra is calculated, and the maximum value is selected; then, the average of the maximum PCC corresponding to all generated samples is taken as the overall similarity evaluation index. Assume that the real spectrum sample set is The generated spectrum sample set is The maximum correlation coefficient of each generated sample is defined as:
[0131]
[0132] Averaging the maximum PCC of all synthetic samples yields:
[0133]
[0134] The final average PCC value is used to represent the overall similarity between the generated sample set and the real sample set.
[0135] MD is an indicator to measure the similarity of high-dimensional data, which is usually used to evaluate the degree of deviation of samples in feature space. In spectral analysis, MD represents the spectral sample xi The distance to the mean spectrum μ of the spectrum set X is calculated as follows:
[0136]
[0137] Where Σ is the covariance matrix of the dataset.
[0138] However, when the dimension of the spectral data is much larger than the sample size, high-dimensional data will make it difficult to estimate the covariance matrix. Therefore, in practice, PCA scores are usually used to replace spectral data. The converted formula is as follows:
[0139]
[0140] where s i is the principal component score of the sample, μ PCA is the mean vector in the PCA space.
[0141] In this study, we used the PCA method to calculate the MD values of the first three principal components (with a cumulative explained variance ratio of at least 95%). The MD values for each generated sample to the center of the original distribution were averaged as a similarity measure between the generated sample set and the true sample set. A smaller MD value indicates a smaller difference between the generated data and the true data, meaning that the spectral data generated by the model is closer to the true distribution.
[0142] The maximum mean difference (MMD) measures the mean difference of samples independently extracted from each distribution in the feature space (mapped by the kernel function). The formula for MMD is:
[0143]
[0144] where x i and y i are the original and generated spectral samples, respectively, and k(·,·) is the kernel function, typically a Gaussian kernel. In this paper, a Gaussian kernel with a width of 1.0 is selected, and the square of the MMD is calculated to provide an intuitive similarity measure. The closer the MMD value is to 0, the closer the distribution of the synthesized sample is to the real sample.
[0145] To verify the applicability of generated samples to deep learning methods, the present invention constructed a one-dimensional convolutional neural network as a leaf nitrogen content prediction model. Furthermore, to address the vanishing and exploding gradient issues often seen in deep CNNs, the present invention introduced residual connections, which not only accelerated network convergence but also improved the model's generalization capabilities. Figure 4The specific network structure is shown, which is mainly composed of the following key components: (1) the initial one-dimensional convolution layer, batch normalization, ReLU activation and maximum pooling layer for preliminary feature extraction and dimensionality reduction; (2) 4 stacked ResBlocks, each containing two convolution layers and residual connections to extract deep local features; (3) Adaptive MaxPool1d, used to map feature sequences to global features of fixed length; (4) Fully connected layer, which performs regression prediction on the flattened feature vector. In the training phase, the present invention uses the SGD optimizer, MSE as the loss function, and the learning rate is set to 1×10 -4 , using regularization, weight_decay is set to 1×10 -3 , the batch size is set to 16, and a total of 2000 epochs are trained.
[0146] In addition, in order to verify that the generated samples are also well adaptable to traditional modeling methods, the present invention uses PLSR and SVR, two classic regression algorithms in the field of spectral analysis, to construct the LNC prediction model. For PLSR, the present invention uses the following parameter settings: the number of latent variables is 11, the maximum number of iterations is 500 by default, and the tolerance is set to 1×10 -6 The hyperparameters of SVR were determined by grid search combined with cross-validation. A Gaussian kernel was used with the kernel coefficient γ set to 0.1 and the initial regularization parameter C set to 20.
[0147] In this paper, the root mean square error (RMSE) and the coefficient of determination (R 2 ) to evaluate the performance of the regression model in predicting nitrogen content in tomatoes. Generally, a lower RMSE and a higher R² indicate a stronger model fit and higher prediction accuracy. The formulas for calculating these performance indicators are as follows:
[0148]
[0149] Where n represents the sample size, Represents the predicted value of the i-th element, y i is the actual value of the i-th element, is the average value.
[0150] In the experiment, it was found that the tomato seedlings in the N2, N3, and N4 groups all grew well, while the tomato seedlings treated with N1 had yellow leaves and were dwarfed. In addition, the tips of the mature leaves of the tomato seedlings in the N5 group showed slight blackening symptoms. Therefore, in the present invention, the treatment of the N1 group was defined as low nitrogen stress, represented by L, the treatment of the N2, N3, and N4 groups was defined as normal concentration, represented by M, and the N5 group was defined as nitrogen excess, represented by H. The spectral curves of all samples are shown in Figure 2. Figure 5As shown, from the overall trend, the spectral curves of each leaf have a high degree of consistency. Among them, the reflection peak near 550nm is due to the strong reflection of chlorophyll on green light, while the troughs at 480-510nm and 650-700nm are caused by the strong absorption of blue light and red light by chlorophyll, respectively. The high reflectivity of 750-914nm is caused by multiple scattering of short-wave near-infrared light inside the leaf. Since nitrogen deficiency can cause pigment deposition and structural deformation of leaves, the average spectral curves of samples under different nitrogen levels show certain differences. Figure 6 As shown, the shaded area represents the standard deviation. In contrast, the L group samples had a higher spectral reflectance due to the lack of nitrogen, which hindered the synthesis of chlorophyll, carotenoids, etc., reducing the leaves' absorption of visible light. In the subsequent model training phase, to avoid overfitting and improve the model's generalization ability for samples cultured at different nitrogen levels, the present invention merged all samples into a unified dataset. Ultimately, the experimental data included the Vis / NIR spectra of 120 tomato leaves and their corresponding leaf nitrogen content.
[0151] The present invention first visualizes the spectral data generated by DDPM at different training steps to evaluate its visual characteristics such as shape and contour. In order to facilitate comparison with the original data, the present invention generates synthetic spectra with the same sample size as the original. Figure 7 As shown in the figure, although the spectra generated by the first 2000 training steps show an overall profile similar to the real data, the noise level is still high. As the number of training steps increases, when it reaches 10,000 steps, the difference in the profile of the generated spectral curve with the real data in the 400nm-700nm band becomes smaller, and the curve becomes gradually smoother. After further increasing the number of training steps to 25,000 steps, the generated spectrum is almost the same as the real spectrum in profile and distribution. However, when the training exceeds 30,000 steps, the similarity between the generated samples and the real samples decreases again, which may be due to overfitting. Figure 8 The spectra generated by the four generative frameworks at the optimal number of training steps are shown. It can be observed that while WGAN-GP is able to learn the distribution of the real spectrum, the generated data has a high level of noise. The spectra generated by DCGAN and GANRAM are relatively smooth, but only cover the main modes of the data and fail to fully reproduce the real data distribution. Notably, the curves generated by DDPM are both smooth and cover the most realistic data distribution. Judging from the waveform characteristics of the generated spectra, it can be concluded that the spectral samples generated by DDPM have a higher physical plausibility. It is important to emphasize that although the data generated by DDPM is visually highly similar to the real data, visual inspection alone cannot determine whether the synthetic data can replace the real data. Therefore, more detailed analysis is required to accurately evaluate the data generated by DDPM.
[0152] Figure 9 The box plots of the actual nitrogen content in tomato leaves and the nitrogen content generated by DDPM are intuitively displayed. These figures clearly reflect the distribution characteristics of LNC. Figure 9 The mean, median, quartiles, and box lengths in the box plots can be used to assess the degree of difference between the generated data and the real data. When the positions of the box plots are close, the differences between the median and mean are small, and there are no outliers, it indicates that the distribution of the generated data is relatively consistent with the real data. If there is a significant deviation between the box plots, it indicates that there is still a large difference between the generated data and the real data. Further observation Figure 9 We can see that the boxplot of the DDPM-generated data at 25,000 training steps is closest to the real data. This indicates that the nitrogen content distribution of the samples generated during this training phase most closely matches the statistical characteristics of the real samples. Combined with the aforementioned analysis of the spectral visualization results, the data generated at 25,000 training steps was ultimately selected as the basis for subsequent data augmentation modeling.
[0153] In order to visualize and analyze the characteristics of data in low-dimensional space, the present invention adopts PCA and t-SNE methods to visualize and analyze the generated spectral samples from the feature space, such as Figure 10 and Figure 11 shown.
[0154] In the qualitative analysis part, PCA analysis was first performed to extract the first three principal components of the spectral data (the cumulative variance contribution rate reached 97.3%). Figure 10 The distribution of synthetic samples (blue) and real samples (red) in the three-dimensional principal component space is shown, and 95% confidence ellipses are plotted, with each ellipse representing the region containing 95% of the samples. The results show that the synthetic samples fall almost completely within the 95% confidence ellipses of the real samples, highly overlapping with the distribution of the real data, and the distribution ranges along the three principal components are basically consistent. This demonstrates that the proposed method can effectively capture the global distribution of the real spectrum and does not generate anomalous samples.
[0155] To further explore the distribution similarity between synthetic samples and real samples, Figure 11 The two-dimensional projection results after dimensionality reduction using t-SNE (perplexity = 40, iterations = 300) are shown. The results show that the synthetic samples (blue) and real samples (red) exhibit a tightly interwoven clustering structure within their local neighborhoods. This demonstrates that this method not only reproduces the global distribution characteristics of real samples but also accurately captures the nonlinear correlations between spectra.
[0156] Qualitative analysis results show that the spectral samples synthesized by the denoised diffusion probability model are highly aligned with the real samples in feature space and accurately correspond to the leaf nitrogen content. This dual consistency verifies the authenticity and rationality of the distribution of the synthesized samples.
[0157] In the quantitative analysis part, the present invention systematically compares the performance differences of four generative models (GANRAM, DCGAN, WGAN-GP and denoising diffusion probability model) in the spectral synthesis task. Based on the same number of synthetic samples, the four models were quantitatively evaluated, and the results are shown in Table 3, where the best values of each indicator are marked in bold. The research results show that the spectral samples generated by the denoising diffusion probability model achieved the best results in all three comprehensive indicators (PCC=0.996, MD=1.604, MMD=0.021). Compared with the traditional generative adversarial network, the spectrum generated based on DDPM can more stably approximate the real data distribution, thereby obtaining higher quality samples, which is consistent with the observation results of visual analysis. The present invention believes that this is because the Diffusion model avoids the traditional GAN's adversarial training mechanism that is susceptible to gradient disappearance or mode collapse, enabling it to more comprehensively cover the real data manifold, avoid ignoring low-probability but regular spectral patterns, and thus improve the rationality of the synthetic samples.
[0158] Table 3 Comparative analysis of index results of different methods
[0159]
[0160] Before model construction, the original spectral dataset was divided using the SPXY algorithm, where the training set contained 90 samples and the validation set contained 30 samples. Subsequently, different numbers of spectral samples were generated using DDPM and combined with the original training set to obtain different numbers of mixed datasets. In order to evaluate the effect of these mixed datasets in nitrogen content prediction, the present invention used PLSR, SVR and 1D-CNN for modeling respectively. Since the present invention mainly focuses on the modeling effect of DDPM generated data, all experiments are based on the same validation set, focusing on comparing the coefficient of determination (R 2 ) and root mean square error (RMSE).
[0161] Table 4 Modeling results of the regression model based on DDPM generated samples
[0162]
[0163]
[0164] Where “+” represents adding generated spectral samples to the training set of real samples. and RMSEP represent the coefficient of determination and root mean square error obtained using the same validation set, respectively.
[0165] In order to explore the effect of the number of generated data on the performance of the prediction model, the present invention gradually introduces the data generated by DDPM to evaluate the performance of the model (the step size is 45, i.e. 50% of the sample size of the original training set). As shown in Table 4, with the increase in the number of synthetic samples, the prediction performance of PLSR, SVR and 1D-CNN has been greatly improved. However, when the number of added generated spectral samples reaches a certain number, the performance of the prediction model begins to decline, which indicates that the addition of more synthetic samples does not have a positive effect on the prediction model. Specifically, PLSR achieves the best performance when 180 synthetic samples are added. When 135 synthetic samples are added, SVR achieves the best prediction performance. 1D-CNN achieves the best prediction performance when adding 225 synthetic samples In contrast, 1D-CNN requires the largest number of synthetic samples to achieve optimal performance, and ultimately achieves better performance than the other two prediction models. The present invention believes that the performance differences between these models are mainly caused by the following factors:
[0166] First, the number of training samples has a direct impact on model performance. Table 4 shows that the more generated data, the better. Experimental results show that when a small amount of generated data is added, the predictive performance of each regression model is significantly improved. This is likely because the generated spectral data expands the distribution range of the original samples, effectively alleviating the overfitting problem caused by insufficient samples, thereby enhancing the model's generalization ability. However, when the number of added synthetic samples exceeds a certain threshold, even if the statistical characteristics of the generated data are similar to those of the real data, slight distribution shifts may still affect model learning. As the proportion of synthetic samples increases, this shift effect gradually accumulates, making the model more likely to learn non-generalized features, thereby reducing prediction accuracy. Therefore, an appropriate amount of noise and data diversity is crucial for enhancing the model's generalization ability, while simply adding a large number of similar samples may be counterproductive. Therefore, when using generated data for training, it is necessary to balance the quality and quantity of the data to ensure that it improves the model's predictive ability.
[0167] Secondly, the complexity of the regression model will directly affect its dependence on the amount of data. In Table 4, it can be observed that the performance of the prediction model trained using only the original spectral samples is not ideal, and the prediction performance of 1D-CNN is the worst. The present invention believes that this is because 1D-CNN, as a deep learning model, is highly dependent on the scale of training data for its performance. Compared with traditional modeling methods, 1D-CNN has a deeper network structure and a larger number of parameters, and requires sufficient data to capture the complex nonlinear relationship between spectral features and target variables. Therefore, when the training samples are too few, the model is very likely to fall into an overfitting state and cannot be generalized to the test set. However, when the number of generated samples gradually increases, the deep learning algorithm shows greater performance improvement. Ultimately, the performance of deep learning is better than that of traditional methods.
[0168] In summary, the present invention verifies the application potential of spectral data enhancement based on DDPM in LNC prediction tasks. It also reveals the key issues that need to be paid attention to in the data enhancement process: on the one hand, the amount of generated data should be reasonably controlled to ensure that it plays a positive role in expanding data distribution and improving dataset quality; on the other hand, the characteristics of specific modeling methods should be combined to formulate targeted data enhancement strategies to give full play to the modeling advantages of deep learning models.
[0169] In summary, in order to solve the problem of limited performance of the prediction model caused by insufficient scale of spectral data, the present invention proposes a spectral data enhancement framework based on DDPM and successfully applies it to tomato LNC prediction. Specifically, the present invention collected spectral samples of a total of 120 tomato leaves under different nitrogen gradient conditions, and used DDPM to simultaneously expand their spectral data and nitrogen content data. Through a large number of iterations, synthetic samples that are very similar to the real data are generated, and various qualitative and quantitative analysis methods are used to verify whether the generated spectral samples can be used as real data. After introducing the generated data, the prediction performance of the regression model has been significantly improved. Experimental results show that the data enhancement method based on DDPM can generate synthetic samples that are highly similar to the distribution of real samples, which can provide richer data support for the regression model, thereby enhancing the prediction performance of the model.
Claims
1. A finite sample spectral data enhancement method based on a denoising diffusion probability model, characterized in that: The steps include: Step 1: Collect leaf spectral data of several plant samples, measure the leaf nitrogen content, and concatenate the leaf spectral data with the corresponding nitrogen content to obtain spectrum-nitrogen content pairs, thereby constructing a leaf spectrum-nitrogen content dataset. Step 2: Input the spectral-nitrogen content data of the leaves into the denoising diffusion probability model for data enhancement; The denoising diffusion probability model, in which the forward diffusion process degenerates the original data distribution into a resolvable distribution by gradually adding Gaussian noise, and in which the reverse denoising process utilizes a learnable Markov chain to inversely transform the noise distribution into the target data distribution, thereby reconstructing synthetic data from random samples of known distribution.
2. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 1 is characterized in that: In step 1, each plant sample is cultured to the seedling stage using the same nitrogen concentration, and then subjected to nitrogen stress treatment with different gradients. The culture medium is regularly replaced until the end of the seedling stage. Subsequently, several leaves at the same leaf position are selected from the plants under the same nitrogen treatment gradient as a sample, and their Vis / NIR spectra and nitrogen content are measured to construct a leaf spectrum-nitrogen content dataset.
3. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 2 is characterized in that: The collected Vis / NIR spectral data are preprocessed by one or more of the following methods: Cut off the starting and ending bands of the spectrum and select the spectral data within the range of 400.717-914.63nm; Smoothing preprocessing was applied to the spectral data; The spectral data were normalized using the maximum-minimum method.
4. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 1, characterized in that: The forward diffusion process uses the spectrum-nitrogen content pair as the original data x0, and gradually adds Gaussian noise to make it degenerate into a standard Gaussian distribution. Noise x T , where I is a diagonal covariance matrix with the same dimension as the data; The reverse denoising process first randomly samples a noise x that obeys the standard Gaussian distribution. t , and then the noise is predicted to be x through the deep neural network t Gradually denoise and restore new data with the same distribution as the original data x0 Finally, Separation into leaf spectra and nitrogen content completes the data generation process.
5. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 4 is characterized in that: The formula for gradually adding Gaussian noise is as follows: Among them, β t ∈(0,1), is a hyperparameter used to control the noise intensity, x t Represents the sample after adding t-step Gaussian noise, corresponding to each time step t, a β is defined t ; Derive the distribution at any time step t from x0: in According to the distribution q(x t |x0), x t The formula for fast sampling from x0 is: When the time step is T, the noise x is obtained T ; where ∈ is standard Gaussian noise.
6. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 4 is characterized in that: The reverse denoising process starts with the sample x after adding t steps of Gaussian noise. t Gradually recover new data with the same distribution as the original data x0 Its conditional probability distribution is: Where θ is the trainable parameter of the neural network, and the mean μ θ (x t ,t) and variance ∑ θ (x t ,t) are parameterized by deep neural networks.
7. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 6, characterized in that: The deep neural network predicts the noise ∈ injected in the forward diffusion process θ (x t ,t), and reconstruct the mean parameter as follows: where α t =1-β t represents the noise scheduling coefficient, variance ∑ θ (x t ,t) adopt a fixed strategy; The training goal of the denoising diffusion probability model is to minimize the forward distribution q(x 1:T |x0) and the reverse distribution p θ (x 1:T ), using variational inference, we get the following optimization objective: Among them, ∈ θ (x t ,t) is the noise estimated by the model when taking the parameter θ.
8. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 5, characterized in that: The reverse denoising process generates data by reversely executing the trained Markov chain. The process is as follows: Given a noise prediction network ∈ θ (x t ,t) and a fixed variance ∑ θ (x t ,t), the inference process is from the standard Gaussian distribution Starting from this, the target spectrum - nitrogen content sample is gradually reconstructed through iterative denoising, and finally 9. The finite sample spectral data enhancement method based on the denoising diffusion probability model according to claim 8, characterized in that: A network architecture for spectral data denoising is constructed based on U-Net. The cosine time scheduling strategy is used to replace the linear time scheduling in U-Net. At the same time, a self-conditioning strategy is adopted. The denoising process is as follows: First, a one-dimensional convolution is applied to the noisy spectral data, and position embeddings associated with the noise level are calculated; Subsequently, a series of downsampling and upsampling blocks are used to extract and restore data features. In the encoding phase, each downsampling block captures multi-scale features through two ResNet blocks, an attention mechanism, a residual connection, and a downsampling operation. In the decoding phase, a symmetrical module is used, and upsampling is used instead of downsampling. The encoder and decoder are connected with a bottleneck layer. Finally, through a Resnet block and a one-dimensional convolutional layer, the noise estimate added to the current input data is output, thus completing the denoising process.
10. A method for inverting physiological and biochemical components based on a denoising diffusion probability model, characterized in that: The steps include: Step 1, obtaining synthetic data using the finite sample spectral data enhancement method based on the denoising diffusion probability model according to any one of claims 1 to 9; Step 2: Combine the original data with the synthetic data to construct an extended training set; Step 3: Using the extended training set to train a regression model for inverting plant physiological and biochemical components, the regression model takes spectral data as input and outputs corresponding component prediction values.