Method for tracing origin of American ginseng based on conditional diffusion model

The near-infrared spectral data of American ginseng are generated through the conditional diffusion model of the cross attention mechanism, which solves the problem of insufficient sample size in the traceability of American ginseng origin, and achieves efficient and lossless traceability of origin, improving the traceability accuracy and model stability.

CN120448908APending Publication Date: 2025-08-08HENAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535271.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The traditional American ginseng origin traceability method has insufficient accuracy and poor real-time performance due to the small sample size, and the operation process is destructive, making it difficult to achieve non-destructive testing.

Method used

The conditional diffusion model based on the cross attention mechanism is used to generate high-quality and highly random near-infrared spectral data of American ginseng. Through the unsupervised data augmentation and reverse generation process, the data set is expanded and the model training effect is improved.

Benefits of technology

It improves the accuracy of the origin traceability of American ginseng, reduces the time and labor cost of data collection, enhances the robustness and efficiency of the model, and solves the shortcomings of traditional methods in real-time, non-destructive testing and small sample prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448908A_ABST
    Figure CN120448908A_ABST
Patent Text Reader

Abstract

The invention provides an American ginseng producing area traceability method based on a conditional diffusion model, which comprises the following steps: preparing an American ginseng sample, carrying out near infrared spectrum data acquisition, dividing the acquired near infrared spectrum data into a training set, a test set and a verification set, and carrying out unsupervised data enhancement on the training set; taking the training set after unsupervised data enhancement as training data, training a conditional diffusion model based on a cross attention mechanism, and obtaining generated American ginseng near infrared spectrum data; fusing the generated American ginseng near infrared spectrum data with the training set before unsupervised data enhancement to obtain a fused training set, and by taking the verification set and the fused training set as training data, performing hyper-parameter adjustment and optimization by using the verification set, training the classification model to obtain a classification model; and after training is completed, testing by using the test set and outputting a classification result of the American ginseng producing areas. According to the method, the defects that a traditional method is insufficient in small sample prediction precision, poor in real-time performance and difficult in nondestructive testing are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural information technology, and in particular to a method for tracing the origin of American ginseng. Background Art

[0002] In modern medicine and the food industry, American ginseng has different uses as both a medicine and a food ingredient depending on its origin. Differences in its ingredients lead to different emphasis on efficacy when used medicinally, and as a food ingredient, its flavor and nourishing properties also differ. Traditional methods for tracing the origin of American ginseng rely primarily on manual judgment and complex chemical analysis processes. While these methods can provide relatively accurate classification information, they often require significant time, manpower, and material resources, resulting in poor real-time performance. Furthermore, manual judgment and complex chemical analysis procedures are somewhat destructive, making them difficult to adapt to the needs of large-scale, rapid origin traceability. Near-infrared spectroscopy (NIRS) technology, due to its non-destructive and rapid nature, has garnered widespread attention and application in American ginseng origin traceability research. Traditional machine learning methods for processing NIRS data suffer from tedious manual feature extraction and difficulty in adequately modeling complex spectral data. In recent years, deep learning methods have gradually emerged in the field of origin traceability, leveraging their powerful automatic feature extraction and pattern recognition capabilities. Convolutional neural networks (CNNs) can automatically extract high-order features from spectral data, improving the accuracy of origin traceability. However, due to the relatively small sample size of American ginseng near-infrared spectral data, directly using it as input for the 1D-CNN model would significantly limit the model's prediction accuracy. Acquiring more American ginseng near-infrared spectral samples through conventional methods would require significant human and material resources, which presents numerous practical difficulties. Summary of the Invention

[0003] In response to the technical problems of existing methods such as insufficient classification model accuracy and poor real-time performance due to the small sample size of American ginseng near-infrared spectral data, and the difficulty of non-destructive detection due to the certain destructiveness of the operation process, the present invention proposes a method for tracing the origin of American ginseng based on a conditional diffusion model. By adopting a conditional diffusion model with a cross-attention mechanism, high-quality and random American ginseng near-infrared spectral data with origin information is generated, and the origin of American ginseng is traced based on the generated infrared spectral data, which solves the shortcomings of traditional methods in insufficient prediction accuracy of small samples, poor real-time performance, and difficulty in non-destructive detection.

[0004] In order to achieve the above object, the technical solution of the present invention is achieved as follows:

[0005] A method for tracing the origin of American ginseng based on a conditional diffusion model, comprising the following steps:

[0006] S1: Prepare American ginseng samples and collect near-infrared spectral data. Divide the collected near-infrared spectral data into training set, test set, and validation set, and perform unsupervised data enhancement on the training set.

[0007] S2: Using the training set after unsupervised data augmentation as training data, the conditional diffusion model based on the cross-attention mechanism is trained and the generated American ginseng near-infrared spectral data is obtained;

[0008] S3: Fuse the generated American ginseng near-infrared spectral data with the training set before unsupervised data enhancement to obtain a fused training set. Use the validation set and the fused training set as training data, use the validation set to tune the hyperparameters, and train the classification model. After the training is completed, use the test set to test and output the classification results of the American ginseng origin.

[0009] Furthermore, in step S1, when collecting near-infrared spectral data, each piece of American ginseng near-infrared spectral data is labeled with text information of the American ginseng origin; the near-infrared spectral data is divided using the KS algorithm; and the training set is subjected to unsupervised data enhancement using random offset, random multiplication, and random slope adjustment.

[0010] Furthermore, the method for training the conditional diffusion model based on the cross attention mechanism and obtaining the generated American ginseng near-infrared spectrum data is as follows:

[0011] The near-infrared spectral data of American ginseng in the training set after unsupervised data enhancement is subjected to noise processing through the forward diffusion process to obtain the near-infrared spectral data of American ginseng with noise x t ;

[0012] The near-infrared spectral data of American ginseng with added noise x t The time step t corresponding to the near-infrared spectral data of American ginseng with added noise and the labels of the origin of American ginseng in the training set after unsupervised data enhancement are used as input to train the noise prediction network;

[0013] The trained prediction noise network was used to perform step-by-step denoising through the reverse generation process to obtain the generated American ginseng near-infrared spectral data.

[0014] Furthermore, the method for performing noise addition processing is:

[0015] The near-infrared spectral data x0 of American ginseng in the training set after unsupervised data enhancement is used as input, and the time step t∈[1,T] is randomly selected. The data with the same shape as the original near-infrared spectral data x0 is randomly sampled from the standard Gaussian distribution as the added noise ∈ t, add noise through the noise addition formula of the forward diffusion process to achieve single-step noise addition at any time step t. The noise addition formula is:

[0016]

[0017] Among them, α t is the control parameter at the t-th time step, is the weight coefficient that controls the noise addition process

[0018] Furthermore, the method for predicting the added noise in the forward diffusion process using the predicted noise network is as follows:

[0019] The time step t corresponding to the near-infrared spectrum data of American ginseng with noise is preprocessed by time embedding, the text information labels of American ginseng origin are converted into semantic vectors, and the near-infrared spectrum data x of American ginseng with noise is preprocessed by time embedding. t Perform initial convolution operation to extract features;

[0020] The Conditional-Unet-1D network based on the cross-attention mechanism is used to realize multi-conditional feature fusion by sequentially performing multi-scale downsampling and multi-scale upsampling operations on the time information after time embedding preprocessing, the American ginseng origin text information after text conversion into semantic vectors, and the American ginseng near-infrared spectral data with added noise after the initial convolution operation, and predict the added noise in the forward diffusion process.

[0021] Furthermore, the multi-scale downsampling operation method is as follows: at each scale other than the minimum scale, for the input feature data, local features are extracted and the number of channels is increased by using convolution block I through sequential convolution operations, activation functions, and normalization operations; then, the dimension and channel number of the American ginseng origin text information after the text is converted into a semantic vector is aligned with the feature data after local feature extraction and channel number increase by expanding the dimension and convolution operations to obtain the American ginseng origin text information I with aligned dimension and channel number, and the American ginseng origin text information I with aligned dimension and channel number is used as the key and value, and the feature data after local feature extraction and channel number increase is used as the query to perform cross attention calculation; The cross attention calculation result I is residually connected with the feature data after local feature extraction and channel number improvement to obtain the residual connection result I. The time information after time embedding preprocessing is aligned with the residual connection result I in terms of dimension and channel number through the fully connected layer I to obtain the dimension and channel number alignment result II. The dimension and channel number alignment result II is feature fused with the residual connection result I to obtain the feature fusion result I. The convolution block II extracts the deep feature I by performing convolution operation, activation function and normalization operation on the feature fusion result I in sequence. The deep feature I is reduced in size by using Conv1d. The reduced-size data is used as the input feature data of the next scale and saved in the next scale.

[0022] In the largest scale of the multi-scale downsampling operation, the input feature data is the near-infrared spectral data of American ginseng with added noise after the initial convolution operation.

[0023] Furthermore, the method of the multi-scale upsampling operation is as follows: at each scale other than the maximum scale, the output feature data of the previous scale is jump-connected with the input feature data of the corresponding scale in the multi-scale downsampling operation as the input feature data of the current scale, and the input feature data is subjected to convolution block III through sequential convolution operations, activation functions and normalization operations to extract local features and reduce the number of channels; then, the dimension and channel number of the American ginseng origin text information after the text is converted into a semantic vector is aligned with the feature data extracted from the local features and with the number of channels reduced by the dimension expansion and convolution operations to obtain the dimension and channel number alignment result II, and the dimension and channel number alignment result II is used as the key and value, and the local feature is used as the key. The feature data of the extracted and reduced channel number is used as Query, and cross attention calculation is performed; the cross attention calculation result II is residually connected with the feature data of the local feature extraction and reduced channel number to obtain the residual connection result II; the time information after time embedding preprocessing is aligned with the residual connection result II in terms of dimension and channel number through the fully connected layer II to obtain the dimension and channel number alignment result III, and the dimension and channel number alignment result III is feature fused with the residual connection result II to obtain the feature fusion result II. The convolution block IV extracts deep features II by performing convolution operations, activation functions and normalization operations on the feature fusion result II in sequence, and deconvolution operations are performed on the deep features II. The data after the deconvolution operation is used as the output feature data;

[0024] In the minimum scale of the multi-scale upsampling operation, the input feature data is the concatenation result of the data stored in the minimum scale of the multi-scale downsampling operation; in the maximum scale of the multi-scale upsampling operation, the output feature data of the previous scale is reduced in number of channels through the convolution operation and then output as the added noise in the forward diffusion process of the prediction.

[0025] Furthermore, when the prediction noise network is trained, loss calculation is performed based on the predicted added noise and the real noise using a mean square error loss function, and the network parameters of the prediction noise network are updated using a back propagation algorithm according to the loss calculation result, so as to obtain a prediction noise network with optimal network parameters as the trained prediction noise network.

[0026] Furthermore, the method of using the trained prediction noise network to perform step-by-step denoising through the reverse generation process is as follows:

[0027] Randomly sample a data with the same shape as the original American ginseng near-infrared spectrum data x0 in the standard Gaussian distribution as the initial noise s T , the denoised data s of the current time step t , time step t, and the text information labels of the origin of American ginseng are input into the trained prediction noise network to obtain the prediction of the current time step and add noise ∈ θ (st ,t,labels), using the denoised data s at the current time step t Subtract the prediction of the current time step from the added noise ∈ θ (s t ,t,labels), and follow the reverse dynamics of the Markov chain by adding randomly sampled Gaussian noise to obtain the denoised data s for the next time step t-1 After T time steps of the reverse generation process, the generated American ginseng near-infrared spectrum data s0 is obtained. The calculation formula of the reverse generation process is:

[0028]

[0029] Among them, σ t is the noise scaling factor, and z~N(0,1) is the randomly sampled Gaussian noise.

[0030] Furthermore, the method of the time embedding preprocessing operation is: performing sinusoidal position encoding, ReLu activation function, and full connection operation in sequence on the time step t corresponding to the near-infrared spectrum data of American ginseng with added noise;

[0031] The classification method of the classification model is as follows: taking the validation set and the fusion training set as input, the near-infrared spectral data of American ginseng in the fusion training set is sequentially extracted through multiple convolution blocks; the features are projected into multiple categories corresponding to the labels of the American ginseng origin text information through sequential flattening operations and multi-layer fully connected layers, and the classification accuracy of the model is verified through the validation set during training.

[0032] The beneficial effects of the present invention are:

[0033] Expanding the data set: The conditional diffusion model based on the cross-attention mechanism can generate corresponding American ginseng near-infrared spectral samples according to the conditional label information when generating data. It can generate high-quality, highly random American ginseng near-infrared spectral data with origin information. It can infinitely expand the near-infrared spectral dataset, providing an effective method to solve the problem of insufficient sample data, and does not rely on manual experience judgment and complex chemical analysis processes, thus avoiding sample damage.

[0034] Improving traceability accuracy: The diffusion model denoising process can generate samples that are difficult to obtain or do not appear in the real data set, which helps to improve the accuracy of American ginseng origin traceability and enhance the precision of traceability technology.

[0035] Reduce costs and improve efficiency: Using diffusion models to generate data significantly reduces the time and labor costs of collecting data. Only a small number of samples are needed to train the model, and data can be generated directly when needed, which improves the convenience and efficiency of data processing and solves the problem of poor real-time performance.

[0036] Improved robustness: Small sample data is susceptible to interference from factors such as noise and outliers, leading to unstable model training. When generating data, the diffusion model, through a deep understanding and simulation of data distribution, can filter out or correct noise and anomalies in the original small sample to a certain extent, generating purer and more reliable data. This enhances the model's resistance to various interference factors, ensures the reliability and effectiveness of model training under small sample conditions, and reduces model bias and errors caused by data quality issues.

[0037] Solve traditional problems: The fusion of generated data and collected data enriches data features, realizes accurate origin traceability prediction, and solves the shortcomings of traditional methods in real-time, non-destructive testing and small sample prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 It is the overall flow chart of the present invention.

[0040] Figure 2 The overall structure diagram of the conditional diffusion model and classification model based on the cross-attention mechanism of the present invention.

[0041] Figure 3 A flow chart is generated for the reverse process of the present invention.

[0042] Figure 4 This is a diagram of the noise prediction network structure of the present invention.

[0043] Figure 5 This is the original near-infrared spectrum data of American ginseng of the present invention.

[0044] Figure 6 This is the near-infrared spectrum data of American ginseng generated by the present invention.

[0045] Figure 7 This is a structural diagram of the classification model of the present invention.

[0046] Figure 8 This is the confusion matrix diagram of the American ginseng near-infrared spectrum test set of the present invention.

[0047] Figure 9 This is an evaluation graph of the classification accuracy, precision, recall rate, and F1-score of the American ginseng near-infrared spectrum test set of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0049] A method for tracing the origin of American ginseng based on the conditional diffusion model, such as Figure 1 As shown, the steps include:

[0050] S1: Prepare American ginseng samples and collect near-infrared spectroscopy (NIR) data. Divide the collected NIR spectroscopy data into training set, test set, and validation set, and perform unsupervised data enhancement on the training set.

[0051] Preparation of American ginseng samples: The American ginseng samples used in the present invention come from five regions (Weihai, Shandong, China, Rongcheng, Shandong, China, Baishan, Jilin, China, Montreal, Canada, and Wisconsin, USA), and the origin of these samples has been verified by experts. All samples are cut into uniform sizes, 5-8 pieces are regarded as one sample, and there are 300 groups in total. All samples are healthy, complete, and stored in a refrigerated environment at 2°C. Before measuring the near-infrared spectral data and evaluation indicators, all samples are taken out and kept at a temperature of 20°C for 8 hours to ensure data consistency during the near-infrared spectral analysis process.

[0052] Near-infrared spectroscopy (NIR) data collection: The near-infrared spectral data of American ginseng samples were collected using a DA7250NIR analyzer in the wavelength range of 950 to 1650 nm with a resolution of 0.5 nm. The device has a built-in tungsten halogen light source and collects spectral data after 30 minutes of warm-up. The sliding window technology is used to reduce the original spectral dimension from 1401 to 272 (i.e., the number of spectral bands is 272) to reduce redundant information. 300 groups of samples were collected, and a total of 2213 pieces of American ginseng near-infrared spectral data were collected, such as Figure 4 As shown in the figure, each piece of American ginseng near-infrared spectrum data is annotated with labels of the American ginseng origin text information.

[0053] Near-infrared spectral preprocessing: The collected near-infrared spectral data was divided into training, test, and validation sets. Unsupervised data augmentation was performed on the training set. Using the KS algorithm, the 2,213 near-infrared spectral data set of American ginseng was divided into a training set of 1,328, a test set of 664, and a validation set of 221. The training set, validation set, and test set were divided in a 6:3:1 ratio. Unsupervised data augmentation techniques such as random offset, random multiplication, and random slope adjustment were used to increase the number of near-infrared spectra in the training set, thereby better training the diffusion model.

[0054] The above preprocessing process did not use the test set and did not cause data leakage, ensuring the independence of the test set and the objectivity of the model evaluation; these unsupervised data augmentation operations only changed the appearance of the spectrum (intensity, offset or slope) at the data level, but did not change the origin information characteristics represented by the spectrum; ensuring that data augmentation would not destroy the true association between the spectrum and the origin, thereby ensuring that the subsequent model can accurately learn the characteristics related to the origin.

[0055] S2: Using the training set after unsupervised data enhancement as training data, the conditional diffusion model based on the cross attention mechanism is trained and the generated American ginseng near-infrared spectral data is obtained, such as Figure 2 shown.

[0056] Specifically, the method for training the conditional diffusion model based on the cross attention mechanism and obtaining the generated American ginseng near-infrared spectrum data is as follows: the American ginseng near-infrared spectrum data in the training set after unsupervised data enhancement is subjected to noise processing through the forward diffusion process to obtain the American ginseng near-infrared spectrum data x with noise added t .

[0057] Specifically, the method for performing noise addition processing is:

[0058] The near-infrared spectral data of American ginseng in the training set after unsupervised data enhancement is recorded as x0, with a shape of (1024, 1, 272), where 1024 represents the batch, 1 represents the feature dimension, that is, the number of channels, and 272 is the length of the one-dimensional data; take the near-infrared spectral data of American ginseng x0 as input, randomly take the time step t∈[1,T], and randomly take the Gaussian noise with the same shape as the original near-infrared spectral data of American ginseng x0 as the added noise ∈ t The noise addition formula of the forward diffusion process is used to directly implement the single-step noise addition of the near-infrared spectrum data of American ginseng from x0 to any time step t, and the near-infrared spectrum data of American ginseng with added noise corresponding to time step t is obtained. t :

[0059]

[0060] Among them, α t is the control parameter at the t-th time step, is the weight coefficient that controls the noise addition process As t increases, the weight coefficient Gradually decrease.

[0061] Furthermore, the near-infrared spectral data of American ginseng with added noise x tThe time step t corresponding to the near-infrared spectral data of American ginseng with added noise and the labels of the origin of American ginseng in the training set after unsupervised data enhancement are used as input to train the noise prediction network.

[0062] Specifically, such as Figure 4 As shown, the method for predicting the added noise in the forward diffusion process using the predicted noise network is:

[0063] The time step t corresponding to the near-infrared spectral data of American ginseng with added noise is subjected to time embedding preprocessing. The shape of the time step t is converted from (1024) to (1024, 4), where 1024 represents the batch and 4 represents the number of channels.

[0064] The method of the time embedding preprocessing operation is as follows: for the time step t corresponding to the near-infrared spectral data of American ginseng with added noise, sinusoidal position encoding, ReLu activation function, and full connection operation are performed in sequence; the time information after the output time embedding preprocessing is passed to the downstream task as the final result to provide key time information representation for the model in time series related tasks.

[0065] The text information labels of the origin of American ginseng are converted into semantic vectors, and the shape becomes (1024, 384), where 1024 represents the batch and 384 represents the length; the specific method of converting the text into the semantic vector operation described in this embodiment is: using the SentenceTransformer library to load the all-MiniLM-L6-v2 pre-trained model, read the txt file containing the text information labels of the origin of American ginseng line by line, remove the leading and trailing spaces of each line of text, and then use the encode() method of the SentenceTransformer library to convert the processed text into a semantic vector, and then move the processing result of the encode() method from the GPU to the CPU and convert it into a NumPy array, and finally return a NumPy array containing the semantic vectors of all texts.

[0066] For the near-infrared spectral data of American ginseng with added noise x t An initial convolution operation is performed, followed by feature extraction and an increase in the number of channels. The shape of the network is transformed from (1024, 1, 272) to (1024, 8, 272). This increase in channels enables the model to extract richer, more multi-dimensional feature information from the input data. Temporal embedding preprocessing, the conversion of text into semantic vectors, and the initial convolution operation are performed in parallel.

[0067] Furthermore, the Conditional-Unet-1D network based on cross-attention is used to realize multi-conditional feature fusion by sequentially performing multi-scale downsampling operations and multi-scale upsampling operations on the time information after time embedding preprocessing, the American ginseng origin text information after text conversion into semantic vectors, and the American ginseng near-infrared spectral data with added noise after the initial convolution operation, and predict the added noise in the forward diffusion process.

[0068] like Figure 4 As shown, this embodiment performs 4 downsampling operations in 5 scales. Figure 4 The fully connected layers, expanded dimensions, and convolution operations shown in the figure are different at each scale and are not fully drawn in the figure.

[0069] Specifically, in the largest scale of the multi-scale downsampling operation, that is, the first scale, the input feature data is the near-infrared spectrum data of American ginseng with noise added after the initial convolution operation, and the shape is (1024, 8, 272); for the input feature data, the convolution block I1 is used to extract local features and increase the number of channels through the convolution operation, ReLu activation function and normalization operation in sequence, and the shape becomes (1024, 16, 272). The I1 in the convolution block I1 represents the convolution block I in the first scale. Similar expressions using Roman numerals plus Arabic numerals in the following description have the same meaning.

[0070] Furthermore, by expanding the dimension and performing convolution operations, the dimension and number of channels of the American ginseng origin text information after the text is converted into a semantic vector are aligned with the feature data after local feature extraction and increasing the number of channels, obtaining the American ginseng origin text information Ⅰ1 with aligned dimensions and channels in the largest scale, and its shape changes from (1024, 384) to (1024, 1, 384) and then to (1024, 16, 384); the expansion dimension operation in the expansion dimension and convolution operation adopts the unsqueeze method in PyTorch.

[0071] Furthermore, the feature data after local feature extraction and channel number improvement is used as Query, and the American ginseng origin text information Ⅰ1 aligned with dimension and channel number is used as Key and Value to perform cross attention calculation; specifically, Query (Q) is generated from the feature data after local feature extraction and channel number improvement, and the shape becomes (1024, 272, 16), and Key (K) and Value (V) are generated from the American ginseng origin text information Ⅰ1 aligned with dimension and channel number, and the shapes are both (1024, 384, 16); then the attention score matrix is calculated, and the shape becomes (1024, 272, 384). The Softmax function is applied to obtain the attention weights, which have a shape of (1024, 272, 384). Finally, the weighted sum is calculated by multiplying the attention weights with the Value to obtain an output matrix with a shape of (1024, 272, 6). The output matrix is transposed to obtain the final cross-attention calculation result I1, which has a shape of (1024, 16, 272). The final data representation that integrates the origin label information is obtained, ensuring the consistency of the data and labels in the feature dimension while handling the difference in their sequence length. The transposition operation is implemented through the permute operation in PyTorch. The formula for cross-attention calculation is:

[0072]

[0073] The cross attention calculation result Ⅰ1 is residually connected with the feature data after local feature extraction and channel number increase; the time information after time embedding preprocessing is aligned with the residual connection result Ⅰ1 in terms of dimension and channel number through the fully connected layer Ⅰ1, that is, the fully connected layer in the first scale in the multi-scale downsampling operation, to obtain the dimension and channel number alignment result Ⅱ1, and the shape becomes (1024, 16, 1). The dimension and channel number alignment result Ⅱ1 is feature fused with the residual connection result Ⅰ1, and the shapes are respectively (1024, 16, 272) and (1024, 16, 1) through the broadcast mechanism in PyTorch. The residual connection result Ⅰ1 and the dimension and channel number alignment result Ⅱ1 are added to realize feature fusion, ensuring that the shape after feature fusion is still (1024, 16, 272); this can retain the original data, that is, all the information in the American ginseng near-infrared spectrum data, while incorporating additional origin information obtained through the cross-attention mechanism and time information; the convolution block Ⅱ1 extracts the deep feature Ⅰ1 by performing convolution operations, ReLu activation functions and normalization operations on the feature fusion result Ⅰ1 in sequence, and then uses Conv1d to reduce the data size of the deep feature Ⅰ1 as the input feature data of the second scale.

[0074] Specifically, for the second scale to the fourth scale, the multi-scale downsampling operation method is as follows: for the input feature data, using convolution block I to extract local features and increase the number of channels through sequential convolution operations, ReLu activation functions, and normalization operations; then, by expanding the dimension and performing convolution operations, the dimension of the American ginseng origin text information after the text is converted into a semantic vector is aligned with the feature data after local feature extraction and channel number increase, to obtain the American ginseng origin text information I with aligned dimension and channel number, and the American ginseng origin text information I with aligned dimension and channel number is used as the key and value, and the feature data after local feature extraction and channel number increase is used as the query, and cross attention calculation is performed; The cross-attention calculation result I is residually connected with the feature data after local feature extraction and channel number improvement to obtain the residual connection result I; the time information after time embedding preprocessing is aligned with the residual connection result I in terms of dimension and channel number through the fully connected layer I to obtain the dimension and channel number alignment result II, and the dimension and channel number alignment result II is feature fused with the residual connection result I to obtain the feature fusion result I. The convolution block II extracts the deep feature I by performing convolution operation, ReLu activation function and normalization operation on the feature fusion result I in sequence, and then uses Conv1d to reduce the data size of the deep feature I. The reduced-size data is used as the input feature data of the next scale and saved in the next scale.

[0075] In the multi-scale downsampling operation, increasing the number of channels and reducing the data size help to extract richer and more complex feature representations while reducing the amount of computation and the risk of overfitting; it enables the network to encode more information in less space by reducing the data size, forcing the network to learn more compact and discriminative features; in addition, this operation expands the receptive field, enabling the network to aggregate information over a larger range and capture the overall structure of the data; downsampling also promotes the fusion of multi-scale features, providing different levels of details and contextual information for subsequent upsampling.

[0076] like Figure 4 As shown, this embodiment performs four upsampling operations at five scales.

[0077] Specifically, in the smallest scale of the multi-scale upsampling operation, that is, the fifth scale, the input feature data of the upsampling operation is the result of splicing the output features in the multi-scale downsampling operation themselves, and the size changes from (1024, 128, 17) to (1024, 256, 17).

[0078] Furthermore, convolution block III5 is used to extract local features and reduce the number of channels through sequential convolution operations, ReLu activation functions and normalization operations, and the size becomes (1024, 64, 17).

[0079] Furthermore, the dimension and channel number of the American ginseng origin text information after the text is converted into a semantic vector are aligned with the feature data after local feature extraction and channel reduction by expanding the dimension and performing convolution operations, and the American ginseng origin text information Ⅱ5 with aligned dimension and channel number is obtained. The American ginseng origin text information Ⅱ5 with aligned dimension and channel number is used as the Key and Value, and the feature data after local feature extraction and channel reduction is used as the Query, and cross-attention calculation is performed, and the result shape is (1024, 64, 17).

[0080] Furthermore, the cross-attention calculation result Ⅱ5 is residually connected with the feature data after local feature extraction and channel reduction to obtain the residual connection result Ⅱ5; the time information after time embedding preprocessing is aligned with the residual connection result Ⅱ5 in terms of dimension and channel number through the fully connected layer Ⅱ5, and the dimension and channel number alignment result Ⅲ5 is feature fused with the residual connection result Ⅱ5, and the result shape is (1024, 64, 17).

[0081] Furthermore, convolution block IV extracts deep features II5 by performing convolution operation, ReLu activation function and normalization operation on the feature fusion result II5 in sequence, and then performs deconvolution operation on the deep features II5. The result shape is (1024, 64, 34). The data after deconvolution operation is used as the output feature data of the fifth scale.

[0082] In the 4th scale in the multi-scale upsampling operation, the output feature data of the 5th scale is jump-connected with the input feature data of the 4th scale in the multi-scale downsampling operation (shape is (1024, 64, 34)) as the input of the current scale, with a shape of (1024, 128, 34).

[0083] Specifically, for the second scale to the fourth scale, the multi-scale upsampling operation method is: the output feature data of the previous scale is jump-connected with the input feature data of the corresponding scale in the multi-scale downsampling operation as the input feature data of the current scale, and the input feature data is subjected to convolution block III through sequential convolution operations, ReLu activation functions and normalization operations to extract local features and reduce the number of channels; then, the dimension and channel number of the American ginseng origin text information after the text is converted into a semantic vector by expanding the dimension and convolution operations are aligned with the feature data after local feature extraction and channel reduction to obtain the American ginseng origin text information II with aligned dimension and channel number, and the American ginseng origin text information II with aligned dimension and channel number is used as K ey and Value, and use the feature data after local feature extraction and channel reduction as Query to perform cross-attention calculation; perform residual connection on the cross-attention calculation result II and the feature data after local feature extraction and channel reduction; align the dimension and channel number of the time information after time embedding preprocessing with the residual connection result II through the fully connected layer II, and perform feature fusion on the dimension and channel number alignment result III and the residual connection result II to obtain the feature fusion result II. The convolution block IV extracts the deep feature II by performing convolution operation, ReLu activation function and normalization operation on the feature fusion result II in sequence, and then performs deconvolution operation on the deep feature II. The data after the deconvolution operation is used as the current scale output feature data in the multi-scale upsampling operation.

[0084] In the maximum scale of the multi-scale upsampling operation, that is, the first scale, the output feature data of the second scale is reduced in number of channels through a convolution operation and then output as a prediction with noise added.

[0085] When training the predicted noise network, the loss is calculated based on the predicted added noise and the real noise through the mean square error loss function and the network parameters are updated. According to the loss calculation result, the network parameters of the predicted noise network are updated through the back propagation algorithm to obtain the predicted noise network with the optimal network parameters as the trained predicted noise network.

[0086] Skip connections effectively retain high-resolution detail information, promote the fusion of features at different levels, improve gradient flow, reduce information loss, improve the prediction accuracy of the model, increase the flexibility of network design, and allow the construction of deeper network structures without worrying about the gradient disappearance problem, thereby learning more complex feature representations.

[0087] Replacing the traditional maximum pooling layer that changes the data size with a convolutional layer (i.e., an operation that reduces the data size) can increase the depth of the network and allow the model to learn more complex feature hierarchies. Compared with traditional splicing and additive fusion methods, using a cross-attention mechanism to embed label information into American ginseng near-infrared spectroscopy data has the following advantages: the cross-attention mechanism allows the model to dynamically focus on the part of the data most relevant to the label, thereby capturing richer contextual information, thus helping to reduce unnecessary feature combinations and reduce the risk of overfitting. Splicing and additive fusion lack this context-awareness capability. The cross-attention mechanism can flexibly handle inputs of different lengths and dimensions, while splicing and additive fusion usually require the inputs to have the same dimensions.

[0088] Furthermore, the trained prediction noise network is used to perform step-by-step denoising through the reverse generation process to obtain the generated American ginseng near-infrared spectrum data. Figure 6 shown.

[0089] Specifically, such as Figure 3 As shown, the method for gradually denoising through the reverse generation process is:

[0090] Randomly sample a data with the same shape as the original American ginseng near-infrared spectrum data x0 in the standard Gaussian distribution as the initial noise s T , the denoised data s of the current time step t , time step t, and the text information labels of the origin of American ginseng are input into the trained prediction noise network to obtain the prediction of the current time step and add noise ∈ θ (s t ,t,labels), using the denoised data s at the current time step t Subtract the prediction of the current time step from the added noise ∈ θ (s t ,t,labels), and follow the reverse dynamics of the Markov chain by adding randomly sampled Gaussian noise to obtain the denoised data s for the next time step t-1 After T time steps of the reverse generation process, the generated American ginseng near-infrared spectrum data s0 is obtained. The calculation formula of the reverse generation process is:

[0091]

[0092] in, is the noise scaling factor, z~N(0,1) is the additional injected randomly sampled Gaussian noise, which is used to maintain the randomness of the generation process to match the Gaussian distribution assumption of forward diffusion.

[0093] S3: Fuse the generated American ginseng near-infrared spectral data with the training set before unsupervised data enhancement to obtain a fused training set. Use the validation set and the fused training set as training data, use the validation set to tune the hyperparameters, and train the classification model. After training, use the test set to test and output the classification results of the American ginseng origin. Figure 7 shown.

[0094] The classification method of the classification model is as follows: taking the validation set and the fusion training set as input, the near-infrared spectral data of American ginseng in the fusion training set is sequentially extracted through multiple convolution blocks; the features are projected into multiple categories corresponding to the labels of the American ginseng origin text information through sequential flattening operations and multi-layer fully connected layers, and the classification accuracy of the model is verified through the validation set during training.

[0095] In this embodiment, the shape of the American ginseng near-infrared spectral data in the input fusion training set is (32, 1, 272), and 4 convolution blocks are used. Convolution operations using different convolution kernels, ReLU activation functions and maximum pooling operations are performed in each convolution block in sequence. From the first convolution block to the fourth convolution block, the convolution kernel sizes of the convolution operations are 1×3, 4×3, 8×3, and 16×3, respectively; by using 4 convolution blocks with different convolution kernels, the data length is reduced, the nonlinear expression ability of the model is increased, and the computational complexity is reduced. The data shape becomes (32, 16, 17), and the pooling window in the maximum pooling is set to 2, the step size is 2, and the padding is 0.

[0096] The 3-dimensional tensor (32, 16, 17) is flattened into a 2-dimensional (32, 272) through a flattening operation. In this embodiment, two fully connected layers are used. The 272-dimensional data is projected to 1024 dimensions through the first fully connected layer to enhance the information expression capability. The ReLU activation function calculation is performed after the first fully connected layer to increase nonlinearity, enabling the model to learn complex features. The 1024-dimensional features are mapped to 5 categories, namely the American ginseng category, through the second fully connected layer. The final data shape becomes (32, 5).

[0097] Finally, the loss function is calculated through the cross-entropy loss function, and the network parameters are updated through the back-propagation algorithm. The classification network after the updated network parameters is verified using the validation set. The training is completed when the preset training rounds are reached or the verification accuracy requirements are met, and the classification model with the optimal parameters is obtained. The classification model with the optimal parameters after training is used to output the classification results of the origin of American ginseng.

[0098] Experimental setup and model evaluation:

[0099] In the embodiment of the present invention, a total of 2213 pieces of near-infrared spectral data of American ginseng covering five different origins and 272 spectral bands are used, and the KS algorithm is used to divide them into training set, validation set and test set in a ratio of 6:3:1. Unsupervised data enhancement processing is carried out on the training set obtained by division to train the diffusion model. After training, the loss function value is about 0.0015. The generation process can randomly sample from the standard Gaussian distribution and generate near-infrared spectral samples of American ginseng from different origins. The generated American ginseng spectral data is shown in the attached figure. Figure 6 As shown in the figure, the generated American ginseng dataset was fused with the original training set, expanding the number of American ginseng samples from each origin in the training set to 600, for a total of 3,000 near-infrared spectral data points for five categories of American ginseng. Finally, the classification model was trained using the fused training set and the original training set. The accuracy of the two American ginseng classification models was tested using the initially partitioned validation and test sets to verify the effectiveness of the generated American ginseng dataset and whether it could improve the classification accuracy of the test set. In this experiment, the validation and test sets did not cause any data leakage.

[0100] The generated near-infrared spectral dataset of American ginseng was fused with the original dataset to train a classification model to trace the origin of American ginseng from different origins. The results showed good accuracy, such as Figure 8 As shown. Let the row of the confusion matrix be i and the column be j. The numbers in the confusion matrix represent the number of samples of class i that are predicted to be class j. Figure 8 The confusion matrix on the left is the experimental result of the test set under the model trained with the original training set. Figure 8 The confusion matrix on the middle right shows the experimental results of the test set using the model trained on the fused training set. The results show that all American ginseng samples from Weihai, Shandong, and Wisconsin, USA, were correctly predicted. There were slight improvements in Rongcheng, Shandong, and Baishan, Jilin, while there was little change in Montreal, Canada, but the overall performance improved.

[0101] like Figure 9 As shown, Figure 9 The left picture shows the accuracy, precision, recall rate, and F1 score of the original American ginseng near-infrared spectroscopy dataset. Figure 9 The right figure in the figure shows the accuracy, precision, recall, and F1 score of the classification using the generated American ginseng near-infrared spectroscopy dataset fused with the original dataset. To evaluate the model's performance in the American ginseng origin traceability task, the support value indicates the number of samples with each origin in the test set.

[0102] The leftmost column represents the origin of American ginseng, with 0 representing Weihai, Shandong, 1 representing Rongcheng, Shandong, 2 representing Baishan, Jilin, 3 representing Montreal, Canada, and 4 representing Wisconsin, USA. Comparison of experimental results revealed that the present invention improved the F1 scores of American ginseng from origins 0, 1, 2, and 4, with 4 achieving a significant improvement. The overall classification accuracy of these samples increased by 3 percentage points, demonstrating the effectiveness of the near-infrared spectra generated for American ginseng using this method.

[0103] The present invention uses the concept of generative artificial intelligence. It can generate high-quality, highly random American ginseng near-infrared spectral data with origin information, and can achieve unlimited expansion of the near-infrared spectral dataset, providing a method for expanding the American ginseng near-infrared spectral dataset.

[0104] In this paper, this noise actually represents data diversity. During the denoising process, the model can generate samples that are difficult to obtain or do not yet appear in real datasets. The emergence of these novel sample data helps improve the accuracy of American ginseng origin traceability, thereby enhancing the precision of traceability technology.

[0105] Generating near-infrared spectral data for American ginseng using a conditional diffusion model significantly reduces the time and labor costs of data collection. Only a certain amount of sample data needs to be collected as a training set for the diffusion model. When additional data is needed, it can be generated using the model, eliminating the need for tedious experimental collection. This greatly improves the convenience and efficiency of data processing.

[0106] In summary, the near-infrared spectral data of American ginseng generated by this method is integrated with the collected near-infrared spectral data set of American ginseng, which enriches the near-spectral data characteristics of American ginseng, realizes the prediction of the origin traceability of American ginseng, improves the accuracy of detection, and solves the shortcomings of traditional methods in real-time, non-destructive detection, and prediction accuracy of small samples.

[0107] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for tracing the origin of American ginseng based on a conditional diffusion model, characterized in that: Including steps: S1: Prepare American ginseng samples and collect near-infrared spectral data. Divide the collected near-infrared spectral data into training set, test set, and validation set, and perform unsupervised data enhancement on the training set. S2: Using the training set after unsupervised data augmentation as training data, the conditional diffusion model based on the cross-attention mechanism is trained and the generated American ginseng near-infrared spectral data is obtained; S3: Fuse the generated American ginseng near-infrared spectral data with the training set before unsupervised data enhancement to obtain a fused training set. Use the validation set and the fused training set as training data, use the validation set to tune the hyperparameters, and train the classification model. After the training is completed, use the test set to test and output the classification results of the American ginseng origin.

2. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 1, characterized in that: In step In S1, when collecting near-infrared spectral data, each piece of American ginseng near-infrared spectral data was annotated with labels of American ginseng origin; the KS algorithm was used to divide the near-infrared spectral data; and random offset, random multiplication, and random slope adjustment were used to perform unsupervised data enhancement on the training set.

3. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 2, characterized in that: The method for training the conditional diffusion model based on the cross attention mechanism and obtaining the generated American ginseng near-infrared spectrum data is as follows: The near-infrared spectral data of American ginseng in the training set after unsupervised data enhancement is subjected to noise processing through the forward diffusion process to obtain the near-infrared spectral data of American ginseng with noise x t ; The near-infrared spectral data of American ginseng with added noise x t The time step t corresponding to the near-infrared spectral data of American ginseng with added noise and the labels of the origin of American ginseng in the training set after unsupervised data enhancement are used as input to train the noise prediction network; The trained prediction noise network was used to perform step-by-step denoising through the reverse generation process to obtain the generated American ginseng near-infrared spectral data.

4. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 3, characterized in that: The method for performing noise addition processing is: The near-infrared spectral data x0 of American ginseng in the training set after unsupervised data enhancement is used as input, and the time step t∈[1,T] is randomly selected. The data with the same shape as the original near-infrared spectral data x0 is randomly sampled from the standard Gaussian distribution as the added noise ∈ t , add noise through the noise addition formula of the forward diffusion process to achieve single-step noise addition at any time step t. The noise addition formula is: Among them, α t is the control parameter at the t-th time step, is the weight coefficient that controls the noise addition process 5. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 3 or 4, characterized in that: The method for predicting the added noise in the forward diffusion process using the predicted noise network is as follows: The time step t corresponding to the near-infrared spectrum data of American ginseng with noise is preprocessed by time embedding, the text information labels of American ginseng origin are converted into semantic vectors, and the near-infrared spectrum data x of American ginseng with noise is preprocessed by time embedding. t Perform initial convolution operation to extract features; The Conditional-Unet-1D network based on the cross-attention mechanism is used to realize multi-conditional feature fusion by sequentially performing multi-scale downsampling and multi-scale upsampling operations on the time information after time embedding preprocessing, the American ginseng origin text information after text conversion into semantic vectors, and the American ginseng near-infrared spectral data with added noise after the initial convolution operation, and predict the added noise in the forward diffusion process.

6. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 5, characterized in that: The multi-scale downsampling operation method is as follows: at each scale other than the minimum scale, for the input feature data, local features are extracted and the number of channels is increased by using a convolution block I through sequential convolution operations, activation functions, and normalization operations; then, the dimension and channel number of the American ginseng origin text information after the text is converted into a semantic vector is aligned with the feature data after local feature extraction and channel number increase by expanding the dimension and convolution operations to obtain the American ginseng origin text information I with aligned dimension and channel number; the American ginseng origin text information I with aligned dimension and channel number is used as the key and value, and the feature data after local feature extraction and channel number increase is used as the query to perform cross attention calculation; the cross attention calculation result I is residually connected with the feature data after local feature extraction and channel number increase to obtain the residual connection result I; The time information after time embedding preprocessing is aligned with the residual connection result I in terms of dimension and channel number through the fully connected layer I to obtain the dimension and channel number alignment result II. The dimension and channel number alignment result II is fused with the residual connection result I to obtain the feature fusion result I. The convolution block II extracts the deep feature I by performing convolution operation, activation function and normalization operation on the feature fusion result I in sequence. The deep feature I is reduced in size by using Conv1d. The reduced-size data is used as the input feature data of the next scale and saved in the next scale. In the largest scale of the multi-scale downsampling operation, the input feature data is the near-infrared spectral data of American ginseng with added noise after the initial convolution operation.

7. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 5, characterized in that: The multi-scale upsampling operation method is as follows: at each scale other than the maximum scale, the output feature data of the previous scale is jump-connected with the input feature data of the corresponding scale in the multi-scale downsampling operation and used as the input feature data of the current scale, and the input feature data is subjected to convolution block III through sequential convolution operations, activation functions and normalization operations to extract local features and reduce the number of channels; then, the dimensionality and channel number of the American ginseng origin text information after the text is converted into a semantic vector is aligned with the feature data of the local feature extraction and the reduced channel number by expanding the dimension and convolution operations to obtain the dimension and channel number alignment result II, and the dimension and channel number alignment result II is used as the key and value, and the local feature extraction is used as the key and value. The feature data with the reduced number of channels is used as Query, and cross attention calculation is performed; the cross attention calculation result II is residually connected with the feature data with local feature extraction and reduced number of channels to obtain the residual connection result II; the time information after time embedding preprocessing is aligned with the residual connection result II in terms of dimension and channel number through the fully connected layer II to obtain the dimension and channel number alignment result III, the dimension and channel number alignment result III is feature fused with the residual connection result II to obtain the feature fusion result II, the convolution block IV extracts the deep feature II by performing convolution operation, activation function and normalization operation on the feature fusion result II in sequence, and deconvolution operation is performed on the deep feature II. The data after deconvolution operation is used as the output feature data; In the minimum scale of the multi-scale upsampling operation, the input feature data is the concatenation result of the data stored in the minimum scale of the multi-scale downsampling operation; in the maximum scale of the multi-scale upsampling operation, the output feature data of the previous scale is reduced in number of channels through the convolution operation and then output as the added noise in the forward diffusion process of the prediction.

8. The method for tracing the origin of American ginseng based on the conditional diffusion model according to any one of claims 3, 4, 6 or 7, characterized in that: When training the predicted noise network, loss calculation is performed based on the predicted added noise and the real noise using a mean square error loss function. The network parameters of the predicted noise network are updated using a back propagation algorithm based on the loss calculation result, and a predicted noise network with optimal network parameters is obtained as the trained predicted noise network.

9. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 3, characterized in that: The method of using the trained prediction noise network to perform step-by-step denoising through the reverse generation process is as follows: Randomly sample a data with the same shape as the original American ginseng near-infrared spectrum data x0 in the standard Gaussian distribution as the initial noise s T , the denoised data s of the current time step t , time step t, and the text information labels of the origin of American ginseng are input into the trained prediction noise network to obtain the prediction of the current time step and add noise ∈ θ (s t ,t,labels), using the denoised data s at the current time step t Subtract the prediction of the current time step from the added noise ∈ θ (s t ,t,labels), and follow the reverse dynamics of the Markov chain by adding randomly sampled Gaussian noise to obtain the denoised data s for the next time step t-1 After T time steps of the reverse generation process, the generated American ginseng near-infrared spectrum data s0 is obtained. The calculation formula of the reverse generation process is: Among them, σ t is the noise scaling factor, and z~N(0,1) is the randomly sampled Gaussian noise.

10. The method for tracing the origin of American ginseng based on the conditional diffusion model according to claim 5, characterized in that: The method of the time embedding preprocessing operation is as follows: for the time step t corresponding to the near-infrared spectrum data of American ginseng with added noise, sinusoidal position encoding, ReLu activation function, and full connection operation are sequentially performed; The classification method of the classification model is as follows: taking the validation set and the fusion training set as input, the near-infrared spectral data of American ginseng in the fusion training set is sequentially extracted through multiple convolution blocks; the features are projected into multiple categories corresponding to the labels of the American ginseng origin text information through sequential flattening operations and multi-layer fully connected layers, and the classification accuracy of the model is verified through the validation set during training.