Soybean producing area traceability detection method based on laser-induced breakdown spectroscopy technology
By analyzing soybean LIBS spectra using a one-dimensional residual network (1D-ResNet), the high dimensionality and nonlinearity of LIBS spectral data processing in existing technologies are solved, enabling rapid and accurate traceability of soybean origins, especially in cases of mixed origins and scarce sample sizes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- OCEAN UNIV OF CHINA
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-08
AI Technical Summary
Existing LIBS-based methods for tracing the origin of agricultural products struggle to accurately extract feature information when processing high-dimensional, nonlinear spectral data. Furthermore, they suffer from low accuracy when dealing with mixed origins and limited sample sizes, failing to meet the complex and ever-changing real-world traceability needs.
A one-dimensional residual network (1D-ResNet) was used to analyze the soybean LIBS spectrum. A deep learning model adapted to the spectral data was constructed through residual modules and skip connections. Combined with simple sample preprocessing and spectral feature extraction, the model's discrimination ability was improved.
It enables rapid and accurate origin tracing of soybean samples from unknown sources, and can identify single-origin and multi-origin mixed samples, improving the accuracy and robustness of the judgment and adapting to complex trade and circulation environments.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural product traceability and testing technology, and more specifically, to a soybean origin traceability and testing method based on laser-induced breakdown spectroscopy technology. Background Technology
[0002] Soybeans are an important global food and oilseed crop. Due to differences in soil, climate, and other natural conditions, soybeans from different producing areas exhibit variations in their nutritional composition and elemental elements. Accurate traceability of soybean origin is crucial for ensuring food safety, protecting consumer rights, and regulating trade order.
[0003] Currently, commonly used origin traceability technologies encompass stable isotope ratio analysis, mineral elemental fingerprinting, molecular biology, and metabolomics. While each method has its advantages, they often suffer from drawbacks such as complex sample pretreatment, long detection cycles, or the need for chemical reagents. Laser-induced breakdown spectroscopy (LIBS), as an emerging elemental analysis technique, utilizes high-energy pulsed lasers to irradiate the sample surface, causing it to ablate and ionize at instantaneous high temperatures, forming a high-temperature plasma. As the plasma rapidly cools, the excited-state elemental atoms and ions relax and emit characteristic spectra, revealing elemental composition information reflecting the sample's origin. This method offers advantages such as speed, non-destructive testing, and simultaneous multi-element detection, and has received widespread attention in the field of agricultural product origin traceability in recent years.
[0004] However, spectral signals obtained by LIBS detection often suffer from background noise interference and overlapping spectral lines, making it difficult to accurately extract origin information from spectral data using traditional spectral analysis methods, thus hindering precise origin identification. Against this backdrop, machine learning, with its powerful data mining, feature extraction, and pattern recognition capabilities, provides an effective way to solve these problems, becoming a crucial support for the precise application of LIBS technology in agricultural product origin traceability. Classification algorithms such as Linear Discriminant Analysis (LDA) and Support Vector Machines (SVM) have been applied to LIBS spectral classification tasks, but these algorithms struggle to effectively handle the high dimensionality, nonlinearity, and overlapping spectral lines of spectral data, leaving room for improvement in traceability accuracy.
[0005] Deep learning algorithms such as Convolutional Neural Networks (CNNs) have demonstrated significant advantages in processing high-dimensional, nonlinear LIBS spectral data and resolving spectral line overlap issues due to their unique convolutional structure and powerful feature extraction capabilities. When applying CNNs to LIBS spectral data processing, a one-dimensional convolutional neural network (1D-CNN) can be constructed by adjusting the sliding dimension of the convolutional kernel to adapt to the high-dimensional tensor characteristics of the spectral data, thereby effectively mining deep feature information in the spectrum. However, research has found that as the number of network layers increases, the input features of the neural network gradually become more abstract with increasing network depth. Although deeper features can be extracted, some shallow feature information that is beneficial to the origin traceability task may be lost. In addition, excessively deep network layers significantly increase model complexity, not only increasing the training difficulty but also easily leading to overfitting, thus affecting the accuracy and stability of soybean origin traceability.
[0006] Of particular concern is that, in practical applications, the soybean trade and distribution process is complex, often involving the mixing of soybeans from different producing areas. Furthermore, soybean samples from specific producing regions are often limited in quantity and difficult to obtain in large quantities. However, most existing machine learning-based traceability models are built for ideal conditions with a single producing area and sufficient sample size, paying insufficient attention to their ability to distinguish mixed samples. Moreover, when dealing with producing areas with sparse sample sizes, these models are prone to overfitting or underfitting due to insufficient training data, leading to a significant decrease in discrimination accuracy and making it difficult to meet the complex and ever-changing practical traceability needs. Summary of the Invention
[0007] To overcome the shortcomings of existing technologies, this invention provides a soybean origin tracing and detection method based on laser-induced breakdown spectroscopy (LIBS). Addressing the complex preprocessing required by traditional traceability methods, this invention aims to improve the traceability model's ability to identify soybean samples from different and mixed origins based on varying training sample conditions by analyzing fingerprint elements in the soybean LIBS spectrum using a one-dimensional residual network (1D-ResNet).
[0008] This invention is achieved through the following technical solution: a soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology, specifically including the following steps: S1: Collect soybean samples from different production areas and screen out impurities and defective soybeans; S2: Weigh soybean samples from different origins in S1 and place them in the same container to ensure even distribution. Different mixed sample models are constructed by adjusting the origin or sample batch. S3: Randomly select a set number of soybeans from the soybean samples in S1 and S2, grind them into powder, and then sieve them. S4: Take out a portion of the soybean powder from S3, mix it evenly with microcrystalline cellulose according to the set weight ratio, and then press it into tablets of uniform size for spectral testing. S5: Use a laser-induced breakdown spectroscopy (LIBS) device to scan the pressed sample from S4 and collect LIBS spectra of different types of soybean samples; S6: Preprocess the LIBS spectra acquired in S5; S7: After preprocessing in S6, the dataset is divided into training and test sets according to the proportions. S8: Construct a 1D-ResNet soybean origin discrimination model using the training set in S7, and evaluate the model performance using the test set.
[0009] As a preferred option, step S1 requires collecting soybean samples from different production areas, with at least 5 soybean samples collected from each production area, and each production area must cover different batches of soybean samples.
[0010] As a preferred option, the construction of each mixed sample model in step S2 should include soybean samples from at least two different origins; when weighing soybean samples from different origins, the weight ratio of soybean samples from any two origins should be controlled between 0.25 and 1.
[0011] As a preferred option, the soybean sample used in step S3 weighs more than 80 g.
[0012] As a preferred embodiment, the weight of soybean powder used in step S4 is not less than 0.5 g, and the mixing weight ratio of microcrystalline cellulose and soybean powder is not higher than 1:2; when compressing tablets, the weight of each tablet does not exceed 0.5 g and the diameter does not exceed 15 mm.
[0013] As a preferred embodiment, in step S5, parallel tests are performed at random points on the tableting plane. The elements included in the LIBS spectrum are one or more of the following: calcium (Ca), potassium (K), magnesium (Mg), sodium (Na), carbon (C), hydrogen (H), and oxygen (O).
[0014] As a preferred option, in step S6, the preprocessing methods include baseline removal, spectral smoothing, and normalization.
[0015] As a preferred option, the 1D-ResNet in step S8 is a variant of the ResNet residual network that adjusts the input structure to suit the characteristics of one-dimensional spectral data. It introduces a residual module, which adds the input identity mapping to the nonlinearly transformed output through a skip connection, thereby changing the network's learning objective from directly fitting the desired mapping H(x) to fitting the residual mapping F(x). The specific calculation is as follows:
[0016] Where x is the input of the residual module, F(x) represents the residual mapping learned by the stacked convolutional layers, and H(x) is the mapping of the expected output of the module.
[0017]
[0018]
[0019]
[0020] .
[0021] By employing the above technical solutions, this invention has the following beneficial effects compared to existing technologies: (1) Based on LIBS spectroscopy, this invention can trace the origin of soybean samples from unknown sources. This method has the advantages of simple pretreatment, no reagent contamination, and convenient operation, providing a good foundation for the rapid evaluation of soybean origin traceability.
[0022] (2) The present invention adopts a one-dimensional residual network (1D-ResNet) architecture. Targeting the one-dimensional sequence characteristics of spectral data, it learns the local features of the spectrum through a one-dimensional convolutional layer and introduces residual modules and skip connections. Based on existing neural networks, it effectively overcomes the information loss of spectral features of small sample size categories in the network transmission and improves the accuracy and robustness of LIBS spectrum in soybean origin identification.
[0023] (3) In the training process, the present invention introduces mixed samples from multiple origins, so that the model can not only accurately identify the LIBS spectrum of soybeans from a single origin, but also make preliminary judgments on mixed origin samples, thereby providing a prompt for the case of soybean origin misnomer.
[0024] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description
[0025] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 Radar charts showing the average accuracy of different models in identifying different origins / categories in the test set after LIBS spectral preprocessing: (a) LDA, (b) SVM, (c) 1D-CNN and (d) 1D-ResNet; Figure 2 Confusion matrix diagrams of different models' discrimination results for each origin / category in the test set after LIBS spectral preprocessing: (a) LDA, (b) SVM, (c) 1D-CNN and (d) 1D-ResNet; Figure 3 Radar chart showing the accuracy of the above four models on soybean samples from different origins / categories on the test set; Figure 4 This is a flowchart of the present invention. Detailed Implementation
[0026] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0027] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0028] The following is combined Figures 1 to 4 The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to embodiments of the present invention will be described in detail.
[0029] like Figure 4 As shown, this invention proposes a soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology, which specifically includes the following steps: S1: Collect soybean samples from different origins and screen out impurities and defective soybeans; soybean samples from different origins need to be collected, with at least 5 soybean samples collected from each origin, covering different batches of soybean samples to ensure that the samples are broadly representative.
[0030] S2: Weigh soybean samples from different origins in S1 and place them in the same container to ensure even distribution. Different mixed sample models are constructed by adjusting the origin or sample batch. Each mixed sample model should include soybean samples from at least two different origins. When weighing soybean samples from different origins, the weight ratio of any two soybean samples should be controlled between 0.25 and 1 to ensure that the origin characteristics of the mixed samples have sufficient distinguishability and to avoid excessive dominance of a particular origin sample.
[0031] S3: Randomly select a set number of soybeans from the soybean samples in S1 (single origin) and S2 (mixed origin), grind them into powder, and then sieve them; S4: Take out a portion of the soybean powder from S3, mix it evenly with microcrystalline cellulose according to the set weight ratio, and then press it into tablets of uniform size for spectral testing; the weight of the soybean sample used should be greater than 80 g.
[0032] S5: Use a laser-induced breakdown spectroscopy (LIBS) device to scan the compressed samples from S4, acquiring LIBS spectra of different types of soybean samples. The weight of the soybean powder used should be no less than 0.5 g, and the mixing weight ratio of microcrystalline cellulose and soybean powder should not exceed 1:2 (w / w). Cellulose, as a binder, can effectively enhance the tablet forming effect, but an excessively high proportion of cellulose may dilute the spectral signal of the soybean sample. During the mixing process, it is necessary to ensure that the soybean sample and cellulose are thoroughly and uniformly mixed. When compressing tablets, the weight of each tablet should not exceed 0.5 g, the diameter should not exceed 15 mm, and the surface of the compressed sample should be flat and uniform in color to obtain stable and reliable spectral data. To ensure the influence of mixing uniformity on the spectral results, parallel tests should be performed at random points on the compressed surface. The elements included in the LIBS spectrum include one or more of the following: calcium (Ca), potassium (K), magnesium (Mg), sodium (Na), carbon (C), hydrogen (H), and oxygen (O).
[0033] S6: Preprocess the LIBS spectra acquired in S5; preprocessing methods include baseline removal, spectral smoothing, and normalization. S7: After preprocessing in S6, the dataset is divided into training and test sets according to the proportions. The ratio of training and test sets should be reasonable to ensure that the model built on the training set has a certain generalization ability and avoid overfitting due to too few samples. At the same time, it should be ensured that the selection of the test set is representative and can effectively evaluate the model training results, ensuring the stability and reliability of the results.
[0034] S8: Construct a 1D-ResNet soybean origin discrimination model using the training set from S7, and evaluate the model performance using the test set. 1D-ResNet is a variant of ResNet that adjusts the input structure for one-dimensional spectral data characteristics. It introduces a residual module that adds the input identity mapping to the nonlinearly transformed output through skip connections, thereby transforming the network's learning objective from directly fitting the desired mapping H(x) to fitting the residual mapping F(x). The specific calculation is as follows:
[0035] Here, x is the input to the residual module, F(x) represents the residual mapping learned by stacked convolutional layers (typically containing convolution, batch normalization, and ReLU activation functions), and H(x) is the mapping of the module's expected output. By introducing identity skip connections, the network only needs to learn the difference F(x) between the output and the input. When the optimal mapping is close to the identity mapping, the residual part F(x) approaches 0. The skip connections directly add the input x and the residual F(x) element by element, and then pass the activation function to achieve the output, thus alleviating the degradation problem caused by gradient decay or amplification layer by layer in deep networks.
[0036] Metrics for evaluating model performance include accuracy, precision, recall, and F1 score. The calculation methods are as follows:
[0037]
[0038]
[0039]
[0040] Spectral samples of soybeans from a specific origin / category that are correctly predicted are called true positives (TP), while spectral samples of soybeans from other origins / categories that are correctly predicted are called true negatives (TN). Correspondingly, there are false positives (FP) and true positives (TN). Precision is a comprehensive assessment of the proportion of correctly predicted soybeans by the model; the closer this metric is to 1, the better the overall model performance. Precision is based on the proportion of correctly predicted soybeans from a specific origin / category out of all predicted soybeans belonging to that category. Recall is the proportion of predicted soybeans out of all actual soybeans belonging to that category. The F1 score is the harmonic mean of precision and recall, and can be used to comprehensively evaluate the model's ability to identify soybean origins.
[0041] This comparative example provides a soybean origin traceability detection method based on laser-induced breakdown spectroscopy, including the following steps: (1) Sample Collection and Construction of Mixed Origin Sample Models: A total of 114 soybean samples from six categories, including Brazil, the United States, Canada, Argentina, Uruguay, and mixed origins, were collected. The specific number is shown in Table 1. The origins of the mixed soybean samples were Brazil, the United States, and Canada. Based on my country's imports of soybeans from different origins, five US-Brazil mixed origin sample models, one US-Canada mixed origin sample model, and one Brazil-Canada mixed origin sample model were designed. The construction method for each multi-origin sample model was consistent: 50 g samples were weighed from soybeans from different origins and combined into a sealed bag, totaling 100 g. The samples were mixed by mechanical stirring to ensure that the mass of soybeans from different origins per unit volume was similar to the mixing ratio.
[0042] (2) Sample pretreatment: Randomly select about 20 soybeans from the soybean sample, grind them into powder and sieve them. Take 0.6g of soybean powder and mix it with 1.2g of microcrystalline cellulose at a ratio of 1:2 (w / w). Take 0.36g of the mixed soybean-microcrystalline cellulose powder for pressing into tablets. Each tablet has a diameter of 12 mm.
[0043] (3) Spectral acquisition: The laser-induced breakdown spectroscopy (LIBS) system was used to acquire the spectrum. 50 different points were collected for each pellet, and 10 pulse tests were accumulated at each point to obtain the average value, thus obtaining LIBS spectral data covering the wavelength range of 200 nm to 900 nm.
[0044] (4) Spectral preprocessing: The original LIBS spectra are subjected to baseline subtraction, Savitzky-Golay smoothing and normalization to reduce the interference of spectral drift, spectral noise and other factors on the analysis results and improve the quality and reliability of spectral data.
[0045] The minimum value method is used for baseline subtraction. Specifically, a sliding window of appropriate width is moved point by point along the wavelength direction, and the minimum value of the spectral intensity within the window is regarded as the baseline level at that point. After performing the above operation for all wavelength points, the baseline curve of the entire spectrum is obtained. Finally, the baseline is subtracted from the original spectrum to achieve correction.
[0046] The Savitzky-Golay algorithm performs local polynomial fitting on the spectrum within a sliding window, replacing the original data with the fitted values. This effectively suppresses high-frequency random noise while preserving characteristic information such as peak shape and peak width. In this embodiment, the window width is 11 and the polynomial order is 3.
[0047] Normalization involves scaling the baseline-subtracted and Savitzky-Golay-smoothed spectral data according to a certain ratio, making the spectral data from different samples or under different measurement conditions comparable, thus further improving the quality and reliability of the spectral data. In this embodiment, area normalization is used to process the baseline-corrected LIBS spectral data. Its core principle can be expressed by the following formula:
[0048] In the formula, x i This represents the intensity value at the i-th wavelength point of the original spectrum. The total area is the integral of the entire spectrum (or characteristic spectral segments). This represents the normalized intensity at that wavelength. This method scales the area under the entire spectral curve for each sample to 1 (i.e., per unit area), making the spectral intensities of all samples comparable on the same scale. After processing, the spectral data retains only the relative proportions of the intensities at each wavelength, rather than the absolute intensity values.
[0049] (5) Dataset partitioning: Soybean LIBS spectra from different origins / categories were divided into training and test sets in a 4:1 ratio, as shown in Table 1. To further evaluate model performance, 5-fold cross-validation was used to comprehensively assess the reliability and stability of the model.
[0050] Table 1. Dataset Division by Origin / Category
[0051] (6) Classification model training: A linear discriminant analysis (LDA) model is constructed to predict the category of soybean samples.
[0052] Specifically, LDA is a supervised dimensionality reduction and classification algorithm. Its core lies in finding the projection direction that maximizes the inter-class divergence and minimizes the intra-class divergence of samples from different categories after projection. In LIBS analysis, spectral data exhibits high dimensionality, multicollinearity, and rich elemental fingerprint information. LDA effectively addresses high-dimensional features by compressing the original spectrum to a low-dimensional space with one fewer category. Furthermore, its linear modeling capability aligns with the approximately linear relationship between element content and intensity in LIBS spectra, enabling the extraction of linear discrimination information for soybeans from different origins. This embodiment selects LDA as a representative of traditional linear classifiers. The solver employs singular value decomposition (SVD), and the prior probabilities of the categories are automatically estimated from the data. The model is optimized through 5-fold cross-validation. The system examines the impact of different preprocessing methods on origin discrimination to determine the optimal preprocessing scheme. Model evaluation methods: The metrics used to evaluate model performance include accuracy, precision, recall, and F1 score. The calculation methods are as follows:
[0053]
[0054]
[0055]
[0056] Furthermore, spectral samples correctly predicted for a specific soybean origin / category are called true positives (TP), while those correctly predicted for other origins / categories are called true negatives (TN). Correspondingly, there are false positives (FP) and true positives (TN). Precision is a comprehensive assessment of the proportion of correctly predicted soybeans by the model; the closer this metric is to 1, the better the overall model performance. Precision is based on the proportion of correctly predicted soybeans from a specific origin / category out of all predicted soybeans belonging to that category. Recall is the proportion of correctly predicted soybeans to the actual number of soybeans belonging to that category. The F1 score is the harmonic mean of precision and recall, and can be used to comprehensively evaluate the model's ability to identify soybean origins.
[0057] Table 2 Results of the LDA soybean origin discrimination model based on LIBS
[0058] The results of LDA in the soybean origin traceability task are shown in Table 2. The model's training and validation accuracy are high, but the accuracy on the test set is significantly lower, possibly due to overfitting on some samples. Specifically, the model has high discrimination accuracy for the three major producing regions of Brazil, the United States, and Canada when facing unknown data. Figure 1 (a) This is because LDA, as a linear classifier, models linear relationships in the spectrum well, but its ability to distinguish between South American origins with similar spectral characteristics, such as Argentina and Uruguay, is limited. See the LDA confusion matrix. Figure 2 (a) The samples from Argentina and Uruguay were severely misclassified from each other, and some were even misclassified as mixed. These results indicate that LDA can effectively capture linear discriminant information for soybeans from different major producing regions, but its discrimination accuracy is limited when dealing with small samples of soybeans from closely spaced producing areas. This suggests that LDA struggles to learn robust discrimination boundaries from limited data when facing producing regions like Argentina and Uruguay with sparse sample sizes. Furthermore, the model's identification of mixed samples relies heavily on spectral similarity to single-origin sources, lacking specific modeling capabilities for mixed characteristics, resulting in low discrimination accuracy.
[0059] This comparative example provides a soybean origin traceability detection method based on laser-induced breakdown spectroscopy. The sample collection and mixed origin sample model construction, sample preprocessing, spectral acquisition, spectral preprocessing, dataset partitioning, and model evaluation methods are the same as those in Comparative Example 1. The difference is that this comparative example uses a support vector machine (SVM) for soybean origin traceability judgment in the classification model training.
[0060] Support Vector Machine (SVM) is a classifier based on statistical learning theory. Its core idea is to map the original data to a high-dimensional feature space using a kernel function, and then find an optimal hyperplane in this space to maximize the classification margin between different classes. SVM can effectively handle nonlinear classification problems and has good generalization ability for high-dimensional data. In LIBS spectral analysis, there is a nonlinear relationship between element content and spectral intensity (such as self-absorption effect, matrix effect, etc.), and the elemental combinations of soybeans from different origins are complex, making it difficult for linear models to fully capture their differences. SVM constructs a nonlinear decision boundary through a kernel function, enabling more precise differentiation of spectral patterns from different origins. This embodiment selects SVM as a representative nonlinear classifier, employing a radial basis function (RBF) kernel function with a penalty coefficient C = 1.0, kernel coefficient gamma = 'scale', and a one-to-many (ovr) decision function. The model is optimized through 5-fold cross-validation to further verify the universality of the preprocessing method. The model results are shown in Table 3.
[0061] Table 3 Results of the LIBS-based SVM soybean origin discrimination model
[0062] See the radar chart of the SVM classification results. Figure 1 (b) The accuracy distribution of the model across categories exhibits clear regional characteristics. Radar charts for North American (Canada, USA) and Brazil show more prominent axis lengths, but these categories are prone to overfitting due to smaller sample sizes. The accuracy for Argentina, Uruguay, and mixed categories is significantly lower, and the confusion matrix also shows a large number of consecutive misclassifications for these categories. (See [link to related data]). Figure 2 (b) Specifically, although SVM introduces a kernel function to handle nonlinear relationships, its nonlinear decision boundary can effectively distinguish between origin categories with large differences in spectral characteristics. However, it misclassifies origins with similar spectral characteristics. Overall, the model has good accuracy and stability in identifying major origins, but it cannot accurately identify samples from Argentina and Uruguay under conditions of limited samples or inappropriate kernel function selection. This further illustrates that when SVM deals with small sample categories, its complex nonlinear boundary is prone to overfitting the noise of a few samples rather than the true origin characteristics; for mixed samples, since their spectral characteristics are a nonlinear superposition of multiple origin element characteristics, the SVM kernel function is difficult to effectively distinguish, leading to samples being misclassified as a single origin.
[0063] This comparative example provides a soybean origin traceability detection method based on laser-induced breakdown spectroscopy. The sample collection and mixed origin sample model construction, sample preprocessing, spectral acquisition, spectral preprocessing, dataset partitioning, and model evaluation methods are the same as those in Comparative Example 1. The difference is that this comparative example uses an end-to-end one-dimensional convolutional neural network (1D-CNN) for soybean origin traceability judgment in the classification model training.
[0064] Specifically, 1D-CNN automatically extracts high-level features such as local spectral peak morphology and inter-peak combination relationships from wavelength sequences by learning convolutional kernel parameters, and outputs the category results through a fully connected classification head to complete multi-class discrimination. In this embodiment, the 1D-CNN network structure and parameters are set as follows: the input spectral intensity is standardized using StandardScaler (Z-score) before entering the network. The 1D-CNN consists of three one-dimensional convolutional layers. The convolutional output is flattened before entering the fully connected classification head. The hidden layer dimension is 64, and Dropout = 0.5 is added to suppress overfitting. In the training configuration, the loss function is cross-entropy loss. During training, the model parameters corresponding to the highest accuracy on the validation set are used as the final model. Specific model results are shown in Table 4.
[0065] Table 4 Results of the LIBS-based 1D-CNN soybean origin discrimination model
[0066] from Figure 1 (c) It can be seen that the CNN model exhibits high discrimination accuracy in the three main producing regions of Brazil, the United States, and Canada, forming a stable high-value region. The vertices corresponding to Argentina, Uruguay, and the mixed region shrink inwards, with Uruguay showing the lowest recognition accuracy. Figure 2 (c) This further confirms that all samples in this category were misclassified and scattered among the prediction results of other categories. This irregular polygonal shape reveals the performance limitations of the model under conditions of uneven sample distribution, that is, the model tends to learn more fully for categories with sufficient sample size, while it is difficult to extract effective discriminative features for categories with small sample size. This indicates that although 1D-CNN can automatically extract features, when faced with categories with small sample size, its convolutional kernels are unable to learn representative deep features from limited data, resulting in the failure of feature extraction for categories with small sample size (such as Uruguay); at the same time, due to the lack of ability to analyze and reconstruct features of different origins in mixed spectra, the model's recognition accuracy for mixed samples is also significantly limited. Example
[0067] This embodiment provides a soybean origin traceability detection method based on laser-induced breakdown spectroscopy. The sample collection and sample model construction of mixed origins, sample preprocessing, spectral acquisition, spectral preprocessing, dataset partitioning and model evaluation methods are the same as those in Comparative Example 1. The difference is that this embodiment uses a one-dimensional residual neural network (1D-ResNet) for soybean origin traceability judgment in the classification model training.
[0068] 1D-ResNet is a skip connection structure that introduces residual networks, primarily addressing the degradation problem in deep network training. This model transforms "directly fitting the mapping" into "fitting the residual" through a residual learning mechanism, thereby achieving deeper feature extraction and more robust classification. The basic principle of this network is that the residual module adds the input identity mapping to the convolutional transformation output through skip connections, ensuring that the desired mapping satisfies the following formula:
[0069] Where x is the module input, and F(x) is the residual mapping learned by the convolutional layer. This structure allows the network to learn only a small residual term when the optimal mapping is close to the identity mapping, effectively mitigating the training degradation and gradient decay problems caused by network deepening. The network structure in this embodiment adapts the convolution and downsampling operations of classic ResNet to a one-dimensional form to match the LIBS spectral sequence input and uses six classifications as the final output. Model training uses cross-entropy loss as the optimization objective, and parameter updates use the AdamW optimizer, with a learning rate of 0.001 and weight decay of 1×10⁻⁶. -4The model incorporates a cosine annealing (LR) strategy to dynamically schedule the learning rate. Additionally, to suppress overfitting and improve generalization ability, a Dropout regularization module can be added before entering the fully connected classification layer. The model training results are shown in Table 5.
[0070] Table 5 Results of the LIBS-based 1D-ResNet soybean origin discrimination model
[0071] 1D-ResNet performs well across all categories, see Figure 1 (d): Brazil, the United States and Canada have excellent recognition accuracy. Soybean samples from Argentina and Uruguay have been effectively identified, and the accuracy for mixed samples has remained stable above 0.73. Figure 1 (d) shows that the difference in length of each axis is significantly smaller than that of other models.
[0072] Misjudgments mainly occurred between production areas with similar spectral characteristics, see Figure 2 (d): Brazil and the US, and mixed samples showed a few misclassifications; Canada was almost entirely correct; Argentina and Uruguay had fewer misclassifications; mixed samples were mainly confused with Brazil and the US. Of particular note is that, compared to LDA, SVM, and 1D-CNN, 1D-ResNet demonstrated stronger stability for small sample sizes (Argentina and Uruguay). Argentina's recall reached 0.92, and Uruguay's recall also improved to 0.65, with its precision (0.867) significantly outperforming other models. This indicates that the residual structure can effectively alleviate the gradient decay problem in deep networks through skip connections, enabling stable training and the extraction of discriminative spectral features even under small sample conditions. For mixed samples, 1D-ResNet achieved a recall of 0.70 and an F1 score of 0.715, significantly outperforming the control model. This is attributed to the residual module's ability to better resolve the superposition effect of spectral features from different origins and to capture the combination patterns of various origin element features in mixed samples, thereby achieving preliminary discrimination of mixed samples. ResNet deepens the network by introducing residual blocks, enabling it to abstract more discriminative deep features layer by layer. Through residual connections and global average pooling, it effectively controls model complexity and exhibits stronger generalization ability even on classes with fewer samples. This demonstrates that 1D-ResNet, on raw LIBS spectra with only simple preprocessing, ensures stable training of deep networks thanks to its deep residual structure, overcoming the bottleneck of traditional models in recognizing minority class samples. Figure 3 By learning the distinguishable features between different categories, it can achieve high-precision prediction results.
[0073] Based on the results of Examples 1, 2, 3, and 4, this invention proposes a method based on laser-induced breakdown spectroscopy (LIBS). Through spectral preprocessing, dataset partitioning, and the construction and training of a 1D-ResNet deep neural network, it is possible to achieve rapid detection of soybean origins with different sample sizes and single or mixed origins based on LIBS spectroscopy. The accuracy of the prediction results is significantly improved compared to traditional machine learning models (LDA, SVM) and 1D-CNN models, indicating that the method of this invention has good beneficial effects and significant inventiveness, and has good deployment capability and application prospects in actual soybean origin traceability detection. In the description of this invention, the term "a plurality of" refers to two or more. Unless otherwise explicitly defined, the terms "upper," "lower," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. The terms "connection," "installation," "fixing," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.
[0074] In the description of this specification, the terms "one embodiment," "some embodiments," "specific embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0075] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for tracing the origin of soybeans based on laser-induced breakdown spectroscopy, characterized in that, Specifically, the steps include the following: S1: Collect soybean samples from different production areas and screen out impurities and defective soybeans; S2: Weigh soybean samples from different origins in S1 and place them in the same container to ensure even distribution. Different mixed sample models are constructed by adjusting the origin or sample batch. S3: Randomly select a set number of soybeans from the soybean samples in S1 and S2, grind them into powder, and then sieve them. S4: Take out a portion of the soybean powder from S3, mix it evenly with microcrystalline cellulose according to the set weight ratio, and then press it into tablets of uniform size for spectral testing. S5: Use a laser-induced breakdown spectroscopy (LIBS) device to scan the pressed sample from S4 and collect LIBS spectra of different types of soybean samples; S6: Preprocess the LIBS spectra acquired in S5; S7: After preprocessing in S6, the dataset is divided into training and test sets according to the proportions. S8: Construct a 1D-ResNet soybean origin discrimination model using the training set in S7, and evaluate the model performance using the test set.
2. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... In step S1, soybean samples from different production areas need to be collected, with at least 5 soybean samples collected from each production area, and each production area needs to cover different batches of soybean samples.
3. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that, In step S2, the construction of each mixed sample model should include soybean samples from at least two different origins; when weighing soybean samples from different origins, the weight ratio of any two soybean samples should be controlled between 0.25 and 1.
4. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... The soybean sample used in step S3 weighs more than 80 g.
5. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... In step S4, the weight of soybean powder used shall not be less than 0.5 g, and the weight ratio of microcrystalline cellulose and soybean powder shall not be higher than 1:2; when compressing tablets, the weight of each tablet shall not exceed 0.5 g and the diameter shall not exceed 15 mm.
6. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... In step S5, parallel tests are performed at random points on the tableting plane. The elements included in the LIBS spectrum are one or more of the following: calcium (Ca), potassium (K), magnesium (Mg), sodium (Na), carbon (C), hydrogen (H), and oxygen (O).
7. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... In step S6, the preprocessing methods include baseline removal, spectral smoothing, and normalization.
8. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... The 1D-ResNet in step S8 is a variant of ResNet, a residual network with an adjusted input structure for one-dimensional spectral data characteristics. It introduces a residual module that adds the input identity mapping to the nonlinearly transformed output via skip connections, thereby transforming the network's learning objective from directly fitting the desired mapping H(x) to fitting the residual mapping F(x). The specific calculation is as follows: , Where x is the input of the residual module, F(x) represents the residual mapping learned by the stacked convolutional layers, and H(x) is the mapping of the expected output of the module.
9. The soybean origin traceability detection method based on laser-induced breakdown spectroscopy technology according to claim 1, characterized in that... In step S8, the metrics for evaluating model performance include accuracy, precision, recall, and F1 score. The calculation methods are as follows: , , , 。
Citation Information
Patent Citations
Method for rapidly and nondestructively discriminating nephrite origin place
CN108535258A
Cereal origin traceability method based on multi-information ticket simulation mechanism and DAN algorithm
CN116228255A
Beef producing area tracing method and device based on Raman spectrum and adaptive attention residual network
CN121093076A
Method and system for predictive classification by mass spectrometry and trained large spectral model
CN121866467A
Method for measuring the concentration of a chemical compound contained in a fluid by means of an optical measurement system
WO2025016729A1