Soybean fingerprint spectrum establishing method and soybean tracing method
Through infrared spectral detection and metabolomics method combined with support vector machine technology, a soybean fingerprint map was established, which solved the problems of high cost and insufficient accuracy of existing soybean traceability technology, and achieved accurate traceability of soybean origin.
Patent Information
- Application Number
- CN202510135099.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
The existing soybean traceability technology has the problems of high testing costs and insufficient evaluation parameters, resulting in inaccurate analysis results.
Through infrared spectroscopy detection combined with metabolomics method and support vector machine recursive feature elimination, a soybean fingerprint map is established to achieve accurate traceability of soybeans.
It reduces the testing cost, improves the accuracy of the analysis results, and achieves accurate traceability of soybean production areas.
Smart Images

Figure CN120064193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soybean traceability, and particularly to a method for establishing a soybean fingerprint spectrum and a method for tracing the origin of soybeans. Background Art
[0002] Soybean is one of the main food crops in China and plays a crucial role in soybean product processing and human daily life. With the continuous deterioration of the climate environment, the yield and quality of Chinese soybeans show a downward trend, and consumers have shown great interest in high-quality soybeans with geographical landmark characteristics. However, unscrupulous merchants take advantage of this consumer mentality and implement fraud on the origin of soybeans by mislabeling and adulterating, etc., so as to seek high profits. As this problem intensifies, it has become a stumbling block restricting the stable and orderly development of the soybean industry and even triggering more serious food safety risks. In order to safeguard the basic interests and legitimate rights and interests of consumers and protect the brand value of geographical indication products, it is necessary to develop a feasible method to solve the problem of tracing the origin of soybeans.
[0003] At present, a number of studies have introduced in detail methods for identifying soybeans from different geographical origins, including stable isotope technology, mineral element fingerprint technology, and characteristic compound methods (such as fatty acids and isoflavones). Although the above methods have all achieved satisfactory results, there are still some defects to be overcome. For example, the test costs of the first two methods are relatively high, and the evaluation parameters of the latter method are insufficient, which may lead to inaccurate analysis results. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for establishing a soybean fingerprint spectrum and a method for tracing the origin of soybeans. The establishment method of the present invention only needs to test the infrared spectrum, and the test cost is low; and the method of the present invention combines metabolomics and support vector machine recursive feature elimination, so that the prediction result accuracy of the established soybean fingerprint spectrum is high.
[0005] In order to achieve the above invention purpose, the present invention provides the following technical solutions:
[0006] The present invention provides a method for establishing a soybean fingerprint spectrum, including the following steps:
[0007] Performing infrared spectrum detection on a target soybean sample and a reference soybean sample respectively to obtain the original infrared spectrum information of the soybean sample;
[0008] Performing calibration and normalization processing on the original infrared spectrum information of the soybean sample to obtain preprocessed infrared spectrum information;
[0009] Performing principal component analysis on the preprocessed infrared spectrum information to obtain the principal component analysis result;
[0010] Based on the results of the principal component analysis, perform cluster analysis on the preprocessed infrared spectral information to obtain the results of the cluster analysis;
[0011] Based on the results of the cluster analysis, perform orthogonal partial least squares discriminant analysis on the preprocessed infrared spectral information to obtain the pre-wave number markers of the target soybean sample;
[0012] Based on the pre-wave number markers of the target soybean sample, perform support vector machine recursive feature elimination on the preprocessed infrared spectral information to obtain the fingerprint of the target soybean sample.
[0013] Preferably, the origin of the target soybean sample is one or more of Xiangyang, Hubei, China; Lvliang, Shanxi, China; Liaocheng, Shandong, China; Jining, Shandong, China; Shangqiu, Henan, China; Mudanjiang, Heilongjiang, China; Harbin, Heilongjiang, China; Suihua, Heilongjiang, China; Baicheng, Jilin, China; and Jinzhou, Liaoning, China.
[0014] Preferably, the origin of the reference soybean sample is one or more of Brazil, the United States, and Argentina.
[0015] Preferably, the parameters of the infrared spectrum detection include: the scanning mode is diffuse reflection, and the scanning range is 400 - 4000 cm -1 , and the spectral resolution is 4 cm -1 .
[0016] Preferably, the calibration is completed in Origin 9.0 software, and the normalization process is performed in SPSS Statistics V17.0 software.
[0017] Preferably, the principal component analysis is performed in SIMCA 14.1 software.
[0018] Preferably, the cluster analysis is performed in Heml 1.0.3.7 software.
[0019] Preferably, the orthogonal partial least squares discriminant analysis is performed in SIMCA 14.1 software.
[0020] Preferably, the support vector machine recursive feature elimination is performed on MATLAB software of R2019b version.
[0021] The present invention also provides a traceability method for soybeans, including the following steps:
[0022] Perform infrared spectrum detection on the soybeans to be tested to obtain the original infrared spectral information of the soybeans to be tested;
[0023] By comparing the original infrared spectrum of the soybeans to be tested with the soybean fingerprint spectrum obtained by the establishment method described in the above technical solution, the origin of the soybeans to be tested can be traced.
[0024] The invention provides a method for establishing a soybean fingerprint.
[0025] The establishment method provided by the present invention combines the metabolomics method (principal component analysis, cluster analysis and orthogonal partial least squares discriminant analysis) with the support vector machine recursive feature elimination to accurately find the infrared spectrum fingerprint of the target soybean. The establishment method of the present invention relies on less sample data to obtain ideal results, which can effectively simplify the workflow. At the same time, the establishment method of the present invention only requires the infrared spectrum of the test sample, and the test cost is low.
[0026] The present invention also provides a soybean traceability method. The traceability method of the present invention can realize the traceability of the soybean to be tested by only measuring the infrared spectrum of the soybean to be tested, and the traceability result has high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 The calibrated infrared spectra of 13 soybean varieties;
[0028] Figure 2 This is the principal component analysis score diagram of 13 soybean varieties;
[0029] Figure 3 This is a cluster analysis diagram of 13 soybean varieties;
[0030] Figure 4 This is the OPLS-DA score diagram of 10 domestic soybeans;
[0031] Figure 5 This is the S-plot of 10 types of domestic soybeans;
[0032] Figure 6 This is a permutation test chart for 10 types of domestic soybeans;
[0033] Figure 7 The infrared spectra of four kinds of soybeans;
[0034] Figure 8 This is the principal component analysis diagram of mixed soybeans;
[0035] Figure 9 This is the cluster analysis diagram of pure soybean;
[0036] Figure 10 Cluster analysis diagram of replacing No. 7 soybean with No. 14 soybean;
[0037] Figure 11 Cluster analysis diagram of replacing No. 9 soybean with No. 15 soybean;
[0038] Figure 12 OPLS-DA score plot for mixed soybeans;
[0039] Figure 13 S-plot for mixed soybeans;
[0040] Figure 14 Permutation test plot for mixed soybeans. Detailed implementation mode
[0041] The present invention provides a method for establishing a fingerprint spectrum of soybeans, comprising the following steps:
[0042] Performing infrared spectroscopy detection on a target soybean sample and a reference soybean sample respectively to obtain the original infrared spectral information of the soybean sample;
[0043] Calibrating and normalizing the original infrared spectral information of the soybean sample to obtain preprocessed infrared spectral information;
[0044] Performing principal component analysis on the preprocessed infrared spectral information to obtain a principal component analysis result;
[0045] Based on the principal component analysis result, performing cluster analysis on the preprocessed infrared spectral information to obtain a cluster analysis result;
[0046] Based on the cluster analysis result, performing orthogonal partial least squares discriminant analysis on the preprocessed infrared spectral information to obtain pre-wave number markers of the target soybean sample;
[0047] Based on the pre-wave number markers of the target soybean sample, performing support vector machine recursive feature elimination on the preprocessed infrared spectral information to obtain the fingerprint spectrum of the target soybean sample.
[0048] Unless otherwise specified, the raw materials used in the present invention are preferably commercially available products.
[0049] The present invention performs infrared spectroscopy detection on a target soybean sample and a reference soybean sample respectively to obtain the original infrared spectral information of the soybean sample.
[0050] In the present invention, the target soybean sample refers to the soybeans for which a fingerprint spectrum is to be established. In the present invention, the origin of the target soybean sample is preferably one or more of Xiangyang, Hubei, China; Lvliang, Shanxi, China; Liaocheng, Shandong, China; Jining, Shandong, China; Shangqiu, Henan, China; Mudanjiang, Heilongjiang, China; Harbin, Heilongjiang, China; Suihua, Heilongjiang, China; Baicheng, Jilin, China; and Jinzhou, Liaoning, China. In the present invention, when the origin of the target soybean sample is multiple, it is preferred to obtain the fingerprint spectra of soybeans from multiple regions simultaneously.
[0051] In the present invention, the place of origin of the reference soybean samples is preferably one or more of Brazil, the United States, and Argentina, and more preferably Brazil, the United States, and Argentina are selected simultaneously.
[0052] In the present invention, each target soybean sample and reference soybean sample are preferably set with 9 parallels. The samples in each parallel of each target soybean sample and reference soybean sample are mixed to form a quality control (QC) sample. In the present invention, the QC sample can monitor the stability of the infrared spectrometer.
[0053] In the present invention, the parameters of the infrared spectrum detection include: the scanning mode is preferably diffuse reflection, and the scanning range is preferably 400 - 4000 cm -1 , and the spectral resolution is preferably 4 cm -1 . In the present invention, the infrared spectrum detection is preferably carried out on an infrared spectrometer.
[0054] In the present invention, the process of the infrared spectrum detection preferably includes: loading the target soybean sample and the reference soybean sample into a sealed plastic bag and storing them in a ventilated room at 25°C; crushing the target soybean sample and the reference soybean sample to obtain soybean powder; mixing the soybean powder and potassium bromide, and pressing them into a tablet to obtain an infrared spectrum detection specimen; using an infrared spectrometer to perform infrared spectrum detection on the infrared spectrum detection specimen. In the present invention, the crushing is preferably carried out on a JFSD-100 type crusher (Jiading Grain and Oil Instrument Co., Ltd., Shanghai, China). In the present invention, the mass ratio of the soybean powder to potassium bromide is preferably 1:100.
[0055] After obtaining the original infrared spectrum information of the soybean sample, the present invention calibrates and normalizes the original infrared spectrum information of the soybean sample to obtain preprocessed infrared spectrum information.
[0056] In the present invention, the calibration preferably includes the following process: the original infrared spectrum information of the soybean sample is composed of a series of data points output by the infrared spectrometer. The entire original infrared spectrum diagram of the soybean sample is translated along the Y-axis until the lowest point reaches the X-axis, and at this time, the absorbance of this point is zero, and this is used as a benchmark for evaluating the infrared absorption of the soybean sample. In the present invention, the calibration is preferably completed in Origin 9.0 software.
[0057] In the present invention, the normalization process is preferably carried out in SPSS Statistics V17.0 software. In the present invention, the normalization process preferably includes: taking the wave number as the abscissa and the soybean sample name as the ordinate to establish a data matrix; performing data normalization processing on the data matrix in SPSS Statistics V17.0 software. In the present invention, the normalization process can eliminate the influence of the order of magnitude on the result evaluation.
[0058] After obtaining the preprocessed infrared spectrum information, the present invention performs principal component analysis on the preprocessed infrared spectrum information to obtain a principal component analysis result.
[0059] In the present invention, the principal component analysis is preferably performed in SIMCA 14.1 software. In the present invention, during the principal component analysis, if the wave number of the relative standard deviation of absorbance in the QC sample or any soybean sample is greater than 30%, it is considered invalid and deleted. In the present invention, the principal component analysis (PCA) as an unsupervised analysis method can naturally classify samples without knowing the sample category and eliminate extreme data.
[0060] After obtaining the principal component analysis result, the present invention performs cluster analysis on the pre-processed infrared spectrum information based on the principal component analysis result to obtain a cluster analysis result.
[0061] In the present invention, the cluster analysis is preferably performed in Heml 1.0.3.7 software. In the present invention, the cluster analysis is also an unsupervised analysis method, which is to cluster samples layer by layer according to the degree of feature similarity. That is to say, in cluster analysis, samples with higher similarity will be clustered together first until all samples are clustered.
[0062] After obtaining the cluster analysis results, the present invention performs orthogonal partial least squares discriminant analysis on the pre-processed infrared spectrum information based on the cluster analysis results to obtain the pre-wavenumber markers of the target soybean sample.
[0063] In the present invention, the orthogonal partial least squares discriminant analysis (OPLS-DA) is preferably performed in SIMCA 14.1 software. In the present invention, unlike PCA and cluster analysis, OPLS-DA is a supervised analysis method, which can design grouping by itself and can better obtain inter-group difference information that PCA and cluster analysis cannot obtain.
[0064] In the present invention, the OPLS-DA is a binary classification model. In a specific embodiment of the present invention, the target soybean sample is set as category 1, and the reference soybean sample as a whole is set as category 2. In the present invention, when the number of samples in category 1 and category 2 of the OPLS-DA is not equal, the present invention preferably uses synthetic minority class oversampling technology (SMOTE) to make the number of samples in category 1 and category 2 equal to avoid introducing bias in the calculation of the decision rule.
[0065] In the present invention, the OPLS-DA includes an S-plot, permutation test analysis, and variable importance in projection (VIP); the S-plot, permutation test analysis, and variable importance in projection are comprehensively used to analyze and screen for "markers". In the present invention, the wavenumbers in the S-plot are described by a series of points, and the importance of these points is reflected by their positions. The points at both ends of the "S-plot" can best distinguish the two major camps in the OPLS-DA plot and are also the most likely candidate "markers". As is well known, the number of variables in metabolomics research far exceeds the number of samples. In the specific embodiment of the present invention, there are a total of 1868 variables and 117 samples (13 types of soybeans × 9 samples / type of soybean), meeting the requirements of metabolomics analysis. Performing metabolomics analysis using an OPLS-DA classification model may result in overfitting, that is, the model can well distinguish the samples in the training set but performs poorly when predicting a new sample set. Therefore, it is necessary to verify the reliability of the model. The present invention uses a permutation test method to judge the overfitting of the OPLS-DA model. VIP reflects the loading weight of each wavenumber and can be used for feature selection. The VIP value is positively correlated with the importance of the wavenumber. In the present invention, VIP > 1.5 is used as the threshold for calculating candidate "wavenumber markers". The paired t-test of univariate analysis is an important step to confirm whether the "marker" is effective, and this process is carried out in the SPSS Statistics V17.0 software.
[0066] After obtaining the pre-wavenumber markers of the target soybean sample, the present invention performs support vector machine recursive feature elimination on the preprocessed infrared spectral information based on the pre-wavenumber markers of the target soybean sample to obtain the fingerprint spectrum of the target soybean sample.
[0067] In the present invention, the support vector machine recursive feature elimination (SVMRFE) is preferably carried out on the MATLAB software of the R2019b version.
[0068] The present invention preferably selects the SVMRFE running program under the radial basis function (RBF), and this process is carried out in MATLAB software (version R2019b). In the present invention, the support vector machine recursive feature elimination preferably includes the following process: randomly divide the data set into 5 subsets, perform 5-fold cross-validation, and calculate the average value of the accuracy rates for 5 times. This process is repeated 10 times to obtain the average accuracy rate as the final model classification accuracy rate. The kernel parameter is selected by the 3-fold cross-validation method of randomly dividing the training data set into three subsets. The classifier is trained on any one of the two subsets and tested on the third subset. With the help of the LIBSVM software package, a set of parameters that can provide the best cross-validation accuracy is used for further analysis. After calculation, the parameter with the highest cross-validation accuracy is selected as the kernel parameter for actual classification. Under the above conditions, the weight values (w) and the squared weight values (w 2 ) of the variables are obtained, and the latter are arranged into a variable weight table according to the numerical size for further analysis.
[0069] The present invention also provides a method for tracing the origin of soybeans, including the following steps:
[0070] Perform infrared spectrum detection on the soybeans to be tested to obtain the original infrared spectrum information of the soybeans to be tested;
[0071] Compare the original infrared spectrum of the soybeans to be tested with the soybean fingerprint spectrum obtained by the establishment method described in the above technical solution, and the origin tracing of the soybeans to be tested can be realized.
[0072] The present invention performs infrared spectrum detection on the soybeans to be tested to obtain the original infrared spectrum information of the soybeans to be tested. In the present invention, the parameters of the infrared spectrum detection are the same as those in the above technical solution and will not be elaborated here.
[0073] After obtaining the original infrared spectrum information of the soybeans to be tested, the present invention compares the original infrared spectrum of the soybeans to be tested with the soybean fingerprint spectrum obtained by the establishment method described in the above technical solution, and the origin tracing of the soybeans to be tested can be realized.
[0074] In the present invention, the comparison is as follows: compare the original infrared spectrum information of the soybeans to be tested with the fingerprint spectrum. If the fingerprint spectrum exists in the original infrared spectrum information of the soybeans to be tested, it indicates that the soybeans to be tested are from this origin, otherwise not.
[0075] The following combines examples to elaborate in detail on the method for establishing the soybean fingerprint spectrum and the method for tracing the origin of soybeans provided by the present invention, but they cannot be understood as limiting the protection scope of the present invention.
[0076] Example 1
[0077] A total of 13 kinds of soybeans were selected in this example, including 10 domestic soybeans and 3 imported soybeans. Their basic information is shown in Table 1.
[0078] Table 1 Basic Information of Soybean Samples
[0079]
[0080] The collected soybean samples were packed in sealed plastic bags and stored in a ventilated room at 25 °C. The samples were ground into powder using a JFSD-100 type grinder (Shanghai Jiading Grain and Oil Instrument Co., Ltd., China). Each soybean powder was mixed with 100 times its own weight of potassium bromide (Sinopharm Chemical Reagent Co., Ltd., China) to form a uniformly distributed sample. The above process was repeated to obtain nine parallel samples. From each of the above uniformly distributed samples, a sub-sample of (0.1615 ± 0.0003) g was weighed, tableted, and another sub-sample was reserved for preparing quality control (QC) samples. All 117 reserved sub-samples (13 kinds of soybeans × 9 parallel samples per kind of soybean) were fully mixed and then a sample of (0.1615 ± 0.0003) g was weighed and tableted for monitoring the stability of the infrared spectrometer.
[0081] Infrared Spectrum Measurement
[0082] The infrared spectrum of soybeans was recorded using a Shimadzu Fourier transform infrared spectrometer (IRTracer-100). The scanning mode was diffuse reflection, and the scanning range was 400 - 4000 cm -1 , and the spectral resolution was 4 cm -1 . According to the requirements of metabolomics analysis, the 13 kinds of soybeans, with 9 parallel samples for each kind, were named as Sample 1-1 to Sample 1-9, Sample 2-1 to Sample 2-9,..., Sample 13-1 to Sample 13-9. Before and after the infrared spectrum scanning of each kind of soybean, the QC sample needed to be scanned three times repeatedly to obtain the original infrared spectrum.
[0083] Data Processing
[0084] Calibration of the original infrared spectrum: The infrared spectra of soybeans from different origins can be accurately distinguished under the same baseline. To achieve this goal, the original infrared spectrum was calibrated: The infrared spectrum consists of a series of data points output by the spectrometer. The entire spectrum was translated along the Y-axis until the lowest point reached the X-axis, at which time the absorbance at this point was zero, and this was used as the benchmark for evaluating the infrared absorption of soybean samples. The above process was completed in Origin 9.0 software. Figure 1 For the calibrated infrared spectra of the 13 kinds of soybeans, from Figure 1It can be seen that all infrared spectra are at the same baseline after calibration, which is conducive to intuitively comparing the absorbance of different soybeans and making preliminary judgments. However, some extremely similar spectral profiles are still difficult to directly distinguish soybeans from different origins (such as No. 8 and No. 10), which is also the key problem to be solved in this embodiment when using metabolomics and support vector machines to conduct in-depth analysis of infrared spectra.
[0085] According to the requirements of metabolomics, the wave numbers with relative standard deviation of absorbance greater than 30% in the QC group or any soybean group were considered invalid and deleted. Then, a data matrix was established with wave number and sample name (such as sample 1-1) as horizontal and vertical coordinates respectively. The matrix was normalized in SPSS Statistics V17.0 software to eliminate the influence of order of magnitude on the result evaluation.
[0086] The principal component analysis (PCA) of the data matrix was processed using SIMCA 14.1 software. As an unsupervised analysis method, PCA can naturally classify sample groups without knowing the sample category and eliminate extreme data. The results are shown in Figure 2 As shown, Figure 2 This is the principal component analysis score diagram of 13 soybeans, such as Figure 2 As shown in the figure, the credibility of all sample groups is within the 95% confidence level, and no outliers or significant extreme data are found. The 42 QC samples are clustered near the origin, indicating that the quality of the available data in this example is high and can be analyzed in the next step. For some soybeans, their similar spectral profiles ( Figure 1 ) may be a PCA score plot ( Figure 2 ) in the root cause of the lack of clear separation. In general, Figure 2 The soybean samples No. 3, 4, 5, 11 and 13 above the red line intersect with each other, indicating that they have high inter-group similarity. Similar phenomena can also be observed in the soybean samples No. 1, 2, 6, 7, 8, 9, 10 and 12 below the red line. The corresponding PCA graph failed to complete the clear sample grouping, and further distinction was required through supervised analysis methods.
[0087] Cluster analysis was performed using Heml 1.0.3.7 software. Cluster analysis is also an unsupervised analysis method that clusters samples layer by layer according to the degree of feature similarity. In other words, in cluster analysis, samples with higher similarity will be clustered together first until all samples are clustered. The results are as follows Figure 3 , Figure 3 This is a cluster analysis diagram of 13 types of soybeans. Figure 3As shown, the first group of soybean samples is the only group in which all 9 parallel samples are clustered together, meaning they have the highest within-group similarity. Followed by groups 5, 10, and 7, all of which consist of 2 subgroups. Among them, the 5th soybean group includes a subgroup composed of 8 samples (sample 5-2 to sample 5-9), and the 10th soybean group includes a subgroup composed of 7 samples (sample 10-1 to sample 10-3, sample 10-5, sample 10-7 to sample 10-9). The number of samples included in the two subgroups of the 7th soybean (subgroup 1: 4 such as sample 7-1 to sample 7-3 and sample 7-7; subgroup 2: 5 such as sample 7-4 to sample 7-6, sample 7-8, and sample 7-9) is close. Other soybean groups are all divided into no less than 3 subgroups, indicating that these soybeans have relatively low within-group similarity. The entire cluster analysis graph ( Figure 3 ) can be divided into two major camps along the red dashed line. Among them, soybeans numbered 1, 2, 6, 7, 8, 9, and 10 form one camp, and soybeans numbered 3, 4, 5, 11, 12, and 13 form another camp. The classification results are very close to those of the principal component analysis. However, both the principal component analysis and the cluster analysis can only achieve a preliminary grouping of soybean samples from different origins and cannot provide specific difference information.
[0088] Orthogonal partial least squares discriminant analysis (OPLS-DA) is carried out on the SIMCA 14.1 software. Different from PCA and cluster analysis, OPLS-DA is a supervised analysis method. This method can design the grouping by itself and can better obtain the between-group difference information that PCA and cluster analysis cannot obtain.
[0089] OPLS-DA is a binary classification model. In this embodiment, any group of domestic soybeans can be set as classification 1, and the imported soybeans numbered 11-13 as a whole are set as classification 2. Since the sample numbers of these two classifications in the OPLS-DA model are not equal (that is, there are only 9 samples in classification 1, while there are 27 samples in classification 2), in order to avoid introducing bias in the calculation of the decision rule, in this embodiment, through the synthetic minority over-sampling technique (SMOTE), the sample number of classification 1 is increased from 9 to 27 (including 9 original samples and 18 SMOTE-expanded samples).
[0090] Figure 4 The OPLS-DA score graph for 10 domestic soybeans is as Figure 4 shown. Classification 1 (green) and classification 2 (blue) are significantly separated on the first principal component axis, meaning there are wavenumbers with significantly different absorbances between the two classifications. R 2 Y and Q 2 in the OPLS-DA model respectively describe the explanatory level and prediction level of the model along the Y-axis. R 2 Y and Q2 The value is closer to 1, indicating that the OPLS-DA model has higher reliability and predictability. From Figure 4 , it can be seen that R 2 Y and Q 2 are both greater than 0.9, indicating that all OPLS-DA models in this embodiment have good credibility. It should be noted that the groups of soybean samples that were not completely separated by PCA ( Figure 2 ) and cluster analysis ( Figure 3 ) were all well-separated in the OPLS-DA model, indicating the advantage of the OPLS-DA model in differentiating between-group samples, even if the differences between these samples are subtle.
[0091] S-plot analysis, permutation test analysis, and variable importance in projection (VIP) analysis embedded in OPLS-DA were used to screen for "markers". The wavenumbers in the S-plot are described by a series of points, and the importance of these points is reflected by their positions. The points at both ends of the "S-plot" can best distinguish the two major camps in the OPLS-DA plot and are also the most likely candidate "markers". As is well known, the number of variables in metabolomics research far exceeds the number of samples. This embodiment has a total of 1,868 variables and 117 samples (13 types of soybeans × 9 samples per type of soybean), meeting the requirements of metabolomics analysis. Performing metabolomics analysis using an OPLS-DA classification model may result in overfitting, that is, the model can well distinguish the samples in the training set but performs poorly in predicting a new sample set. Therefore, it is necessary to verify the reliability of the model. In this embodiment, the permutation test method was used to judge the overfitting of the OPLS-DA model. VIP reflects the loading weight of each wavenumber and can be used for feature selection. The VIP value is positively correlated with the importance of the wavenumber. In this study, VIP > 1.5 was used as the threshold for calculating candidate "wavenumber markers". The paired t-test of univariate analysis is an important step to confirm whether the "marker" is effective, and this process is carried out in the SPSS Statistics V17.0 software.
[0092] Figure 5 is the S-plot of 10 domestic soybeans, Figure 5 and each point in it represents a wavenumber, and its contribution and confidence level are determined by its position along the X-axis and Y-axis respectively. Points farther from the origin indicate that they are more important in differentiating the soybean samples classified as 1 and 2 in the OPLS-DA plot, and it also well explains why the most different components (i.e., the most qualified candidate "markers") are searched for at both ends of the S-plot. For soybean samples No. 3 to No. 5, their S-plot only covers the third quadrant, meaning that these 3 types of soybeans only have absorbance significantly higher than that of soybeans No. 11 to No. 13, while soybeans No. 1, 2, 6 to 10 have both significantly higher and significantly lower absorbance than soybeans No. 11 to No. 13.
[0093] The OPLS-DA model was examined for overfitting through 200 iterations of permutation testing, and evaluated by the significance test (P-value) of the index Q 2 of Y. If P < 0.05, the model is not overfitted, and at this time the intercept of the corresponding Q 2 regression line is less than 0.05, and the results are as Figure 6 shown Figure 6 is the permutation test chart of 10 domestic soybeans, as Figure 6 shown. All intercepts are significantly lower than 0.05, indicating that none of the OPLS-DA models are overfitted.
[0094] Table 2 lists the wavenumber markers selected for each soybean in the variable weight table.
[0095] Wavenumber markers selected for each soybean in Table 2 of variable weights
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108] In Table 2, soybean No. 14 is composed of 80% soybean No. 7 + 20% soybean No. 6, and this soybean is used to replace soybean No. 7 for metabolomics analysis. Similarly, soybean No. 15 is composed of 80% soybean No. 9 + 20% soybean No. 10, and is used to replace soybean No. 9 for further analysis.
[0109] According to the VIP>1.5 principle, 53, 15, 152, 22, 178, 81, 189, 96, 244, and 225 "markers" were respectively screened out from 1 to 10 groups of soybeans. In this example, the sequence composed of the longest continuous variable numbers was selected as the "marker" characteristic band to distinguish specific soybeans from other soybeans. As shown in Table 2, 49, 15, 93, 17, 129, 77, 177, 50, 80, and 98 "markers" were respectively screened out from 10 groups of soybeans, and the corresponding wavenumber ranges were 777.313 - 954.764, 3971.432 - 3998.436, 962.479 - 1139.930, 1506.405 - 1537.266, 2935.658 - 3182.546, 3853.774 - 4000.364, 3660.893 - 4000.364, 723.306 - 817.818, 671.228 - 823.604, and 671.228 - 858.323 cm -1 .
[0110] In this example, the SVMRFE running program under the radial basis function (RBF) was selected, and this process was carried out in MATLAB software (version R2019b). The dataset was randomly divided into 5 subsets for 5-fold cross-validation, and the average value of the accuracy rates calculated 5 times was calculated. This process was repeated 10 times to obtain the average accuracy rate as the final model classification accuracy rate. The kernel parameter was selected by the 3-fold cross-validation method of randomly dividing the training dataset into three subsets. The classifier was trained on any one of the two subsets and tested on the third subset. With the help of the LIBSVM software package, a set of parameters that could provide the best cross-validation accuracy was used for further analysis. After calculation, the parameter with the highest cross-validation accuracy was selected as the kernel parameter for actual classification. Under the above conditions, the weight value (w) and the weighted square value (w 2 ) of the variable were obtained, and the latter was arranged into a variable weight table according to the numerical size for further analysis.
[0111] After running the support vector machine, the classification accuracy of 5-fold cross-validation was calculated to exceed 90%, indicating that the model has good generalization ability. All variables with VIP>1.5 have corresponding positions in the variable weight table, and the sequences of these variables in the weight table are continuous, meaning that the two methods of VIP and weight ranking to find eligible variables have exactly the same results. However, when the sequences of these variables in the weight table are not continuous, it means that some variables with VIP<1.5 are also included in the weight table. Generally speaking, SVMRFE has better predictive ability than OPLS-DA, that is, the method of weight ranking to obtain eligible variables is more reliable than the VIP method. Therefore, in this embodiment, this part of variables with VIP<1.5 is still retained in the weight table for analysis to prevent the possible loss of some effective variables during the VIP calculation process. According to the above method, in 10 groups of soybean samples, 144, 15, 382, 23, 292, 102, 475, 327, 428, and 227 "markers" were screened out respectively. Similarly, the sequence composed of the longest continuous variable numbers was selected as the "marker" characteristic band (Table 2), including 57, 15, 102, 18, 175, 77, 177, 97, 80, and 98 "markers" respectively, and the corresponding wavenumber ranges were 761.882~954.764, 3971.432~3998.436, 943.191~1139.930, 1504.476~1537.266, 2933.729~3269.343, 3853.774~4000.364, 3660.893~4000.364, 673.1568~858.323, 671.228~823.604, and 671.228~858.323 cm -1 By comparing the "marker" characteristic bands found by the two methods of VIP>1.5 and SVMRFE, it was found that the latter completely covered the former. Among them, the characteristic bands of soybeans No. 2, 6, 7, 9, and 10 were exactly the same. The overlapping rates of the "marker" characteristic bands found by the two methods for the remaining soybeans No. 1, 3, 4, 5, and 8 were 85.96%, 91.17%, 94.4%, 73.7%, and 51.5% respectively, indicating that there are still certain differences between the two methods in finding the "marker" characteristic bands.
[0112] Since the "wavenumber marker" characteristic bands selected by the SVMRFE method completely cover the "wavenumber marker" characteristic bands selected by the VIP>1.5 method, it shows that the former is more comprehensive than the latter in the screening results. Therefore, in this embodiment, the absorbance of the "wavenumber marker" characteristic bands selected by the SVMRFE method is used for paired t-test to verify whether the wavenumber results selected by the SVMRFE method are more comprehensive. The results show that for any domestic soybean (No. 1-10), the absorbance of its "wavenumber marker" characteristic bands is significantly different from the overall absorbance of imported soybeans No. 11-13 (P<0.05). The results of univariate analysis further support the rationality of the "markers" selected by multivariate analysis. On the corresponding "marker" characteristic bands, the soybean groups of No. 2, 3, 4, 5, 6, and 7 have significantly higher absorbance, while the soybean groups of No. 1, 8, 9, and 10 have significantly lower absorbance. The above results show that the SVMRFE method is indeed more comprehensive and the data is more reliable than the VIP method in screening the "wavenumber marker" characteristic bands.
[0113] Practicality test
[0114] To evaluate the practicality of the method of the present invention, a new soybean similar to the soybeans No. 1-10 in this embodiment is considered to be introduced to examine whether this soybean can also be accurately distinguished by the method of the present invention. As shown in Table 2, the "marker" characteristic bands of soybeans No. 6 and 9 are respectively included in the "marker" characteristic bands of soybeans No. 7 and 10. In addition, soybeans No. 6 and 7 have obvious inter-group similarity ( Figure 2 ), and they have significantly high absorbance on the "marker" characteristic bands. Soybeans No. 9 and 10 also have relatively high inter-group similarity ( Figure 2 ), but have significantly low absorbance. The inter-group similarity characteristics of the above samples are beneficial to evaluating the practicality of the method of the present invention. After consideration, two similar soybeans are mixed evenly to form a new soybean, such as the new soybean No. 14 (composed of 80% soybean No. 7 + 20% soybean No. 6), and this soybean is used to replace soybean No. 7 for metabolomics analysis. Similarly, a new soybean No. 15 (composed of 80% soybean No. 9 + 20% soybean No. 10) is prepared for replacing soybean No. 9 for further analysis.
[0115] Figure 7 are the infrared spectra of 4 kinds of soybeans. As shown in Figure 7 (c) and (d) of, the infrared spectrum of soybean No. 15 is significantly different from that of soybean No. 9, while the infrared spectrum difference between soybean No. 14 and soybean No. 7 is not obvious ( Figure 7(a) and (b) above. As mentioned above, Soybean No. 14 is composed of Soybean No. 7 and Soybean No. 6, both of which are from Heilongjiang Province, while Soybean No. 15 is a mixture of Soybean No. 9 and Soybean No. 10 from different provinces. It is speculated that geographical location may be an important reason why the spectral difference between Soybean No. 15 and Soybean No. 9 is more obvious than that between Soybean No. 14 and Soybean No. 7.
[0116] Figure 8 It is the principal component analysis diagram of the mixed soybeans, Figure 9 It is the cluster analysis diagram of the pure soybeans, Figure 10 It is the cluster analysis diagram with Soybean No. 14 replacing Soybean No. 7, Figure 11 It is the cluster analysis diagram with Soybean No. 15 replacing Soybean No. 9. From Figure 8 it can be seen that: the overall aggregation position of Soybean No. 14 is higher than that of Soybean No. 7, and Soybean No. 15 is closer to the origin than Soybean No. 9, indicating that there are indeed obvious differences between the mixed soybeans and the pure soybeans, and this difference can also be observed from the cluster analysis diagram ( Figures 9 - 11 ).
[0117] Figure 12 It is the OPLS-DA score diagram of the mixed soybeans, Figure 13 It is the S-plot diagram of the mixed soybeans, Figure 14 It is the permutation test diagram of the mixed soybeans. After Figures 12 - 14 and univariate analysis, the sequence composed of the longest continuous variable numbers is selected as the "marker" characteristic band (Table 2). Among them, 167 "markers" are screened out for Soybean No. 14 by the SVMRFE method (characteristic bands 3323.350 - 3643.533 cm -1 ), which are completely inconsistent with the 177 "markers" (characteristic bands 3660.893 - 4000.364 cm -1 ) screened out for Soybean No. 7. There is also a similar phenomenon of non-overlapping "markers" between Soybean No. 15 and Soybean No. 9. The SVMRFE method screens out 330 "markers" for the former (characteristic bands 3365.784 - 4000.364 cm -1 ), and only 80 "markers" for the latter (characteristic bands 671.228 - 823.604 cm -1 ). In addition, the "marker" characteristic bands of Soybean No. 9 have significantly lower absorbance, while Soybean No. 15 has significantly higher absorbance. The results show that even if the spectral difference is very small, metabolomics based on infrared spectroscopy combined with the SVMRFE method can be used for accurate traceability of soybean origin, further supporting the practicality of the present invention.
[0118] Previous studies have used chemometrics or machine learning algorithms to process infrared spectra and achieved good results. However, these methods require a large number of experimental samples and advanced algorithms to improve the prediction ability of the model. In the present invention, ideal results can be obtained with fewer samples, which is attributed to permutation tests and SVMRFE, both of which have been proven suitable for processing small sample data, which is a significant advantage of the present invention. In actual work, it is possible to preliminarily determine whether the new soybean has the same origin as the soybeans numbered 1 to 10 by evaluating whether the target soybean and the new soybean have the same absorbance significance on the same "marker" characteristic band. If the observation results are inconsistent, it can be directly determined that the origins of the two soybeans are different, thus omitting further metabolomics and support vector machine analyses and improving the screening efficiency.
[0119] The present invention utilizes the advantages of metabolomics and support vector machines in data analysis to develop a feasible method for tracing the origin of soybeans by finding the infrared spectrum "wavenumber markers" of each type of soybean, and it is expected that this method can more accurately improve the accuracy of origin tracing. This is the first time to combine infrared spectroscopy with metabolomics and support vector machines and apply them to the origin tracing of soybeans. The present invention can obtain ideal results relying on less sample data, which can effectively simplify the work process and improve the screening efficiency. Although the present invention only considers 10 types of soybeans with similar infrared spectra from different cities in China, the present invention can completely distinguish them, indicating that this method has broad application prospects, is expected to be applied to the origin tracing of products, and also provides a new analysis perspective for infrared spectrum data processing.
[0120] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for establishing soybean fingerprint, characterized in that: The following steps are involved: Perform infrared spectrum detection on the target soybean sample and the reference soybean sample respectively to obtain the original infrared spectrum information of the soybean sample; The original infrared spectrum information of the soybean sample is calibrated and normalized in sequence to obtain pre-processed infrared spectrum information; Performing principal component analysis on the preprocessed infrared spectrum information to obtain a principal component analysis result; Based on the principal component analysis result, cluster analysis is performed on the preprocessed infrared spectrum information to obtain a cluster analysis result; Based on the cluster analysis results, orthogonal partial least squares discriminant analysis is performed on the pre-processed infrared spectrum information to obtain the pre-wave number marker of the target soybean sample; Based on the pre-wave number markers of the target soybean sample, the pre-processed infrared spectrum information is subjected to support vector machine recursive feature elimination to obtain a fingerprint of the target soybean sample.
2. The establishment method according to claim 1, characterized in that: The target soybean samples are produced from one or more of Xiangyang, Hubei, China, Luliang, Shanxi, China, Liaocheng, Shandong, China, Jining, Shandong, China, Shangqiu, Henan, China, Mudanjiang, Heilongjiang, China, Harbin, Heilongjiang, China, Suihua, Heilongjiang, China, Baicheng, Jilin, China, and Jinzhou, Liaoning, China.
3. The establishment method according to claim 1, characterized in that: The reference soybean samples are produced from one or more of Brazil, the United States and Argentina.
4. The establishment method according to claim 1, characterized in that: The parameters of the infrared spectrum detection include: the scanning mode is diffuse reflection, the scanning range is 400-4000cm -1 , the spectral resolution is 4cm -1 .
5. The establishment method according to claim 1, characterized in that: The calibration was completed in Origin 9.0 software, and the normalization process was performed in SPSS Statistics V17.0 software.
6. The establishment method according to claim 1, characterized in that: The principal component analysis was performed in SIMCA 14.1 software.
7. The establishment method according to claim 1, characterized in that: The cluster analysis was performed in Heml1.0.3.7 software.
8. The establishment method according to claim 1, characterized in that: The orthogonal partial least squares discriminant analysis was performed in SIMCA 14.1 software.
9. The establishment method according to claim 1, characterized in that: The support vector machine recursive feature elimination was performed on the R2019b version of MATLAB software.
10. A soybean traceability method, characterized in that: The following steps are involved: Perform infrared spectrum detection on the soybean to be tested to obtain original infrared spectrum information of the soybean to be tested; By comparing the original infrared spectrum of the soybean to be tested with the soybean fingerprint spectrum obtained by the establishment method described in any one of claims 1 to 9, the origin of the soybean to be tested can be traced.