A method for quantitatively detecting DON toxin in wheat grains
Through ATR-FTIR combined with neural network and recursive feature elimination technology, the problem of high cost and low accuracy of DON toxin detection in wheat grains is solved, and low-cost, fast and accurate detection effect is achieved.
Patent Information
- Application Number
- CN202410488088.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-04-23
AI Technical Summary
The prior art wheat grain DON toxin detection method is costly, complex in operation, and is not suitable for large-scale sample detection, and the ATR-FTIR detection feature wavenumber extraction is incomplete and the model prediction stability is poor.
ATR-FTIR combined with neural network and recursive feature elimination technology is used to extract feature wavenumbers and construct a linear regression model to achieve low-cost and rapid detection of DON toxins in wheat grains.
The detection cost is reduced, the detection accuracy and efficiency are improved, the correlation coefficient reaches 0.97, and the cost is between 1-5 yuan/sample, which is significantly reduced compared with traditional methods.
Smart Images

Figure CN118392816B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for quantitatively detecting DON toxin in wheat grains, belonging to the field of food safety. Background Art
[0002] Deoxynivalenol (DON), also known as vomitoxin, is a toxin produced by certain Fusarium fungi and is commonly found in grains such as wheat and corn. DON poses a threat to human and animal health. Especially after consuming contaminated grains, it may cause vomiting, diarrhea, immune system suppression, and poor nutrient absorption. Therefore, accurately detecting the content of DON toxin in wheat is crucial for food safety.
[0003] Common methods for traditionally detecting the content of DON toxin in wheat include liquid chromatography and enzyme-linked immunosorbent assay. These detection methods require professional instruments, are costly, have high technical requirements for operators, and also have the disadvantages of a long detection period and being unsuitable for large-scale sample detection. Therefore, it is particularly important to construct a technical system for accurately, efficiently, and low-costly detecting DON toxin in wheat.
[0004] Attenuated Total Reflectance Fourier Transform Infrared Spectroscopy (ATR-FTIR) is a variant of Fourier transform infrared spectroscopy that combines attenuated total reflection technology. This technology has unique advantages in sample analysis and is particularly suitable for the analysis of liquid, paste, thin film, and solid samples. Researchers have made certain explorations in using ATR-FTIR to detect DON toxin in wheat grains, but there are problems such as incomplete extraction of characteristic wave numbers and poor stability of model prediction. Therefore, establishing a low-cost and rapid detection technical system for DON toxin in wheat grains using ATR-FTIR technology has important theoretical and practical significance. Summary of the Invention
[0005] Object of the Invention: Aiming at the deficiencies of the prior art, the present invention provides a method for quantitatively detecting DON toxin in wheat grains to achieve low-cost, rapid, and accurate detection.
[0006] Technical Solution: To achieve the above object, the present invention adopts the following technical solution:
[0007] A method for quantitatively detecting DON toxin in wheat grains, comprising the following steps:
[0008] Step 1: After grinding the sample wheat grains into powder, extract DON toxin with an acetonitrile aqueous solution, and reserve it after centrifugation and filtration;
[0009] Step 2: Absorb the sample solution from Step 1, and obtain spectral information through scanning by an ATR-FTIR instrument. According to the DON molecular structure and its absorption peaks in the infrared spectrum, select the spectral information within a specified wavenumber range from the full spectral wavelength range as the model input;
[0010] Step 3: Feed the spectral information screened in Step 2 into a sequential neural network model composed of two fully connected layers to extract characteristic wavenumbers. The first fully connected layer of this neural network is a hidden layer containing 10 neurons, using a linear activation unit, and the second fully connected layer is an output layer containing one neuron, which outputs a predicted value;
[0011] Step 4: Use the characteristic wavenumbers extracted in Step 3 to train the first linear regression model. During the training process, use the absolute value of the coefficient of the first linear regression model to evaluate the importance of each feature, delete the feature with the smallest absolute value of the coefficient, and repeat this process until the specified condition is met to obtain the final feature set. The i-th characteristic wavenumber is denoted as w i , and the corresponding coefficient is β i ;
[0012] Step 5: Use the feature set obtained in Step 4 to train the second linear regression model, y pred =b + β1v1 + β2ν2 + … + β i v i , where b is a constant, and v i is the spectral measurement value at the characteristic wavenumber w i ;
[0013] Step 6: Use the second linear regression model, substitute the spectral measurement value obtained by the method of Steps 1 and 2 for the wheat grains to be measured, and output the quantitative detection result of DON toxin.
[0014] Further, in Step 2, select the spectral information within the following wavenumber ranges from the 400 - 4000 cm -1 full spectral wavelength range: 850 - 1300 cm -1 , 1500 - 1750 cm -1 , 2800 - 3100 cm -1 , 3200 - 3600 cm -1 as the model input.
[0015] Further, in Step 3, the neural network model is obtained through training. During the training process, use the DON content of the sample in Step 1 determined by the immunoaffinity chromatography purification high performance liquid chromatography method according to the national food safety standard GB5009.111 - 2016 as the true value, use the Adam optimizer to calculate the first moment and the second moment of the gradient to adjust the learning rate of each parameter, and use the mean square error as the loss function.
[0016] Furthermore, during the training process of the neural network model, the weights are extracted from the first fully connected layer, and the cumulative sum of their absolute values is calculated. Among them, S is the cumulative sum of the absolute values of the weights of the entire first fully connected layer, M corresponds to the number of wavenumbers in the spectral information, |w ij | is the absolute value of the weight corresponding to the j-th neuron in the input layer of the i-th neuron in the first fully connected layer. The double summation represents the accumulation of the absolute values of all weights in the first fully connected layer; the top 10% of the characteristic wavenumbers with the highest importance are selected based on the S value.
[0017] Furthermore, in step 4, 43 characteristic wavenumbers are obtained.
[0018] Furthermore, in step 5, the performance of the second linear regression model is evaluated, including the coefficient of determination, root mean square error, and relative prediction error.
[0019] Beneficial effects:
[0020] 1. Due to the influence of different varieties and qualities of wheat, there are large detection errors when scanning the spectra of wheat flour, and it is more difficult to perform accurate quantification. The present invention uses the extracted liquid sample of DON toxin for spectral scanning, which can reduce errors to a certain extent and improve the detection accuracy. Compared with the existing liquid chromatography-mass spectrometry method, the sample processing method of the present invention is simpler and faster, and only requires one chemical reagent.
[0021] 2. The present invention combines the molecular structure characteristics of DON toxin, applies the recursive feature elimination technology to the screening of spectral characteristic wavenumbers, identifies the characteristic wavenumbers significantly related to DON toxin, and constructs a linear regression model with very good prediction effect, and the correlation coefficient reaches 0.97.
[0022] 3. The present invention uses a neural network to obtain features, performs feature screening and then establishes a linear regression model. The method is simple and convenient. The cost of detecting with the constructed model is 1 - 5 yuan per sample. Compared with liquid chromatography-mass spectrometry and ELISA methods, the cost is greatly reduced. Compared with other ATR-FTIR detection methods, the prediction accuracy is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flow chart of the method for quantitatively detecting DON toxin in wheat grains of the present invention;
[0024] Figure 2 It is a diagram showing the comparison result of the predicted value and the true value obtained by the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0025] The technical solutions of the present invention will be further described below with reference to the accompanying drawings.
[0026] Reference Figure 1 , in the embodiment of the present invention, a method for quantitatively detecting DON toxin in wheat grains is provided, including the following steps:
[0027] Step 1: Accurately weigh 5 g of the pulverized wheat sample into a 50 mL centrifuge tube, add 20 mL of an acetonitrile aqueous solution with a volume fraction of 84%, oscillate and extract for 30 min, then centrifuge. After filtering through a 0.22 μm filter membrane, the liquid chromatography is used for spectral acquisition.
[0028] Step 2: Determine the DON content in the sample in Step 1 according to the determination steps of immunoaffinity chromatography purification high performance liquid chromatography in the national food safety standard GB5009.111-2016. After detection, the change range of DON toxin content is 2.4 ppm - 141 ppm, and this value is used as the true value in the following neural network learning process.
[0029] Step 3: Collect the spectral information of the liquid sample in Step 1. Turn on the ATR-FTIR instrument. In the embodiment, the Fourier transform infrared spectrometer Nicolet iS50 equipped with an ATR module is used. The experimental temperature is controlled at about 20 °C, and let it preheat to a stable state. Conduct a background scan once without a sample to eliminate the influence of the environment and the instrument. Use a pipette to suck 10 μL of the liquid sample and directly drop it onto the surface of the ATR crystal, and ensure that the sample spreads evenly on the crystal surface to obtain the best measurement results. Conduct a sample scan, collect the infrared spectral data, and save it in the CSV file format. The scanned spectral wavelength range: 4000 - 400 cm-1; resolution: 4 cm-1. The number of scans: 32. Since the DON molecule carries specific structures, these structures have specific absorption ranges in this spectrogram. Therefore, according to the DON molecular structure and its absorption peaks in the infrared spectrum, the present invention screens and extracts the spectral information within the following wavenumber ranges from the above CSV file: 850 - 1300 cm -1 , 1500 - 1750 cm -1 , 2800 - 3100 cm -1 , 3200 - 3600 cm -1 , a total of 2909 spectral information is extracted, which is integrated with the DON toxin content in Step 2 and used as the input data in Step 4. The samples are divided into a training set and a test set to ensure that the samples can be evenly divided to improve the generalization ability of the model.
[0030] Step 4: Screen and export important characteristic wavenumbers from the spectral data.
[0031] 4.1. Construct a sequential neural network model consisting of two fully connected layers. The first layer is the hidden layer, with 10 neurons selected to provide sufficient model complexity to capture the features of the input data while avoiding overfitting. Use the ReLU activation function (Rectified Linear Unit), whose main advantage lies in its simplicity and computational efficiency. Compared with other activation functions such as sigmoid or tanh, it is computationally simpler and more efficient. The size of the input layer matches the number of wavenumbers in the spectral information of step 3. The input data contains 2909 spectral information extracted and the DON content detected by traditional methods, specifically in matrix form. Each row corresponds to a sample, the first column is the true value, and the columns from the second onwards are the measured values. The first row is the wavenumber w, and starting from the second row are the spectral measurement values v under w. The second layer of the neural network is the output layer, which contains only one neuron for outputting the predicted value of the DON content.
[0032] 4.2. Use the Adam optimizer, which combines the advantages of momentum and adaptive learning rate (RMSprop) and is considered an optimizer for efficiently handling large datasets and deep learning models. Adam adjusts the learning rate of each parameter by calculating the first moment (mean) and second moment (unbiased variance) of the gradient. The specific calculation process is as follows: Calculate the first moment (mean) and second moment (unbiased variance) of the gradient: Perform bias correction on the first moment and second moment: Update the parameters: where θ is the model parameter, g t is the gradient at time step t, β1 and β2 are values close to 1, set here to 0.999, η is the learning rate, and are the estimates of the first and second moments after bias correction respectively, and ε is a very small number used to prevent division by zero. Use the Mean Squared Error (MSE) as the loss function to measure the difference between the predicted value and the true value of the model output. The goal of the optimizer is to minimize the loss function by adjusting the model parameters.
[0033] 4.3. The model will be iteratively trained. Extract the weights from the first layer of the model and calculate the cumulative sum of their absolute values to evaluate the importance of characteristic wavenumbers. The cumulative sum S of the absolute values of the weights in the first layer can be expressed as: where S is the cumulative sum of the absolute values of all the weights in the first layer, M corresponds to the number of wavenumbers in the spectral information, |w ij| is the absolute value of the weight corresponding to the j-th neuron in the input layer for the i-th neuron in the first layer. The double summation represents the accumulation of the absolute values of all weights in the first layer. Based on the S score, the top 10% of the feature wavenumbers in terms of importance are selected for the analysis in step 5.
[0034] Step 5: Refine spectral feature wavenumbers. The Recursive Feature Elimination (RFE) technique is used to refine the feature wavenumbers preliminarily screened by the neural network, and linear regression is used as the estimator to implement RFE. During the execution of RFE, first, a linear regression model (the first linear regression model) is trained using the feature set in step 4. According to the feature wavenumbers selected in step 4, for each feature wavenumber of each sample, there is a corresponding spectral measurement value, and the first linear regression model is trained. Then, by analyzing the absolute values of the coefficients of this model, the importance of each feature is evaluated, where the absolute value of the coefficient reflects the contribution degree of the feature to the model output. Based on this evaluation, the feature with the lowest contribution, that is, the feature with the smallest absolute value of the coefficient, is removed one by one to ensure that the final feature set is both refined and effective. Finally, 43 feature wavenumbers are obtained, as shown in Table 1. Where i is the feature serial number, W is the feature wavenumber, and β is the corresponding coefficient.
[0035] Table 1 Feature wavenumbers and related coefficients
[0036]
[0037] Step 6: Construction and evaluation of the prediction model. The prediction model uses a linear regression model (the second linear regression model), prediction model: y pred = b + β1ν1 + β2v2 + … + β i ν i , where b = 19819.2581333114, i is the serial number, there are a total of 43 feature wavenumbers, w i in Table 1 is the finally screened feature wavenumber, ν i is the spectral measurement value at the w i wavenumber. For example, for a new sample to be measured, scanning the ATR-FTIR spectrum, v1 is the measurement value of w1 (w1 here is 857.2036). β i is the coefficient for each w i . Substituting the spectral measurement value of the wheat grain to be measured into the formula, the predicted DON content can be calculated.
[0038] In the embodiment, the performance of the model is evaluated using a test set, and the evaluation parameters include the coefficient of determination (R-squared), the root mean squared error (RMSE), and the relative prediction difference (RPD). The fitting curve of the sample true value and the predicted value is as Figure 2 shown, with a slope of 1.04, an intercept of -1.44, and the R 2 of the fitting curve being 0.97, indicating a very good prediction effect. The RMSE value is 4.16 and the RPD value is 6.16, indicating that the model has high prediction accuracy and can be used for actual quantitative analysis.
[0039] The present invention uses a neural network to obtain features, performs feature screening, and then establishes a linear regression model. The method is simple and convenient, and the cost of detecting using the constructed model is 1 - 5 yuan / sample. Compared with liquid chromatography-mass spectrometry and ELISA methods, the cost is greatly reduced. Compared with other ATR-FTIR detection methods, the prediction accuracy is higher.
[0040] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A method for quantitatively detecting DON toxin in wheat grains, characterized in that, It includes the following steps: Step 1: After grinding the sample wheat grains into flour, extract DON toxin with an acetonitrile aqueous solution, and set aside after centrifugation and filtration; Step 2: Absorb the sample solution in Step 1, obtain spectral information through scanning with an ATR-FTIR instrument, and select the spectral information within a specified wavenumber range from the full spectral wavelength range as the model input according to the DON molecular structure and its absorption peaks in the infrared spectrum; Step 3: Feed the spectral information screened in Step 2 into a sequential neural network model composed of two fully connected layers to extract characteristic wave numbers. The first fully connected layer of this neural network is a hidden layer containing 10 neurons, using a linear activation unit, and the second fully connected layer is an output layer containing one neuron, which outputs a predicted value. The neural network model is obtained through training. During the training process, the DON content of sample D in Step 1 determined by the immunoaffinity chromatography purification high performance liquid chromatography method according to the national food safety standard GB5009.111-2016 is used as the true value, and the Adam optimizer is used to calculate the first moment and the second moment of the gradient to adjust the learning rate of each parameter. The mean square error is used as the loss function. During the training process of the neural network model, the weights are extracted from the first fully connected layer, and the sum of the absolute values is calculated. , where S is the sum of the absolute values of all the weights of the entire first fully connected layer, M corresponds to the number of wave numbers in the spectral information, is the absolute value of the weight corresponding to the i th neuron in the first fully connected layer corresponding to the j th neuron in the input layer. The double summation represents the accumulation of the absolute values of all the weights of the first fully connected layer. Based on the S value, the top 10% of the characteristic wave numbers with the highest importance are selected. Step 4: Train the first linear regression model using the characteristic wave numbers extracted in Step 3. During the training process, use the absolute value of the coefficients of the first linear regression model to evaluate the importance of each feature, delete the feature with the smallest absolute value of the coefficient, and repeat this process until the specified condition is met to obtain the final feature set. The \(i\)-th characteristic wave number is denoted as , and the corresponding coefficient is ; Step 5: Use the feature set obtained in Step 4 to train a second linear regression model. , where b is a constant. is the characteristic wavenumber. is the spectrometric value at Step 6: Use the second linear regression model, input the spectral measurement values obtained by the method in Steps 1 and 2 for the wheat grains to be tested, and output the quantitative detection result of DON toxin.
2. The method according to claim 1, wherein In the said step 2, spectral information within the following wavenumber ranges is selected from the full-spectrum wavelength range of 400 - 4000 cm -1 as the model input: 850 - 1300 cm -1 , 1500 - 1750 cm -1 , 2800 - 3100 cm -1 , 3200 - 3600 cm -1 3. The method according to claim 1, wherein In the said Step 4, 43 characteristic wavenumbers are obtained.
4. The method according to claim 1, wherein In the said Step 5, performance evaluation is carried out on the second linear regression model, including the coefficient of determination, root mean square error, and relative prediction error.
Citation Information
Patent Citations
Amino acid adsorption mechanism research method based on infrared spectroscopy and DFT calculation
CN104749129A
Quantitative detection method for vomitoxin in flour
CN110057777A