Corn production place tracing method based on combination of open-type mass spectrometry technology and machine learning

By integrating machine learning with open-source mass spectrometry, the method addresses environmental interference in AMS, enabling rapid and precise pesticide residue detection in corn, enhancing detection accuracy and throughput.

CN120316643APending Publication Date: 2025-07-15CHINA INNOVATION INSTR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386106.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-15
Filing Date
2025-03-30
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing food pesticide residue detection methods are complex and cumbersome and costly. The open mass spectrometry technology is susceptible to environmental interference to false positive results, and lacks broad spectrum.

Method used

Combining open mass spectrometry technology and machine learning, through data screening, model construction and optimization, rapid detection without complex preprocessing is achieved, and machine learning models are used to identify and eliminate environmental interference to improve detection accuracy.

Benefits of technology

It realizes rapid pesticide residue detection without complex sample pretreatment, improves detection throughput and accuracy, supports online and real-time analysis, and has a wide range of crop origin traceability capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316643A_ABST
    Figure CN120316643A_ABST
Patent Text Reader

Abstract

The invention provides a corn origin tracing method based on combination of an open-type mass spectrometry technology and machine learning. The method comprises the following steps: (A1) identifying and testing pesticide residues in corn; (A2) removing background noise; (A3) constructing a mass spectrum secondary analysis model; (A4) optimizing a mass spectrum method; (A5) data preprocessing; (A6) statistical analysis: importing the data into SPSS statistical analysis software, comparing differences among substances in different regions by adopting a significance method, and generating a box plot to visually display data distribution and differences; (A7) establishing an algorithm model, and selecting a classification algorithm model according to the characteristics of the data and the number of samples; performing hyper-parameter tuning by using a grid search method, and selecting a parameter combination when the classification accuracy is the highest as an optimal parameter of the algorithm; and (A8) cross validation. The method has the advantages of accurate detection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to mass spectrometry technology, and particularly to a method for tracing the origin of corn based on the combination of open mass spectrometry technology and machine learning. Background Art

[0002] The detection of pesticide residues in food is of great significance for protecting public health and ensuring food safety. Although the extensive use of pesticides has effectively increased crop yields, it has also brought the risk of pesticide residues. These residues enter the human body through the food chain, and long-term accumulation may cause a series of health problems, such as poisoning, endocrine disorders, chronic diseases such as cancer, etc. Therefore, regular detection of pesticide residues in food can timely detect excessive pesticides, trace the origin of food, and protect the lives and health of consumers.

[0003] Food safety detection mainly relies on laboratory detection methods, such as liquid chromatography / gas chromatography tandem mass spectrometry (LC / GC-MS). For different detection items, the detection limit can reach the ng / g or ng / mL level. However, it is usually accompanied by complex and cumbersome pretreatment processes (such as extraction, enrichment, separation, etc.), which are quite time-consuming, laborious and costly. Currently, the existing rapid detection methods mainly include chemical detection methods, enzyme inhibition methods, immunoassay methods, spectroscopic methods, etc. For example, supermarkets or markets mainly use immunoassay reagent colorimetry for pesticide residue detection. Its principle mainly utilizes the heterogeneity of pesticides on the activity of acetylcholinesterase, and uses a spectrophotometer to measure the activity of acetylcholinesterase to calculate the residue amount of pesticides. Although this method has the advantage of rapid detection, it is only applicable to organophosphorus and carbamate pesticides, and the types of detected substances are limited and it does not have broad-spectrum properties.

[0004] Ambient ionization mass spectrometry (AMS) can directly perform rapid detection and analysis on complex matrix samples such as gases, liquids, or solids without complex sample pretreatment. This technology simplifies the analysis process, increases the sample detection throughput, and supports on-line and real-time analysis. AMS technology has the advantages of simplicity, economy, and high throughput. It plays a very important role in improving the traceability and certification of food. However, ambient mass spectrometry technology may be interfered by environmental factors such as air, solvent vapor, and dust. Especially in complex samples or when high-precision analysis is required, the influence of impurities may lead to an increase in signal noise, the appearance of additional peaks or peak changes in the mass spectrum, which may be misjudged as target compounds, resulting in false positive results. To overcome this challenge, it is decided to combine machine learning technology to improve the detection performance of ambient mass spectrometry technology. Machine learning algorithms can learn and extract useful feature information from a large amount of data, helping to better understand and distinguish the differences between target compounds and environmental interference factors. By training a machine learning model, the model learns from the existing mass spectrometry data and judges the detection results accordingly. In this way, even if environmental factors interfere during the actual detection process, the machine learning model can accurately identify and eliminate these interferences, thereby reducing the false positive phenomenon. Summary of the Invention

[0005] To address the deficiencies in the above prior art solutions, the present invention provides a method for tracing the origin of corn based on the combination of ambient mass spectrometry technology and machine learning.

[0006] The object of the present invention is achieved through the following technical solutions: A method for tracing the origin of corn based on the combination of ambient mass spectrometry technology and machine learning, comprising the following steps: (A1) Identification test of pesticide residues in corn. Analyze the sample solution obtained after pretreatment in the positive ion mode to obtain the primary mass spectrometry data of all samples; (A2) Remove background noise. If the highest peak intensity of the primary mass spectrum is greater than 105 and the number of mass spectrum peaks is greater than 400, retain this spectrum. Screen the primary mass spectrometry data to select target substances that may contain pesticide residues; (A3) Construct a secondary mass spectrometry analysis model; (A4) Optimize the mass spectrometry method. Perform secondary mass spectrometry detection on the target substances in the primary mass spectrometry data to improve the accuracy of mass spectrometry detection, and obtain a secondary mass spectrometry data matrix containing detected ions and their relative intensities; (A5) Data preprocessing: Fill in the missing values in the data with zero values; perform normalization processing on the secondary mass spectrometry data matrix; Statistical analysis: Import the data into the SPSS statistical analysis software, and use the significance method to compare the differences between substances in different regions, and generate a box plot to visually display the data distribution and differences. (A7) Establish an algorithm model: Select a classification algorithm model according to the characteristics of the data and the number of samples; use the grid search method to optimize the hyperparameters, and select the parameter combination with the highest classification accuracy as the optimal parameters of the algorithm. (A8) Cross-validation.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: The advantage of the present invention is that it does not require complex sample pretreatment, can complete the mass spectrometry data acquisition of a single sample within 1 minute, combines machine learning to establish a classification model, and can be used for the rapid identification of pesticide residues in corn and the origin traceability. Compared with the traditional method, this method not only has simple operation, simplifies the analysis process, increases the detection throughput of samples, but also realizes the online and real-time analysis of samples, and can quickly trace the origin of crops by combining machine learning, and has broad application prospects. Description of the Drawings

[0008] Referring to the accompanying drawings, the disclosure of the present invention will become more understandable. It is easy for those skilled in the art to understand that these drawings are only used to illustrate the technical solutions of the present invention and are not intended to limit the protection scope of the present invention. In the drawings: Figure 1 is a schematic structural diagram of a method for tracing the origin of corn based on the combination of open mass spectrometry technology and machine learning according to an embodiment of the present invention; Figures 2 - 7 is the collected spectrogram. Specific Embodiments

[0009] Figures 1 - 7 The following description and illustration describe alternative specific embodiments of the present invention to teach those skilled in the art how to implement and reproduce the present invention. Some conventional aspects have been simplified or omitted to teach the technical solutions of the present invention. Those skilled in the art should understand that variations or substitutions derived from these specific embodiments will fall within the scope of the present invention. Those skilled in the art should understand that the following features can be combined in various ways to form multiple variations of the present invention. Thus, the present invention is not limited to the following alternative specific embodiments, but is defined only by the claims and their equivalents. Embodiment

[0010] The method for tracing the origin of corn based on the combination of open mass spectrometry technology and machine learning according to Embodiment 1 of the present invention, as Figure 1 shown, includes the steps: Sample collection: Freeze and store the samples purchased from various places, and mark the sampling location and sampling date for later use.

[0011] Sample pretreatment: After thawing the sample, homogenize it. Weigh 10 g of the sample and place it in a 50 mL centrifuge tube, add 10 mL of methanol. Weigh 10 g of the sample in a 50 mL centrifuge tube, add 10 mL of methanol, and then add a QuEChERS extraction packet (containing 4 g of MgSO4, 1 g of NaCl, 1 g of sodium citrate, and 0.5 g of disodium hydrogen citrate). Vortex mix for 5 min, let it stand for 3 min, then centrifuge at 8000 r / min for 5 min. Take the supernatant and transfer it to a 10 mL centrifuge tube. Repeat the extraction of the residue with 10 mL of methanol once, combine the supernatants in a 10 mL centrifuge tube, and filter through a 0.22 μm organic filter membrane for further analysis.

[0012] Dielectric barrier discharge mass spectrometry analysis: Perform thermal desorption on the sample solution, and the volatilized substances enter the mass spectrometer for detection and identification. The analysis conditions of the mass spectrometry are as follows: The ion source uses a DBDI-200 source; the discharge voltage and frequency are 3.9 kV and 20 kHz respectively; the ion source temperature is 274 °C; the scanning mode is positive ion mode; the discharge gas is air; the injection method is that 10 μL of the sample solution is heated and volatilized and directly inhaled into the ion source; the mass spectrometry scanning range is m / z 100 - 500; the MS / MS collision energy is 30 V.

[0013] The mass spectrometry acquisition range is m / z 100 - 500. Samples are collected from various locations (A - F), and the collected spectra are as Figures 2 - 7 shown.

[0014] (A1) Identification test for pesticide residues in corn. Analyze the sample solution obtained after pretreatment in the positive ion mode to obtain the primary mass spectrometry data of all samples; (A2) Remove background noise. If the peak intensity of the highest peak in the primary mass spectrometry is greater than 105 and the number of mass spectrometry peaks is greater than 400, retain this spectrum. Screen the primary mass spectrometry data to select the target substances that may contain pesticide residues; (A3) Construct a secondary mass spectrometry analysis model; (A4) Optimize the mass spectrometry method. Perform secondary mass spectrometry detection on the target substances in the primary mass spectrometry data to improve the accuracy of mass spectrometry detection and obtain a secondary mass spectrometry data matrix containing detection ions and their relative intensities; (A5) Data preprocessing: Fill the missing values in the data with zero values; perform normalization processing on the secondary mass spectrometry data matrix; After collecting data by a mass spectrometer, the data is preprocessed to extract the effective information in each sample. First, the detected secondary mass spectrometry peaks (i.e., m / z values) and their corresponding ion intensities are extracted from the mass spectrometry graph to form a mass spectrometry peak list. Second, the total intensity of each detected ion intensity is normalized, and the sampling solvent background signal is removed to obtain a secondary mass spectrometry data matrix containing detected ions and their relative intensities. In all samples, missing values are filled with zeros.

[0015] (A6)Statistical analysis: Import the data into the SPSS statistical analysis software, and use the significance method to compare the differences between substances in different regions, and generate a box plot to visually display the data distribution and differences; (A7)Establish an algorithm model: Select a classification algorithm model according to the characteristics of the data and the number of samples; use the grid search method to optimize the hyperparameters, and select the parameter combination with the highest classification accuracy as the optimal parameters of the algorithm; (A8)Cross-validation, the specific method is as follows: The entire data set is used in the classification model. First, the data set is randomly divided into five uniform subsets; In each round of validation, four subsets are selected as the training set, and the remaining one subset is used as the test set; by training the model and evaluating its performance on the test set, and then cycle this process five times, each time changing the different test sets; In each cross-validation, the performance indicators of the model, such as accuracy, precision, and recall, are calculated; Finally, calculate the average performance of all five validations as the overall evaluation index of the model.

[0016] In addition, the ROC curve (Receiver Operating Characteristic Curve) can be drawn and the AUC value (Area Under the Curve) can be calculated to further judge the performance of the classification model in distinguishing different categories.

[0017] The constructed classification models specifically include Logistic Regression, Support Vector Machine, and Random Forest. The specific parameter settings for each method are as follows: The number of RF selection trees is 100; SVM selects the poly function as the kernel function; LR selects C = 1 and solver = lbfgs. The data is randomly divided into 5 equal parts. Four of them are taken in turn as the training set to train the classification model, and the last one is used as the test set to test the performance of the finally obtained classification model and verify the accuracy of the model. The model is verified according to the five-fold cross-validation method, and the classification accuracy of the model is calculated using the following formula: 。

[0018] 。

[0019] 。

Claims

1. A method for tracing the origin of corn based on the combination of open mass spectrometry technology and machine learning, comprising the steps: (A1) Identification test of pesticide residues in corn. Analyze the sample solution obtained through pretreatment in the positive ion mode to obtain the first-order mass spectrometry data of all samples; (A2) Remove background noise. If the peak intensity of the highest peak in the first-order mass spectrometry is greater than 105 and the number of mass spectrometry peaks is greater than 400, retain this spectrum, and perform data screening on the first-order mass spectrometry data to screen out target substances that may contain pesticide residues; (A3) Construct a second-order mass spectrometry analysis model; (A4) Optimize the mass spectrometry method. Perform second-order mass spectrometry detection on the target substances in the first-order mass spectrometry data to improve the accuracy of mass spectrometry detection, and obtain a second-order mass spectrometry data matrix containing detection ions and their relative intensities; (A5) Data preprocessing: Fill in the missing values in the data with zero values; Normalize the second-order mass spectrometry data matrix; (A6) Statistical analysis. Import the data into the SPSS statistical analysis software, and use the significance method to compare the differences between substances in different regions, and generate a box plot to visually display the data distribution and differences; (A7) Establish an algorithm model. Select a classification algorithm model according to the characteristics of the data and the number of samples; use the grid search method to optimize the hyperparameters, and select the parameter combination with the highest classification accuracy as the optimal parameters of the algorithm; (A8) Cross-validation.

2. The method for tracing the origin of corn based on the combination of open mass spectrometry technology and machine learning according to claim 1, wherein The way of cross-validation is: The entire data set is used in the classification model. First, randomly divide the data set into five uniform subsets; In each round of validation, select four subsets as the training set, and the remaining one subset as the test set; evaluate its performance on the test set by training the model, and then cycle this process five times, each time changing the different test set; In each cross-validation, the performance indicators of the accuracy, precision, and recall rate of the model will be calculated; Finally, calculate the average performance of all five validations as the overall evaluation index of the model.