A method for safety detection and analysis of pesticide residues in food based on machine learning
Through the combination of two-dimensional liquid chromatography-high resolution mass spectrometry combined technology and machine learning model, the complexity and instability of pesticide residue detection are solved, and high-precision and high-sensitivity pesticide residue detection are achieved.
Patent Information
- Application Number
- CN202411646904.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The existing pesticide residue detection methods are complex in operation, long detection time, low sensitivity, and environmental factors and instrument response deviations lead to unstable detection results.
Combining two-dimensional liquid chromatography-high resolution mass spectrometry combined technology and machine learning algorithms, machine learning models are established to identify and classify pesticide residue components through environmental factor correction and baseline drift correction.
It improves the accuracy and sensitivity of detection, reduces the impact of environmental factors and instrument response deviations, and achieves fast and accurate pesticide residue detection.
Smart Images

Figure CN119580877B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food detection, and in particular, to a method for detecting and analyzing food pesticide residues based on machine learning. Background Technique
[0002] With the increasing severity of global food safety issues, the detection of pesticide residues has become an important part of food safety monitoring. Pesticide residues can not only pose a threat to the health of consumers but also affect the export and market competitiveness of agricultural products. Therefore, quickly and accurately detecting pesticide residues in food is the key to ensuring food safety. Existing methods for detecting pesticide residues mainly include gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), enzyme-linked immunosorbent assay (ELISA), high-performance liquid chromatography (HPLC), etc. Although these traditional methods have certain advantages in pesticide residue detection, they usually have limitations such as complex operation, long detection time, and low sensitivity.
[0003] In recent years, with the rapid development of big data technology and machine learning algorithms, machine learning-based methods for detecting pesticide residues have gradually attracted attention. Machine learning can automatically learn and extract potential patterns from a large amount of experimental data, thereby achieving precise identification and quantification of pesticide residues. By combining with high-performance chromatography-mass spectrometry analysis technology, machine learning models can identify different types of pesticides in complex food matrices and quantify their concentrations. However, in practical applications, the complexity of food samples and changes in the experimental environment often lead to unstable detection results. Therefore, how to effectively preprocess and correct experimental data to eliminate the interference of environmental factors and instrument responses is the key to improving detection accuracy and reliability.
[0004] Chromatography-mass spectrometry coupling technology (such as liquid chromatography-high-resolution mass spectrometry coupling technology) is widely used in pesticide residue detection due to its high resolution, high selectivity, and sensitivity. This technology can accurately obtain the mass spectrometry map of each pesticide by separating and qualitatively analyzing the extract. However, chromatographic and mass spectrometric instruments may exhibit response deviations or baseline drifts under different environmental conditions, such as temperature and humidity, which will affect the accuracy of detection results. Therefore, when performing pesticide residue analysis, it is necessary to correct the mass spectrometry data to ensure the reliability of the data.
[0005] In recent years, in order to improve the accuracy and efficiency of pesticide residue detection, research has begun to combine machine learning with data correction techniques. By obtaining environmental information (such as temperature, humidity, etc.) in real time and correcting experimental data, the interference of environmental factors on mass spectrometry data can be effectively reduced. At the same time, based on the corrected data, a machine learning model is constructed, which can automatically identify the types and concentrations of pesticide residues, greatly improving the detection efficiency and accuracy. Using feature extraction methods and matrix effect correction techniques can further optimize the data quality and reduce the impact of sample matrix on the detection results.
[0006] Therefore, combining chromatography-mass spectrometry technology and machine learning algorithms, and adding environmental factor correction and data correction steps has become an important research direction in the field of food pesticide residue detection. Through this innovative technical method, not only can the accuracy and efficiency of detection be improved, but also more reliable data support can be provided in large-scale food safety monitoring. Summary of the Invention
[0007] In view of this, the present invention proposes a machine learning-based food pesticide residue safety detection and analysis method, aiming to solve the problems existing in the prior art and provide a fast, accurate and highly adaptable food pesticide residue detection method.
[0008] The present invention proposes a machine learning-based food pesticide residue safety detection and analysis method, including:
[0009] Collect a number of food samples, perform pretreatment on all food samples to obtain extracts for detection; use two-dimensional liquid chromatography-high resolution mass spectrometry technology to analyze the extracts, obtain mass spectrometry maps, and convert the map data into mass spectrometry numerical data;
[0010] Correct the mass spectrometry numerical data to obtain corrected data; establish a machine learning model based on the corrected data, and the machine learning model is used to identify and classify the pesticide residue components in the food samples to be tested;
[0011] Extract the pretreatment extracts of the food samples to be tested, and use the same two-dimensional liquid chromatography-high resolution mass spectrometry technology for analysis to obtain the mass spectrometry maps of the food samples to be tested; convert the map data of the food samples to be tested into mass spectrometry numerical data to be tested; perform the same correction on the mass spectrometry numerical data to be tested to obtain corrected data to be tested;
[0012] Input the corrected data to be tested into the established machine learning model for identification and classification of pesticide residue components; according to the results output by the model, perform quantitative analysis on the types and concentrations of pesticide residues to obtain analysis results;
[0013] Compare the analysis results with food safety standards to determine whether the food sample to be tested meets the safety standards and generate corresponding test reports.
[0014] Preferably, the mass spectrometry numerical data includes ion intensity data, peak shape parameters, ion current curve shape characteristics, and retention time.
[0015] Preferably, the correction includes environmental factor correction and baseline drift correction;
[0016] The environmental factor correction includes temperature correction and humidity correction. Calculate the environmental correction response value based on the actual temperature and actual humidity; then calculate the final baseline correction value based on the environmental correction response value;
[0017] Use the final baseline correction value to correct the ion intensity data and peak shape parameters.
[0018] Preferably, the correction formula for the environmental factor correction is:
[0019] Rt = Rm·(1 + kT·(Ta - Tr));
[0020] Rh = Rt·(1 + kH·(Ha - Hr));
[0021] Wherein, kT and kH respectively represent the temperature and humidity correction coefficients; Ta and Ha respectively represent the actual temperature and actual humidity; Tr and Hr respectively represent the standard temperature and standard humidity; Rt and Rh respectively represent the temperature correction value and the environmental correction response value;
[0022] The baseline drift correction formula is:
[0023]
[0024] Rbase = Rh - Bd;
[0025] Wherein, N represents the size of the moving average window; Rh represents the environmental correction response value; Rbase represents the final baseline correction value.
[0026] Preferably, when using the final baseline correction value to correct the ion intensity data, it includes:
[0027] Ic(t) = Ir(t) - Rbase(t);
[0028] Wherein, Ic(t) represents the ion intensity after baseline correction; Ir(t) represents the original ion intensity data.
[0029] Preferably, when using the final baseline correction value to correct the peak shape parameters, it includes:
[0030] The peak shape parameters include peak area, peak height, and peak width;
[0031]
[0032] Hc = Hr - Rbase(tpeak);
[0033] Wc = t(Hc / 2,left) - t(Hc / 2,right);
[0034] tleft = t(Hc / 2,left);
[0035] tright = t(Hc / 2,right);
[0036] Among them, Ac represents the peak area after baseline correction; Ic(t) represents the ion intensity after baseline correction; tpeak represents the time point at which the peak height appears; Hc represents the peak height after baseline correction; Hr represents the maximum ion intensity value of the peak; tleft and tright respectively represent the left and right time points at half the height of the peak height after baseline correction; Wc represents the peak width after baseline correction.
[0037] Preferably, based on the corrected ion intensity, the shape features of the ion current curve are extracted, including the calculation of the curve symmetry coefficient and the curve slope ratio:
[0038]
[0039] Among them, S represents the curve symmetry coefficient; Rslope represents the curve slope ratio; slopea and slopeb respectively represent the average slopes of the rising and falling segments.
[0040] Preferably, the correction also includes standardization
[0041] The formula for performing the standardization is:
[0042]
[0043] Among them, X represents the corrected data; μ and σ respectively represent the mean and standard deviation of the corresponding data; Xnor represents the data after standardization.
[0044] Preferably, according to the corrected data, a machine learning model is established. When the machine learning model is used to identify and classify the pesticide residue components in the food sample to be tested, feature selection is also included;
[0045] The feature selection uses the recursive feature elimination method to screen the features that contribute the most to the model performance from the peak area, peak height, peak width, curve symmetry coefficient, and curve slope ratio.
[0046] Preferably, the pretreatment includes washing, crushing, homogenization, extraction and concentration, and filtration.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] 1. Improve detection accuracy and reliability:
[0049] By introducing environmental factor correction and baseline drift correction technologies, the present invention solves the problem of detection instability in traditional pesticide residue detection methods caused by changes in the experimental environment (such as temperature, humidity) and instrument response deviation. The correction of environmental factors is achieved by real-time collecting data such as temperature and humidity and correcting the mass spectrometry numerical data, effectively eliminating the errors caused by environmental changes, thereby improving the accuracy and reliability of the detection results.
[0050] 2. Higher detection sensitivity:
[0051] Combined with two-dimensional liquid chromatography-high resolution mass spectrometry technology, the present invention can achieve precise separation and quantitative analysis of pesticide residues in complex food samples. At the same time, based on the automated analysis of machine learning models, more types of pesticide residues can be identified and their concentrations accurately quantified. Compared with traditional detection methods, the present invention has higher sensitivity and can detect pesticide residues in a low concentration range, adapting to more complex food matrices.
[0052] 3. Automation and intelligence:
[0053] The present invention uses a machine learning model for data analysis, which can automatically identify and classify different types of pesticide residues. Compared with traditional manual analysis methods, the machine learning model not only improves the analysis efficiency but also reduces human operation errors. The automatic learning and optimization ability of the model makes the detection process more intelligent, capable of continuously improving according to new data and adapting to new pesticide types or changing detection requirements.
[0054] 4. Comprehensive correction method:
[0055] The correction process of the present invention includes multiple aspects such as baseline drift correction, environmental factor correction, and matrix effect correction. By correcting ion intensity data, peak shape parameters, etc., it not only solves the influence caused by instrument deviation and environmental changes but also optimizes the feature extraction process through data preprocessing. This comprehensive correction method greatly improves the quality of data and ensures the accuracy of the machine learning model in analysis.
[0056] 5. Stronger result interpretability:
[0057] When identifying and classifying pesticide residue components, feature selection and importance analysis methods such as principal component analysis (PCA) and feature importance scoring are adopted, making the analysis results of the model more interpretable. This not only increases the transparency of the detection process but also provides a basis for subsequent optimization and debugging, improving the operability and credibility of the model in practical applications.
[0058] 6. Efficient quantitative analysis:
[0059] Through the quantitative analysis results output by the machine learning model, the present invention can accurately calculate the concentration of pesticide residues and compare it with food safety standards. In this process, by combining the corrected data and concentration calculation methods, the accuracy of the pesticide residue concentration can be ensured, effectively judging whether the sample meets the safety standards and reducing false positives or false negatives.
[0060] 7. Strong adaptability and easy to expand:
[0061] The method of the present invention can not only be used for conventional pesticide residue detection but also be extended to more food sample types and pesticide species according to needs. With the continuous addition of new data, the automatic update and retraining capabilities of the machine learning model endow it with strong adaptability and can continuously improve the detection ability for new types of pesticides.
[0062] In summary, the present invention has significant advantages compared with the prior art, can improve the accuracy, sensitivity and efficiency of pesticide residue detection, and at the same time solves problems such as environmental interference and data deviation existing in the existing methods through correction and optimization technologies, providing a more reliable and intelligent solution for food safety detection. Brief Description of the Drawings
[0063] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0064] Figure 1 It is a flowchart of a food pesticide residue safety detection and analysis method based on machine learning. Detailed Description of the Preferred Embodiments
[0065] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0066] This embodiment provides a method for safety detection and analysis of pesticide residues in food based on machine learning, including:
[0067] Collect a number of food samples, perform pretreatment on all food samples to obtain extracts for detection; use two-dimensional liquid chromatography-high resolution mass spectrometry technology to analyze the extracts, obtain mass spectrometry maps, and convert the map data into mass spectrometry numerical data;
[0068] Correct the mass spectrometry numerical data to obtain corrected data; establish a machine learning model based on the corrected data, and the machine learning model is used to identify and classify pesticide residue components in food samples to be tested;
[0069] Extract the pretreatment extract of the food sample to be tested, and use the same two-dimensional liquid chromatography-high resolution mass spectrometry technology for analysis to obtain the mass spectrometry map of the food sample to be tested; convert the map data of the food sample to be tested into mass spectrometry numerical data to be tested; perform the same correction on the mass spectrometry numerical data to be tested to obtain corrected data to be tested;
[0070] Input the corrected data to be tested into the established machine learning model for identification and classification of pesticide residue components; perform quantitative analysis on the types and concentrations of pesticide residues according to the results output by the model to obtain analysis results;
[0071] Compare the analysis results with food safety standards to determine whether the food samples to be tested meet the safety standards and generate corresponding test reports.
[0072] It can be understood that this embodiment provides a method for safety detection and analysis of pesticide residues in food based on machine learning, and this method can efficiently and accurately detect and analyze pesticide residues in food.
[0073] The specific steps are as follows:
[0074] First, collect a number of food samples, which can be vegetables, fruits, grains, etc. Perform pretreatment on all food samples, including steps such as cleaning, homogenization, extraction, etc., to remove impurities in the samples and extract pesticide residues, obtaining extracts for detection.
[0075] Next, use two-dimensional liquid chromatography-high resolution mass spectrometry coupling technology to analyze the above extracts. Two-dimensional liquid chromatography technology can provide more complex separation capabilities, while high resolution mass spectrometry can provide high-precision molecular weight information and structural information. Through this technology combination, a mass spectrometry map containing rich information can be obtained.
[0076] Convert the obtained mass spectrometry map data into mass spectrometry numerical data for subsequent computer processing. Perform necessary corrections on the mass spectrometry numerical data, such as correcting baseline drift, eliminating noise, etc., to improve data quality.
[0077] Establish a machine learning model based on the corrected mass spectrometry numerical data. The model can be a support vector machine, random forest, neural network, etc., for identifying and classifying pesticide residue components in food samples to be tested. The training of the model requires the use of sample data with known pesticide residue components to ensure the accuracy and reliability of the model.
[0078] Extract the pretreatment extract of the food sample to be tested, and use the same two-dimensional liquid chromatography-high resolution mass spectrometry coupling technology as when training the model for analysis to obtain the mass spectrometry map of the food sample to be tested. Convert the map data into mass spectrometry numerical data to be tested, and perform the same corrections as the training data to obtain the corrected data to be tested.
[0079] Input the corrected data to be tested into the established machine learning model for identifying and classifying pesticide residue components. The model will output the identification results, including the types and concentrations of pesticide residues.
[0080] Finally, based on the results output by the model, perform quantitative analysis on the types and concentrations of pesticide residues to obtain the analysis results. Compare the analysis results with food safety standards to determine whether the food sample to be tested meets the safety standards. If it meets the standards, generate a qualified test report; if it does not meet the standards, generate an unqualified test report and propose corresponding treatment suggestions.
[0081] Through the method provided in this embodiment, rapid detection and accurate analysis of pesticide residues in food can be achieved, providing strong technical support for food safety supervision.
[0082] In some embodiments of the present application, the mass spectrometry numerical data includes ion intensity data, peak shape parameters, ion flow curve shape characteristics, and retention time.
[0083] It can be understood that the mass spectrometry numerical data described in this embodiment includes ion intensity data, peak shape parameters, ion current curve shape characteristics, and retention time. Specifically, the ion intensity data provides information on the relative or absolute quantity of ions with different mass-to-charge ratios (m / z) at specific time points or time periods, reflecting the concentration of each component in the sample. The peak shape parameters describe the shape characteristics of each peak in the mass spectrometry diagram, such as peak height, peak width, peak area, etc. These parameters help to evaluate the accuracy and repeatability of the analysis. The ion current curve shape characteristics involve the pattern of ion current change over time and can be used to identify and distinguish compounds with similar mass-to-charge ratios but different dynamic behaviors. The retention time refers to the time required for a specific compound to travel from injection to detection in the chromatographic column and is a key parameter in chromatographic analysis for qualitative analysis of compounds. By comprehensively analyzing these mass spectrometry numerical data, a comprehensive analysis and identification of the chemical components in the sample can be achieved.
[0084] In some embodiments of the present application, the correction includes environmental factor correction and baseline drift correction;
[0085] The environmental factor correction includes temperature correction and humidity correction. The environmental correction response value is calculated based on the actual temperature and actual humidity; and then the final baseline correction value is calculated based on the environmental correction response value;
[0086] The ion intensity data and peak shape parameters are corrected using the final baseline correction value.
[0087] It can be understood that the correction process described in this embodiment includes two main steps: environmental factor correction and baseline drift correction. First, the environmental factor correction aims to adjust the possible impact of environmental temperature and humidity changes on the measurement results. Specifically, the system will continuously monitor the current temperature and humidity and calculate the corresponding environmental correction response value according to a pre-set algorithm or look-up table. This response value reflects the specific impact degree of temperature and humidity changes on the measurement data.
[0088] Next, the calculated environmental correction response value will be used to further correct the baseline drift. Baseline drift refers to the slow change of the baseline (i.e., zero point) during a long-term measurement process due to various factors (such as temperature fluctuations, instrument aging, etc.). By applying the environmental correction response value, a final baseline correction value can be obtained, which comprehensively considers the impact of environmental factors on the baseline.
[0089] Finally, the ion intensity data and peak shape parameters are corrected using this baseline-corrected final value. The ion intensity data reflects the concentration of ions in the sample, while the peak shape parameters are related to the peak shape in chromatographic or mass spectrometric analysis. These parameters are crucial for the accuracy and reliability of the analysis results. By correcting these data, it can be ensured that the analysis results more accurately reflect the true situation of the sample and reduce errors caused by environmental factors and instrument drift.
[0090] In some embodiments of the present application, the correction formula for environmental factor correction is:
[0091] Rt = Rm·(1 + kT·(Ta - Tr));
[0092] Rh = Rt·(1 + kH·(Ha - Hr));
[0093] Wherein, kT and kH respectively represent the temperature and humidity correction coefficients; Ta and Ha respectively represent the actual temperature and actual humidity; Tr and Hr respectively represent the standard temperature and standard humidity; Rt and Rh respectively represent the temperature correction value and the environmental correction response value;
[0094] The baseline drift correction formula is:
[0095]
[0096] Rbase = Rh - Bd;
[0097] Wherein, N represents the size of the moving average window; Rh represents the environmental correction response value; Rbase represents the baseline-corrected final value.
[0098] It can be understood that this embodiment aims to provide an effective method for correcting environmental factors (temperature and humidity) and baseline drift. This method includes two key steps: correction of environmental factors and correction of baseline drift. First, the influence of temperature and humidity is adjusted using the environmental factor correction formula to obtain the temperature correction value Rt and the environmental response correction value Rh. Next, by applying the baseline drift correction formula and using the moving average window to calculate the average value of a series of environmental response correction values Rh, the final value Rbase of the baseline correction is determined. By adopting this method, the interference of environmental factors and baseline drift on the measurement results can be significantly reduced, thereby improving the precision and credibility of the measurement data. By implementing these correction formulas, the corrected ion intensity data and peak shape parameters can be obtained, ensuring that the analysis results more accurately reflect the actual state of the sample and reducing errors caused by environmental factors and instrument drift.
[0099] In some embodiments of the present application, when correcting the ion intensity data using the baseline-corrected final value, it includes:
[0100] Ic(t) = Ir(t) - Rbase(t);
[0101] Wherein, Ic(t) represents the ion intensity after baseline correction; Ir(t) represents the original ion intensity data.
[0102] It can be understood that the purpose of this embodiment is to correct the original ion intensity data Ir(t) through the baseline correction final value Rbase to eliminate or reduce the influence of baseline drift on ion intensity measurement. The specific operation is to subtract the baseline correction final value Rbase from the original ion intensity data Ir(t) to obtain the ion intensity Ic(t) after baseline correction. In this way, it can be ensured that the ion intensity data more accurately reflects the actual ion concentration in the sample, thereby improving the reliability and accuracy of the analysis results. In practical applications, this correction method is crucial for improving the performance of detection techniques such as mass spectrometry and chromatography, especially when long-term monitoring is required or measurements are carried out under changing environmental conditions.
[0103] In some embodiments of the present application, when using the baseline correction final value to correct the peak shape parameters, it includes:
[0104] The peak shape parameters include peak area, peak height, and peak width;
[0105]
[0106] Hc = Hr - Rbase(tpeak);
[0107] Wc = t(Hc / 2,left) - t(Hc / 2,right);
[0108] tleft = t(Hc / 2,left);
[0109] tright = t(Hc / 2,right);
[0110] Wherein, Ac represents the peak area after baseline correction; Ic(t) represents the ion intensity after baseline correction; tpeak represents the time point when the peak height appears; Hc represents the peak height after baseline correction; Hr represents the maximum ion intensity value of the peak; tleft and tright respectively represent the left and right time points at half of the peak height after baseline correction; Wc represents the peak width after baseline correction.
[0111] It is understandable that this embodiment aims to correct the peak shape parameters through the baseline-corrected final value Rbase to ensure that the peak shape parameters can more accurately reflect the actual components in the sample. The specific operations include recalculating the peak area Ac, peak height Hc, and peak width Wc using the baseline-corrected ion intensity Ic(t). For example, the peak area Ac can be obtained by integrating the baseline-corrected ion intensity Ic(t) within a specific time range. The peak height Hc can be determined by the baseline-corrected ion intensity Ic(tpeak) at the time point tpeak where the peak height appears. And the peak width Wc can be obtained by determining the left and right time points tleft and tright at half of the peak height and calculating the time difference between these two time points. In this way, the accuracy of the peak shape parameters can be ensured, thereby improving the reliability and accuracy of the analysis results. This is particularly important for occasions that require precise analysis of sample components, such as in the fields of drug development, environmental monitoring, and food safety detection.
[0112] In some embodiments of the present application, extracting the shape characteristics of the ion current curve based on the corrected ion intensity includes calculating the curve symmetry coefficient and the curve slope ratio:
[0113]
[0114] Wherein, S represents the curve symmetry coefficient; Rslope represents the curve slope ratio; slopea and slopeb respectively represent the average slopes of the rising section and the falling section.
[0115] It is understandable that this embodiment aims to extract the shape characteristics that can reflect the sample characteristics by analyzing the ion current curve after baseline correction. The curve symmetry coefficient S can be used to describe the symmetry degree of the peak shape, and its calculation can be achieved by comparing the average slopes of the rising section and the falling section of the peak. Specifically, in the formula of the curve symmetry coefficient S, slopea and slopeb respectively represent the average slopes of the rising section and the falling section. If the value of S is close to 1, it indicates that the peak shape is symmetric; if the value of S is greater than or less than 1, it indicates that the peak shape is asymmetric, and the magnitude of the S value can reflect the degree of asymmetry.
[0116] The curve slope ratio Rslope can be used to describe the relative relationship between the rising and falling rates of the peak. When the value of the curve slope ratio Rslope is greater than 1, it indicates that the rising rate is faster than the falling rate; when the value of Rslope is less than 1, it indicates that the falling rate is faster than the rising rate. By analyzing the value of Rslope, the dynamic characteristics of the sample can be further understood.
[0117] By extracting these shape features, a more in-depth analysis of the chemical composition of the samples can be carried out, thereby providing more accurate data support in fields such as drug development, environmental monitoring, and food safety detection.
[0118] In some embodiments of the present application, the correction further includes a normalization process.
[0119] The formula for performing the normalization process is:
[0120]
[0121] Wherein, X represents the corrected data; μ and σ respectively represent the mean and standard deviation of the corresponding data; Xnor represents the data after the normalization process.
[0122] It can be understood that the purpose of this embodiment is to eliminate the difference in data magnitude between different samples through the normalization process, so that the shape features of the ion flow curves of different samples are comparable. The normalized data Xnor, by subtracting the mean μ and dividing by the standard deviation σ, can convert the data into a distribution with a mean of 0 and a standard deviation of 1, so that the curve features between different samples can be compared and analyzed on the same scale.
[0123] In actual operation, first, the ion flow curve data of each sample needs to be baseline corrected to eliminate the influence of non-sample factors such as baseline drift. Then, according to the above formula, the normalization process is carried out to obtain the normalized data Xnor. Based on these normalized data, the curve symmetry coefficient S and the curve slope ratio Rslope of each sample can be calculated, and then an in-depth analysis of the chemical composition and dynamic characteristics of the samples can be carried out.
[0124] Through this comprehensive analysis method, the chemical components in different samples can be more accurately identified and distinguished, providing more reliable data support for research and applications in related fields. For example, in drug development, the changes in drug components can be monitored more precisely; in environmental monitoring, pollutants can be detected and identified more effectively; in food safety detection, the content of additives and harmful substances in food can be evaluated more accurately.
[0125] In some embodiments of the present application, when establishing a machine learning model based on the corrected data, and the machine learning model is used to identify and classify the pesticide residue components in the food samples to be tested, feature selection is further included;
[0126] The feature selection adopts the recursive feature elimination method to screen the features that contribute the most to the model performance from the peak area, peak height, peak width, curve symmetry coefficient, and curve slope ratio.
[0127] It is understandable that this embodiment aims to optimize the feature set of the model through the Recursive Feature Elimination (RFE) method to improve the prediction accuracy and efficiency of the model. The Recursive Feature Elimination method is a feature selection technique that works by repeatedly building models and selecting or excluding the least important features. In each iteration, the model ranks the features according to their importance and then removes the least important features. This process is repeated until a predetermined number of features is reached or a certain performance metric is achieved.
[0128] In this embodiment, the purpose of feature selection is to identify the most critical features for the identification and classification of pesticide residues. Through the Recursive Feature Elimination method, the features that contribute the most to the model performance can be screened out from the original feature set, such as peak area, peak height, peak width, curve symmetry coefficient, and curve slope ratio. After these features are optimized, a more concise and efficient machine learning model can be constructed, thereby improving the identification accuracy and classification efficiency of pesticide residues.
[0129] In practical applications, such a model can be used for the rapid detection and monitoring of pesticide residues in food samples, providing guarantee for food safety. Through automated and intelligent analysis means, the detection efficiency can be significantly improved, the time and cost required for manual detection can be reduced, and at the same time, the accuracy and reliability of the detection results can be ensured. This has important practical significance and application value for food production, processing, and regulatory agencies.
[0130] In some embodiments of the present application, the pretreatment includes washing, crushing, homogenization, extraction and concentration, and filtration.
[0131] The following are the pretreatment steps for food samples, involving washing, crushing, homogenization, and extraction, aiming to provide clean, uniform, and effective extracts for subsequent chromatography-mass spectrometry analysis:
[0132] 1. Washing
[0133] Steps:
[0134] Rinse the collected food samples (such as fruits, vegetables, grains, etc.) with pure water to remove the dust, soil, and impurities on the surface.
[0135] If necessary, use a detergent or surfactant to clean the sample, especially when there are pesticide residues on the sample surface.
[0136] Replace the clear water repeatedly to ensure thorough washing and avoid introducing any other contaminants.
[0137] For some solid foods (such as nuts, meats, etc.), alcohol or other appropriate solvents can be used for cleaning.
[0138] Purpose: To remove visible surface contaminants or impurities and reduce background interference.
[0139] 2. Crushing
[0140] Steps:
[0141] Crush the washed food sample (such as using a crusher, grinder, or manual mashing).
[0142] Select an appropriate crushing equipment and method according to the physical properties of the sample (such as hardness, texture, etc.).
[0143] For hard samples (such as dried fruits, seeds, etc.), low-temperature treatment (such as cryogenic grinding) can be selected to reduce the denaturation of the sample caused by heat.
[0144] Purpose: To process the sample into small particles or powder, increase the effective contact area, and facilitate subsequent homogenization and extraction.
[0145] 3. Homogenization
[0146] Steps:
[0147] Homogenize the crushed sample, usually using a homogenizer, ultrasonic treatment, or high-speed shear mixer.
[0148] For some hard or difficult-to-homogenize samples, treatment can be carried out at low temperature to avoid component loss caused by heating of the sample.
[0149] During homogenization, an appropriate amount of solvent (such as water or buffer solution) can be added to assist in the homogenization of the sample.
[0150] Purpose: To uniformly mix different components in the sample, ensure that each part of the sample is similar in chemical composition, and facilitate extraction.
[0151] 4. Extraction
[0152] Steps:
[0153] Mix the homogenized sample with an appropriate amount of solvent (such as methanol, ethanol, water, ethyl acetate, n-hexane, etc.). Common solvent combinations include ethanol / water (70 / 30) and ethyl acetate / n-hexane, etc.
[0154] Adopt liquid-liquid extraction or solid-liquid extraction methods, and use rotary evaporation, ultrasonic extraction, or other extraction methods to transfer the target pesticide components from the sample matrix to the solvent.
[0155] If the liquid-liquid extraction method is used, the solvent and aqueous phase can be separated through a separatory funnel, and the organic phase solution containing the pesticide is collected.
[0156] During the extraction process, the selection of temperature, time, and solvent can affect the extraction effect. Therefore, these parameters need to be optimized to obtain the best extraction effect.
[0157] Objective: To extract the pesticide components from food samples from the matrix and provide effective extracts for subsequent analysis.
[0158] 5. Concentration and Filtration
[0159] Steps:
[0160] Concentrate the extract. The solvent can be removed using a rotary evaporator or by blowing dry with nitrogen until it is concentrated to an appropriate volume.
[0161] In the concentrated extract, use a filter membrane (such as a 0.45 μm filter membrane) to filter out the remaining solid substances to obtain a clear sample.
[0162] Objective: To increase the concentration of the extract and remove solid impurities that may affect the analysis.
[0163] 6. Preparation for Analysis
[0164] Steps:
[0165] Dissolve the processed extract in an appropriate solvent to ensure that its concentration is suitable for subsequent analysis.
[0166] If necessary, the sample can be further diluted or an internal standard can be added according to chromatographic requirements.
[0167] After preparing the sample, place it in a chromatographic injection vial and wait for analysis by a gas chromatography-mass spectrometry system.
[0168] Objective: To ensure that the sample has an appropriate concentration and stability during gas chromatography-mass spectrometry analysis.
[0169] Summary:
[0170] The order and specific operations of these steps can be appropriately adjusted according to the type of food sample, the types of pesticides to be detected, equipment availability, and laboratory conditions. The key objective of the pretreatment is to retain the pesticide components as much as possible, remove impurities and interfering substances that may affect the detection, and ensure the accuracy and reliability of subsequent gas chromatography-mass spectrometry analysis results.
[0171] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0172] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0173] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A food pesticide residue safety detection and analysis method based on machine learning, characterized in that, Including: Collect a number of food samples, perform pretreatment on all food samples to obtain extracts for detection; Use two-dimensional liquid chromatography-high resolution mass spectrometry technology to analyze the extract, obtain a mass spectrometry map, and convert the map data into mass spectrometry numerical data; Correct the mass spectrometry numerical data to obtain corrected data; Establish a machine learning model based on the corrected data, and the machine learning model is used to identify and classify pesticide residue components in food samples to be tested; Extract the pretreatment extract of the food sample to be tested, and use the same two-dimensional liquid chromatography-high resolution mass spectrometry technology for analysis to obtain the mass spectrometry map of the food sample to be tested; convert the map data of the food sample to be tested into mass spectrometry numerical data to be tested; Perform the same correction on the mass spectrometry numerical data to be tested to obtain corrected data to be tested; Input the corrected data to be tested into the established machine learning model for identification and classification of pesticide residue components; according to the results output by the model, perform quantitative analysis on the types and concentrations of pesticide residues to obtain analysis results; Compare the analysis results with food safety standards to determine whether the food sample to be tested meets the safety standards and generate corresponding test reports; The mass spectrometry numerical data includes ion intensity data, peak shape parameters, ion current curve shape characteristics, and retention time; The correction includes environmental factor correction and baseline drift correction; The environmental factor correction includes temperature correction and humidity correction, and the environmental correction response value is calculated according to the actual temperature and actual humidity; Then calculate the final baseline correction value according to the environmental correction response value; Use the final baseline correction value to correct the ion intensity data and peak shape parameters.
2. The food pesticide residue safety detection and analysis method according to claim 1, wherein The correction formula for the environmental factor correction is: ; ; Where kT and kH represent the temperature and humidity correction coefficients respectively; Ta and Ha represent the actual temperature and actual humidity respectively; Tr and Hr represent the standard temperature and standard humidity respectively; Rt and Rh represent the temperature correction value and the environmental correction response value respectively; The baseline drift correction formula is: ; ; Where N represents the size of the moving average window; Rh represents the environmental correction response value; Rbase represents the final baseline correction value.
3. The food pesticide residue safety detection and analysis method according to claim 2, characterized in that When using the final baseline correction value to correct the ion intensity data, it includes: ; Where Ic(t) represents the ion intensity after baseline correction; Ir(t) represents the original ion intensity data.
4. The food pesticide residue safety detection and analysis method according to claim 3, characterized in that When using the final baseline correction value to correct the peak shape parameters, it includes: The peak shape parameters include peak area, peak height, and peak width; ; ; ; ; ; Where Ac represents the peak area after baseline correction; Ic(t) represents the ion intensity after baseline correction; tpeak represents the time point when the peak height appears; Hc represents the peak height after baseline correction; Hr represents the maximum ion intensity value of the peak; tleft and tright represent the left and right time points when the peak height is half of the height after baseline correction; Wc represents the peak width after baseline correction.
5. The food pesticide residue safety detection and analysis method according to claim 4, characterized in that, Extract the ion current curve shape characteristics based on the corrected ion intensity, including the calculation of the curve symmetry coefficient and the curve slope ratio: ; ; Among them, S represents the curve symmetry coefficient; Rslope represents the curve slope ratio; slope a and slope b respectively represent the average slopes of the rising segment and the falling segment.
6. The food pesticide residue safety detection and analysis method according to claim 1, characterized in that, The correction also includes a standardization process, and the formula for performing the standardization process is as follows: ; Wherein, X represents the corrected data; μ and σ respectively represent the mean and standard deviation of the corresponding data; Xnor represents the data after the standardization process.
7. The food pesticide residue safety detection and analysis method according to claim 1, characterized in that When establishing a machine learning model based on the corrected data, and the machine learning model is used to identify and classify the pesticide residue components in the food sample to be tested, feature selection is also included; The feature selection adopts the recursive feature elimination method to screen the features that contribute the most to the model performance from the peak area, peak height, peak width, curve symmetry coefficient, and curve slope ratio.
8. The food pesticide residue safety detection and analysis method according to claim 1, characterized in that The pretreatment includes washing, crushing, homogenization, extraction and concentration, and filtration.
Citation Information
Patent Citations
Crop pesticide residue detection method and system based on convolutional neural network
CN116884512A
Automatic pesticide residue detection system for food safety
CN118243827A