Method for identifying shelf life of yam based on machine learning metabolite marker model

By using a machine learning-based metabolite biomarker model, the subjectivity and instability issues in determining the shelf life of yam have been resolved, enabling rapid and accurate determination of yam shelf life. This model is suitable for real-time monitoring during yam production and processing, improving the objectivity and reliability of the determination.

CN117288865BActive Publication Date: 2026-03-31INST OF AGRI QUALITY STANDARDS & TESTING TECH HENAN ACAD OF AGRI SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-08
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies for determining the shelf life of yams suffer from high subjectivity, instability, and are not conducive to industrial application. Traditional visual, olfactory, and tactile methods cannot accurately determine changes in nutrients, while chemical analysis methods suffer from high time costs and poor real-time performance.

Method used

A machine learning-based metabolite biomarker model was adopted. By collecting metabolite data of yam, LASSO regression was applied for feature selection and metabolite biomarker model construction to quickly and accurately determine whether the main nutrients of yam have changed. The metabolite biomarker model was constructed and analyzed using machine learning algorithms.

Benefits of technology

It enables rapid and accurate determination of yam shelf life, eliminating the subjectivity and uncertainty of traditional methods, improving the objectivity and reliability of the determination, and is suitable for real-time monitoring of shelf life changes during yam production and processing, reducing waste and losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117288865B_ABST
    Figure CN117288865B_ABST
Patent Text Reader

Abstract

The application provides a method for identifying the storage period of Chinese yam based on a metabolite marker model of machine learning, and the steps are as follows: collecting Chinese yam samples of different storage periods for crushing and drying treatment; placing the dried Chinese yam samples into extraction solvents to obtain different sample injection liquids; performing high-throughput analysis on the different sample injection liquids by using a liquid chromatograph-mass spectrometer to obtain metabolite data; performing data cleaning on the collected metabolite data, and selecting metabolite markers that have a greater influence on the storage period judgment; screening characteristic markers of the optimal storage period of Chinese yam from potential differential metabolites based on a machine learning algorithm, and constructing a metabolite marker model; calculating the model score of a Chinese yam sample to be judged according to the metabolite marker model, and judging the storage period according to the model score. The application is more objective and accurate, the influence of subjective factors is reduced, the operation is simple, the Chinese yam can be judged in real time, and a new solution is provided for the storage and consumption of Chinese yam.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of yam shelf-life identification, and more particularly to a method for identifying yam shelf-life based on a machine learning-based metabolite biomarker model. Background Technology

[0002] In the food industry, accurately determining shelf life is crucial for ensuring food safety and reducing food waste. Yam, an important edible and medicinal plant resource, is a popular food, widely used in daily life and traditional Chinese medicine due to the rich nutrients and functional compounds in its underground tubers. Yam is widely used in traditional Chinese medicine, offering numerous benefits and supporting people's health and well-being. Whether yam is still edible after a period of storage, and whether its main nutrients have changed or decreased, is currently determined using visual, olfactory, and tactile methods. However, there are no relevant standards or documents regarding whether the main nutrients have changed. Traditional visual, olfactory, and tactile detection methods alone cannot accurately determine this, and these methods suffer from high subjectivity, instability, and limitations in industrial application, leading to inconsistent and inaccurate results. No patents or journal articles have been reported domestically or internationally regarding the use of metabolite marker models to identify shelf life in yam.

[0003] Currently, visual, olfactory, and tactile identification methods can only determine whether yams are edible, whether they have exceeded their shelf life, and whether their nutritional content has decreased. Visual inspection involves observing the yam's appearance. Fresh yams typically have smooth skin without obvious rot or spots. If rot, mold, discoloration, or other abnormalities are found, the yam is no longer suitable for consumption. Olfactory inspection involves smelling the yam. Normal, fresh yams have a faint sweet aroma. If there is an off-odor or foul smell, it indicates that the yam has spoiled and should not be eaten. Tactile inspection involves touching the surface and texture of the yam. Fresh yams are firm and elastic. If the yam becomes soft or hard, it indicates that it is no longer fresh and should not be eaten. A cooking test is an option if other methods are inaccurate; the yam can be cooked before consumption. If the cooked yam has an off-odor or tastes unpleasant, it indicates that it has exceeded its shelf life. While the above methods can determine to some extent whether yams are edible, sensory perception and experience cannot determine whether yams have exceeded their shelf life, or whether their main nutrients have changed or decreased.

[0004] This invention aims to address the following shortcomings of existing technologies in determining the shelf life of yams:

[0005] 1. Unable to determine changes in nutrients

[0006] Currently, widely used identification methods such as vision, smell, and touch can only determine whether yams are edible, but cannot determine changes in the main nutrients, especially whether the main nutrients have been reduced.

[0007] 2. Determining the shelf life of yams is highly subjective.

[0008] Traditional methods for determining the shelf life of yams typically rely on manual observation or experience, which leads to a high degree of subjectivity in the results. Different people may arrive at different conclusions based on their own experience and intuition, thus affecting the accuracy and consistency of the yam's shelf life.

[0009] 3. Assessing the instability of yam shelf life

[0010] Judging whether yams have exceeded their shelf life based on appearance or experience is highly unreliable. For example, different people or different external environmental factors often lead to inconsistent judgments.

[0011] 4. Applications that are not conducive to industrialization:

[0012] In the production and preservation of yam, a great deal of stability assessment and identification is required. Relying solely on human experience is often insufficient and even less conducive to industrial application.

[0013] Therefore, existing technologies for determining the shelf life of yams have several major problems, including high subjectivity, instability, and limitations on industrial application. Traditional methods rely on manual observation or experience, leading to high subjectivity and inconsistent conclusions from different individuals. Furthermore, judging whether yams have exceeded their shelf life based on appearance or experience is highly unstable, significantly influenced by external environmental factors and individual differences. Relying on human experience in yam production and preservation is detrimental to industrial application, necessitating more stable and objective methods for assessment.

[0014] In recent years, the application of new technologies combining metabolomics and machine learning algorithms in food preservation has gradually attracted attention. Metabolomics is a science that systematically studies the metabolites of organisms. By analyzing the overall changes in metabolic products produced by an organism under specific conditions and the relationships between metabolites, a comprehensive understanding of the organism's state can be obtained. Metabolites can be biomolecules such as proteins, lipids, and carbohydrates, or their metabolic products, such as metabolic enzymes. In food science, metabolomics can be used for food quality evaluation, food safety testing, and food ingredient traceability. Machine learning, a branch of artificial intelligence, can automatically improve and optimize algorithms by learning and adapting to data, thereby achieving accurate task execution. Combining metabolomics technology and machine learning algorithms, using metabolites as potential biomarkers, can quickly and accurately predict the shelf life of food, eliminating the subjectivity of traditional judgment methods, improving the objectivity and reliability of judgments, and providing better product quality control and food safety assurance for food processing enterprises and consumers.

[0015] Currently, no patents or journal articles have been reported on identifying shelf-life using metabolite biomarker models for yam. There are reports of applying analytical techniques from different omics disciplines (such as proteomics, metabolomics, lipidomics, nutrigenomics, metagenomics, and transcriptomics) to the field of food preservation. There are also numerous reports on the application of different omics technologies in analyzing food components, food identification, and assessing food safety and quality. Clearly, applying advanced omics research techniques to food and nutrition science allows researchers to study these fields from a broader perspective. Combining different omics disciplines provides a comprehensive and holistic approach to research related to food quality, safety, and nutrition. This multidisciplinary approach opens new avenues for exploring and understanding the complexity of food components and their impact on human health.

[0016] However, the reported literature does not address the determination of the shelf life of yams, but only describes the application of metabolomics technology in the general food industry. Therefore, the criteria for determining the shelf life of yams remain an urgent problem to be solved in yam production. The metabolite biomarker model based on machine learning provided by this invention can effectively solve the problem of the lack of a method for determining the shelf life of yams. Currently, there are some methods for determining the shelf life of food; however, the most similar related reports are based on traditional sensory observation and chemical analysis methods.

[0017] Traditional sensory observation is the most common method for determining whether food has passed its expiration date. For yams, this often involves observing their appearance, including the color of the skin, and whether there are cracks or signs of rot. Then, smell and touch are used to check for any off-odors or changes in texture. However, this method relies on personal senses and experience, is easily influenced by subjective factors, and the results are not objective or accurate enough. Chemical analysis methods: Modern science and technology can determine whether food has passed its expiration date by chemically analyzing its components. For yams, this can involve measuring their moisture content and changes in metabolites. Although existing technologies can help determine the expiration date of food to some extent, they still fall short of meeting the food safety requirements of modern society due to their high subjectivity, high time cost, and poor real-time performance. Summary of the Invention

[0018] To address the technical problems of existing methods for identifying the shelf life of yams, such as high subjectivity, high time cost, and poor real-time performance, this invention proposes a method for identifying the shelf life of yams based on a machine learning-based metabolite biomarker model. By collecting metabolite data from yams and applying LASSO regression for feature selection and metabolite biomarker model construction, this method can quickly and accurately determine whether the main nutrients in yams have changed. Furthermore, this invention is more objective and accurate, reduces the influence of subjective factors, and is simpler to operate, enabling real-time judgment and providing a new solution for the storage and consumption of yams.

[0019] To achieve the above objectives, the technical solution of the present invention is as follows: a method for identifying the shelf life of yam based on a machine learning-based metabolite biomarker model, comprising the following steps:

[0020] Step 1: Sample collection and pretreatment: Collect yam samples from different storage periods and perform crushing and drying processes;

[0021] Step 2, Sample extraction: Yam dried samples from different storage periods were placed in the extraction solvent to obtain different sample injection solutions;

[0022] Step 3, Instrumental Analysis: High-throughput analysis of different sample injection solutions was performed using liquid chromatography-mass spectrometry to obtain metabolite data;

[0023] Step 4: Screening for potential differential metabolites: Clean the collected metabolite data, select metabolite biomarkers that have a significant impact on the determination of shelf life, and scale the selected metabolite biomarker data to a uniform scale.

[0024] Step 5: Machine learning model construction and feature biomarker identification: Based on machine learning algorithms, potential differential metabolites are screened to identify feature biomarkers for the optimal storage period of yam, and a metabolite biomarker model is constructed.

[0025] Step 6: Calculate the model score of the metabolite biomarker model and determine the shelf life: Calculate the model score of the yam sample to be judged based on the metabolite biomarker model, and determine the shelf life based on the model score.

[0026] Preferably, no fewer than 15 yam samples are selected for each storage period; the extraction solvent is any one of methanol-water, ethanol-water, or acetonitrile-water; the concentration of the extraction solvent is in the range of 30%-90%; and the extraction method is one of ultrasonic extraction, hot reflux extraction, or vortex extraction.

[0027] Preferably, the liquid chromatography conditions for analysis using liquid chromatography-mass spectrometry include: the chromatographic column used is a reversed-phase column; the mobile phase consists of acid-containing ultrapure water as the aqueous phase and acid-containing acetonitrile as the organic phase; and an elution gradient is used to improve the separation effect.

[0028] The liquid chromatography conditions are as follows: the mobile phase is ultrapure water containing 0.02-2% formic acid, and the organic phase is acetonitrile containing 0.02-2%; the mobile phase flow rate is 0.2-0.4 mL·min⁻¹, the column temperature is 35-45 ℃, and the injection volume is 2-4 μL.

[0029] Preferably, the mass spectrometry conditions for analysis using liquid chromatography-mass spectrometry include: the electrospray ion source temperature is set to 500~600 ℃, the ion spray voltage for positive ion mode is set to 4500~6500 V, the ion spray voltage for negative ion mode is set to -4000~-6000 V, the ion source gas I, gas II and curtain gas are set to 50, 60 and 25 psi respectively, and the collision-induced ionization parameter is set to high.

[0030] Preferably, the screening of potential differential metabolites is performed using the R software platform and includes the following steps:

[0031] ① Perform baseline filtering, peak identification, retention time correction, peak alignment, and mass spectrometry fragment structure analysis on the metabolite data acquired by mass spectrometry, and convert the spectral data into two-dimensional matrix data;

[0032] ② Perform unsupervised principal component analysis on all data to identify differences between groups and regroup them;

[0033] ③ Use supervised orthogonal partial least squares discriminant analysis to redefine the new groups and identify potential differential metabolites.

[0034] Preferably, potential differential metabolites are screened by comprehensively analyzing the importance of variable weights, VIP value, p-value, and fold change value, and selecting variables with VIP>1.0, p-value<0.05, and |log2(fold change)|>1 as important features and final potential differential metabolites.

[0035] Preferably, the screened potential differential metabolites are analyzed using LASSO regression in a machine learning algorithm. LASSO regression uses cross-validation curves and coefficient of variation path analysis to find the optimal model for variable screening.

[0036] Preferably, the metabolite biomarker model is: model score = β0 + Σ(βi * Xi);

[0037] Where β0 is the intercept of the model, βi is the coefficient of the model, Xi is the content data of metabolite biomarkers, and modelscore is the model score;

[0038] The method for determining the shelf life is as follows:

[0039] When the model score is greater than 0, the yam is within its expiration date and can be consumed.

[0040] When the model score is less than 0, the yam has expired and should not be eaten.

[0041] Preferably, the metabolite biomarker model is validated using receiver operating characteristic (ROC) curves. The ROC curves are calculated by changing the classification threshold to obtain the sensitivity and specificity at different classification thresholds. These sensitivity and specificity values ​​are plotted as ROC curves, with the horizontal axis representing 1-specificity and the vertical axis representing sensitivity. An area under the ROC curve (AUC) greater than 0.95 indicates that the metabolite biomarker model has good accuracy, sensitivity, and specificity.

[0042] Preferably, the metabolite markers obtained through screening are: hydroxyacetone, L-valine-L-phenylalanine, 1,7-diphenyl-4,6-diene-3-heptanone, pantothenic acid thioethylamine, leboadenosine, and 3'-adenine nucleotide, totaling 6 metabolite markers.

[0043] Using the six screened metabolite biomarkers as indicators, the formula for the constructed metabolite biomarker model is:

[0044] Model score = 55.4710 + hydroxyacetone * 0.1243 +

[0045] L-Valyl-L-phenylalanine*(-0.2013)+

[0046] 1,7-Diphenyl-4,6-diene-3-heptanone*(-2.2379)+

[0047] Pantothenic thioethylamine*(-0.267)+

[0048] Lipoadenosine*0.6788+

[0049] 3'-Adenine nucleotide*(-1.0823).

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] Determining changes in nutrients: This invention can not only determine whether yam is edible, but also accurately determine changes in its main nutrients.

[0052] High efficiency and speed: This invention uses a metabolite biomarker model based on machine learning, which can predict the shelf life of yams in a shorter time. This allows food processing companies to understand the freshness of yams more promptly, so as to make corresponding processing and arrangements and reduce waste and losses caused by exceeding the shelf life.

[0053] More accurate and reliable: This invention utilizes machine learning algorithms for analysis and modeling, eliminating the subjectivity and uncertainty of traditional methods and improving the objectivity and accuracy of judgment. Compared to manual visual, olfactory, and tactile examinations, the metabolite biomarker model of this invention can more accurately predict the shelf life of yams, reducing judgment errors.

[0054] Highly practical and applicable: The method of this invention is suitable for the storage of yams during the production and processing of yams, and has practical application value for food processing enterprises. It is easy to use; simply inputting the metabolite data of the yams is sufficient to obtain the judgment result, facilitating real-time monitoring of changes in shelf life during the production, processing, and sales of yams.

[0055] The metabolite biomarker model employed in this invention overcomes the shortcomings of existing technologies in determining the shelf life of yams, such as high subjectivity and instability. It provides a new, more scientific, and reliable method for determining the shelf life of yams, effectively addressing the deficiencies of existing technologies. This invention collects metabolite data from yams and applies machine learning algorithms to build and analyze a model, enabling rapid and accurate determination of whether yams have passed their shelf life. This eliminates the subjectivity and uncertainty of traditional methods, improving the objectivity and reliability of the assessment. The machine learning-based metabolite biomarker model of this invention has significant practical value, identifying whether yams are within their shelf life and whether their main nutrients have decreased, providing consumers with more reliable food safety assurance. Users simply input the yam metabolite data into the model to obtain the results, providing a clearer answer as to whether the yams are still edible. This invention accurately determines changes in the nutritional content of yams, ensuring their nutritional safety and providing technical support for people to consume safe and nutritious yams.

[0056] In summary, this invention solves the problems of high subjectivity, instability, and unfavorable industrial application of traditional methods for determining the shelf life of yams. By introducing machine learning technology and combining it with yam metabolite data, a metabolite biomarker model is constructed for accurate determination, providing a scientific, efficient, and reliable method for determining the shelf life of yams, which is expected to be promoted and applied by food processing enterprises and consumers. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a total ion flow chromatogram of samples from different storage periods according to the present invention.

[0059] Figure 2 The PCA score chart of this invention.

[0060] Figure 3 This is a schematic diagram of the screening of potential differential metabolites in this invention, where A represents VIP>1.0, B represents P value<0.05 and |log2(fold change)|>1, and C represents the intersection of the three.

[0061] Figure 4 This is a graph of the LASSO regression of the present invention, where A is the cross-validation curve and B is the LASSO coefficient path graph.

[0062] Figure 5 This is the ROC analysis curve for Example 2 of the present invention.

[0063] Figure 6 This is a box plot of the mode score values ​​of this invention. Detailed Implementation

[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example

[0065] This example presents a method for identifying the shelf life of yam based on a machine learning-based metabolite biomarker model, comprising the following steps:

[0066] Step 1: Sample Collection: Collect yam samples from different storage times to ensure that the yam samples from each storage period are representative. Select no fewer than 15 yam samples from each storage period to better reflect the characteristics and differences of yams from different storage periods and increase the reliability of the samples.

[0067] Step 2: Sample Extraction: The extraction solvent can be any one of methanol-water, ethanol-water, or acetonitrile-water; the concentration of the extraction solvent is 30%-90%. The extraction method is one of ultrasonic extraction, hot reflux extraction, or vortex extraction. Solvent extraction is performed on dried yam samples stored for different times to obtain different sample injection solutions.

[0068] Step 3, Instrumental Analysis: High-throughput analysis of sample injection solutions from different storage periods (storage time) was performed using liquid chromatography-mass spectrometry to obtain metabolite data.

[0069] The liquid chromatography conditions for analysis using liquid chromatography-mass spectrometry include: a reversed-phase column; an aqueous phase of acid-containing ultrapure water and an organic phase of acid-containing acetonitrile; and an elution gradient to improve separation efficiency.

[0070] In the liquid chromatography conditions, the mobile phase is ultrapure water containing 0.02–2% formic acid, and the organic phase is acetonitrile containing 0.02–2%. The mobile phase flow rate is 0.2–0.4 mL·min⁻¹, the column temperature is 35–45 °C, and the injection volume is 2–4 μL.

[0071] The mass spectrometry conditions for analysis using liquid chromatography-mass spectrometry (LC-MS) included: electrospray ion source temperature set to 500–600 °C, ion spray voltage for positive ion mode set to 4500–6500 V, ion spray voltage for negative ion mode set to -4000–-6000 V, ion source gas I, gas II, and curtain gas set to 50, 60, and 25 psi, respectively, and collision-induced ionization parameter set to high.

[0072] Step 4: Screening for Potential Differential Metabolites: The collected metabolite data is cleaned to remove outliers and noise; feature selection is performed: metabolite biomarkers that significantly influence storage period determination are selected; data normalization is performed: the selected metabolite biomarker data are scaled to a uniform scale; based on metabolomics technology, potential differential metabolites in yam samples from different storage periods are screened using mass spectrometry analysis.

[0073] Screening for potentially differentially expressed metabolites using omics technologies was performed on the R software platform, employing the following steps:

[0074] ① The mass spectrometry data is processed by baseline filtering, peak identification, retention time correction, peak alignment and mass spectrometry fragment structure analysis, and the spectral data is converted into two-dimensional matrix data for analysis of potential differential substances.

[0075] ② Perform unsupervised principal component analysis on all data to identify differences between groups and regroup them.

[0076] Principal Component Analysis (PCA): PCA is a commonly used dimensionality reduction algorithm used to transform high-dimensional data into low-dimensional data while retaining as much information as possible. PCA projects the original data onto a low-dimensional space spanned by the principal components by identifying the most prominent features (principal components). This reduces the dimensionality of the data, facilitating data analysis and visualization, while preserving the data's structure and differences as much as possible.

[0077] ③ Use supervised orthogonal partial least squares discriminant analysis (OPLS-DA) to redefine the new groups and identify potential differential metabolites.

[0078] OPLS-DA is an algorithm for classifying and distinguishing samples, particularly suitable for comparing multiple datasets and performing differential analysis. OPLS-DA combines the ideas of Orthogonal Partial Least Squares (OPLS) and discriminant analysis. By finding the linear combinations (latent structures) that best distinguish different sample classes, it maps the data to a new space, making similar samples as close together as possible and dispersing dissimilar samples as much as possible, thus achieving effective classification and discrimination of samples.

[0079] The VIP (Variable Importance in Projection) value is an indicator used in OPLS-DA models to assess the importance of variables (metabolites). A higher VIP value indicates a greater contribution of the variable to distinguishing between groups. The VIP value is calculated by comprehensively considering information such as the variable's weights, loadings, and variances in the model. Generally, metabolites with a VIP value greater than 1.0 are considered to have significant differences between different groups. The p-value, or p-value, is usually calculated using statistical methods (such as t-tests and ANOVA) to assess whether the difference in metabolites between two groups is significant. A smaller p-value indicates a more significant difference. The fold change value compares the folding change in metabolite abundance between two groups; the fold change is usually calculated as the ratio of the means between the two groups.

[0080] Here, the screening of potential differentially expressed metabolites is based on a comprehensive analysis of the importance ranking of variable weights, VIP value, p-value, and fold change value. Variables with VIP > 1.0, p-value < 0.05, and |log2(fold change)| > 1 are selected as important features for model screening and the final potential differentially expressed metabolites.

[0081] Step 5: Machine Learning Model Construction and Feature Biomarker Identification: Using the LASSO (Least Absolute Shrinkage and Selection Operator) algorithm of machine learning, potential differential metabolites are screened based on machine learning algorithms to identify feature biomarkers of the optimal storage period of yam, constructing a metabolite biomarker model, and further validating the model through receiver operating characteristic curves (ROC curves).

[0082] The method for screening characteristic biomarkers of yam shelf life using machine learning algorithms includes the following steps: First, the potential differentially expressed metabolites screened in step four are analyzed using LASSO regression, a linear regression algorithm used for feature selection and variable screening. LASSO regression introduces an L1 regularization term on top of ordinary linear regression, which makes certain feature coefficients zero, effectively filtering out features with little impact on the results. Ultimately, metabolites significantly correlated with differences between sample groups are identified, eliminating irrelevant features, identifying target metabolites for feature selection, and improving the model's generalization ability and interpretability. The LASSO regression algorithm uses cross-validation curves and coefficient of variation path analysis to determine the optimal model for variable screening.

[0083] The model for constructing metabolite biomarkers using characteristic biomarkers is: model score = β0 + Σ(βi * Xi).

[0084] Where β0 is the intercept of the model, βi is the coefficient of the model, and Xi is the content data of metabolite markers.

[0085] In this invention, PCA, OPLS-DA, and LASSO regression are applied to the analysis of yam metabolite data to construct a metabolite biomarker model, enabling the prediction and determination of yam shelf life. Their combined use can fully extract information from the metabolite data, improving the accuracy and stability of yam shelf life determination.

[0086] ROC curves are used to evaluate the accuracy and stability of metabolite biomarker models. By changing the classification threshold (usually ranging from 0 to 1), the sensitivity and specificity at different thresholds can be calculated. These sensitivity and specificity values ​​are then plotted on an ROC curve, with the horizontal axis representing 1-specificity and the vertical axis representing sensitivity. The area under the ROC curve is called the AUC (Area Under the ROC Curve). The AUC value ranges from 0.5 to 1; the closer the AUC is to 1, the better the model performance.

[0087] Preferably, the model established by the characteristic biomarkers screened by the LASSO regression method is evaluated based on the area under the ROC curve (AUC). If the corresponding area AUC is greater than 0.95, it indicates that the established metabolite biomarker model has good accuracy, sensitivity, and specificity.

[0088] Step Six: Calculate the model score and determine the shelf life: Calculate the model score of the yam sample to be evaluated based on the trained metabolite biomarker model, and determine the shelf life based on the calculated model score.

[0089] When the model score is greater than 0, the yam is within its expiration date and can be consumed.

[0090] When the model score is less than 0, the yam has expired and should not be eaten.

[0091] Based on the result of the shelf life determination, the corresponding prompt message is output, that is, greater than 0 or less than 0, to inform whether the yam is within the shelf life. Example

[0092] This example demonstrates a machine learning-based metabolite biomarker model for identifying whether yams have exceeded their shelf life, comprising the following steps:

[0093] Step 1: Sample collection and pretreatment: Collect at least 15 yam samples from different storage periods. After drying, crush and mix five yams at a time, and pass them through an No. 8 sieve to obtain one yam analysis sample. There should be no less than three yam analysis samples from each storage period.

[0094] Step 2, Sample extraction: Weigh 1.0 g of each yam analysis sample, add 5 mL of 70% ethanol solution, and extract by ultrasonication for 70 minutes. Centrifuge at 12000 r / min for 5 minutes to collect the supernatant, and filter through a 0.22 µm filter membrane to obtain the sample injection solution.

[0095] Step 3: Instrumental Analysis: Metabolic fingerprints of yam samples from different storage periods were acquired using chromatography-mass spectrometry (GC-MS). The HPLC conditions were as follows: C18 column (Agilent SB-C18, 1.8 μm, 2.1 mm × 100 mm); the mobile phase consisted of ultrapure water containing 0.05% formic acid and acetonitrile containing 0.05% formic acid; the elution gradient was water / acetonitrile (95:5, V / V) at 0 min, 5:95 at 9.0 min, 5:95 at 10.0 min, 95:5 at 11.1 min, and 95:5 at 14.0 min; and the mobile phase flow rate was 0.2 mL / min. -1 The column temperature was 40 ℃, and the injection volume was 2 μL. The mass spectrometry conditions were as follows: electrospray ionization source temperature was set to 550 ℃, ion spray voltage for positive ion mode was set to 4500 V, ion spray voltage for negative ion mode was set to -5500 V, ion source gas I, gas II, and curtain gas were set to 50, 60, and 25 psi, respectively, and the collision-induced ionization parameter was set to high. The total ion chromatogram is shown below. Figure 1 This demonstrates the overall distribution of all compounds detected in the sample and strong information about all ion peaks.

[0096] Step 4: Screening for potential differential metabolites: After processing the raw data obtained from mass spectrometry, including baseline filtering, peak identification, retention time correction, peak alignment, and mass spectrometry fragment structure analysis, the data is converted into two-dimensional matrix data for potential differential metabolite analysis. Figure 2 As shown, the PCA score plot divides the six samples into two large groups. It is clear that the data separation between these two new groups is good. Further screening for potential differentially expressed metabolites was conducted using VIP > 1.0, P-value < 0.05, and |log2(fold change)| > 1. (See [link to PCA score plot]). Figure 3 As shown, there are 200 metabolites that meet these three conditions, which are important characteristics and potential differential metabolites selected by OPLS-DA discrimination screening.

[0097] Step 5: Machine learning model construction and feature identification: Use LASSO regression in machine learning algorithms to screen out the most representative potential differential metabolites. Figure 4 The cross-validation curves in LASSO regression are shown, where the horizontal axis represents the value of the regularization parameter λ, and the vertical axis represents the cross-validation mean squared error (CV-MSE). In LASSO regression, the range of λ values ​​typically decreases gradually from a series of large values. The smaller the CV-MSE, the better the model fit. The figure shows that when λ is set to 0.005, the CV-MSE reaches its minimum value, and after screening, six metabolite markers are obtained: hydroxyacetone, L-valine-L-phenylalanine, 1,7-diphenyl-4,6-diene-3-heptanone, pantothenic acid thioethylamine, lipoadenosine, and 3'-adenine nucleotide. Detailed information is shown in Table 1.

[0098] Table 1. Information on 6 metabolic biomarkers

[0099] Using the six selected metabolite biomarkers as indicators, a LASSO regression model was constructed. The formula for the constructed model is:

[0100] Model score = 55.4710 + hydroxyacetone * 0.1243 +

[0101] L-Valyl-L-phenylalanine*(-0.2013)+

[0102] 1,7-Diphenyl-4,6-diene-3-heptanone*(-2.2379)+

[0103] Pantothenic thioethylamine*(-0.267)+

[0104] Lipoadenosine*0.6788+

[0105] 3'-Adenine nucleotide*(-1.0823)

[0106] The model was further evaluated using the AUC of the ROC curve, such as... Figure 5 As shown, the results show that the corresponding AUC is 1.

[0107] Step 6: Calculate the model score and determine the retention period. Figure 6 The model shows that the mode score values ​​of 20 samples are located around the y=0 axis, further demonstrating that the established metabolite biomarker model has good accuracy and specificity, as well as high stability and predictive ability.

[0108] When the model score is less than 0, the yam is within its shelf life and can be consumed.

[0109] When the model score is greater than 0, the yam has expired and should not be eaten. Example

[0110] A method based on machine learning and metabolite biomarker models to identify the shelf life of yams is used to determine whether yams have exceeded their shelf life. The method includes the following steps:

[0111] Ten batches of yam samples from unknown preservation periods were collected as the research object. After drying, five samples from each batch of yam were mixed together to form a yam analysis sample, which was then passed through an No. 8 sieve. At least three yam samples from each period were analyzed.

[0112] Weigh 1.0 g of each sample, add 5 mL of 70% ethanol solution, and extract by ultrasonication for 70 minutes. Centrifuge at 12000 r / min for 5 minutes to collect the supernatant, and filter through a 0.22 µm filter membrane to obtain the injection solution.

[0113] Metabolic fingerprints were acquired using chromatography-mass spectrometry (GC-MS). The HPLC conditions were as follows: C18 column (Agilent SB-C18, 1.8 μm, 2.1 mm × 100 mm); the mobile phase consisted of ultrapure water containing 0.05% formic acid and acetonitrile containing 0.05% formic acid; the elution gradient was water / acetonitrile (95:5, V / V) at 0 min, 5:95 at 9.0 min, 5:95 at 10.0 min, 95:5 at 11.1 min, and 95:5 at 14.0 min; and the mobile phase flow rate was 0.2 mL / min. -1 The column temperature was 40 ℃, and the injection volume was 2 μL. The mass spectrometry conditions were as follows: the electrospray ion source temperature was set to 550 ℃, the ion spray voltage for positive ion mode was set to 4500 V, the ion spray voltage for negative ion mode was set to -5500 V, the ion source gas I, gas II and curtain gas were set to 50, 60 and 25 psi, respectively, and the collision-induced ionization parameter was set to high.

[0114] Baseline filtering, peak identification, retention time correction, and peak alignment were performed on the fingerprint spectra acquired by mass spectrometry. The data were then standardized and normalized. Mass spectrometry data for hydroxyacetone, L-valine-L-phenylalanine, 1,7-diphenyl-4,6-diene-3-heptanone, pantothenic acid, lipoadenosine, and 3'-adenine nucleotide were extracted. Modelscores were calculated using a metabolite biomarker model.

[0115] Based on the calculated model score, determine the retention period and output the results:

[0116] Eight batches had a model score < 0, indicating that the yams were still within their shelf life and could be consumed.

[0117] Two batches had a model score greater than 0; the yams had exceeded their shelf life and were not suitable for consumption.

[0118] This invention employs a combination of machine learning algorithms and metabolomics technology, utilizing metabolite data for modeling and analysis to predict the shelf life of yams. This invention improves the objectivity and accuracy of the prediction, eliminating the subjectivity and uncertainty of traditional methods. The model formula is:

[0119] Model score = 55.4710 + hydroxyacetone * 0.1243 +

[0120] L-Valyl-L-phenylalanine*(-0.2013)+

[0121] 1,7-Diphenyl-4,6-diene-3-heptanone*(-2.2379)+

[0122] Pantothenic thioethylamine*(-0.267)+

[0123] Lipoadenosine*0.6788+

[0124] 3'-Adenine nucleotide*(-1.0823)

[0125] Rapid and accurate prediction of yam shelf life: The machine learning model of this invention can quickly predict the shelf life of yams by analyzing their metabolite data in a short time. Compared with traditional time-consuming methods, this invention provides an efficient and rapid method for judgment.

[0126] Improving the quality control and food safety of yam: Through the method of this invention, yam growers and processing enterprises can better control product quality, reduce waste and losses caused by product expiration, and improve the safety and reliability of food.

[0127] Practical for industrial applications: The technical method of this invention is applicable to the preservation process of yam in yam production and processing enterprises, providing enterprises with a practical and reliable judgment tool, which facilitates real-time monitoring of the freshness and shelf life of yam during storage and processing.

[0128] This invention introduces machine learning and metabolomics technologies to establish a method based on a metabolite biomarker model, thereby addressing the shortcomings of existing technologies in determining the shelf life of yams. By collecting metabolite data from yams and applying machine learning algorithms to construct a metabolite biomarker model for analysis, it is possible to quickly and accurately determine whether yams have exceeded their shelf life, eliminating the subjectivity and uncertainty of traditional methods and improving the objectivity and reliability of the judgment.

[0129] The innovative method provided by this invention allows for a more objective assessment of the safety of yam consumption, preventing waste of yam resources and food safety issues, promoting the full utilization of yam resources, and enhancing public confidence in food quality. This will provide yam producers and processing enterprises with a convenient and practical judgment standard. This invention aims to advance the technology of yam preservation and processing, improve the quality and safety of yams, and contribute to the development of the yam industry. In summary, existing technologies have shortcomings in terms of subjectivity, accuracy, and difficulty in industrial application. The metabolite model method of this invention solves these problems through machine learning algorithms, offering advantages such as ease of operation, objectivity, accuracy, and ease of industrial application, providing a more efficient and reliable judgment standard for determining the shelf life of yams.

[0130] The core innovation of this invention is the metabolite biomarker model based on machine learning. While this method has significant advantages in improving objectivity and accuracy, there may still be some alternatives, some of which may be more suitable for certain specific situations or varieties of yam. To broaden the scope of patent protection and prevent others from circumventing this technology to achieve the same inventive purpose, the following alternatives may be considered:

[0131] Metabolite Models Based on Deep Learning: In addition to machine learning algorithms, deep learning technology is also a powerful artificial intelligence method. Deep learning algorithms, such as deep neural networks, can be considered for constructing metabolite models to further improve the accuracy and stability of yam shelf-life prediction.

[0132] Multi-source data fusion models: In addition to metabolite data, data from multiple sources can be considered, such as data on the growth environment and storage conditions of yam, to build a more comprehensive predictive model. Fusion of multiple data sources allows for a more comprehensive consideration of factors affecting the shelf life of yam, improving the reliability of predictions.

[0133] Real-time monitoring based on sensor technology: During the production, processing and sales process, sensor technology can be used to monitor specific indicators of yams in real time, such as temperature and humidity. Combined with metabolite data, a real-time monitoring and prediction model can be built to promptly remind consumers whether the yams have exceeded their shelf life.

[0134] A model combining image recognition technology: It is possible to combine image recognition technology with metabolite data, and by analyzing the appearance characteristics and morphological changes of yam, together with metabolite data, to make predictions and judgments, thereby improving the accuracy and stability of predictions.

[0135] The above alternative solutions are all improvements and extensions based on the core innovation, which can further enhance the patent protection scope of this invention, ensure that innovations in different variations and applications are fully protected, and prevent others from bypassing this technology to achieve the same inventive purpose. By combining different technical solutions, different application scenarios and needs can be better met, thereby improving the value and practicality of the invention.

[0136] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying the shelf life of Chinese yam based on a machine learning-based metabolite marker model, characterized by, Comprising the following steps: Step one, sample collection and pretreatment: collect different storage period of yam samples, and carry out crushing and drying treatment; Step two, sample injection liquid extraction: put the dried yam samples of different storage periods into the extraction solvent to obtain different sample injection liquids; Step three, instrument analysis: use liquid chromatography-mass spectrometry to analyze different sample injection liquids to obtain metabolite data; Step four, screening of potential differential metabolites: clean the collected metabolite data, select metabolite markers that have greater influence on storage period, and scale the selected metabolite marker data to a unified scale; Step five, machine learning model construction and feature marker determination: screen out feature markers related to yam storage period based on machine learning algorithm, and construct metabolite marker model; The screened metabolite markers are six metabolite markers of hydroxypropanone, L-valyl-L-phenylalanine, 1,7-diphenyl-4,6-diene-3-heptanone, pantetheine, liposaccharide and 3'-adenine nucleotide; The constructed metabolite marker model formula is: Model score = 55.4710+hydroxypropanone*0.1243+ L-valyl-L-phenylalanine*(-0.2013)+ 1,7-diphenyl-4,6-diene-3-heptanone*(-2.2379)+ Pantetheine*(-0.267)+ Liposaccharide*0.6788+ 3'-adenine nucleotide*(-1.0823); Step six, calculate the model score of the metabolite marker model, and judge the storage period: calculate the model score of the yam sample to be judged according to the metabolite marker model, and judge the storage period according to the model score; The storage period judgment method is: When model score < 0, the yam is in the effective period and can be eaten; When model score > 0, the yam has exceeded the storage period and is not suitable for eating.

2. The method of claim 1, wherein the machine learning-based metabolite marker model identifies the shelf life of the Chinese yam. Not less than 15 yam samples are selected for each storage period; the extraction solvent is selected from any one of methanol-water, ethanol-water or acetonitrile-water; the concentration of the extraction solvent is in the range of 30%-90%; the extraction method is one of ultrasonic extraction, hot reflux extraction or vortex extraction.

3. The method for identifying the shelf life of yam using a machine learning-based metabolite biomarker model according to claim 1 or 2, characterized in that, The liquid chromatography conditions for analysis using liquid chromatography-mass spectrometry include: the used chromatographic column is a reversed-phase chromatographic column; in order to improve the separation effect, elution gradient is adopted; In the liquid chromatography conditions, the mobile phase water phase is selected from ultrapure water containing 0.02-2% formic acid, and the organic phase is selected from acetonitrile containing 0.02-2% formic acid; in the liquid chromatography conditions, the flow rate of the mobile phase is 0.2-0.4 mL·min-1, the column temperature is 35-45 ℃, and the injection amount is 2-4 μL.

4. The method of claim 3, wherein the machine learning-based metabolite marker model identifies the shelf life of the Chinese yam. The mass spectrometry conditions for analysis using liquid chromatography-mass spectrometry include: the temperature of the electrospray ion source is set to 500-600 ℃, the ion spray voltage is set to 4500-6500 V in positive ion mode and -4000--6000 V in negative ion mode, the ion source gas I, gas II and curtain gas are set to 50, 60 and 25 psi respectively, and the collision-induced ionization parameters are set to high.

5. The method for identifying the shelf life of Chinese yam by a metabolite marker model based on machine learning according to claim 1 or 4, characterized in that, The screening of potential differential metabolites is implemented under the R software platform, including the following steps: ① The metabolite data collected by mass spectrometry are processed by baseline filtering, peak identification, retention time correction, peak alignment and mass spectrometry fragment structure analysis, and the spectrum data are converted into two-dimensional matrix data; ② All data are subjected to unsupervised principal component analysis to find differences between groups and re-grouping; ③ The newly defined groups are subjected to supervised orthogonal partial least squares discriminant analysis to find potential differential metabolites.

6. The method of claim 5, wherein the machine learning-based metabolite marker model identifies the shelf life of the Chinese yam. The potential differential metabolites are comprehensively analyzed according to the importance of variable weight, VIP value, difference significance P value and difference fold change value, and the variables with VIP>1.0, P value<0.05 and |log2(fold change)|>1 are screened as important features and final potential differential metabolites.

7. The method for identifying the shelf life of yam using a machine learning-based metabolite biomarker model according to claim 6, characterized in that, The screened potential differential metabolites are analyzed by LASSO regression in machine learning algorithm, and the best model is analyzed by cross-validation curve and coefficient of variation path to screen variables.

8. The method for identifying the shelf life of Chinese yam by a metabolite marker model based on machine learning according to claim 7, characterized in that, The metabolic marker model is verified by the receiver operating characteristic curve, which calculates the sensitivity and specificity at different classification thresholds by changing the classification threshold, and plots the values of sensitivity and specificity as an ROC curve, with the horizontal axis being 1-specificity and the vertical axis being sensitivity; the area AUC under the ROC curve is greater than 0.95, indicating that the accuracy, sensitivity and specificity of the metabolic marker model are good.

Citation Information

Patent Citations

  • Mung bean production place tracing method based on characteristic difference metabolites

    CN116381085A