Method for predicting optimal harvest time of yam based on machine learning-based marker metabolite model

By constructing a yam harvesting period prediction model using metabolomics and machine learning, the problem of relying on experience to determine the yam harvesting period has been solved, enabling scientific and accurate determination of the harvesting period, improving yam yield and quality, and promoting the development of the yam industry.

CN117169388BActive Publication Date: 2026-04-07INST OF AGRI QUALITY STANDARDS & TESTING TECH HENAN ACAD OF AGRI SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for determining the harvest period of yam rely on subjective experience, lack scientific rigor and accuracy, and are greatly affected by the external environment, resulting in unstable yam yield and quality.

Method used

Metabolomics data of yam samples were collected using metabolomics technology. A marker metabolite model was constructed by combining it with machine learning algorithms. Characteristic metabolites were screened using the LASSO regression method to establish a predictive model. The accuracy of the model was verified by ROC curves, providing a scientific basis for determining the harvest period.

Benefits of technology

It eliminates subjectivity and reliance on experience, improves the accuracy and stability of yam harvesting time, ensures yam yield and quality, and promotes the sustainable development of the yam industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117169388B_ABST
    Figure CN117169388B_ABST
Patent Text Reader

Abstract

The present application provides a kind of based on machine learning's mark metabolite model prediction method of optimal harvest period of Chinese yam, steps are as follows: collecting Chinese yam samples of different harvest periods, obtain metabolomics data by analyzing Chinese yam samples through metabolomics technology;Metabolomics data are preprocessed;The feature related to the growth period of Chinese yam is obtained by using machine learning algorithm to select potential marker metabolite;LASSO regression method is used to screen potential marker metabolite to construct marker metabolite prediction model;The area under ROC curve is used to verify the constructed marker metabolite prediction model;The metabolomics data of new Chinese yam are input into the marker metabolite prediction model to obtain model score, and whether Chinese yam is suitable for harvesting is judged according to model score.The present application can accurately predict the optimal harvest period of Chinese yam, eliminate subjectivity and experience dependence, improve scientificity, reduce external environmental influence, realize Chinese yam production capacity maximization, and provide reliable technical support for agricultural production.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of yam harvest period prediction, and particularly relates to a method for predicting the optimal harvest period of yam based on a machine learning-based marker metabolite model. BACKGROUND

[0002] Yam is a rhizome plant of the Dioscoreaceae family and Dioscorea genus, and its rhizome is the edible part. It is not only a food but also has medicinal and health care effects, and has been listed as a Chinese medicinal and edible plant resource. Its unique nutritional composition and rich bioactive substances make it have wide application value in traditional Chinese medicine and modern health fields. Yam is rich in amino acids, organic acids, sugars and other nutrients, and also contains rich secondary metabolites such as saponins, dioscorin, flavonoids and alkaloids, which have biological activities. Yam is flat and sweet in taste, and has the effects of enhancing immunity, invigorating the spleen and stopping diarrhea, tonifying the lung and kidney, etc. The digging of yam usually starts at the end of September, around the Mid-Autumn Festival and National Day, and lasts until early December, lasting for three months. Harvesting is an important link in the production process of yam, which directly affects its quality and yield. However, there is no research on the correlation between different harvest periods and the changes of yam chemical components, and there is no report on the identification of related maturity markers. In actual production, whether yam is mature is still judged by traditional experience. As a medicinal and edible plant resource, the optimal harvest period can ensure the quality and efficacy of yam as a medicinal material, and the content of functional nutrients is high when it is used as food, and the taste is good, which can improve the market competitiveness of yam industry and promote the sustainable and healthy development of yam industry.

[0003] The existing schemes for yam harvest period are usually based on traditional experience and external observation. The following are some common existing schemes:

[0004] Growth cycle observation method: through regular observation of yam plants in the planting area, mainly the growth state, leaf color, stem state, etc., combined with the experience of producers, to judge the growth cycle and collection period of yam. This method is simple and intuitive, but the accuracy is affected by the personal experience and subjective judgment of the producer, and there may be some errors.

[0005] Subterranean organ observation method: by digging part of yam plants, observing the size, shape and color of underground organs (tubers), to judge the growth stage and collection period of yam. This method is more direct in reflecting the growth state of yam than the growth cycle observation method, but is also affected by the personal experience and subjective judgment of the producer.

[0006] Growth model prediction method: based on historical planting data and meteorological data of yam, a mathematical model is established to predict the growth cycle and collection period of yam. This method is relatively scientific, but the model needs a large amount of data to support, and is greatly affected by external factors such as climate.

[0007] It should be noted that the above existing solutions have certain limitations, mainly in accuracy and dependence on experience. Therefore, establishing a more scientific and accurate prediction method has an important role in promoting the development of yam planting and improving yield and quality.

[0008] The existing technology mainly has the following shortcomings:

[0009] Subjectivity and experience dependence: Traditional yam harvest period determination methods often rely on the planting and harvesting experience of farmers, or on market prices, lacking scientificity, leading to a lack of scientific criteria for yam harvest period determination, with large price differences between different years, and individual differences between different farmers, resulting in different judgments based on personal experience.

[0010] Data deficiency: Traditional harvest period determination methods usually rely on limited observation data, ignoring other potential important factors, limiting the reliability and accuracy of the prediction model, especially in the face of complex natural environment and climate change.

[0011] Low accuracy and low prediction precision: The traditional method of determining the harvest period of yam often has large errors, leading to unreasonable harvest time and affecting the yield and quality of yam.

[0012] Waste of time and resources: Due to the need for long-term observation and data accumulation and a large number of trial and error processes in traditional methods, time and resources are wasted.

[0013] Affected by external environment: Traditional methods do not consider external environmental factors (such as climate, temperature, etc.), but these factors have important influence on the growth and development of yam, limiting the accuracy of the judgment. Traditional visual inspection method is difficult to accurately determine the best harvest period of yam, resulting in a part of yam not reaching the best harvest state at the time of harvest, reducing the quality and yield of yam.

[0014] The chemical components contained in yam of different harvest periods are different, which further affects its quality and medicinal value. The maturity of yam directly affects the accumulation of metabolites, quality and medicinal value, and the appropriate harvest period has important influence on the accumulation of yam quality and nutrients. Determining the best harvest period of yam helps to improve the quality of yam and ensure the nutritional and medicinal value of yam.

[0015] Metabolomics and machine learning are two independent but combined technical fields, which have important application value in predicting the best harvest period of yam.

[0016] Metabolomics reflects the metabolic characteristics of organisms at different growth stages by analyzing the overall changes in metabolites under different growth conditions. In predicting the harvest period of yam, metabolomics can be used to obtain the metabolic profiles of yam at different growth stages, and to identify specific metabolites related to the growth period. By collecting yam samples at different growth stages, extracting and analyzing metabolites, full-spectrum metabolite data can be obtained. By comparing the metabolite composition at different growth stages, some characteristic metabolites closely related to the growth stage of yam can be found. These characteristic metabolites can be used as indicators to reflect the growth status of yam and provide a basis for predicting the optimal harvest period of yam.

[0017] Machine learning is a branch of artificial intelligence that enables computers to learn and adapt to data, thereby achieving prediction or decision-making for specific tasks. In predicting the harvest period of yam, machine learning can help build a prediction model to correlate metabolomics data with the growth period of yam. Using algorithms such as PCA, OPLS-DA, LASSO, etc., feature selection and pattern recognition can be performed on the omics data. By establishing a marker metabolite model, the selected metabolite features are correlated with the growth period of yam, thereby predicting the optimal harvest period of yam.

[0018] By integrating metabolomics and machine learning methods, metabolite data of yam samples can be collected, and then a prediction model can be built using machine learning algorithms to identify characteristic metabolite features related to the growth period of yam, thereby achieving accurate prediction of the harvest period of yam. SUMMARY

[0019] To address the technical problems of subjectivity and experience dependence, low accuracy, and greater influence from external environment in existing methods for predicting the harvest period of yam, the present invention proposes a method for predicting the optimal harvest period of yam based on a marker metabolite model using machine learning. By using large-scale full-spectrum metabolite data and advanced machine learning algorithms, a marker metabolite model based on machine learning is introduced, which can better explore the change patterns during the growth of yam, provide basic data for yam producers to scientifically determine the harvest time, and provide a more accurate, scientific, and efficient method for determining the harvest period of yam, increasing yam yield while improving the quality of yam, and providing a basis for the scientific and reasonable harvesting of yam.

[0020] To achieve the above-mentioned purpose, the technical solution of the present invention is as follows: a method for predicting the optimal harvest period of yam based on a marker metabolite model using machine learning, comprising the following steps:

[0021] Step 1: Data collection: Collect yam samples at different harvest periods, analyze the yam samples using metabolomics technology, and obtain metabolomics data of yam at different harvest periods;

[0022] Step two, data preprocessing: preprocessing the collected metabolomics data;

[0023] Step three, selection of potential marker metabolites: using machine learning algorithms to select features related to the growth period of yam from the preprocessed metabolomics data to obtain potential marker metabolites;

[0024] Step four, construction and verification of prediction model: using LASSO regression method to screen potential marker metabolites to construct marker metabolite prediction model; based on the area under the ROC curve as the evaluation index, the constructed marker metabolite prediction model is verified;

[0025] Step five, input the new yam metabolomics data into the verified marker metabolite prediction model to obtain the model score, and determine whether the yam is suitable for harvesting according to the model score.

[0026] Preferably, the yam sample of each harvesting period is selected not less than 15;

[0027] The metabolomics technology adds an extraction solvent to the yam sample after freeze-drying, ultrasonic extraction, centrifugal supernatant, and filters the membrane to obtain the sample solution;

[0028] The metabolomics technology collects metabolite fingerprint by different harvesting period yam sample solution.

[0029] Preferably, the chromatography-mass spectrometry instrument is used to collect the metabolite fingerprint, and the chromatography conditions of the chromatography-mass spectrometry instrument are: the flow rate of the mobile phase is 0.2-0.4 mL·min -1 , the column temperature is 35-45℃, and the sample injection amount is 2-4 μL; the mass spectrometry conditions of the chromatography-mass spectrometry instrument are: the temperature of the electrospray ion source is set to 500-600℃, the ion spray voltage is set to 4500-6500V in positive ion mode, the ion spray voltage is set to -4000--6000V in negative ion mode, the ion source gas I, gas II and curtain gas are set to 50, 60, 25 psi respectively, and the collision-induced ionization parameter is set to high.

[0030] Preferably, the preprocessing includes data cleaning, removing outliers and data normalization.

[0031] Preferably, the selection of potential marker metabolites is implemented on the R software platform, including the following steps:

[0032] ①Baseline filtering, peak identification, retention time correction, peak alignment and mass spectrum fragment structure analysis are performed on the data collected by mass spectrometry, and the spectrum data is converted into two-dimensional matrix data;

[0033] ② Perform unsupervised principal component analysis on all two-dimensional matrix data to identify differences between groups and regroup them;

[0034] ③ Use supervised orthogonal partial least squares discriminant analysis to redefine the new group and identify potential marker metabolites.

[0035] Preferably, the potential biomarker metabolites are selected according to importance ranking (VIP value), significance level (P value), and fold change value; variables with VIP > 1.0, P value < 0.05, and |log2(fold change)| > 1 are selected as potential biomarker metabolites.

[0036] Preferably, the marker metabolite prediction model uses LASSO regression analysis to identify potential marker metabolites. LASSO regression minimizes the mean square error of the cross-validation curve and eliminates irrelevant potential marker metabolites through cross-validation curve and coefficient of variation path analysis. The optimal penalty and penalty coefficient values ​​are determined by the lowest point of the cross-validation curve. Metabolites with non-zero coefficients are identified as marker metabolites through the coefficient path corresponding to the optimal values. The marker metabolite data are then fitted to construct the marker metabolite prediction model.

[0037] Preferably, the area under the ROC curve is greater than 0.95, indicating that the marker metabolite prediction model has good accuracy, sensitivity, and specificity.

[0038] Preferably, the marker metabolites screened are allantoin, 5-oxo-L-proline, 4-hydroxymandelinniol, L-methionine, 6,7-dihydroxy-2,4-dimethoxyphenanthrene, N-feruloylguanidine, and glucosylsyringic acid.

[0039] Preferably, using the seven screened biomarker metabolites as indicators, the biomarker metabolite prediction model constructed by the LASSO regression method is as follows:

[0040] Model score = -117.2510 + allantoin * (-2.7340) +

[0041] 5-Oxo-L-proline*(-0.8350)+

[0042] 4-Hydroxymandenitrile*5.5810+

[0043] L-methionine*1.7373+

[0044] 6,7-Dihydroxy-2,4-Dimethoxyphenanthrene*0.4416+

[0045] N-Feruloylguanidine*1.9960+

[0046] Glucosyl eugenol* (-0.7759);

[0047] Wherein, model score is the model score;

[0048] A threshold is used to determine the model score. When the model score is greater than zero, the yam sample is suitable for harvesting; when the model score is less than zero, the yam sample is not suitable for harvesting.

[0049] Compared with existing technologies, the beneficial effects of this invention are: it solves the problems of subjectivity, insufficient data, and low prediction accuracy in traditional yam harvesting period determination methods, providing yam growers with a more scientific, accurate, and reliable determination method, and offering the yam industry an advanced and efficient solution for harvesting period determination, thus promoting the sustainable development and modernization of the yam industry; this invention can provide yam growers with a scientific and reasonable harvesting period, increasing yam yield while ensuring its high nutritional quality, increasing the economic benefits of yam growers, improving the market competitiveness of the yam industry, and promoting the sustainable and healthy development of the yam industry. It has the following advantages:

[0050] Eliminating subjectivity and reliance on experience: Machine learning algorithms are used to establish predictive models based on metabolite characteristics. Machine learning algorithms can learn and summarize patterns from a large amount of data, eliminating the subjectivity and reliance on experience in traditional judgment methods. Yam producers no longer rely on personal experience, but obtain more scientific and accurate harvesting periods with the support of scientific data and algorithms.

[0051] Improved accuracy: By adopting metabolomics analysis technology, we can comprehensively obtain information on the metabolite composition of yam at different growth stages, avoiding the problem of lack of scientific data support in traditional methods; machine learning algorithms, combined with metabolite feature selection and pattern recognition, can more accurately predict the optimal harvesting period of yam, thereby optimizing the harvesting time and improving the yield and quality of yam.

[0052] The influence of the external environment is relatively small: When establishing the prediction model, this invention increases the stability of the prediction by reasonably screening features and eliminating factors related to fluctuations in the external environment, thus ensuring the accuracy of the prediction. Yam producers can harvest yams more reliably based on the prediction results, thereby maximizing the yield and quality of yams.

[0053] Maximizing production capacity: By fully utilizing the advantages of machine learning algorithms based on metabolite characteristics, we can ensure that yams are harvested at the optimal harvest time, thereby increasing yam yield and improving yam quality, thus enhancing the economic benefits of the crop. In addition, accurate prediction of the optimal harvest time for yams can also help to rationally plan the planting cycle, formulate production plans, maximize resource utilization, and improve the overall economic benefits of yam cultivation.

[0054] In summary, this invention provides a scientific and accurate prediction method that eliminates subjectivity and reliance on experience, improves scientific rigor, and reduces the impact of the external environment, thereby maximizing yam production capacity and providing reliable technical support for agricultural production. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 This is the total ion chromatogram of the yam sample collected according to the present invention.

[0057] Figure 2 This is the PCA score chart of the present invention.

[0058] Figure 3 This is the OPLS-DA score graph of the present invention.

[0059] Figure 4 This diagram illustrates the screening process for potential differential metabolites in this invention, where A represents VIP > 1.0, B represents P value < 0.05 and |log2(fold change)| > 1, and C represents the intersection of the three.

[0060] Figure 5 This is a graph of the LASSO regression of the present invention, where A is the cross-validation curve and B is the LASSO coefficient path graph.

[0061] Figure 6 The results are the ROC curve analysis results of this invention.

[0062] Figure 7 This is a score box plot of the prediction model of this invention. Detailed Implementation

[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] Example 1

[0065] like Figure 1 As shown, a method for predicting the optimal harvesting period of yam based on a machine learning-based marker metabolite model includes the following steps:

[0066] Step 1: Data Collection: First, yam samples were collected at different harvest stages, and then analyzed using metabolomics technology to obtain metabolomics data of yam at different harvest stages. At least 10 samples were selected from each harvest stage.

[0067] The metabolomics data were obtained by collecting metabolomics fingerprints from yam samples at different harvest times. The chromatographic conditions were: mobile phase flow rate of 0.2–0.4 mL / min. -1 The column temperature was 35–45 °C, and the injection volume was 2–4 μL. The mass spectrometry conditions were as follows: the electrospray ion source temperature was set to 500–600 °C, the ion spray voltage for positive ion mode was set to 4500–6500 V, the ion spray voltage for negative ion mode was set to -4000–-6000 V, the ion source gas I, gas II, and curtain gas were set to 50, 60, and 25 psi, respectively, and the collision-induced ionization parameter was set to high.

[0068] Step 2: Data Preprocessing: The collected metabolomics data are preprocessed, including data cleaning, outlier removal, and normalization, to ensure data quality and consistency.

[0069] Step 3: Feature selection of metabolite biomarkers: Using metabolomics and machine learning algorithms, features related to the growth stage of yam are selected from the preprocessed metabolomics data to obtain potential differential (biomarker) metabolites. These features will serve as input to the predictive model to build the biomarker metabolite model.

[0070] The selection of characteristics related to the growth stage of yam was carried out using the R software platform, following these steps:

[0071] ① The mass spectrometry data is processed by baseline filtering, peak identification, retention time correction, peak alignment and mass spectrometry fragment structure analysis, and the spectral data is converted into two-dimensional matrix data for analysis of potential differential substances;

[0072] ② Perform unsupervised principal component analysis (PCA) on all data to identify differences between groups and regroup them;

[0073] ③ Use supervised orthogonal partial least squares discriminant analysis (OPLS-DA) to analyze the redefined new groups and identify potential marker metabolites.

[0074] PCA is a dimensionality reduction algorithm used to map high-dimensional data to a low-dimensional space, thereby reducing the number of features. OPLS-DA is a multivariate statistical analysis method used to find discriminative metabolite features. Features related to the growth stage of yam are selected according to the variable importance in projection (VIP) value, the p-value of the difference, and the fold change value. Variables with VIP > 1.0, p-value < 0.05, and |log2(fold change)| > 1 are selected as important features and final potential biomarker metabolites for the model.

[0075] In addition to fold change and P-value, VIP value is also a very important indicator for identifying differentially differentiated metabolites. VIP is the variable weight value of the variables in the OPLS-DA model, which can be used to measure the strength and explanatory power of the accumulation difference of each metabolite on the classification of each group of samples. VIP ≥ 1 is a common screening criterion for differentially differentiated metabolites.

[0076] Step 4: Predictive Model Construction: After feature selection, the Least Absolute Shrinkage and Selection Operator (LASSO) regression algorithm is used to construct a marker metabolite prediction model. This algorithm will learn the correlation between the optimal harvesting period of yam and metabolite features based on the input metabolite features.

[0077] The machine learning algorithm establishes a characteristic biomarker model for the optimal harvest time of yam by screening potential biomarker metabolites through LASSO regression analysis. LASSO regression eliminates irrelevant features and screens target metabolites through cross-validation curves and coefficient of variation path analysis, constructing a biomarker metabolite prediction model to predict the optimal harvest time of yam.

[0078] LASSO is a linear regression method that can be used for feature selection, reducing irrelevant feature coefficients to zero. The application of machine learning algorithms provides the metabolite model of this invention with efficient and accurate feature selection and pattern recognition capabilities.

[0079] Step 5: Predictive Model Validation and Optimization: To validate and optimize the established biomarker metabolite prediction model, the area under the receiver operating characteristic (ROC) curve (AUC) is used as an evaluation metric to validate the model. ROC AUC is a commonly used metric for evaluating classification models, measuring their performance and discriminative ability.

[0080] The ROC curve is a tool used to evaluate the performance of classification models. It plots the true positive rate (Sensitivity) on the vertical axis and the false positive rate (1-Specificity) on the horizontal axis, showing the model's performance at different thresholds. AUC, the area under the ROC curve, measures the overall performance of the classification model. The closer the AUC value is to 1, the better the model's performance. In this invention, AUC is used as an evaluation metric to comprehensively assess and optimize the model's classification accuracy and robustness.

[0081] The predictive model of biomarker metabolites established by the characteristic biomarkers screened by the LASSO regression method was evaluated based on the area under the ROC curve (AUC). The corresponding AUC was greater than 0.95, indicating that the established model had good accuracy, sensitivity and specificity.

[0082] Step Six: Model Application: After establishing and validating the predictive model, it can be applied to actual yam production. Input the new yam metabolomics data into the valid predictive model to obtain a model score. Based on the model score, determine whether the yam is suitable for harvesting.

[0083] When new metabolomics data is input, the model outputs a numerical value representing the correlation between a yam sample and its optimal harvest time, known as the "model score." A threshold is applied to the "model score": a score greater than zero indicates the yam sample is suitable for harvesting; a score less than zero indicates it is not suitable. Farmers can use the "model score" to scientifically determine the optimal harvest time, thereby increasing both yam yield and quality.

[0084] Example 2

[0085] A method for predicting the optimal harvest time of yam using a marker metabolite model based on machine learning was proposed. Yam samples were collected from six different harvest periods: late September (S1), early October (S2), late October (S3), early November (S4), late November (S5), and early December (S6). Fifteen samples were collected from each period. After freeze-drying, five samples were mixed to form one yam analysis sample, which was then passed through a 150-mesh sieve. Three yam analysis samples were collected from each period. Quality control samples were prepared by mixing equal volumes of 5g of yam analysis samples from each of the six periods. Three 1.0g portions of each quality control sample were weighed, and 5mL of 70% ethanol solution was added for ultrasonic extraction for 70 minutes. The samples were then centrifuged at 12000r / min for 5 minutes, and the supernatant was collected and filtered through a 0.22μm filter membrane to obtain the injection solution.

[0086] Metabolic fingerprints of yam samples from different harvest stages were collected using chromatography-mass spectrometry (GC-MS). The HPLC conditions were as follows: C18 column (Agilent SB-C18, 1.8 μm, 2.1 mm × 100 mm); mobile phase consisted of ultrapure water containing 0.05% formic acid as the aqueous phase and acetonitrile containing 0.05% formic acid as the organic phase; the elution gradient was water / acetonitrile (95:5, V / V) at 0 min, 5% aqueous phase at 9.0 min, 5% aqueous phase at 10.0 min, 95% aqueous phase at 11.1 min, and 95% aqueous phase at 14.0 min; the mobile phase flow rate was 0.2 mL / min. -1 The column temperature was 40℃, and the injection volume was 2μL. The mass spectrometry conditions were as follows: electrospray ionization source temperature was set to 550℃, ion spray voltage for positive ion mode was set to 4500V, and for negative ion mode was set to -5500V. Ion source gases I, II, and curtain gas were set to 50, 60, and 25 psi, respectively. The collision-induced ionization parameter was set to high. The total ion chromatogram is shown below. Figure 1 As shown.

[0087] After baseline filtering, peak identification, retention time correction, peak alignment, and mass spectrometry fragment structure analysis, the fingerprint spectra acquired by mass spectrometry were converted into two-dimensional matrix data for potential differential analysis. The PCA score chart is shown below. Figure 2 As can be seen, the six groups of samples are clearly separated into two major categories, namely S1-S3 and S4-S6. Meanwhile, the clustering of the QC samples is good, indicating that the instrument is stable and the method has good repeatability.

[0088] Figure 3 The OPLS-DA score plot clearly shows good separation between the two new groups. Further screening for potential differentially expressed metabolites was conducted using criteria of VIP > 1.0, P < 0.05, and |log2(fold change)| > 1. (See attached image.) Figure 4 As shown, a total of 41 metabolites meet these three conditions, which are the important features and final potential differential metabolites screened by OPLS-DA. Preferably, the selection of potential differential metabolites is based on a comprehensive selection of variable weights, including the importance value (VIP), the significance value (P-value), and the fold change value. More preferably, variables with VIP > 1.0, P-value < 0.05, and |log2(fold change)| > 1 are selected as the important features and final potential differential metabolites screened by the model.

[0089] LASSO regression, a machine learning algorithm, was used to screen 41 differentially expressed metabolites for feature metabolite identification. A smaller mean squared error of the cross-validation curve in the LASSO regression indicates a better LASSO fit. The penalty value λ was determined based on the lowest point of the cross-validation curve. The penalty coefficient compressed the coefficients of variables with insignificant impact on the prediction results to 0. Figure 5 The data shows that when the penalty value λ is 0.005, the mean squared error reaches its minimum, identifying seven metabolites: allantoin, 5-oxo-L-proline, 4-hydroxymandelinniol, L-methionine, 6,7-dihydroxy-2,4-dimethoxyphenanthrene, N-feruloylguanidine, and glucosylsyringic acid. Using these seven metabolites as indicators, a LASSO regression model was constructed. The resulting model formula is as follows:

[0090] Model score = -117.2510 + allantoin * (-2.7340) +

[0091] 5-Oxo-L-proline*(-0.8350)+

[0092] 4-Hydroxymandenitrile*5.5810+

[0093] L-methionine*1.7373+

[0094] 6,7-Dihydroxy-2,4-Dimethoxyphenanthrene*0.4416+

[0095] N-Feruloylguanidine*1.9960+

[0096] Glucosyl eugenol* (-0.7759)

[0097] Based on the calculated model score, the shelf life is determined:

[0098] When the model score is less than 0, harvesting is not suitable.

[0099] When the model score is greater than 0, it is suitable for harvesting.

[0100] Model Evaluation: Ten samples each of yams from unsuitable harvest periods (september-grown yams) and suitable harvest periods (December-grown yams) were collected. Mass spectrometry analysis and data standardization were performed using the method described above. Peak values ​​of six metabolites were extracted and substituted into the prediction model's model score formula to calculate the model score for each group. The model was evaluated using ROC AUC, and the results showed a corresponding ROC of 1. Figure 6 The Model score values ​​of yams unsuitable for harvesting are all less than 0, while the Model score values ​​of yams suitable for harvesting are all greater than 0.Figure 7 The results showed that the model had an accuracy of 100%, good accuracy and specificity, and high stability and predictive ability. The model was able to effectively identify whether yam met the harvesting conditions.

[0101] Example 3

[0102] A method for predicting the optimal harvesting time of yam based on a machine learning-based marker metabolite model was proposed. Ten batches of yam samples with unknown growth times were collected as the research object. After freeze-drying, five samples were mixed to form one yam analysis sample, with each sample weighing 1.0g. 5mL of 70% ethanol solution was added, and the sample was extracted by ultrasonic extraction for 70 minutes. The sample was then centrifuged at 12000r / min for 5 minutes, and the supernatant was collected and filtered through a 0.22μm filter membrane to obtain the injection solution.

[0103] Metabolic fingerprints were acquired using chromatography-mass spectrometry (GC-MS). The HPLC conditions were as follows: C18 column (Agilent SB-C18, 1.8 μm, 2.1 mm × 100 mm); mobile phase consisted of ultrapure water containing 0.05% formic acid as the aqueous phase and acetonitrile containing 0.05% formic acid as the organic phase; the elution gradient was water / acetonitrile (95:5, V / V) at 0 min, 5% aqueous phase at 9.0 min, 5% aqueous phase at 10.0 min, 95% aqueous phase at 11.1 min, and 95% aqueous phase at 14.0 min; the mobile phase flow rate was 0.2 mL / min. -1 The column temperature was 40℃, and the injection volume was 2μL. The mass spectrometry conditions were as follows: the electrospray ion source temperature was set to 550℃, the ion spray voltage for positive ion mode was set to 4500V, the ion spray voltage for negative ion mode was set to -5500V, the ion source gas I, gas II and curtain gas were set to 50, 60 and 25psi respectively, and the collision-induced ionization parameter was set to high.

[0104] Baseline filtering, peak identification, retention time correction, and peak alignment were performed on the fingerprint spectra acquired by mass spectrometry. The data were then standardized and normalized. Metabolite mass spectrometry data for allantoin, 5-oxo-L-proline, 4-hydroxymandelinniol, L-methionine, 6,7-dihydroxy-2,4-dimethoxyphenanthrene, N-feruloylguanidine, and glucosylsyringic acid were extracted and calculated using the Modelscore formula. The Modelscore formula is:

[0105] Model score = -117.2510 + allantoin * (-2.7340) +

[0106] 5-Oxo-L-proline*(-0.8350)+

[0107] 4-Hydroxymandenitrile*5.5810+

[0108] L-methionine*1.7373+

[0109] 6,7-Dihydroxy-2,4-Dimethoxyphenanthrene*0.4416+

[0110] N-Feruloylguanidine*1.9960+

[0111] Glucosyl eugenol* (-0.7759)

[0112] Based on the calculated model score, the shelf life is determined and the result is output:

[0113] There are 6 batches with a model score < 0, indicating that the yams in 6 plots are not yet suitable for harvesting during their growth period.

[0114] There are 4 batches with a model score > 0, indicating that the yams in the 4 plots are ready for harvest during their growth period.

[0115] This invention relates to the establishment of a marker metabolite model using various machine learning algorithms, such as PCA, OPLS-DA, and LASSO, to predict the optimal harvest time of yam. By leveraging the advantages of machine learning algorithms, subjectivity and reliance on experience are eliminated, improving accuracy and stability. The machine learning model of this invention can rapidly predict the harvest time of yam by analyzing metabolite data in a short time. Compared to traditional time-consuming methods, this invention provides an efficient and rapid method for judgment. This invention is applicable to the harvesting of yam in large-scale production, providing a practical and reliable tool for determining the harvest time. Yam is an important economic crop; accurate prediction and rational management of its optimal harvest time can improve yield and quality while reducing production costs. Therefore, the technology of this invention has practical application value for agricultural production.

[0116] Metabolomics is a technique for studying the overall composition and changes of all metabolites in an organism. By analyzing metabolites in samples, a large amount of metabolic profile data can be obtained, thereby understanding the state and changes of the metabolic network in the organism. Machine learning is a type of artificial intelligence algorithm that can automatically build models and make predictions by learning and summarizing large amounts of data. In this invention, metabolomics technology provides data on the composition of metabolites during the growth of yam, while machine learning algorithms use this data to build predictive models, achieving accurate prediction of the optimal harvest time for yam.

[0117] Metabolite feature selection methods: In addition to the LASSO regression method used in this invention, other feature selection algorithms can be considered, such as random forests and recursive feature elimination (RFE). Each method has its advantages and disadvantages, and the appropriate method can be selected based on the actual data.

[0118] This invention employs multiple machine learning algorithms, including PCA, OPLS-DA, and LASSO, to construct metabolite models. Besides these, many other machine learning algorithms can be used for model building, such as decision trees, support vector machines, and neural networks. Different algorithms may have varying effects on data processing and pattern recognition; therefore, it is worthwhile to try different algorithms to obtain better predictive performance.

[0119] Besides AUC, which is used in this invention as a model performance evaluation metric, other metrics can be used for model validation and optimization, such as accuracy, recall, and F1-score. Different metrics are suitable for evaluating different problems, and the appropriate metric can be selected according to the specific circumstances.

[0120] The metabolite data used in this invention may be obtained from specific experiments or samples, or data from other sources may be considered, such as data from public databases or data collected by other laboratories. Different data sources may affect the training and prediction results of the model.

[0121] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning, characterized in that, Includes the following steps: Step 1: Data Collection: Collect yam samples from different harvesting periods, and analyze the yam samples using metabolomics technology to obtain metabolomics data of yam at different harvesting periods; Step 2, Data Preprocessing: Preprocess the collected metabolomics data; Step 3: Selection of potential marker metabolites: Using machine learning algorithms, features related to the growth stage of yam are selected from the preprocessed metabolomics data to obtain potential marker metabolites; Step 4: Construction and validation of the prediction model: The LASSO regression method was used to screen potential marker metabolites and construct a marker metabolite prediction model; the constructed marker metabolite prediction model was validated based on the area under the ROC curve as an evaluation index. The marker metabolites screened were allantoin, 5-oxo-L-proline, 4-hydroxymandelinnisone, L-methionine, 6,7-dihydroxy-2,4-dimethoxyphenanthrene, N-feruloylguanidine, and glucosylsyringic acid. Using the seven selected biomarker metabolites as indicators, the biomarker metabolite prediction model constructed using the LASSO regression method is as follows: Model score = -117.2510 + allantoin * (-2.7340) + 5-Oxo-L-proline* (-0.8350) + 4-Hydroxymandelinnitrile* 5.5810 + L-Methionine* 1.7373 + 6,7-Dihydroxy-2,4-Dimethoxyphenanthrene* 0.4416 + N-Feruloylguanidine* 1.9960+ Glucosyl eugenol* (-0.7759); Wherein, model score is the model score; A threshold is used to determine the model score. When the model score is greater than zero, the yam sample is suitable for harvesting; when the model score is less than zero, the yam sample is not suitable for harvesting. Step 5: Input the new yam metabolomics data into the validated marker metabolite prediction model, obtain the model score, and determine whether the yam is suitable for harvesting based on the model score.

2. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to claim 1, characterized in that, No fewer than 15 yam samples should be selected from each harvest period; Before implementing the metabolomics technology, the yam sample was freeze-dried, then an extraction solvent was added, ultrasonic extraction was performed, the supernatant was collected by centrifugation, and the sample solution was obtained by filtration through a membrane. The metabolomics data were obtained by collecting metabolite fingerprints from the injection solutions of yam samples from different harvesting periods.

3. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to claim 2, characterized in that, Metabolite fingerprints were acquired using chromatography-mass spectrometry (GC-MS). The chromatographic conditions for the GC-MS were: mobile phase flow rate of 0.2–0.4 mL / min. -1 The column temperature was 35–45 °C, and the injection volume was 2–4 μL. The mass spectrometry conditions for the chromatography-mass spectrometry system were as follows: the electrospray ion source temperature was set to 500–600 °C, the ion spray voltage for positive ion mode was set to 4500–6500 V, the ion spray voltage for negative ion mode was set to -4000–-6000 V, the ion source gas I, gas II, and curtain gas were set to 50, 60, and 25 psi, respectively, and the collision-induced ionization parameter was set to high.

4. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to any one of claims 1-3, characterized in that, The preprocessing includes data cleaning, outlier removal, and data normalization.

5. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to claim 4, characterized in that, The selection of potential marker metabolites is performed using the R software platform and includes the following steps: ① Perform baseline filtering, peak identification, retention time correction, peak alignment, and mass spectrometry fragment structure analysis on the mass spectrometry data, and convert the spectral data into two-dimensional matrix data; ② Perform unsupervised principal component analysis on all two-dimensional matrix data to identify differences between groups and regroup them; ③ Redefine the new group using supervised orthogonal partial least squares discriminant analysis to identify potential marker metabolites.

6. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to claim 5, characterized in that, The potential biomarker metabolites were selected based on importance (VIP value), significance (P value), and fold change (fold change value); variables with VIP > 1.0, P value < 0.05, and |log2(fold change)| > 1 were selected as potential biomarker metabolites.

7. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to any one of claims 1-3 and 5, characterized in that, The marker metabolite prediction model uses LASSO regression analysis to identify potential marker metabolites. LASSO regression minimizes the mean square error of the cross-validation curve and eliminates irrelevant potential marker metabolites through cross-validation curve and coefficient of variation path analysis. The remaining marker metabolite data are then fitted to construct the marker metabolite prediction model.

8. The method for predicting the optimal harvesting period of yam using a marker metabolite model based on machine learning according to claim 7, characterized in that, The area under the ROC curve is greater than 0.95, indicating that the marker metabolite prediction model has good accuracy, sensitivity, and specificity.

Citation Information

Patent Citations

  • Mung bean production place tracing method based on characteristic difference metabolites

    CN116381085A