Methods for predicting the origin and traceability of offspring

CN121678874BActive Publication Date: 2026-09-01江西省 中国科学院庐山植物园
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511908069.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-09-01
Estimated Expiration
2045-12-17

AI Technical Summary

Technical Problem

[0005]为解决现有技术存在的问题,本发明提供一种预知子产地溯源的方法,基于非靶向代谢组学技术,从预知子中筛选并验证了一组由7个特定代谢物组成的标志物组合;该组合能够作为"化学指纹"实现对预知子产地的精准判别,解决了现有技术无法准确溯源预知子产地的技术难题

Benefits of technology

1、高准确性:基于标志物组合构建的判别模型准确率高达100%。特别地,仅使用重要性排名前三或前五的代谢标志物子集,模型仍能保持100%的判别准确率,这为权利要求提供了强有力的、多层次的数据支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121678874B_ABST
    Figure CN121678874B_ABST
Patent Text Reader

Abstract

This invention discloses a method for predicting the origin of a sample and tracing its origin, comprising: step S1, extracting metabolites from the sample to be tested and detecting the content of a combination of metabolic markers using ultra-high performance liquid chromatography-mass spectrometry; step S2, inputting the content data of the metabolic marker combination into a pre-trained origin discrimination model; and step S3, the discrimination model outputting the predicted origin of the sample. The technical solution of this invention solves the problem that existing technologies cannot accurately trace the origin of a sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of traditional Chinese medicine identification technology, specifically involving a method for predicting the origin and tracing the source of medicinal materials. Background Technology

[0002] Yuzhizi is a traditional Chinese medicinal herb, the dried, nearly mature fruit of *Akebia quinata*, *Akebia trifoliata*, or *Akebia quinata* (all belonging to the Lardizabalaceae family). It is bitter and cold in nature, and enters the liver, gallbladder, stomach, and bladder meridians. Yuzhizi has the effects of soothing the liver and regulating qi, promoting blood circulation and relieving pain, dispersing nodules, and promoting diuresis. It is mainly used for abdominal distension and pain, dysmenorrhea, amenorrhea, phlegm nodules, and difficulty urinating. Modern pharmacological studies have shown that it possesses various pharmacological activities, including antidepressant, antithrombotic, antioxidant, and antitumor effects.

[0003] Traditional Chinese medicinal herbs are characterized by their distinct "origin," with their place of origin being a key factor determining their growing environment and intrinsic quality. However, the current market is rife with problems such as confusion regarding origin and the substitution of inferior goods for superior ones, severely impacting the quality, safety, and clinical efficacy of these herbs. Therefore, establishing precise origin traceability technology is of paramount importance.

[0004] Existing traceability technologies, such as morphological identification, are highly subjective; DNA barcoding is insufficient to reflect changes in medicinal components caused by environmental stress; and elemental fingerprinting is easily affected by human interventions such as fertilization. Metabolomics, on the other hand, can comprehensively and unbiasedly detect endogenous small-molecule metabolites in organisms, thereby capturing the overall metabolic response caused by differences in the production environment. Its metabolite profile can directly reflect the chemical quality of medicinal materials, making it an ideal technology for origin traceability. Currently, there are no reports on precise origin traceability of medicinal materials based on specific metabolic markers. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a method for tracing the origin of predicted progeny. Based on non-targeted metabolomics technology, a set of biomarkers consisting of seven specific metabolites was screened and verified from the predicted progeny. This set can serve as a "chemical fingerprint" to accurately identify the origin of the predicted progeny, thus solving the technical problem that the prior art cannot accurately trace the origin of predicted progeny.

[0006] To achieve the above objectives, the present invention provides the following solution: A method for predicting the origin and tracing the origin of offspring includes: Step S1: Extract metabolites from the sample to be tested and detect the content of the combination of metabolic markers using ultra-high performance liquid chromatography-mass spectrometry. Step S2: Input the content data of the metabolic biomarker combination into the pre-trained origin discrimination model; Step S3: The discrimination model outputs the predicted origin of the sample to be tested.

[0007] As a preferred combination of metabolic markers, the combination includes: norsine, (+ / -)-2-hydroxy-2-methylbutyrate ethyl ester, histidine-alanine-methionine, 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxo}oxacyclohexane-2-carboxylic acid, 2-(2,4-dihydroxyphenyl)-5,7-dihydroxy-3-(4-hydroxy-3,7-dimethyloct-2,6-dien-1-yl)-6-(3-methylbut-2-en-1-yl)-4H-benzopyran-4-one, lysophosphatidic acid (i-19:0 / 0:0), and undecanoic acid.

[0008] Preferably, in step S3, the origin of the sample to be tested is obtained based on the fact that the combination of metabolic markers includes at least three of the seven compounds.

[0009] As a preferred option, the origin discrimination model is obtained by training a known sub-training set of known origins on the combination of the seven metabolic biomarkers using machine learning algorithms such as random forest, support vector machine, K-nearest neighbor, linear discriminant analysis or logistic regression.

[0010] Preferably, the chromatographic column used in step S1 for ultra-high performance liquid chromatography-mass spectrometry is a Waters ACQUITY UPLC HSS T3 column, with the aqueous phase (A) being a 0.1% formic acid aqueous solution and the organic phase (B) being an acetonitrile solution containing 0.1% formic acid; the gradient elution is as follows: 0–5 min, 95% A → 35% A; 5–6 min, 35% A → 1% A; 6–7.5 min, 1% A; 7.5 min–7.6 min, 1% A → 95% A; 7.6–10 min, 95% A; the flow rate is 0.4 mL / min; the column temperature is 40 °C; and the injection volume is 4 μL.

[0011] Preferably, the mass spectrometer used in step S1 for the ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS) technique is an ABTripleTOF 6600; the ion source is an ESI, and positive and negative ion modes are used for scanning respectively. The source parameters are set as follows: spray gas 50psi; auxiliary heating gas 60psi; curtain gas 35psi; ion source temperature 550℃; declustering voltage in positive and negative modes are 80V and -80V respectively; ionization voltage in positive and negative modes are 5500V and -4500V respectively. The time-of-flight mass spectrometry scanning parameters were set as follows: mass range of 50 to 1250 Da; cumulative time of 200 ms. The product ion scanning parameters were set as follows: mass range of 50 to 1250 Da; cumulative time of 40 milliseconds; collision energy step size of 15 V; and collision energies of 30 V and -30 V in positive and negative modes, respectively.

[0012] Preferably, the method for obtaining the combination of metabolic biomarkers includes: The pre-training set of known origins is preprocessed, and a standardized data matrix containing the relative content of each metabolite is obtained by comparing with the metabolite database and combining secondary mass spectrometry information. The preprocessing includes: chromatographic peak identification and extraction, peak alignment, and retention time correction. The standardized data matrix underwent univariate and multivariate statistical analyses, including fold change analysis (FC) and orthogonal partial least squares discriminant analysis (OPLS-DA). After establishing the OPLS-DA model, the model was validated and overfitting was assessed using 200 permutation tests. The screening criterion for differentially expressed metabolites was VIP > 1.0. And |logFC|≥1; count the metabolites that meet the above conditions in each pairwise comparison; then take the common differential metabolites that meet the above conditions in all 10 pairwise comparisons, and obtain 7 as metabolic markers, namely: norsine, (+ / -)-2-hydroxy-2-methylbutyrate ethyl ester, histidine-alanine-methionine, 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxacyclohexane-2-carboxylic acid, 2-(2,4-dihydroxyphenyl)-5,7-dihydroxy-3-(4-hydroxy-3,7-dimethyloct-2,6-dien-1-yl)-6-(3-methylbut-2-en-1-yl)-4H-benzopyran-4-one, lysophosphatidic acid (i-19:0 / 0:0), and undecanoic acid.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. High accuracy: The discrimination model built based on biomarker combinations achieves an accuracy of up to 100%. In particular, the model can still maintain 100% discrimination accuracy even when using only a subset of metabolic biomarkers ranked by importance (top three or top five), which provides strong, multi-layered data support for the claims.

[0014] 2. Strong robustness: The discrimination ability of the marker combination does not depend on a specific algorithm. It has been verified by five machine learning models with different principles and all have achieved 100% accuracy, demonstrating strong universality and robustness.

[0015] 3. High practicality: Based on the combination of biomarkers, rapid detection kits or portable detection devices can be developed, providing reliable technical support for the supervision of the Chinese medicinal materials market and the quality control of raw materials for enterprises, and has broad market application prospects. Attached Figure Description

[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of a method for predicting the origin of offspring in an embodiment of the present invention; Figure 2 A heatmap showing the relative abundance of seven metabolic biomarkers in five production areas reveals the specific metabolic patterns of each production area. Figure 3 This is a performance comparison chart of five machine learning models, including accuracy comparisons and confusion matrices for each model. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Example 1 like Figure 1 As shown, the present invention provides a method for predicting the origin and tracing the source of offspring, comprising: Step S1: Extract metabolites from the sample to be tested and detect the content of the combination of metabolic markers using ultra-high performance liquid chromatography-mass spectrometry. Step S2: Input the content data of the metabolic biomarker combination into the pre-trained origin discrimination model; Step S3: The discrimination model outputs the predicted origin of the sample to be tested: As one embodiment of the present invention, step S1 includes: Step S11, Sample Collection Predictive sub-samples were collected from five production areas: Sichuan, Chongqing, Henan, Jiangxi, and Guizhou. Ten batches were collected from each production area, for a total of 50 samples.

[0021] Step S12: Extraction of sample metabolites The sesame seeds from different origins were pulverized into powder using a grinder and passed through a 40-mesh sieve. 50 mg of sample powder was weighed using an electronic balance (MS105DM), and 1200 μL of 70% methanol-water internal standard extraction solution pre-cooled at -20 ℃ was added. The mixture was vortexed once every 30 minutes for 30 seconds each time, for a total of 6 vortexes. After centrifugation (12000 rpm, 3 minutes), the supernatant was collected, the sample was filtered through a microporous membrane (0.22 μm), and stored in a sample vial for UPLC-MS / MS analysis.

[0022] Step S13, UPLC-MS / MS detection and analysis Chromatographic acquisition conditions: The column type was Waters ACQUITY UPLC HSS T3 column; the aqueous phase (A) was 0.1% formic acid aqueous solution, and the organic phase (B) was acetonitrile solution containing 0.1% formic acid; the gradient elution was as follows: 0–5 min, 95% A → 35% A; 5–6 min, 35% A → 1% A; 6–7.5 min, 1% A; 7.5 min–7.6 min, 1% A → 95% A; 7.6–10 min, 95% A; the flow rate was 0.4 mL / min; the column temperature was 40℃; and the injection volume was 4 μL.

[0023] Mass spectrometry conditions: Mass spectrometer model: AB TripleTOF 6600; Ion source: ESI; positive and negative ion modes were used for scanning respectively. The source parameters are set as follows: spray gas 50psi; auxiliary heating gas 60psi; curtain gas 35psi; ion source temperature 550℃; declustering voltage in positive and negative modes are 80V and -80V respectively; ionization voltage in positive and negative modes are 5500V and -4500V respectively. The time-of-flight mass spectrometry scanning parameters were set as follows: mass range of 50 to 1250 Da; cumulative time of 200 ms. The product ion scanning parameters were set as follows: mass range of 50 to 1250 Da; cumulative time of 40 milliseconds; collision energy step size of 15 V; and collision energies of 30 V and -30 V in positive and negative modes, respectively.

[0024] Step S14, Quality Control Quality control (QC) samples are prepared by mixing sample extracts and are used to analyze the reproducibility of samples under the same processing methods. During instrumental analysis, a QC sample is typically inserted for every 10 samples analyzed to monitor the reproducibility of the analytical process. By overlaying and analyzing the total ion current (TIC) plots of mass spectrometry analyses of different QC samples, the reproducibility of metabolite extraction and detection, i.e., technical reproducibility, can be determined. Pearson correlation analysis of the QC samples, along with coefficient of variation (CV) distribution plots and principal component analysis of all samples, are used to assess the reliability and stability of the combined analysis.

[0025] As one embodiment of the present invention, the metabolic marker combination includes: norsine, (+ / -)-2-hydroxy-2-methylbutyrate ethyl ester, histidine-alanine-methionine, 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxo}oxacyclohexane-2-carboxylic acid, 2-(2,4-dihydroxyphenyl)-5,7-dihydroxy-3-(4-hydroxy-3,7-dimethyloct-2,6-dien-1-yl)-6-(3-methylbut-2-en-1-yl)-4H-benzopyran-4-one, lysophosphatidic acid (i-19:0 / 0:0), and undecanoic acid; Furthermore, in step S3, based on the discriminant model, the origin of the sample to be tested is obtained according to the fact that the combination of metabolic biomarkers includes at least three of the seven compounds; wherein, the origin discrimination model is obtained by training a known sub-training set of known origins based on the combination of the seven metabolic biomarkers using machine learning algorithms such as random forest, support vector machine, K-nearest neighbor, linear discriminant analysis or logistic regression.

[0026] Furthermore, methods for obtaining combinations of metabolic biomarkers include: Preprocessing of the known sub-training set from known origins includes steps such as chromatographic peak identification and extraction, peak alignment, and retention time correction. Appropriate algorithms are used to handle missing values ​​and correct data to ensure data quality. Detected metabolites are identified by comparing with a metabolite database and combining this with secondary mass spectrometry information. Finally, a standardized data matrix containing the relative amounts of each metabolite is obtained.

[0027] The standardized data matrix underwent univariate and multivariate statistical analyses, including fold change (FC) analysis and orthogonal partial least squares-discriminant analysis (OPLS-DA). After implementing the OPLS-DA model, it was validated and overfitting was assessed using 200 permutation tests. The screening criteria for differentially expressed metabolites were: VIP > 1.0 and |logFC| ≥ 1 (i.e., FC ≥ 2 or FC ≤ 0.5). Metabolites meeting these criteria were counted in each pairwise comparison. Seven metabolic biomarkers were identified as common differential metabolites that met the above conditions in all 10 pairwise comparisons: norsine, (+ / -)-2-hydroxy-2-methylbutyrate ethyl ester, histidine-alanine-methionine, 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxacyclohexane-2-carboxylic acid, 2-(2,4-dihydroxyphenyl)-5,7-dihydroxy-3-(4-hydroxy-3,7-dimethyloctyl-2,6-dien-1-yl)-6-(3-methylbut-2-en-1-yl)-4H-benzopyran-4-one, lysophosphatidic acid (i-19:0 / 0:0), and undecanoic acid, as shown in Table 1. These seven metabolic biomarkers exhibit specific metabolic patterns in different origins, such as... Figure 2 As shown.

[0028] Table 1

[0029] As one embodiment of the present invention, the construction and verification of a predictive progeny origin discrimination model based on seven metabolic biomarkers are as follows: 1. Data Source Using the seven metabolic biomarkers screened in Example 1 as the core, predicted sub-samples from five producing areas—Jiangxi (JX), Hunan (HN), Sichuan (SC), Chongqing (CQ), and Guizhou (GZ)—were collected, with 10 batches from each producing area, totaling 50 samples. The relative peak area data of the seven metabolic biomarkers in each sample were extracted to form a feature variable matrix (50×7), with the corresponding known producing area labels serving as response variables.

[0030] 2. Data Processing Logarithmic transformation was performed on the original relative peak area data to improve the data distribution characteristics and reduce the impact of extreme values. Statistical analysis of the transformed data showed significant differences in various metabolites across different production areas, providing a good characteristic basis for model discrimination.

[0031] 3. Model Building A multi-class classification model was constructed using the random forest algorithm. Random forest is an ensemble learning method that can effectively handle high-dimensional data and evaluate feature importance, making it particularly suitable for classification problems in metabolomics data. Key parameters were set as follows: 1000 decision trees, maximum tree depth of 5, and 42 random seeds to ensure reproducible results. The model was evaluated using rigorous leave-one-out cross-validation.

[0032] 4. Model Performance Validation The model achieved 100% accuracy in leave-one-out cross-validation, with all 50 samples correctly classified to their true origin. Detailed performance metrics are shown in Table 2, and the confusion matrix is ​​shown in Table 3.

[0033] Table 2

[0034] Table 3

[0035] 5. Importance analysis and optimal combination verification of metabolic biomarkers Feature importance scores were extracted from the trained random forest model. The importance distribution of the seven metabolic biomarkers was relatively balanced, indicating that they worked synergistically as a whole chemical fingerprint. The importance ranking is shown in Table 4.

[0036] Table 4

[0037] To verify the robustness of the model and identify a more advantageous subset of metabolic biomarkers, a feature subset validation test was conducted. The results show: The model accuracy remained 100% when only the top 3 most important metabolites (i.e., undecanoic acid, norsine, and lysophosphatidylcholine (i-19:0 / 0:0)) were used. The model accuracy was also 100% when using the top 5 most important metabolites (i.e., undecanoic acid, norsine, lysophosphatidic acid (i-19:0 / 0:0), 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxo}oxacyclohexane-2-carboxylic acid, and histidine-alanine-methionine).

[0038] This discovery demonstrates that a specific subset of the metabolic biomarker combination also possesses perfect origin discrimination capability, providing a solid experimental basis for the hierarchical and preferred solutions in the claims of this invention.

[0039] 6. Practical Application Verification The constructed random forest model was used to test three batches of market-purchased unknown origin predictors. All prediction results were confirmed to be accurate by supply chain tracing, which verified the practical application value of the model. The results are shown in Table 5.

[0040] Table 5

[0041] This example demonstrates that: The random forest discriminant model constructed based on the seven metabolic biomarkers can accurately distinguish the predictors from the five production sites with 100% accuracy.

[0042] The importance of each metabolic marker is evenly distributed, forming a synergistic "chemical fingerprint" rather than relying on a single marker.

[0043] The model exhibits extremely high stability and reproducibility, and demonstrates reliable performance in actual market sample testing, fully proving the practical value and technical advantages of this invention in predicting the origin of offspring.

[0044] As one embodiment of the present invention, multi-model validation of the universal discriminative ability of combinations of metabolic biomarkers specifically includes: This embodiment verifies the universality and robustness of the combination of the seven metabolic biomarkers in predicting the origin of offspring by constructing five machine learning discrimination models based on different principles.

[0045] 1. Data Materials The relative peak area data of seven metabolic biomarkers from 50 predicted subsamples were used.

[0046] 2. Model Selection and Principles Five classification algorithms representing different machine learning paradigms were selected: Random Forest (ensemble learning), Support Vector Machine (kernel method), K-Nearest Neighbors (instance-based), Linear Discriminant Analysis (statistical learning), and Logistic Regression (linear model).

[0047] 3. Experimental Design All models employ the same leave-one-out cross-validation strategy to ensure consistency in evaluation criteria.

[0048] 4. Multi-model performance analysis All five machine learning models based on different principles achieved 100% discrimination accuracy. Specific results are shown in Table 6 and [example missing]. Figure 3 As shown. The confusion matrices of all models exhibit a perfect diagonal pattern, as... Figure 3 As shown in Table 7, the single-prediction accuracy analysis using leave-one-out cross-validation shows that all models exhibit perfect stability with a standard deviation of 0.

[0049] Table 6

[0050] Table 7

[0051] The consistent perfect performance (100% accuracy) of five different machine learning paradigms demonstrates that the metabolic biomarker combination screened in this invention possesses inherent and powerful discriminative capabilities, and its effectiveness is independent of any specific algorithm. This discovery has significant patent implications, meaning that in practical applications, appropriate classification algorithms can be flexibly selected according to needs, greatly enhancing the practical value and commercial prospects of this invention. The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for predicting the origin and tracing the source of offspring, characterized in that, include: The place of origin is traced back to Sichuan, Chongqing, Henan, Jiangxi, or Guizhou; Step S1: Extract metabolites from the sample to be tested and detect the relative content of the combination of metabolic markers using ultra-high performance liquid chromatography-mass spectrometry. Step S2: Input the relative content data of the metabolic biomarker combination into the pre-trained origin discrimination model; Step S3: The origin discrimination model outputs the predicted origin of the sample to be tested; The metabolic marker combination consisted of norsine, (+ / -)-2-hydroxy-2-methylbutyrate ethyl ester, histidine-alanine-methionine, 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxo}oxacyclohexane-2-carboxylic acid, 2-(2,4-dihydroxyphenyl)-5,7-dihydroxy-3-(4-hydroxy-3,7-dimethyloct-2,6-dien-1-yl)-6-(3-methylbut-2-en-1-yl)-4H-benzopyran-4-one, lysophosphatidic acid (i-19:0 / 0:0), and undecanoic acid; The origin discrimination model is based on the combination of the seven metabolic biomarkers and is trained on the predicted sub-training sets of five origin regions: Sichuan, Chongqing, Henan, Jiangxi, and Guizhou using machine learning algorithms such as random forest, support vector machine, K-nearest neighbor, linear discriminant analysis, or logistic regression. Step S1 involves extracting metabolites from the sample of the predictor molecule as follows: The predictor molecule is pulverized into powder using a grinder and passed through a 40-mesh sieve; 50 mg of sample powder is weighed using an electronic balance, and 1200 μL of 70% methanol-water internal standard extraction solution pre-cooled at -20 ℃ is added. The sample is vortexed once every 30 minutes for 30 seconds each time, for a total of 6 vortexes; after centrifugation, the supernatant is collected, the sample is filtered through a microporous membrane, and stored in a sample vial for ultra-high performance liquid chromatography-mass spectrometry analysis; The ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS / MS) technique described in step S1 uses a Waters ACQUITY UPLCHSS T3 column. Aqueous phase A is a 0.1% formic acid aqueous solution, and organic phase B is an acetonitrile solution containing 0.1% formic acid. The gradient elution is as follows: 0–5 min, 95% A → 35% A; 5–6 min, 35% A → 1% A; 6–7.5 min, 1% A; 7.5 min–7.6 min, 1% A → 95% A; 7.6–10 min, 95% A. The flow rate is 0.4 mL / min, the column temperature is 40 °C, and the injection volume is 4 μL. The mass spectrometer used in step S1 for the ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS) technique is an AB TripleTOF 6600; the ion source is an ESI, and positive and negative ion modes are used for scanning respectively. The source parameters are set as follows: spray gas 50psi; auxiliary heating gas 60psi; curtain gas 35psi; ion source temperature 550℃; declustering voltage in positive and negative modes are 80V and -80V respectively; ionization voltage in positive and negative modes are 5500V and -4500V respectively. The time-of-flight mass spectrometry scanning parameters were set as follows: mass range of 50 to 1250 Da; cumulative time of 200 ms. The product ion scanning parameters were set as follows: mass range of 50 to 1250 Da; cumulative time of 40 ms; collision energy step size of 15 V; and collision energies of 30 V and -30 V in positive and negative modes, respectively.

2. The method for predicting the origin and tracing the source of offspring as described in claim 1, characterized in that, Methods for obtaining combinations of metabolic biomarkers include: The predicted sub-training sets of the five origins were preprocessed, and a standardized data matrix containing the relative contents of each metabolite was obtained by comparing with the metabolite database and combining secondary mass spectrometry information. The preprocessing included: chromatographic peak identification and extraction, peak alignment, and retention time correction. The standardized data matrix underwent univariate and multivariate statistical analyses, including fold change analysis and OPLS-DA. After implementing the OPLS-DA model, the model was validated and overfitting was assessed using a permutation test with 200 iterations. The screening criteria for differentially expressed metabolites were: VIP > 1.0 and |logFC| ≥ 1.

0.

1. Metabolites that meet the above screening criteria in pairwise comparisons across different production areas are statistically analyzed. Then, common differential metabolites that meet the above screening criteria in all pairwise comparisons are selected, resulting in 7 metabolic markers: norsine, (+ / -)-2-hydroxy-2-methylbutyrate ethyl ester, histidine-alanine-methionine, 3,4,5-trihydroxy-6-{[3-(3,4,5-trihydroxyphenyl)propionyl]oxacyclohexane-2-carboxylic acid, 2-(2,4-dihydroxyphenyl)-5,7-dihydroxy-3-(4-hydroxy-3,7-dimethyloct-2,6-dien-1-yl)-6-(3-methylbut-2-en-1-yl)-4H-benzopyran-4-one, lysophosphatidic acid (i-19:0 / 0:0), and undecanoic acid.

Citation Information

Patent Citations

  • Dried ginger producing area tracing method based on characteristic metabolites

    CN115586283A