Metabolic marker combination for identifying okra related species, identification method and application of metabolic marker combination
By using UPLC-MS/MS detection and OPLS-DA analysis, common differential metabolites among closely related species of the genus *Abelmoschus* were screened, solving the accuracy problem of species identification in the genus *Abelmoschus* using traditional methods. This method achieves efficient and accurate interspecific identification and is applicable to the management of Chinese medicinal materials and germplasm resources.
Patent Information
- Application Number
- CN202511217966.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, traditional morphological methods are insufficient to accurately identify closely related species of the Okra genus, leading to mixed germplasm and misidentification of medicinal origins. Current metabolomics studies have failed to achieve accurate cross-species identification.
Untargeted metabolomics detection was performed using ultra-high performance liquid chromatography-tandem mass spectrometry (UPLC-MS/MS) to screen out common differential metabolites among okra, hibiscus, romaine lettuce, and coffee okra. An OPLS-DA analysis model was constructed to screen out 47 characteristic metabolic biomarkers, including flavonoids and phenolic acids, and accurate identification was achieved through metabolite combinations.
It enables accurate identification of morphologically similar species in the genus *Abelmoschus*, improves the accuracy of identification of medicinal materials and management of germplasm resources, provides highly specific molecular tools, and the detection method is accurate and reliable, suitable for large-scale sample analysis.
Smart Images

Figure CN120992803A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of drug detection technology, and relates to closely related species of the genus *Okra*, specifically to a combination of metabolic markers for the identification of closely related species of the genus *Okra*, identification methods, and their applications. Background Technology
[0002] okra ( Abelmoschus manihot (L.) Medicus), Abelmoschus manihot ( Abelmoschus moschatus Medicus), Arrow Leaf Okra ( Abelmoschus sagittifolius (Kurz) Merr.) and coffee hibiscus ( Abelmoschus esculentus (L.) Moench. are all herbaceous plants belonging to the genus Abelmoschus in the Malvaceae family. The stem bark fiber of Abelmoschus hupehensis can replace hemp raw materials, and its root mucilage is used as a papermaking paste. The whole plant can be used medicinally and has the effects of clearing heat and cooling blood. The whole plant of Abelmoschus hupehensis can be used medicinally. It tastes slightly sweet and is cold in nature. It enters the heart and lung meridians and has the effects of clearing heat and detoxifying, promoting lactation and relieving constipation. The root of Abelmoschus salsa is used medicinally to treat stomach pain and neurasthenia. Externally, it is used to remove blood stasis and reduce swelling, treat sprains and bruises, and promote bone healing. The tender fruit of Abelmoschus hupehensis can be used as a vegetable and has certain economic and edible value.
[0003] However, the four plants mentioned above exhibit highly similar phenotypic characteristics in the seedling stage, such as leaf morphology and trichome density. Traditional morphological identification methods have a high error rate, which can easily lead to germplasm contamination, misjudgment of medicinal origins, and difficulties in seedling quality control. In existing technologies, Yang Xiaonan et al. (Yang Xiaonan et al. Differential analysis of components in different parts of Abelmoschus mandshurica based on global non-targeted metabolomics [J], Modern Chinese Materia Medica, 2023) analyzed the differences in metabolites in different parts of Abelmoschus mandshurica, such as roots, stems, and leaves. However, their research was limited to the distribution of metabolites within a single species and did not involve the screening of characteristic metabolic markers and the construction of identification models among closely related species of the Abelmoschus genus. Therefore, it is impossible to achieve accurate cross-species identification and quality correlation evaluation.
[0004] Metabonomics is an emerging "omics" technology that has developed in recent years. It is a high-throughput, qualitative, and quantitative technique for measuring small molecule metabolites in the metabolome. It studies metabolic pathways in biological systems by examining the dynamic, multi-parameter responses of metabolite profiles over time before and after stimulation or disturbance (such as mutations in a specific gene or environmental changes). The metabolome refers to the collection of all endogenous and exogenous small molecule metabolites in cells, tissues, or the entire organism. These small molecules include peptides, amino acids, nucleic acids, carbohydrates, organic acids, vitamins, polyphenols, alkaloids, and inorganic substances. Advances in the technology for separating, analyzing, and identifying these small molecules have made metabolomics research possible. These technologies primarily include: high-resolution mass spectrometry (MS) for precise mass determination, high-resolution, high-throughput magnetic resonance (NMR), capillary electrophoresis (CE), and high-pressure liquid chromatography (HPLC) and ultra-high-pressure liquid chromatography (UPLC). Supported by these technological platforms, a series of datasets suitable for metabolite identification can be generated. However, to accurately identify the metabolites represented by each signal, powerful bioinformatics tools are needed to automate data analysis. Building metabolomics databases based on various spectroscopic techniques, building upon experimental data, will make metabolite identification more convenient and feasible.
[0005] The prior art CN117607325A discloses a method for identifying daylilies from the Datong region based on metabolomics, including: freeze-drying samples; extraction; UPLC-MS / MS analysis; qualitative and quantitative analysis of metabolites; and comparison of differentially expressed metabolites in different samples. Specifically, ultra-high performance liquid chromatography and tandem mass spectrometry are used for sequencing; unsupervised principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (OPLS-DA) are performed based on the metabolite composition distribution of daylilies from Datong and other regions; and KEGG enrichment analysis is used to identify key differentially expressed metabolites in daylilies from Datong and other regions.
[0006] To identify metabolites with medicinal and edible value in cotton flowers, Zhen Junbo et al. conducted metabolomics analysis on flowers of upland cotton and hibiscus on the day of flowering, detecting a total of 465 metabolites. Differential metabolites were screened using a combination of fold change and VIP value. The results showed that 269 significantly different metabolites were identified in the flowers of upland cotton and hibiscus, while 196 metabolites showed no significant difference in content between the two species (Zhen Junbo, Song Shijia, Liu Linlin, et al. Screening and analysis of differential metabolites in flowers of upland cotton and hibiscus (also known as yellow hibiscus) [J]. Journal of China Agricultural University, 2021). Summary of the Invention
[0007] This invention addresses the lack of existing metabolomics analysis methods for identifying closely related species in the genus *Abelmoschus*, providing a combination of metabolic markers, an identification method, and its application for this purpose. The identification method provided by this invention includes the following steps: [The text abruptly shifts to a different topic] Abelmoschus manihot (L.) Medik.), Abelmoschus manihot ( Abelmoschus moschatus Medicus), Arrow Leaf Okra ( Abelmoschus sagittifolius (Kurz)Merr.) and coffee hibiscus ( Abelmoschus esculentus Dry powder samples of *Okra* (Linn.) Moench were analyzed using ultra-high performance liquid chromatography-tandem mass spectrometry (UPLC-MS / MS) for non-targeted metabolomics detection to obtain metabolite data. The metabolite data were then subjected to qualitative and quantitative mass spectrometry analysis to screen out common differential metabolites among the four plant species as characteristic metabolites, enabling varietal identification of closely related species within the *Okra* genus. (See schematic diagram). Figure 1 This invention can effectively distinguish morphologically similar species of the Okra genus and has application value in the identification of medicinal materials, screening of germplasm resources, and detection of variety purity in the seedling stage.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: On the one hand, the present invention provides an identification method for identifying closely related species of the genus *Okra*, comprising the following steps: S1. Sample preparation: Collect mixed stem and leaf samples of closely related species of the genus *Abelmoschus*, freeze-dry them separately under vacuum, grind them into powder, add the extraction solution, centrifuge after extraction, take the supernatant, filter and obtain the solution to be tested; the extraction solution includes methanol solution, and the closely related species of the genus *Abelmoschus* include *Abelmoschus quinquefolius*, *Abelmoschus quinquefolius*, *Abelmoschus quinquefolius*, and *Abelmoschus quinquefolius*. S2. Metabolomics Data Detection: The test solution obtained in step S1 was subjected to non-targeted metabolomics detection using UPLC-MS / MS to obtain metabolomics data of closely related species of the Okra genus mentioned in step S1. The UPLC-MS / MS liquid chromatography conditions included: separation of the sample extract using an Agilent SB-C18 1.8µm, 2.1mm*100mm column; mobile phase A consisting of formic acid solution, mobile phase B consisting of acetonitrile containing formic acid, and column temperature of 30-50°C. The mass spectrometry conditions for the UPLC-MS / MS include: an electrospray ion source temperature of 400-600°C; and an ion spray voltage of 5000-6000V in positive ion mode. S3. Processing and Analysis of Metabolic Data: The metabolomics data obtained in step S2 are first assessed by principal component analysis to evaluate the overall metabolic differences among samples and the variability within groups; an OPLS-DA analysis model is constructed and variable weights (VIPs) are calculated to screen for variables with differences among groups; finally, the fold change (FC) is calculated. S4. Screening of Common Differential Metabolites: Using the OPLS-DA analysis model from step S3, common differential metabolites of the closely related species of the *Abelmoschus* genus described in step S1 are screened as a combination of metabolic markers to distinguish these closely related species. This combination of metabolic markers includes N-ethylleucine, trimethyllysine, 1-O-caffeoyl-6-O-glucosyl-β-D-glucose, 1-O-coumaroyl-β-D-glucose, 2-phenylethylβ-primrosein, and 2-(3-... 4-Dihydroxyphenethoxy)6-feruloyl-O-glucosyl-D-xylose, 2'-hydroxygenistein, 3',4'-dihydroxy-5,6,7,8,5'-pentamethoxyflavone, 3,5-dihydroxy-7,4'-dimethoxyflavone, 5,2',5'-trihydroxy-3,7,4'-trimethoxyflavone-2'-oxo-β-D-glucoside, 5,4'-dihydroxy-6,7,8,3'-tetramethoxyflavone, 5,7-dihydroxy-3,4',6 8-Tetramethoxyflavonoids, 5,7-dihydroxy-3,6,8,3',4'-pentamethoxyflavonoids, 5,7-dihydroxy-6,3',4',5'-tetramethoxyflavonoids, 5-hydroxy-3',4',7,8-tetramethoxyflavonoids, 7,4'-dihydroxy-3,5,3'-trimethoxyflavonoids, 7,8-dihydroxy-5,6,4'-trimethoxyflavonoids, kaempferol-3-O-mannoside, andrographolide D-aglycone, Brickellin, and cat's eye grass Flavonoids, Methylcitriol, Hemigerminol, Gardeniaflavin C, Gardeniain E Malonyl Glucoside, 5,7-Dihydroxy-6,8,3',4'-Tetramethoxyflavonoids, Isorhamnetin-3-O-rutinoside, Kaempferol-3-O-(2''-O-acetyl)glucoside, Kaempferol-3-O-rutinoside, Sophora flavescens, Scutellaria baicalensis Flavonoids II, Syringin-3-O-(6''-acetyl)glucoside, Tamarixanthin-3-O-rutinoside, Tupichinol E, tamarind, 5,4'-dihydroxy-6,7,8-trimethoxyflavone, 5,6-dihydroxyrucidin-3β-alizarin, safflower yellow A, 8-hydroxycoumarin, schisandrin I, 7-methyl-5,8-dioxododecyl hydrogen sulfate, dehydrodieugenol, galactitol, glucosyl 5,8-dihydroxy-2,6-dimethyloctacarbon-2,6-dienoic acid, indole-3-lactic acid, tetrahydrofolate, and gingerol C; S5. Identification of Unknown Okra Species: After obtaining the metabolite mass spectrometry data of the unknown sample through measurement and analysis, the peak area of all substances in the chromatogram is integrated and corrected, and merged with the peak area database of the closely related species of the Okra genus measured in steps S1-S4; after statistical analysis of the data, the relative content of the common differential metabolites screened in step S4 is obtained, a heatmap is generated, and the aggregation of unknown samples in the graph is observed; the known species that aggregates with the unknown sample first is identified as the unknown sample species.
[0009] Preferably, the grinding in step S1 includes grinding for 1-3 minutes using a grinder at 20-50 Hz; the extraction solution includes a 60%-80% methanol-water internal standard extraction solution pre-cooled at -20℃; the vortexing conditions include vortexing once every 30 minutes for 30 seconds each time, for a total of 4-8 vortexes; the centrifugation conditions include 10000-15000 rpm for 2-5 minutes; and the filtration includes filtering the sample using a microporous membrane.
[0010] Preferably, the grinding in step S1 is performed by grinding at 30 Hz for 1.5 minutes using a grinder; the extraction solution is a 70% methanol-water internal standard extraction solution pre-cooled at -20℃; the vortexing is performed once every 30 minutes, each time lasting 30 seconds, for a total of 6 vortexes; the centrifugation is performed at 12000 rpm for 3 minutes; and the filtration is performed by filtering the sample using a 0.22 μm microporous membrane.
[0011] Preferably, the HPLC conditions for UPLC-MS / MS in step S2 include: separating the sample extract using an Agilent SB-C18 1.8µm, 2.1mm*100mm column; mobile phase A being ultrapure water containing 0.1% formic acid, and mobile phase B being acetonitrile containing 0.1% formic acid; the elution gradient is as follows: at 0.00 min, the proportion of phase B is 5%; within 9.00 min, the proportion of phase B increases linearly to 95% and is maintained at 95% for 1 min; from 10.00 to 11.10 min, the proportion of phase B decreases to 5% and equilibrates to 5% for 14 min; the flow rate is 0.35 mL / min; the column temperature is 40°C; and the injection volume is 2 μL.
[0012] Preferably, the mass spectrometry conditions for UPLC-MS / MS in step S2 include: an electrospray ion source temperature of 500°C; an ion spray voltage of 5500V in positive ion mode and -4500V in negative ion mode; ion source gas I, gas II, and curtain gas set to 50, 60, and 25 psi, respectively; a collision-induced ionization parameter set to high; QQQ scan using MRM mode; and nitrogen as the collision gas, set to medium.
[0013] Preferably, step S3 further includes qualitative and quantitative analysis of metabolic data; the qualitative analysis includes analysis based on the existing database MWDB using secondary spectral information; the quantitative analysis includes analysis using the multi-reaction monitoring mode of triple quadrupole mass spectrometry, after obtaining the signal intensity of characteristic ions in the detector, integrating and correcting all chromatographic peaks of substances using MultiQuant software, characterizing the relative content of metabolites by peak area, and finally exporting and saving all chromatographic peak area integrated data.
[0014] Preferably, in step S4, the OPLS-DA analysis model selects metabolites that meet the conditions of VIP≥1.0 and |Log2FC|≥1.0 as differential metabolites.
[0015] On the other hand, the present invention provides a combination of metabolic markers obtained by the above-described identification method, the combination of metabolic markers including N-ethylleucine, trimethyllysine, 1-O-caffeoyl-6-O-glucosyl-β-D-glucose, 1-O-p-coumaryl-β-D-glucose, 2-phenylethyl β-primrosein, 2-(3,4-dihydroxyphenethoxy)6-feruloyl-O-glucosyl-D-xylose, 2'-hydroxygenistein, 3 ',4'-dihydroxy-5,6,7,8,5'-pentamethoxyflavone, 3,5-dihydroxy-7,4'-dimethoxyflavone, 5,2',5'-trihydroxy-3,7,4'-trimethoxyflavone-2'-oxo-β-D-glucoside, 5,4'-dihydroxy-6,7,8,3'-tetramethoxyflavone, 5,7-dihydroxy-3,4',6,8-tetramethoxyflavone, 5,7-dihydroxy-3,6,8,3', 4'-Pentamethoxyflavonoids, 5,7-Dihydroxy-6,3',4',5'-Tetramethoxyflavonoids, 5-Hydroxy-3',4',7,8-Tetramethoxyflavonoids, 7,4'-Dihydroxy-3,5,3'-Trimethoxyflavonoids, 7,8-Dihydroxy-5,6,4'-Trimethoxyflavonoids, Kaempferol-3-O-Mannoside, Andrographolide D-Aglycone, Brickellin, Erythroflavonoidin, Methylcitriol, and Scutellarin Lansu, Gardenia Flavonoid C, Gardeniae E Malonyl Glucoside, 5,7-Dihydroxy-6,8,3',4'-Tetramethoxyflavonoid, Isorhamnetin-3-O-rutinoside, Kaempferol-3-O-(2''-O-acetyl)glucoside, Kaempferol-3-O-rutinoside, Lithocarpin, Scutellaria Flavonoid II, Syringin-3-O-(6''-acetyl)glucoside, Tamarixanthin-3-O-rutinoside, Tupichinol E, teosin, 5,4'-dihydroxy-6,7,8-trimethoxyflavone, 5,6-dihydroxyrucidin-3β-alizarin, safflower yellow A, 8-hydroxycoumarin, schisandrin I, 7-methyl-5,8-dioxododecyl hydrogen sulfate, dehydrodieugenol, galactitol, glucosyl 5,8-dihydroxy-2,6-dimethyloctacarbon-2,6-dienoic acid, indole-3-lactic acid, tetrahydrofolate, and gingerol C.
[0016] On the other hand, the present invention provides the application of the above-described identification method or the above-described combination of metabolic markers, the application of which includes the identification of closely related species of the genus *Okra*.
[0017] Preferably, the application also includes the identification of the origin of Chinese medicinal materials, the screening of germplasm resources, or the detection of variety purity during the seedling stage.
[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. For the first time, a unique metabolic biomarker profile for okra, hibiscus, okra sagittatum, and okra coffee was constructed based on broad-targeted metabolomics technology. By screening key differential metabolites such as flavonoids and phenolic acids, precise interspecific identification was achieved, breaking through the dependence of traditional methods on morphological characteristics. This provides a highly specific molecular tool for the identification of the origin of Chinese medicinal materials and the management of germplasm resources.
[0019] 2. The UPLC-MS / MS analytical method used in this study has the advantages of accuracy and reliability, small number of analytes, and short detection cycle, providing an efficient solution for rapid analysis of large-scale samples. Attached Figure Description
[0020] Figure 1 This is a schematic diagram illustrating the principle of the method of the present invention.
[0021] Figure 2 A circular diagram showing the composition of total metabolite categories.
[0022] Figure 3 This is a superimposed image of the total ion flow map (TIC map) of the quality control sample (QC) under positive ion mode.
[0023] Figure 4 This is an overlay of the total ion flow map (TIC map) of the quality control sample (QC) under negative ion mode.
[0024] Figure 5 The coefficient of variation (CV) distribution is shown for each group of samples; among them, okra, yellow okra, arrow-leaf okra and coffee okra are named Ama, Amo, Asa and Aes, respectively.
[0025] Figure 6 PCA score plots of mass spectrometry data for each group of samples and quality control samples (QC); among them, okra, abelmoschus mandshurica, okra, and abelmoschus mandshurica are named Ama, Amo, Asa, and Aes, respectively.
[0026] Figure 7 The control chart for the overall sample PC1 is shown; among them, okra, yellow okra, arrow-leaf okra, and coffee okra are named Ama, Amo, Asa, and Aes, respectively.
[0027] Figure 8 The OPLS-DA score plots for Aes and Ama are shown; the okra and coffee okra are named Ama and Aes, respectively.
[0028] Figure 9 The OPLS-DA score plots for Aes and Amo are shown; where Okra and Amo are named Aes and Amo, respectively.
[0029] Figure 10The OPLS-DA score plots for Amo and Ama are shown; the yellow okra and yellow hollyhock are named Amo and Ama, respectively.
[0030] Figure 11 The OPLS-DA score plots for Asa and Aes are shown; among them, okra and hibiscus are named Asa and Aes, respectively.
[0031] Figure 12 The OPLS-DA score plots for Asa and Ama are shown; among them, okra and yellow hibiscus are named Asa and Ama, respectively.
[0032] Figure 13 The OPLS-DA score plots for Asa and Amo are shown; among them, okra and hibiscus are named Asa and Amo, respectively.
[0033] Figure 14 Cluster heatmaps of 47 common differential metabolites were generated; among them, okra, hibiscus, arrow-leaf okra and coffee okra were named Ama, Amo, Asa and Aes, respectively.
[0034] Figure 15 Cluster heatmaps of common differential metabolites were generated using only VIP≥1.0 as the criterion for screening differential metabolites; among them, okra, hibiscus, and coffee okra were named Ama, Amo, Asa, and Aes, respectively.
[0035] Figure 16 Heatmap generated to identify an unknown okra plant; among them, okra, okra, okra salsa, and okra cambodiana and the unknown okra were named Ama, Amo, Asa, Aes and XAma1, respectively. Detailed Implementation
[0036] Unless otherwise specified, all raw materials and reagents used in this invention were purchased from commercial suppliers, and experiments were conducted in accordance with the operating instructions. Unless otherwise specified, all instruments, equipment, and apparatus used in this invention are conventional instruments, equipment, and apparatus, and experiments were conducted in accordance with the operating instructions and the accompanying reagents.
[0037] Data analysis was performed using data processing software. Significance analysis was conducted using t-tests and one-way ANOVA tests, with P < 0.05 indicating a significant difference.
[0038] Example 1: Identification of closely related species in the genus Okra (1) Sample collection and solution preparation A mixture of stem and leaf samples from four species—Ama, Amo, Asa, and Aes—was collected (samples were taken approximately one week after sowing the seeds during the germination and growth stage, using stems and leaves). These samples were then freeze-dried separately under vacuum, followed by grinding into powder using a grinder (30 Hz, 1.5 min). 50 mg of the powder was weighed using an electronic balance and 1200 μL of pre-cooled (-20°C) 70% methanol-water internal standard extract was added. The mixture was vortexed for 30 seconds every 30 minutes, for a total of 6 times. After centrifugation (12000 rpm, 3 min), the supernatant was aspirated, filtered through a microporous membrane (0.22 μm pore size), and stored in syringes for UPLC-MS / MS analysis.
[0039] (2) Metabolomics data detection Four species of okra were analyzed using ultra-high performance liquid chromatography-tandem mass spectrometry (UPLC-MS / MS, Agilent 8890-7000D).
[0040] The main liquid chromatography conditions included: separation of the sample extract using an Agilent SB-C18 (1.8 µm, 2.1 mm * 100 mm) column; mobile phase A being ultrapure water containing 0.1% formic acid, and mobile phase B being acetonitrile containing 0.1% formic acid; the elution gradient was: 5% for phase B at 0.00 min, linearly increasing to 95% within 9.00 min and maintaining at 95% for 1 min, decreasing to 5% from 10.00 to 11.10 min, and equilibrating to 5% at 14 min; the flow rate was 0.35 mL / min; the column temperature was 40°C; and the injection volume was 2 μL.
[0041] The mass spectrometry conditions mainly included: an electrospray ionization (ESI) source temperature of 500°C; an ion spray voltage (IS) of 5500 V (positive ion mode) / -4500 V (negative ion mode); ion source gas I (GSI), gas II (GSII), and curtain gas (CUR) were set to 50, 60, and 25 psi, respectively; and the collision-induced ionization parameter was set to high. QQQ scans used MRM mode with the collision gas (nitrogen) set to medium. DP and CE of each MRM ion pair were completed through further optimization of the declustering voltage and collision energy. A specific set of MRM ion pairs was monitored at each epoch based on the metabolites eluted within each epoch.
[0042] (3) Qualitative and quantitative analysis by mass spectrometry and instrument stability test Mass spectrometry data were processed using Analyst 1.6.3 software. Based on the existing metabolic database MWDB (metware database), qualitative and quantitative mass spectrometry analyses of the sample metabolites were performed. Qualitative analysis was based on the MWDB database and performed according to secondary spectrum information. Isotope signals, repetitive signals containing K+, Na+, and NH4+ ions, as well as repetitive signals from fragment ions that are themselves larger molecular weight substances, were removed during analysis. Quantitative analysis was performed using the multiple reaction monitoring (MRM) mode of triple quadrupole mass spectrometry. In MRM mode, the quadrupole first screens for precursor ions of the target substance, excluding ions corresponding to other molecular weight substances to initially eliminate interference. After the precursor ions are induced to ionize in the collision chamber, they break into many fragment ions. These fragment ions are then filtered by the triple quadrupole to select a characteristic fragment ion, eliminating interference from non-target ions (making quantification more accurate and reproducible). After obtaining the signal intensity (CPS) of the characteristic ion in the detector, the chromatographic peaks of all substances were integrated and corrected using MultiQuant software, and the peak area (Area) was used to characterize the relative content of the metabolites. Finally, all chromatographic peak area integral data were exported and saved. To compare the differences in the content of each metabolite in different samples, the chromatographic peaks of each metabolite detected in different samples were corrected to ensure the accuracy of qualitative and quantitative analysis.
[0043] The composition of metabolites is sample-specific; different types of samples contain different types and proportions of metabolites. Furthermore, the composition of metabolites can change under different treatments or biological processes. Metabolite composition ratio analysis allows for an overall assessment of the distribution of major metabolites in a sample. Based on the metabolite information collected by UPLC / MS-MS in step (2) above from four Okra species (Abelmoschus manihot, Abelmoschus quinquefolius, Abelmoschus quinquefolius, and Abelmoschus quinquefolius), a total of 1825 metabolites were detected, mainly including flavonoids, phenolic acids, lipids, etc. (see [link to article]). Figure 2 (This includes 1274 differential metabolites.)
[0044] To verify the reliability of the analytical system, this study systematically evaluated the instrument stability using multi-dimensional indicators. TIC analysis showed high curve overlap for total ion current detection of metabolites, meaning that retention time and peak intensity were consistent, indicating that the mass spectrometer exhibited good signal stability when detecting the same sample at different times (see [link to TIC analysis]). Figure 3 and Figure 4 To further quantify data repeatability, the coefficient of variation (CV) statistics of QC sample metabolites showed that over 75% of the substances had a CV value below 0.3, confirming that the experimental system had low dispersion and repeatability that met metabolomics standards. Figure 5The instrument signal exhibits excellent stability. However, principal component analysis of the samples (including quality control samples) revealed significant aggregation in the QC samples. Figure 6 ), and the standard deviation of its PC1 score is within plus or minus 2 standard deviations ( Figure 7 This further confirms that the instrument remains stable during continuous testing and that technical deviations are controllable. Based on these indicators, the analytical system provides high reproducibility and reliable data quality assurance for metabolite detection.
[0045] (4) Processing and analysis of metabolic data First, PCA analysis was performed on the samples to preliminarily understand the overall metabolite differences among the groups and the magnitude of variability within each group. The PCA results showed metabolomic separation trends between groups, indicating whether there were metabolomic differences within the sample groups. Figure 6 The PCA results showed that the variance contribution rates of the first and second principal components were 32.1% and 28.13%, respectively, explaining a total of 60.23% of the data variation. This indicates that the first two principal components can effectively characterize the overall features of the dataset, and there is a clear trend of metabolomic separation among the groups. The okra and coffee okra groups were completely separated along the PC1 axis, while the okra and sword-leaf okra groups were separated along the PC2 axis.
[0046] While PCA effectively extracts key information, it is insensitive to variables with low correlation, a problem that OPLS-DA addresses. Orthogonal partial least squares discriminant analysis combines orthogonal signal correction (OSC) and partial least squares discriminant analysis (PLS-DA) to decompose the X matrix information into two classes: those correlated with and those uncorrelated with Y. By removing uncorrelated differences, differential variables are filtered out. Metabolomics data are analyzed using the OPLS-DA model, and score plots are generated for each group to further illustrate the differences between them. Figures 8-13 ).
[0047] Finally, to more clearly and intuitively demonstrate the overall metabolic differences, VIP and FC values of metabolites in the comparison group were calculated.
[0048] (5) Screening for common differential metabolites Metabolomics data are characterized by their high dimensionality and massive volume, thus requiring a combination of univariate and multivariate statistical analysis methods. Analysis should be conducted from multiple perspectives based on the data characteristics to accurately identify differential metabolites. The differential metabolite screening criteria in this patent are: VIP ≥ 1.0 and |Log2FC| ≥ 1.0. Further screening was performed on common differential metabolites among samples of *Abelmoschus humilis*, *Abelmoschus quinquefolius*, *Abelmoschus salsa*, and *Abelmoschus aubergine* as characteristic metabolites. Finally, 47 common differential metabolites that can distinguish *Abelmoschus humilis*, *Abelmoschus quinquefolius*, *Abelmoschus quinquefolius*, and *Abelmoschus aubergine* were obtained, including 30 flavonoids, 4 phenolic acids, 2 quinones, 2 amino acids and their derivatives, 2 lignans and coumarins, 1 alkaloid, 1 organic acid, 1 lipid, and 4 other compounds (see Table 1).
[0049] Based on the 47 common differential metabolites selected above, a cluster heatmap was generated, which clearly shows the differences in metabolites among different Okra species, allowing for comparison of Okra samples to be identified (see...). Figure 14 ).
[0050] Table 1 Summary of Common Differential Metabolites
[0051] (6) Identification of unknown Okra species After obtaining the metabolite mass spectrometry data of the unknown sample, the peak areas of all substances were integrated and corrected, and then merged with the peak area database of the known okra varieties measured above. After statistical analysis of the data, the relative contents of the 47 differential metabolites screened in step (5) were obtained, and a heatmap was generated. If the unknown sample clusters together with any of the okra species measured above in the clustering heatmap, then that species is identified.
[0052] Comparative Example 1: The impact of differential metabolite screening criteria on identification specificity (1) Experimental Procedure: The same sample sources (Abelmoschus manihot, Abelmoschus manihot, Abelmoschus manihot, and Abelmoschus manihot) as in Example 1, sample pretreatment methods, ultra-high performance liquid chromatography-tandem mass spectrometry (UPLC-MS / MS) detection conditions, and preliminary raw metabolomics data were used. During the screening of differentially expressed metabolites, the variable weight value VIP ≥ 1.0 calculated by the OPLS-DA model was used as the sole criterion for screening differentially expressed metabolites between groups. Subsequent steps remained consistent with Example 1.
[0053] (2) Results and Problems: The number of differentially expressed metabolites screened increased significantly. These additional metabolites were mainly noise or non-specific changes, lacking clear distinguishing patterns among species. A clustering heatmap generated from the common differentially expressed metabolites screened from these metabolites ( Figure 15 The results showed that the cluster boundaries of samples from different species were blurred and overlapped. Furthermore, the accuracy of identification was significantly reduced, while the false positive rate increased.
[0054] (3) Conclusion: Using only VIP≥1.0 for screening introduces a large amount of noise, resulting in insufficient specificity of feature combinations and failure to achieve high-accuracy identification. This proves the necessity of the present invention using the dual standard of VIP≥1.0 and |Log2FC|≥1.0, which can accurately locate the core markers and ensure the accuracy and reliability of the method.
[0055] Verification Example 1: Accuracy of Identification of Unknown Okra Genus Plants In addition, this invention purchased 10 plants each of *Abelmoschus humilis*, *Abelmoschus humilis*, *Abelmoschus salvia*, and *Abelmoschus cambodianus* from 10 different cities. Following the methods described in steps (1)-(5) of Example 1, sample processing, data collection, and statistical analysis were performed to obtain the relative contents of the 47 differential metabolites screened in step (5), generating a heatmap. For each of the 10 plants, it was determined whether they clustered with the datasets of their respective *Abelmoschus* species measured in Example 1, thus matching them. This was to analyze the identification accuracy of *Abelmoschus humilis*, *Abelmoschus humilis*, *Abelmoschus salvia*, and *Abelmoschus cambodianus*. The results are shown in Table 2. It can be seen that the identification accuracy of the method provided by this invention is no less than 80%. The heatmap generated when identifying one unknown *Abelmoschus humilis* plant is shown in Table 2. Figure 16 .
[0056] Table 2. Results of the accuracy test for identifying unknown Okra species.
[0057] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.
Claims
1. A method for identifying closely related species of the genus *Okra*, characterized in that, Includes the following steps: S1. Sample preparation: Collect mixed stem and leaf samples of closely related species of the genus *Abelmoschus*, freeze-dry them separately, grind them into powder, add the extraction solution, extract, centrifuge, take the supernatant, filter and obtain the solution to be tested; the extraction solution includes methanol solution, and the closely related species of the genus *Abelmoschus* include *Abelmoschus quinquefolius*, *Abelmoschus quinquefolius*, *Abelmoschus quinquefolius*, and *Abelmoschus quinquefolius*. S2. Metabolomics data detection: The test solution obtained in step S1 was subjected to metabolomics detection using UPLC-MS / MS to obtain the metabolomics data of the closely related species of the Okra genus mentioned in step S1. The UPLC-MS / MS liquid chromatography conditions include: separation of the sample extract using an Agilent SB-C18 column; mobile phase A comprising formic acid solution, mobile phase B comprising acetonitrile containing formic acid, and column temperature of 30-50°C; the UPLC-MS / MS mass spectrometry conditions include: electrospray ionization source temperature of 400-600°C, and ion spray voltage of 5000-6000V in positive ion mode. S3. Processing and Analysis of Metabolic Data: The metabolomics data obtained in step S2 are first assessed by principal component analysis to evaluate the overall metabolic differences among samples and the variability within groups; an OPLS-DA analysis model is constructed and variable weights (VIPs) are calculated to screen for variables with differences among groups; finally, the fold change (FC) is calculated. S4. Screening of Common Differential Metabolites: Using the OPLS-DA analysis model from step S3, common differential metabolites of the closely related species of the *Abelmoschus* genus described in step S1 are screened as a combination of metabolic markers to distinguish these closely related species. This combination of metabolic markers includes N-ethylleucine, trimethyllysine, 1-O-caffeoyl-6-O-glucosyl-β-D-glucose, 1-O-coumaroyl-β-D-glucose, 2-phenylethylβ-primrosein, and 2-(3-... 4-Dihydroxyphenethoxy)6-feruloyl-O-glucosyl-D-xylose, 2'-hydroxygenistein, 3',4'-dihydroxy-5,6,7,8,5'-pentamethoxyflavone, 3,5-dihydroxy-7,4'-dimethoxyflavone, 5,2',5'-trihydroxy-3,7,4'-trimethoxyflavone-2'-oxo-β-D-glucoside, 5,4'-dihydroxy-6,7,8,3'-tetramethoxyflavone, 5,7-dihydroxy-3,4',6 8-Tetramethoxyflavonoids, 5,7-dihydroxy-3,6,8,3',4'-pentamethoxyflavonoids, 5,7-dihydroxy-6,3',4',5'-tetramethoxyflavonoids, 5-hydroxy-3',4',7,8-tetramethoxyflavonoids, 7,4'-dihydroxy-3,5,3'-trimethoxyflavonoids, 7,8-dihydroxy-5,6,4'-trimethoxyflavonoids, kaempferol-3-O-mannoside, andrographolide D-aglycone, Brickellin, and cat's eye grass Flavonoids, Methylcitriol, Hemigerminol, Gardeniaflavin C, Gardeniain E Malonyl Glucoside, 5,7-Dihydroxy-6,8,3',4'-Tetramethoxyflavonoids, Isorhamnetin-3-O-rutinoside, Kaempferol-3-O-(2''-O-acetyl)glucoside, Kaempferol-3-O-rutinoside, Sophora flavescens, Scutellaria baicalensis Flavonoids II, Syringin-3-O-(6''-acetyl)glucoside, Tamarixanthin-3-O-rutinoside, Tupichinol E, tamarind, 5,4'-dihydroxy-6,7,8-trimethoxyflavone, 5,6-dihydroxyrucidin-3β-alizarin, safflower yellow A, 8-hydroxycoumarin, schisandrin I, 7-methyl-5,8-dioxododecyl hydrogen sulfate, dehydrodieugenol, galactitol, glucosyl 5,8-dihydroxy-2,6-dimethyloctacarbon-2,6-dienoic acid, indole-3-lactic acid, tetrahydrofolate, and gingerol C; S5. Identification of Unknown Okra Species: After obtaining the metabolite mass spectrometry data of the unknown sample through measurement and analysis, the peak area of all substances in the chromatogram is integrated and corrected, and merged with the peak area database of the closely related species of the Okra genus measured in steps S1-S4; after statistical analysis of the data, the relative content of the common differential metabolites screened in step S4 is obtained, a heatmap is generated, and the aggregation of unknown samples in the graph is observed; the known species that aggregates with the unknown sample first is identified as the unknown sample species.
2. The identification method according to claim 1, characterized in that, The grinding in step S1 includes grinding for 1-3 minutes using a grinder at 20-50 Hz; the extraction solution includes a 60%-80% methanol-water internal standard extraction solution pre-cooled at -20℃; the vortexing conditions include vortexing once every 30 minutes for 30 seconds each time, for a total of 4-8 vortexes; the centrifugation conditions include 10000-15000 rpm for 2-5 minutes; and the filtration includes filtering the sample using a microporous membrane.
3. The identification method according to claim 2, characterized in that, The grinding in step S1 is performed by grinding at 30 Hz for 1.5 minutes using a grinder; the extraction solution is a 70% methanol-water internal standard extraction solution pre-cooled at -20℃; the vortexing is performed once every 30 minutes, each time lasting 30 seconds, for a total of 6 vortexes; the centrifugation is performed at 12000 rpm for 3 minutes; and the filtration is performed by filtering the sample using a 0.22 μm microporous membrane.
4. The identification method according to claim 1, characterized in that, The HPLC conditions for UPLC-MS / MS described in step S2 include: using an Agilent SB-C18 1.8µm, 2.1mm*100mm column to separate the sample extract; mobile phase A is ultrapure water containing 0.1% formic acid, and mobile phase B is acetonitrile containing 0.1% formic acid; the elution gradient is as follows: at 0.00 min, the proportion of phase B is 5%; within 9.00 min, the proportion of phase B increases linearly to 95% and is maintained at 95% for 1 min; from 10.00 to 11.10 min, the proportion of phase B decreases to 5% and equilibrates to 5% for 14 min; the flow rate is 0.35 mL / min; the column temperature is 40°C; and the injection volume is 2 μL.
5. The identification method according to claim 1, characterized in that, The mass spectrometry conditions for UPLC-MS / MS in step S2 include: an electrospray ion source temperature of 500°C; an ion spray voltage of 5500V in positive ion mode and -4500V in negative ion mode; ion source gas I, gas II, and curtain gas set to 50, 60, and 25 psi, respectively; a collision-induced ionization parameter set to high; QQQ scan using MRM mode; and nitrogen as the collision gas set to medium.
6. The identification method according to claim 1, characterized in that, Step S3 also includes qualitative and quantitative analysis of metabolic data; the qualitative analysis includes analysis based on the existing database MWDB using secondary spectral information; the quantitative analysis is performed using the multi-reaction monitoring mode of triple quadrupole mass spectrometry. After obtaining the signal intensity of characteristic ions in the detector, the chromatographic peaks of all substances are integrated and corrected using MultiQuant software. The peak area is used to characterize the relative content of metabolites, and finally, the integrated data of all chromatographic peak areas are exported and saved.
7. The identification method according to claim 1, characterized in that, In step S4, the OPLS-DA analysis model selects metabolites that meet the conditions of VIP≥1.0 and |Log2FC|≥1.0 as differential metabolites.
8. The combination of metabolic markers obtained by the identification method according to any one of claims 1-7, characterized in that, The combination of metabolic markers includes N-ethylleucine, trimethyllysine, 1-O-caffeoyl-6-O-glucosyl-β-D-glucose, 1-O-p-coumaryl-β-D-glucose, 2-phenylethyl-β-primroseside, 2-(3,4-dihydroxyphenethoxy)6-feruloyl-O-glucosyl-D-xylose, 2'-hydroxygenistein, and 3',4'-dihydroxy-5,6,7,8,5'-pentamethyl Oxyflavonoids, 3,5-dihydroxy-7,4'-dimethoxyflavonoids, 5,2',5'-trihydroxy-3,7,4'-trimethoxyflavonoid-2'-oxo-β-D-glucoside, 5,4'-dihydroxy-6,7,8,3'-tetramethoxyflavonoids, 5,7-dihydroxy-3,4',6,8-tetramethoxyflavonoids, 5,7-dihydroxy-3,6,8,3',4'-pentamethoxyflavonoids, 5,7- Dihydroxy-6,3',4',5'-tetramethoxyflavonoids, 5-hydroxy-3',4',7,8-tetramethoxyflavonoids, 7,4'-dihydroxy-3,5,3'-trimethoxyflavonoids, 7,8-dihydroxy-5,6,4'-trimethoxyflavonoids, kaempferol-3-O-mannoside, andrographolide D-aglycone, Brickellin, chrysanthin, methyl cypermethrin, hemi-glyphosate, genistein C. Gardeniae E malonyl glucoside, 5,7-dihydroxy-6,8,3',4'-tetramethoxyflavone, isorhamnetin-3-O-rutinoside, kaempferol-3-O-(2''-O-acetyl)glucoside, kaempferol-3-O-rutinoside, scutellarin, baicalin II, syringin-3-O-(6''-acetyl)glucoside, tamariscin-3-O-rutinoside, Tupichinol E, teosin, 5,4'-dihydroxy-6,7,8-trimethoxyflavone, 5,6-dihydroxyrucidin-3β-alizarin, safflower yellow A, 8-hydroxycoumarin, schisandrin I, 7-methyl-5,8-dioxododecyl hydrogen sulfate, dehydrodieugenol, galactitol, glucosyl 5,8-dihydroxy-2,6-dimethyloctacarbon-2,6-dienoic acid, indole-3-lactic acid, tetrahydrofolate, and gingerol C.
9. The application of the identification method according to any one of claims 1-7 or the combination of metabolic markers according to claim 8, characterized in that, The application includes the identification of closely related species of the genus Okra.
10. The application according to claim 9, characterized in that, The applications also include the identification of medicinal herb origins, the screening of germplasm resources, or the detection of variety purity during the seedling stage.
Citation Information
Patent Citations
Method for identifying day lily in Datong region based on metabonomics
CN117607325A