Identification method of red-pulp kiwi fruit variety
By using metabolomics technology and machine learning algorithms, a combination of highly specific metabolic biomarkers was screened out, and a discriminative model was constructed. This solved the problem of distinguishing red-fleshed kiwifruit varieties, enabling rapid and accurate variety identification, which is suitable for market supervision and enterprise testing.
Patent Information
- Application Number
- CN202511816338.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies make it difficult to quickly and accurately distinguish between the red-fleshed kiwifruit varieties 'Donghong', 'Hongyang', and 'Xiaohongguo'. Conventional morphological identification is highly subjective, DNA molecular marker technology is costly and time-consuming, and physicochemical index analysis cannot accurately differentiate them.
Using metabolomics technology, a combination of metabolic biomarkers was constructed. The metabolite spectra of red-fleshed kiwifruit were analyzed using ultra-high performance liquid chromatography-triple quadrupole mass spectrometry. Combined with multivariate statistical analysis and machine learning algorithms, highly specific metabolic biomarkers were screened out, and a discriminant model was constructed to distinguish varieties.
It enables accurate and reliable identification of red-fleshed kiwifruit varieties, overcomes the subjectivity of appearance identification, and has a fast testing speed, making it suitable for market supervision and rapid batch testing by enterprises.
Smart Images

Figure CN121385237A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of agricultural product identification, and particularly relates to a method for identifying a red-fleshed kiwifruit variety. BACKGROUND
[0002] With the increasing demand of consumers for fruit quality and health attributes, red-fleshed kiwifruit has become a representative of the high-end kiwifruit market due to its unique red flesh, sweet and smooth taste, and rich antioxidant substances such as anthocyanins. Among them, 'Donghong', 'Hongyang' and 'Xiaohongguo' (Ruby Red) are the main red-fleshed kiwifruit varieties in the market, which are independently bred in China. However, these three varieties are extremely similar in appearance, especially after peeling or processing, but their market positioning, price and flavor taste are significantly different.
[0003] Currently, there are many limitations in the identification method of such agricultural product varieties. Conventional morphological identification is highly dependent on experience and is greatly affected by fruit maturity, production area and cultivation conditions, and is highly subjective. Although DNA molecular marker technology is highly accurate, the detection process is complicated, time-consuming and costly, and usually needs to be completed by technical personnel in a professional laboratory, which is difficult to meet the needs of rapid and batch detection in market supervision, enterprise raw material inspection and other scenarios. Conventional physicochemical index analysis: such as soluble solids (sugar content), titratable acidity, etc. These indexes have high overlap between varieties and large dynamic change range, and cannot achieve accurate differentiation of the three highly similar varieties.
[0004] Metabolomics, as a new technology that can systematically qualitatively and quantitatively analyze all small molecule metabolites in the body, directly reflects the physiological and biochemical state of the organism and its response to the external environment. By comparing the metabolite "fingerprint" of different varieties, stable differential metabolites closely related to the genetic background of the variety can be discovered, which can serve as an irrefutable "chemical identity card".
[0005] Although metabolomics technology has been widely used in crop quality research, there is currently no research report that can provide specific and stable metabolic markers for rapid and accurate differentiation of 'Donghong', 'Hongyang' and 'Xiaohongguo', which are three high-value red-fleshed kiwifruit varieties. Developing such a method is not only a key technical requirement to solve the current market pain points, but also a powerful tool to strengthen the protection of new plant varieties and promote the high-quality development of the industry. SUMMARY
[0006] To solve the problems in the prior art, the application provides a red-fleshed kiwifruit variety identification method, which can realize identification of red-fleshed kiwifruit varieties 'Donghong', 'Hongyang' and 'Ruby Red' based on specific metabolic markers, and in particular, can realize variety differentiation by using relative and absolute contents of single or multiple metabolic markers.
[0007] To achieve the above-mentioned object, the application provides the following solutions. A red-fleshed kiwifruit variety identification method, comprising: a) sample preparation and metabolite data acquisition are performed on a to-be-tested kiwifruit sample to obtain metabolite spectrum data of the sample; b) the metabolite spectrum data of the to-be-tested sample are compared with content information of metabolic markers or each marker in a metabolic marker group to obtain relative abundance; c) a specific marker in the to-be-tested sample is compared with relative abundance of a corresponding metabolic marker in a reference sample group to construct a discrimination model or judge the variety of the to-be-tested sample.
[0008] As a preferred solution, the metabolic marker combination specifically comprises: The metabolic marker combination comprises at least three of the following 19 substances: glutamine-alanine-glycine, thiamethoxam, 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine, norpinoresinol, L-isoleucine-L-glutamic acid-L-valine, dehydrovicianoside, galangin-7-glucoside, mediodrosin glucoside, ranunculin, 3,4-dihydro-4-(4'-hydroxyphenyl)-5,7-dihydroxycoumarin glucoside, 3-hydroxy-1,2-dimethoxyxanthone-glucoside, isohypsoicin, hydroxyisomperatorin glucoside, 1,7-dihydroxy-3,4-dimethoxyxanthone-glucoside, dihydrokaempferol-3-O-glucoside, 3,4,5,2',4',6'-hexahydroxychalcone 2'-glucoside, 6-O-caffeoylajugol, cyanidin-3,5-di-O-glucoside. The content information of each marker in the metabolic marker group specifically comprises: The metabolic marker group is composed of at least one of the following three substances: 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine and cyanidin-3,5-di-O-glucoside.
[0009] As preferred, when using the metabolic marker combination, step c) is: comparing the relative content of the metabolic markers in the sample to be tested with the average relative content of the reference sample group consisting of the domestic red-fleshed kiwifruit cultivars ‘Donghong’ and ‘Hongyang’; if the relative content of 1-O-p-coumaroyl lysine in the sample to be tested is more than 15 times the reference value, or the relative content of N-p-coumaroyl hydroxyguanidinobutyrin is more than 15 times the reference value, or the relative content of cyanidin-3,5-di-O-glucoside is more than 50 times the reference value, then the sample is determined to be ‘Xiaohongguo’; otherwise, it is determined to be a domestic red-fleshed kiwifruit cultivar.
[0010] As preferred, the “more than 10 times” specifically refers to: the relative content of 1-O-p-coumaroyl lysine is 15-105 times, or the relative content of N-p-coumaroyl hydroxyguanidinobutyrin is 15-110 times, or the relative content of cyanidin-3,5-di-O-glucoside is 50-300 times As preferred, the screening of the metabolic markers includes the following steps: a) Sample preparation and metabolite extraction: Fruit samples of ‘Donghong’, ‘Hongyang’ and ‘Xiaohongguo’ were collected respectively, at least 3 biological replicates were set for each cultivar, and pre-cooled methanol water extraction solution was used for extraction; b) Metabolomics data acquisition: Ultra-high performance liquid chromatography-triple quadrupole mass spectrometry was used for analysis; c) Data processing and multivariate statistical analysis: The original data were preprocessed by peak extraction, alignment and normalization to obtain an information table containing all samples and all metabolite peaks; unsupervised principal component analysis was used to observe the natural clustering of samples and determine whether there was a separation trend between groups; supervised orthogonal partial least squares discriminant analysis (OPLS-DA) was used to construct a model to find the variables that contributed most to the separation between groups; d) Screening and identification of differential metabolites: Orthogonal partial least squares discriminant analysis (OPLS-DA) was performed on the normalized data, and metabolites with a variable projection importance (VIP) value greater than 1.0 were selected; one-way analysis of variance (ANOVA) was performed on the metabolites with a VIP value greater than 1.0, and metabolites with a P value less than 0.01 and a fold change (FC) greater than 2 or less than 0.5 in the comparison between groups were selected; common differential metabolites that simultaneously meet the conditions of VIP>1.0, P<0.01 and FC>2 or FC<0.5 in any pairwise comparison were selected as candidate metabolites; the above candidate metabolites were simultaneously input into the random forest and LASSO models, and then the intersection of the two was taken as the final metabolic marker combination for identifying red-fleshed kiwifruit cultivars; As preferred, the screening condition for screening the metabolic marker group for quickly identifying New Zealand small red fruits and domestic red-fleshed kiwifruit varieties is that the fold difference of the metabolic markers of the small red fruits and the Donghong and Hongyang is more than 10, and the metabolic markers with P<0.05 are taken as the metabolic marker group.
[0011] As preferred, the discriminant model is at least one of gradient boosting, random forest, support vector machine, K-nearest neighbor, linear discriminant analysis and naive Bayes.
[0012] As preferred, the chromatographic condition in the ultra-high performance liquid chromatography-high resolution mass spectrometry technology is that a C18 chromatographic column is used, the mobile phase is water (containing 0.1% formic acid) and acetonitrile (containing 0.1% formic acid), and gradient elution is performed.
[0013] As preferred, the mass spectrometry condition in the ultra-high performance liquid chromatography-high resolution mass spectrometry technology is that an electrospray ionization source is adopted, and full scanning is performed in positive and negative ion modes.
[0014] Compared with the prior art, the present application has the following beneficial effects: Precise and reliable: based on internal metabolites for identification, overcoming the subjectivity of appearance identification, high accuracy, and scientific and reliable results.
[0015] High specificity: the provided metabolic markers have significant content differences between varieties (up to tens to hundreds of times), which can clearly distinguish the easily confused varieties.
[0016] Flexible application: both complete marker combination-based accurate identification and single key marker-based rapid screening can be used to meet different application scenarios.
[0017] Convenient and efficient: simple pretreatment and fast detection. Especially suitable for developing detection reagent kits, convenient for market supervision and rapid batch detection of enterprises, providing technical support for industrial standardization and brand protection. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present application, the following briefly introduces the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 The flow chart of the red-fleshed kiwifruit variety identification method of the embodiments of the present application; Figure 2 The LDA model based on 19 metabolic markers; Figure 3 The LDA model confusion matrix based on 5 important metabolic markers; Figure 4 For performance comparison of different models. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Example 1 like Figure 1 As shown, this invention provides a method for identifying red-fleshed kiwifruit varieties, comprising: a) Sample preparation and metabolite data collection were performed on the kiwifruit samples to be tested to obtain their metabolite spectral data; b) Compare the metabolite spectrum data of the sample to be tested with the content information of each marker in the combination of metabolic markers or the group of metabolic markers to obtain the relative abundance. c) By constructing a discriminant model or comparing the relative abundance of a specific biomarker in the sample with the corresponding metabolic biomarker in the reference sample group, the species classification of the sample to be tested is determined; wherein the discriminant model is at least one of gradient boosting, random forest, support vector machine, K-nearest neighbor, linear discriminant analysis and Naive Bayes.
[0023] Furthermore, the combination of metabolic biomarkers specifically comprises: The metabolic marker combination comprises at least three of the following 19 substances: glutamyl-alanine-glycine, thiamethoxam, 1-O-coumaroyl lysine, N-coumaroyl hydroxyguanidine, norpinaside, L-isoleucine-L-glutamic acid-L-valine, vinorelbine, galangin-7-glucoside, medroxypyr glucoside, euryptic acid, 3,4-dihydro-4-(4'-hydroxyphenyl)-5,7-dihydroxy Coumarin glucoside, 3-hydroxy-1,2-dimethoxyxanthonone-glucoside, isocosmosin, hydroxyisoglycyrrhizin glucoside, 1,7-dihydroxy-3,4-dimethoxyxanthonone-glucoside, dihydrokaempferol-3-O-glucoside, 3,4,5,2',4',6'-hexahydroxychalcone 2'-glucoside, 6-O-caffeoyl sago alcohol, cyanidin-3,5-di-O-glucoside; The specific content information of each marker in the metabolic marker group is as follows: The metabolic marker group consists of at least one of the following three substances: 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine, and cyanidin-3,5-di-O-glucoside.
[0024] As preferred, when using the metabolic marker group, step c) is: comparing the relative content of the metabolic marker in the sample to be tested with the average relative content of the reference sample group consisting of the domestic red and red sun varieties. If the relative content of 1-O-p-coumaroyl lysine in the sample to be tested is more than 15 times the reference value, or the relative content of N-p-coumaroyl hydroxyagmatine is more than 15 times the reference value, or the relative content of cyanidin-3,5-di-O-glucoside is more than 50 times the reference value, it is determined that the sample is New Zealand small red fruit; otherwise, it is determined to be a domestic red-fleshed kiwifruit variety.
[0025] Further, the screening of the metabolic markers includes the following steps: a) Sample preparation and metabolite extraction: Fruit samples of the 'Donghong', 'Hongyang' and 'Xiaohongguo' varieties were collected respectively, with at least 3 biological replicates for each variety, and pre-cooled methanol water extraction solution was used for extraction; b) Metabolomics data acquisition: Ultra-high performance liquid chromatography-triple quadrupole mass spectrometry was used for analysis; in which, the chromatographic conditions of the ultra-high performance liquid chromatography-triple quadrupole mass spectrometry were: using a C18 chromatographic column, the mobile phase was water (containing 0.1% formic acid) and acetonitrile (containing 0.1% formic acid), and gradient elution was performed; the mass spectrometry conditions of the ultra-high performance liquid chromatography-triple quadrupole mass spectrometry were: using an electrospray ionization source, and full scan was performed in positive and negative ion modes respectively; c) Data processing and multivariate statistical analysis: The original data were preprocessed by peak extraction, alignment and normalization to obtain an information table containing all samples and all metabolite peaks; unsupervised principal component analysis was used to observe the natural aggregation of samples and determine whether there was a separation trend between groups; supervised partial least squares discriminant analysis (OPLS-DA) was used to construct a model to find the variables that contributed most to the separation between groups; d) Screening and identification of differential metabolites: The orthogonal partial least squares discriminant analysis (OPLS-DA) is performed on the standardized data, and metabolites with a variable projection importance (VIP) value greater than 1.0 are selected; single factor variance analysis (ANOVA) is performed on the metabolites with a VIP value greater than 1.0, and metabolites with a P value less than 0.01 and a difference multiple (Fold Change) greater than 2 or less than 0.5 in the comparison between groups are selected; common differential metabolites that simultaneously meet the conditions of VIP>1.0, P<0.01 and FC>2 or FC<0.5 in any two comparisons are selected as candidate metabolites; the above candidate metabolites are simultaneously input into the random forest and LASSO models for running, and then the intersection of the two is taken as the final metabolic marker combination for identifying red kiwifruit varieties; Further, the screening condition of the metabolic marker group for quickly identifying New Zealand small red fruits and domestic red meat kiwifruit varieties is that the difference multiples of the small red fruits and Donghong and Hongyang are all above 10, and the metabolites with P<0.05 are taken as the metabolic marker group.
[0026] As an embodiment of the embodiment of the present application, the screening and identification of the metabolic marker are specifically: 1. Sample collection Fresh fruits of Donghong, Hongyang and small red fruits are collected, and each species comes from 8 different batches or origins, and the pulp is taken for detection.
[0027] 2. Sample metabolite extraction The three varieties of kiwifruit pulp are quickly frozen in liquid nitrogen and then placed in a freeze dryer (Scientz-100F) for vacuum freeze drying for about 60 hours, and then the dried pulp is ground (30 Hz, 1.5 minutes) into powder by a grinder (MM 400, Retsch). 50 mg of sample is weighed by an electronic balance (MS105DM), and 1200 μL of pre-cooled 70% methanol containing internal standard extraction solution is added. Vortex every 30 minutes, each for 30 seconds, a total of 6 times. After centrifugation (12000 rpm for 3 min), the supernatant is aspirated, the sample is filtered with a microporous filter membrane (0.22 μm pore size), and is stored in a sample bottle for UPLC-MS / MS analysis.
[0028] 3. UHPLC-MS / MS detection and analysis Chromatographic acquisition conditions: The chromatographic column type is Agilent SB-C18 chromatographic column, 1.8 μm, 2.1 mm * 100 mm, the aqueous phase (A) is 0.1% formic acid aqueous solution, the organic phase (B) is acetonitrile solution containing 0.1% formic acid, gradient elution: 0-9 min, 95% A→5% A; 9-10 min, 5% A; 10-11 min, 5% A→95% A; 11 min-14 min, 95% A, flow rate: 0.35 mL / min, column temperature 40°C, injection volume 2 μL.
[0029] Mass spectrometry conditions: electrospray ion source (ESI) temperature 500°C; ion spray voltage (IS) 5500 V (positive ion mode) / -4500 V (negative ion mode); ion source gas I (GSI), gas II (GSII) and curtain gas (CUR) are set to 50, 60 and 25 psi respectively, and the collision-induced ionization parameter is set to high. QQQ scanning uses MRM mode, and the collision gas (nitrogen) is set to medium. Through further optimization of declustering potential (DP) and collision energy (CE), the DP and CE of each MRM ion pair are completed. According to the metabolites eluted in each period, a specific group of MRM ion pairs is monitored in each period.
[0030] 4. Quality control Quality control samples (QCs) are prepared by mixing sample extracts for analyzing the reproducibility of samples under the same processing method. During instrument analysis, generally one quality control sample is inserted in every 10 detection and analysis samples to monitor the reproducibility of the analysis process. By overlapping and analyzing the total ion chromatogram (TIC) of different quality control QC sample mass spectrometry detection and analysis, the reproducibility of metabolite extraction and detection, i.e. technical repeatability, can be judged. The reliability and stability of the analysis are evaluated together using Pearson correlation analysis of QC samples and coefficient of variation (CV) distribution chart of all samples, principal component analysis.
[0031] 5. Processing and analysis of metabolomics data The obtained mass spectrometry raw data is preprocessed, including chromatographic peak recognition and extraction, peak alignment, retention time correction and other steps. Appropriate algorithms are used to process and correct missing values to ensure data quality. The detected metabolites are identified by comparing the metabolite database and combining the secondary mass spectrometry information. Finally, a standardized data matrix containing the relative content of each metabolite is obtained.
[0032] 6. Screening of metabolic markers Step 1: Multivariate statistical analysis of data matrix. Before performing orthogonal partial least squares discriminant analysis (OPLS-DA), the data is logarithmically transformed with base 2 and centered, the variable projection importance (VIP) is calculated, and metabolites with VIP values greater than 1.0 are selected.
[0033] Step 2: Univariate statistical analysis. Perform one-way analysis of variance (ANOVA) on the metabolites with VIP values greater than 1.0, select metabolites with P values less than 0.01 and fold change (FC) greater than 2 or less than 0.5 for comparison between groups.
[0034] Step 3: Candidate metabolites. Select common differential metabolites that simultaneously satisfy the conditions of VIP > 1.0, P < 0.01, and FC > 2 or FC < 0.5 in pairwise comparisons of the three groups as candidate metabolites.
[0035] Step 4: Input the above candidate metabolites into the random forest and LASSO models, then take the intersection of the two as the final combination of marker metabolites. The present application finally screens out 19 metabolic markers, including the following substances: Glutamine-alanine-glycine, thiamethoxam, 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine, norpinoresinol, L-isoleucine-L-glutamic acid-L-valine, dehydrovinaloside, galangin-7-glucoside, mediodrosin glucoside, ranunculin, 3,4-dihydro-4-(4'-hydroxyphenyl)-5,7-dihydroxycoumarin glucoside, 3-hydroxy-1,2-dimethoxyxanthone-glucoside, isovectorol glucoside, hydroxyisomelitoxin glucoside, 1,7-dihydroxy-3,4-dimethoxyxanthone-glucoside, dihydrokaempferol-3-O-glucoside, 3,4,5,2',4',6'-hexahydroxychalcone 2'-glucoside, 6-O-caffeoylerythrocentaurin, cyanidin-3,5-di-O-glucoside.
[0036] Step 5: Further screening of the 19 metabolites obtained in the fourth step. The screening condition is that the difference fold of small red fruit and east red and red sun is more than 10, and the metabolite with P < 0.05 is selected as the metabolic marker group. Finally, three substances are obtained as the marker metabolic group for distinguishing domestic and foreign red kiwifruit varieties: 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine, and cyanidin-3,5-di-O-glucoside.
[0037] As an embodiment of the present application, the LDA identification model based on 19 metabolic markers is constructed and applied, specifically: 1. Model construction A linear discriminant analysis (LDA) model was constructed using the 19 metabolite relative peak area data of the 24 samples (8 samples per variety) obtained in Example 1. Before modeling, all metabolite features were standardized by Z-score.
[0038] 2. Discriminant function and model performance Based on the standardized data, the following linear discriminant function was established. Using leave-one-out cross-validation, the model training set accuracy reached 100%.
[0039] As shown in Table 1, a linear discriminant function based on standardized metabolite data was established: Discriminant function of Donghong: Score_DH = -13056.643167 - 62.198581 × X1 + 356.861081 × X2 -34177.761326 × X3 - 4109.871271 × X4 - 5088.150988 × X5 + 825.048523 × X6- 4037.396605 × X7 + 1561.182832 × X8 + 1561.182832 × X9 + 1887.397678 ×X10 + 403.140103 × X11 + 403.140103 × X12 + 10463.512761 × X13 -5413.329762 × X14 - 1662.022051 × X15 + 5255.117592 × X16 - 7584.581821 ×X17 + 1004.688895 × X18 + 7935.569361 × X19 Discriminant function of Hongyang: Score_HY = -8540.336847 - 464.673503 X1 + 296.572006 X2 - 23925.466742 X3 - 2633.605431 X4 - 1006.275817 X5 + 279.739372 X6 - 2382.022906 X7 + 1592.388934 X8 + 1592.388934 X9 + 2089.182200 X10 + 579.918356 X11 + 579.918356 X12 + 6600.525289 X13 - 7941.202543 X14 - 1770.098317 X15 + 2056.291227 X16 - 7596.355817 X17 + 914.060092 X18 + 6095.411042 X19 Discriminant function for small red fruits: Score_XHG = -41025.533272 + 526.872084 X1 - 653.433087 X2 + 58103.228068 X3 + 6743.476702 X4 + 6094.426805 X5 - 1104.787895 X6 + 6419.419511 X7 - 3153.571766 X8 - 3153.571766 X9 - 3976.579878 X10 - 983.058459 X11 - 983.058459 X12 - 17064.038050 X13 + 13354.532305 X14 + 3432.120368 X15 - 7311.408818 X16 + 15180.937638 X17 - 1918.748987 X18 - 14030.980403 X19 Table 1
[0040] 3. Analysis of the contribution of key markers By analyzing the standardized coefficients of each metabolite in the LDA model, the five markers with the greatest contribution to the variety discrimination and their contribution direction were identified, as shown in Table 2. Figure 2 Table 2
[0041] 1-O-p-Coumaroyl lysine: 74128.35 (positive contribution) Isoorientin: -23504.57 (negative contribution) 3,4,5,2',4',6'-Hexahydroxychalcone 2'-glucoside: 22097.09 (positive contribution) Hydroxyisomelin glycoside: 17607.60 (positive contribution) Cyanidin-3,5-di-O-glucoside: -16503.00 (negative contribution) 4. Unknown sample identification An unknown kiwifruit sample of a variety was detected, and the standardized data was substituted into the discriminant function to calculate the score, and the probability belonging to each category was calculated by the Softmax function. The results showed that the sample was accurately identified as 'Xiaohongguo', and the classification probability was 1.000.
[0042] The LDA identification method provided in this embodiment not only realizes 100% accurate variety identification, but also establishes a complete discriminant function system, providing a reliable method for rapid and automatic identification of three red-fleshed kiwifruit varieties.
[0043] As an embodiment of the present application, the simplified LDA model based on 5 key metabolic markers was verified, specifically: This example aims to verify whether using fewer key markers can still maintain high identification accuracy.
[0044] Model construction: only the data of the 5 key metabolic markers identified in embodiment 2 (1-O-p-coumaroyl lysine, isoorientin, etc.) was used to reconstruct the LDA model.
[0045] As shown in Table 2, linear discriminant functions based on standardized metabolite data were established: East red variety classification function: Score_DH = -1126.690040 - 2204.370918 × Y1 + 1297.853251 × Y2 -447.376309 × Y3 - 362.345331 × Y4 - 1223.635574 × Y5 Red sun variety classification function: Score_HY = -819.664069 - 1811.004391 × Y1 + 1161.863917 × Y2 -393.013495 × Y3 - 268.535349 × Y4 - 1194.024339 × Y5 Small Red Fruit Variety Classification Function: Score_XHG = -3843.375072 + 4015.375309 × Y1 - 2459.717168 × Y2 +840.389804 × Y3 + 630.880680 × Y4 + 2417.659913 × Y5 Table 2
[0046] Model performance: Leave-one-out cross-validation shows that the simplified model also achieves 100% accuracy on the training set, and the confusion matrix shows that all samples are correctly classified (corresponding to...). Figure 3 ).
[0047] Application Validation: Predictions were made on the same unknown sample, and the results were completely consistent with those in Example 2, accurately identifying it as 'small red fruit'. This demonstrates that the simplified model significantly reduces detection costs and complexity while maintaining accuracy.
[0048] As one embodiment of the present invention, the general identification method based on multiple machine learning models and the identification of key markers specifically includes: 1. Data and Models Using data from 19 metabolic biomarkers across 24 samples, six algorithms were employed to construct classifiers: Gradient Boosting, Random Forest, Support Vector Machine (SVM), K-Nearest Neighbors (K-NN), Linear Discriminant Analysis (LDA), and Naive Bayes.
[0049] 2. Model Performance Comparison All models were evaluated using 5-fold stratified cross-validation. Figure 4 As shown, all six models achieved 100% discrimination accuracy, and the confusion matrices all exhibited a perfect diagonal distribution, indicating that the metabolic biomarker combination of the present invention has extremely strong robustness and discriminative ability under different algorithms.
[0050] 3. Screening of key biomarkers Eight key metabolites were successfully screened, and these metabolites all showed high importance in the six machine learning models, as shown in Table 3. This result is consistent with the simplified model in Example 3. Table 3
[0051] 3. Prediction of unknown samples When predicting the same unknown sample, all models produce completely consistent results, accurately identifying it as 'small red fruit'.
[0052] As an embodiment of the present application, rapid identification based on a single metabolic marker, specifically: Method and principle: rapid identification based on the stable fold relationship identified in Example 4. Judgment and verification: calculate the ratio of the relative content of 1-O-p-coumaroyl lysine in the sample to be tested to the reference value (average relative content of domestic varieties). When the ratio is more than 15 times, it can be determined that the sample belongs to the 'Xiaohongguo' variety. The validation set samples of known varieties were detected, and the fold relationship was used for judgment. The results showed that the screening accuracy for 'Xiaohongguo' was more than 98%. This method is suitable for scenarios that require rapid and low-cost differentiation of domestic and foreign varieties.
[0053] The above-described embodiments are only descriptions of the preferred modes of the present application and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements to the technical solutions of the present application made by those of ordinary skill in the art shall fall within the scope of protection determined by the claims of the present application.
Claims
1. A method of identifying a red-fleshed kiwifruit variety, characterized by, Comprise: a), sample preparation and metabolite data collection on the test kiwifruit sample, to obtain its metabolite spectrum data; b), the metabolite spectrum data of the test sample is compared with the content information of the metabolic marker combination or each marker in the metabolic marker group, to obtain the relative abundance; c) By constructing a discriminant model or comparing the relative abundance of specific markers in the sample with the corresponding metabolic markers in the reference sample group, the variety of the test sample is judged.
2. The red-fleshed kiwifruit variety identification method according to claim 1, wherein the metabolic marker combination comprises: The metabolic marker combination comprises at least three of the following 19 substances: glutamine-alanine-glycine, thiamethoxam, 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine, norpiopodophyllin, L-isoleucine-L-glutamic acid-L-valine, dehydrovinaloin, galangin-7-glucoside, mediodrosin glucoside, chamaecypanin, 3,4-dihydro-4-(4'-hydroxyphenyl)-5,7-dihydroxy coumarin glucoside, 3-hydroxy-1,2-dimethoxy xanthone-glucoside, isohypsoicin, hydroxyisomichemica glucoside, 1,7-dihydroxy-3,4-dimethoxy xanthone-glucoside, dihydrokaempferol-3-O-glucoside, 3,4,5,2',4',6'-hexahydroxychalcone 2'-glucoside, 6-O-caffeoyl acteoside, cyanidin-3,5-di-O-glucoside; The content information of each marker in the metabolic marker group is: The metabolic marker group consists of at least one of the following three substances: 1-O-p-coumaroyl lysine, N-p-coumaroyl hydroxyagmatine, and cyanidin-3,5-di-O-glucoside. When using the metabolic marker group, step c) is: comparing the relative content of the metabolic markers in the test sample with the average relative content of the reference sample group composed of the domestic Donghong and / or Hongyang varieties; if the relative content of any of the metabolic markers in the test sample is more than 10 times the average relative content of the reference sample group, it is determined that the sample is New Zealand Xiaohongguo.
3. The method of claim 2, wherein the red-fleshed kiwifruit variety is identified by the presence of the nucleotide sequence of SEQ ID NO:
1. The "more than 10 times" specifically refers to: the relative content of 1-O-p-coumaroyl lysine is 15-105 times, or the relative content of N-p-coumaroyl hydroxyagmatine is 15-110 times, or the relative content of cyanidin-3,5-di-O-glucoside is 50-300 times.
4. The method of claim 3, wherein, The screening of the metabolic markers comprises the following steps:
5. The method of claim 4, wherein the red-fleshed kiwifruit variety is identified by the presence of the nucleotide sequence of SEQ ID NO:
1. a) Sample preparation and metabolite extraction: Collect fruit samples of 'Donghong', 'Hongyang' and 'Xiaohongguo' varieties, with at least 3 biological replicates for each variety, and use pre-cooled methanol water extraction solution for extraction; b) Metabolomics data collection: Use ultra-high performance liquid chromatography-triple quadrupole mass spectrometry for analysis; c) Data processing and multivariate statistical analysis: The original data is pre-processed by peak extraction, alignment and normalization to obtain an information table containing all samples and all metabolite peaks; unsupervised principal component analysis is used to observe the natural aggregation of samples and determine whether there is a separation trend between groups; supervised orthogonal partial least squares discriminant analysis (OPLS-DA) is used to construct a model to find the variables that contribute most to the distinction between groups; d) Differential metabolite screening and identification: Orthogonal partial least squares discriminant analysis (OPLS-DA) is performed on the normalized data, and metabolites with a variable projection importance (VIP) value greater than 1.0 are selected; single-factor analysis of variance (ANOVA) is performed on the metabolites with a VIP value greater than 1.0, and metabolites with a P value less than 0.01 and a difference ratio (Fold Change) greater than 2 or less than 0.5 in the comparison between groups are selected; common differential metabolites that simultaneously satisfy the conditions of VIP>1.0, P<0.01 and FC>2 or FC<0.5 in any two-way comparison are selected as candidate metabolites; the above candidate metabolites are simultaneously input into the random forest and LASSO models, and then the intersection of the two is taken as the final metabolic marker combination for identifying red kiwifruit varieties.
6. The method of claim 5, wherein the red-fleshed kiwifruit variety is identified by the presence of the nucleotide sequence of SEQ ID NO:
1. The screening condition for screening the metabolic marker group for rapidly identifying New Zealand small red fruit and domestic red-fleshed kiwifruit varieties is that the metabolites with a difference ratio of more than 10 and a P value of less than 0.05 between small red fruit and Donghong and Hongyang are selected as the metabolic marker group.
7. The method of claim 6, wherein the red-fleshed kiwifruit variety is identified by the presence of the nucleotide sequence of SEQ ID NO:
1. The discriminant model is at least one of gradient boosting, random forest, support vector machine, K-nearest neighbor, linear discriminant analysis and naive Bayes.
8. The method of claim 7, wherein the red-fleshed kiwifruit variety is identified by the presence of the nucleotide sequence of SEQ ID NO:
1. In the ultra-high performance liquid chromatography-high resolution mass spectrometry technology, the chromatographic conditions are as follows: a C18 chromatographic column is used, the mobile phase is water containing 0.1% formic acid and acetonitrile containing 0.1% formic acid, and gradient elution is performed.
9. The method of claim 8, wherein the red-fleshed kiwifruit variety is identified by the presence of the nucleotide sequence of SEQ ID NO:
1. In the ultra-high performance liquid chromatography-triple quadrupole mass spectrometry technology, the mass spectrometry conditions are as follows: an electrospray ionization source is used, and full scan is performed in positive and negative ion modes, respectively.