Prediction model of garlic variety and production place and training method and related application thereof

By detecting the content of sulfide compounds in garlic and combining machine learning algorithms to establish a prediction model, the accuracy and economic problems of traceability of garlic varieties and origins are solved, and garlic quality monitoring and industrial upgrading are achieved.

CN120432029APending Publication Date: 2025-08-05INST OF QUALITY STANDARD & TESTING TECH FOR AGRO PROD OF CAAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493398.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and economically trace garlic varieties and origins. Sensory assessment depends on expert experience and is easily subjectively affected, DNA detection costs are high, and chemical composition analysis lacks specificity.

Method used

Thioether compounds are used as markers and a prediction model is established in combination with machine learning algorithms. By detecting the content of sulfioether compounds in garlic and updating parameters using machine learning models, accurate traceability of garlic varieties and origin is achieved.

Benefits of technology

It has achieved rapid and accurate traceability of garlic varieties and origins, provided a new way to monitor garlic quality, and improved the specificity and operability of the garlic industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432029A_ABST
    Figure CN120432029A_ABST
Patent Text Reader

Abstract

The invention discloses a garlic variety and production place prediction model and a training method and related application thereof, and relates to the field of biology. According to the method, the thioether compound is used as a marker, a machine learning algorithm is combined, a model capable of tracing the garlic variety and / or the producing area is established, compared with a traditional method, the scheme has better specificity and operability, rapid and accurate tracing of garlic products can be achieved, a new way is provided for garlic quality monitoring, and the method is suitable for large-scale popularization and application. The positive influence is also brought to the variety protection and industrial upgrading of the garlic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the biological field, and in particular, to a prediction model for garlic varieties and origins, a training method thereof, and related applications. Background Art

[0002] Garlic (Allium sativum L.) is the underground bulb of a plant in the Liliaceae family and the Allium genus. As a widely used herbaceous plant, garlic has important application values in the fields of food, medicine, and health care. Due to the large variety and wide distribution of garlic, garlic from different varieties and origins varies in its flavor, nutrition, and medicinal activity. These differences not only affect the market positioning and economic value of garlic but also arouse consumers' concerns about the quality and authenticity of garlic. Therefore, how to accurately trace the variety and origin of garlic has become an urgent problem to be solved in the industry.

[0003] Currently, garlic traceability mainly relies on means such as sensory evaluation, DNA detection, and chemical composition analysis. However, sensory evaluation usually depends on the experience of experts and is easily affected by subjective factors; while DNA detection is relatively accurate, it is costly and difficult to be widely used in the industrial chain; existing chemical composition analysis methods lack specificity and often have difficulty effectively identifying garlic from different origins and varieties.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a prediction model for garlic varieties and origins, a training method thereof, and related applications.

[0006] The present invention is implemented as follows:

[0007] In a first aspect, an embodiment of the present invention provides the use of a reagent for detecting markers in garlic in the preparation of a product for predicting garlic varieties and / or origins, wherein the markers include sulfide compounds.

[0008] In a second aspect, an embodiment of the present invention provides a training method for a prediction model for garlic varieties and / or origins, which includes: obtaining the detection results of the marker contents in the training samples and their corresponding annotation results; the markers are the markers described in the foregoing embodiments, and the annotation results are labels representing garlic varieties and / or origins; inputting the detection results of the training samples into a pre-constructed machine learning model to obtain a prediction result; and updating the parameters of the model based on the prediction result and the annotation result to obtain a prediction model.

[0009] In a third aspect, an embodiment of the present invention provides a method for training a prediction model for garlic varieties and / or origins, comprising: obtaining detection results of marker contents in training samples and corresponding annotation results thereof; the markers are the markers described in the aforementioned embodiments, and the annotation results are labels representing garlic varieties and / or origins; inputting the detection results of the training samples into a pre-built machine learning model to obtain prediction results; and updating model parameters based on the prediction results and the annotation results to obtain a prediction model.

[0010] In a fourth aspect, an embodiment of the present invention provides a device for predicting garlic varieties and / or origins, comprising: an acquisition module for obtaining a detection result of a marker content in a test sample; the marker is the marker described in the aforementioned embodiment; and a prediction module for inputting the detection result of the test sample into a prediction model trained by the training method described in the aforementioned embodiment to obtain a prediction result for the test sample.

[0011] In a fifth aspect, an embodiment of the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a program. When the program is executed by the processor, the processor implements the method for tracing garlic varieties and / or origins described in the aforementioned embodiment.

[0012] In a sixth aspect, an embodiment of the present invention provides a computer-readable medium having a computer program stored thereon. When the computer program is executed by a processor, the method for tracing the garlic varieties and / or origins described in the aforementioned embodiment is implemented.

[0013] The present invention has the following beneficial effects:

[0014] The embodiments of the present invention use sulfide compounds as markers and combine them with machine learning algorithms to establish a model capable of tracing garlic varieties and / or origins. Compared with traditional methods, this solution has better specificity and operability, can achieve rapid and accurate traceability of garlic products, provides a new approach to garlic quality monitoring, and also has a positive impact on garlic variety protection and industrial upgrading. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 The corresponding diagram of the DAS content in garlic training set and test set; green is the training set sample, and red is the test set sample;

[0017] Figure 2 Corresponding diagram of the training set and test set for the AMDS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0018] Figure 3 Corresponding diagram of the training set and test set for the MPDS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0019] Figure 4 Corresponding diagram of the training set and test set for the DMTS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0020] Figure 5 Corresponding diagram of the training set and test set for the DADS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0021] Figure 6 Corresponding diagram of the training set and test set for the DATTS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0022] Figure 7 Corresponding diagram of the training set and test set for the AMS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0023] Figure 8 Corresponding diagram of the training set and test set for the DATS content in garlic; among them, the green ones are the training set samples and the red ones are the test set samples;

[0024] Figure 9 Confusion matrix of the random forest model (stratified test set);

[0025] Figure 10 Confusion matrix of the KNN model (stratified test set);

[0026] Figure 11 Heat map of feature importance and relative concentration;

[0027] Figure 12 Confusion matrix of the random forest model (stratified test set) with DAS as the feature variable;

[0028] Figure 13 Confusion matrix of the random forest model (stratified test set) with DAS and AMDS as the feature variables. Specific implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described clearly and completely below. For those not specified in the embodiments, they are carried out according to conventional conditions or conditions recommended by the manufacturer. For reagents or instruments not specified by the manufacturer, they are all conventional products that can be obtained through commercial purchase.

[0030] In view of the existing technical problems, the inventors of the present application found that sulfide compounds can be used as traceability markers for garlic varieties and / or origins. By detecting the content of sulfide compounds in garlic, accurate traceability of garlic varieties and / or origins can be achieved.

[0031] On the one hand, the embodiments of the present invention provide an application of a reagent for detecting markers in garlic in the preparation of a product for predicting garlic varieties and / or origins, where the markers include sulfide compounds.

[0032] In some embodiments, the sulfide compounds include any one or more of: diallyl sulfide (DAS), methyl allyl disulfide (AMDS), methyl propyl disulfide (MPDS), dimethyl trisulfide (DMTS), diallyl disulfide (DADS), diallyl tetrasulfide (DATTS), methyl allyl sulfide (AMS), and diallyl trisulfide (DATS).

[0033] In some embodiments, the sulfide compounds include: DAS.

[0034] In some embodiments, the sulfide compounds include: DAS and DATTS. In some embodiments, the sulfide compounds further include any one or more of: AMDS, MPDS, DMTS, DADS, AMS, and DATS.

[0035] In some embodiments, the sulfide compounds include: DAS, AMDS, MPDS, DMTS, DADS, DATTS, AMS, and DATS.

[0036] On the other hand, the embodiments of the present invention provide a method for training a prediction model for garlic varieties and / or origins, which includes:

[0037] Obtaining the detection results of the marker content in the training samples and their corresponding annotation results; the markers are the markers described in any of the foregoing embodiments, and the annotation results are labels representing garlic varieties and / or origins;

[0038] Inputting the detection results of the training samples into a pre-constructed machine learning model to obtain prediction results;

[0039] The model parameters are updated based on the prediction results and the annotation results to obtain a prediction model.

[0040] In some embodiments, the machine learning model includes any one of logistic regression, support vector machine (SVM), random forest, gradient boosting classifier, K-nearest neighbor (KNN), decision tree and naive Bayes.

[0041] It is understandable that the categories and number of training samples can be routinely selected by those skilled in the art, and the number of total training samples and training samples of various categories (such as different varieties and origins) can be ≥ any value among 10, 50, 100, 200, 300, 400 and 500 or a range between any two of them.

[0042] In some embodiments, the label may be a character or a string of characters. The content of the predicted result corresponds to the content of the labeled result.

[0043] In another aspect, an embodiment of the present invention provides a method for tracing garlic varieties and / or origins, comprising:

[0044] Obtaining a detection result of a marker content in a sample to be tested; the marker is a marker described in any of the aforementioned embodiments;

[0045] The detection result of the sample to be tested is input into the prediction model trained by the training method described in any of the above embodiments to obtain the prediction result of the sample to be tested.

[0046] In another aspect, an embodiment of the present invention provides a device for predicting garlic varieties and / or origins, comprising:

[0047] An acquisition module, configured to obtain a detection result of a marker content in a sample to be tested; the marker is a marker described in any of the aforementioned embodiments;

[0048] The prediction module is used to input the detection result of the sample to be tested into the prediction model trained by the training method described in any of the above embodiments to obtain the prediction result of the sample to be tested.

[0049] The modules described in the embodiments of the present invention may be stored in a memory in the form of software or firmware or embedded in the operating system (OS) of the electronic device provided herein, and may be executed by a processor in the electronic device. Furthermore, the data and program code required to execute the modules may be stored in the memory.

[0050] On the other hand, an embodiment of the present invention provides an electronic device, which includes a processor and a memory. The memory is used to store a program, and when the program is executed by the processor, the processor implements the traceability method for garlic varieties and / or origins described in any of the foregoing embodiments.

[0051] The electronic device may include a memory, a processor, a bus, and a communication interface. The memory, the processor, and the communication interface are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more buses or signal lines.

[0052] The memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0053] The processor can be an integrated circuit chip with signal processing capabilities. The processor 120 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0054] The electronic device can be a server, a cloud platform, a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, etc. Therefore, the embodiments of the present application do not limit the types of electronic devices.

[0055] In addition, an embodiment of the present invention also provides a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the traceability method for garlic varieties and / or origins described in any of the foregoing embodiments.

[0056] In some embodiments, the computer-readable medium can be a general storage medium, such as a removable disk, a hard disk, etc.

[0057] The features and performance of the present invention will be further described in detail below in conjunction with embodiments.

[0058] Embodiment 1

[0059] 1. Sample collection method:

[0060] In the embodiments of the present invention, a total of 23 varieties of garlic samples were collected, from Henan Province (HN, 4), Yunnan Province (YN, 2), Shandong Province (SD, 3), Gansu (GS, 3), Heilongjiang (HLJ, 3), Shaanxi Province (SX, 1), Jiangsu Province (JS, 5) and Jiangxi Province (JX, 2). All samples were harvested from July to September 2024. 30 garlic bulbs were taken from each sample.

[0061] Take garlic bulbs, remove the dust on the surface, wrap them with tin foil after peeling off one layer of the bulb skin, and store them in a -20°C refrigerator for later use. Before detection, peel off the bulb skin and the inner skin of the bulb bud, and use a grinder for homogenization to make garlic homogenate for later use.

[0062] Pretreatment: Accurately weigh 1.000 g (±0.001 g) of garlic homogenate into a 50 mL centrifuge tube, add 10 mL of extraction solvent (dichloromethane: petroleum ether = 5:5), vortex for 1 min with a parallel vortex mixer, shake on a shaker at 25°C (300 rpm) for extraction for 2.5 h, centrifuge the mixed solution at 4°C and 4000 rpm for 10 min, take 5 mL of the supernatant, add 1.5 g of anhydrous magnesium sulfate for drying treatment, centrifuge the mixed solution at 4°C and 4000 rmp for 10 min, and filter through a 0.20 μm PTFE filter membrane into an injection vial to make a test sample for on-machine analysis.

[0063] 2. Instrument conditions

[0064] The chromatographic conditions are as follows:

[0065] Chromatographic column: SH-Rxi-5Sil MS (Shimadzu Corporation, Japan), column length 30 m, inner diameter 0.25 mm, film thickness 0.25 μm; inlet temperature: 280 °C; carrier gas: helium; carrier gas flow rate: 1.59 mL / min; carrier gas pressure: 94.1 kPa; split mode: splitless; injection mode: high-pressure injection, pressure 250 kPa, injection time: 1 min; temperature program: first hold at 50 °C for 1 min, increase the temperature at a rate of 10 °C / min to 100 °C, then increase the temperature at a rate of 15 °C / min to 200 °C, and finally increase the temperature at a rate of 30 °C / min to 280 °C; injection volume: 1 μL.

[0066] The mass spectrometry conditions are as follows:

[0067] Ion source: electron impact (EI) source; scan mode: multiple reaction monitoring (MRM) mode; ion source temperature (Temperature): 230 °C; solvent delay: 1.05 min; the ion pairs and collision energy (CE) and other plasma characteristic mass spectrometry conditions of each compound are shown in Table 1.

[0068] Table 1 Names, retention times, MRM ion pairs and mass spectrometry detection parameters of 8 sulfide compounds

[0069]

[0070] 3. Data analysis:

[0071] Use the Shimazu GC-QQQ-MS post-run analysis software for data acquisition and processing. Compare the retention time and mass spectrometry with those in the NIST 17 mass spectrometry library (2017 version), and use a matching score higher than 70% as the acceptable standard to identify sulfide compounds, and perform quantitative analysis through a standard curve.

[0072] 4. Establishment and application of the prediction model

[0073] 4.1. Verify whether the random forest is applicable to predict based on the contents of DAS, AMDS, MPDS, DMTS, DADS, DATTS, AMS and DATS.

[0074] Take each compound as the target variable and the other 7 compounds as the characteristic variables respectively, and establish a random forest regression model with DAS, AMDS, MPDS, DMTS, DADS, DATTS, AMS and DATS as the target variables. Take the output DAS random forest regression model as an example, use AMDS, MPDS, DMTS, DADS, DATTS, AMS and DATS as the characteristic variables and DAS as the target variable. Split the data of the entire dataset, and split the dataset into a training set and a test set respectively.

[0075] Training-test split ratio: The data is divided in an 80 / 20 ratio, which means 80% of the data is used to train the model and the remaining 20% is used for testing. Random state: A random state is used to ensure reproducibility. This means that each time the model is run, the same split is applied, resulting in consistent results. Data splitting process: The training-test split is performed randomly within this ratio, ensuring a mix of data points in the training and test sets. In the figure, the green dots represent the training set predictions, showing how well the model fits the data it was learned from. The red dots represent the test set predictions, demonstrating the model's ability to predict data not encountered during training. Red dashed line: Represents the ideal line where the predictions exactly match the actual values. The model is evaluated using the Mean Squared Error (MSE) and R 2 The model is evaluated. MSE measures the average squared difference between the actual and predicted values. A lower MSE indicates that the predictions are closer to the actual values, reflecting higher model accuracy. A value of MSE close to 0 is ideal as it means the predictions are very close to the actual values. For example, if the MSE on the test set is much lower compared to another model or a baseline, it indicates good prediction performance. R 2 Represents the proportion of the variance in the target variable that the model can explain. It ranges from 0 to 1, and a value closer to 1 indicates that the model explains more of the data variability. An R 2 value close to 1 indicates a good fit, meaning the model captures most of the variation in the data. For example, an R 2 of 0.9 means that 90% of the variance in the target variable is explained by the model. The correspondence diagram between the training set and the test set in the prediction models for the contents of 8 sulfide compounds in garlic is shown in Figures 1 to 8 , and the MSE and R 2 are shown in Table 2.

[0076] Table 2 Evaluation indicators related to the random forest regression prediction model for the contents of 8 sulfide compounds in garlic

[0077] Compound Test set MSE <![CDATA[Test set R 2 > Training set MSE <![CDATA[Training set R 2 > DAS 1.063 0.811 0.138 0.978 AMDS 9.316 0.951 2.510 0.983 MPDS 0.0001286 0.979 0.0000785 0.988 DMTS 0.0441 0.979 0.0203 0.985 DADS 8.495 0.892 2.495 0.972 DATTS 1012.717 0.884 199.736 0.989 AMS 103.99 0.990 58.458 0.991 DATS 6.130 0.921 1.851 0.983

[0078] The data in the table shows that the overall performance of the prediction results using the random forest regression model based on the contents of 8 sulfide compounds in garlic is good. Among them, the predictions of MPDS and DMTS are the most accurate. Their MSEs for the test set and prediction set are only 0.0001286, 0.0000785 and 0.0441, 0.0203 respectively, and the R 2 is greater than 0.979, indicating that the prediction error of the model for these two compounds is very small and the explanatory ability is very strong. For DAS, DATS, AMDS, and AMS, the model also shows high prediction accuracy, and the R 2The values are all above 0.811. However, the prediction effects of DATTS and DADS are relatively poor. Although the MSE value of DATTS is relatively high, this does not affect the overall evaluation of the model. In model evaluation, although MSE is an important indicator to measure the prediction error of the model, R 2 provides important information on the model's interpretability. According to the data in the table, the R 2 values of DATTS are 0.884 and 0.989, indicating that the model can still explain most of the variability. In addition, the model performs well in predicting other several sulfide compounds. Especially, the R 2 values of MPDS and DMTS are both greater than 0.979, showing the strong interpretability of the model as a whole. Therefore, despite the relatively high MSE of DATTS, the model can still provide reliable predictions for the nutritional value of sulfide compounds in garlic. The R 2 values of the contents of 8 sulfide compounds in the prediction models of 23 garlic varieties are between 0.811 and 0.991, indicating good prediction ability of the prediction models. At the same time, models can also be established for individual compounds. It is necessary to simultaneously meet the conditions that there is no strong correlation between compounds, and when the MSE is closer to 0 and the R2 is closer to 1, single-compound models can be established. For the establishment of multi-compound models, it is considered when there is a strong correlation between compounds and when the MSE is closer to 0 and the R 2 is closer to 1.

[0079] 4.2. Compare the performance of two machine learning models, random forest and KNN, in predicting the contents of 8 key sulfide compounds in garlic.

[0080] Based on the random forest model and the KNN model respectively, taking the contents of 8 key sulfide compounds (DAS, AMDS, MPDS, DMTS, DADS, DATTS, AMS and DATS) as markers, prediction models for predicting garlic varieties and / or origins are constructed.

[0081] Evaluate the performance of the models in different regions through indicators such as accuracy, confusion matrix, and classification report on the test set. The classification reports of the random forest model and the KNN model are shown in Table 3 and Table 4.

[0082] Table 3 Classification report of the random forest model

[0083] Precision Recall F1-score Support GS 1 1 1 3.0 HLJ 1 1 1 3.0 HN 1 1 1 3.0 JS 1 1 1 4.0 JX 1 1 1 2.0 SD 1 1 1 3.0 SX 1 1 1 1.0 YN 1 1 1 2.0 Accuracy 1 1 1 1.0 Macro-average 1 1 1 21.0 Weighted average 1 1 1 21.0

[0084] Table 4 Classification report of the KNN model

[0085] Precision Recall F1-score Support GS 1 0.33 0.5 3.0 HLJ 1 0.33 0.5 3.0 HN 0.6 1 0.75 3.0 JS 0.67 1 0.80 4.0 JX 0.67 1 0.80 2.0 SD 1 1 0.80 3.0 SX 0 0 1 1.0 YN 1 1 1 2.0 Accuracy 0.76 0.76 0.76 0.76 Macro-average 0.74 0.71 0.67 21.0 Weighted average 0.80 0.76 0.72 21.0

[0086] Among them, there are indicators such as precision, recall, and F1 score. The accuracies of the random forest and KNN models on the test set are 100% and 76% respectively. It can be seen that the random forest model has a good recognition effect on all regional labels, while the KNN model has a relatively poor recognition effect on some labels. Through further analysis of the confusion matrix and feature importance, the confusion matrix can show the classification errors of the model on each regional label, which helps to identify which labels are more likely to be misclassified. The confusion matrices of the random forest and KNN models can be generated respectively to analyze the classification accuracy of different labels. The confusion matrices of the random forest and KNN models are shown in Figure 9 and Figure 10 . The random forest is correctly classified in almost all regions without obvious misclassification. The confusion matrix shows that the model has a very good recognition effect on each regional label and basically has no false predictions. The KNN model has some misclassifications in [region names]. For example, some samples are misclassified as [wrong classification 1], while some samples are misclassified as other regions (GS, HN). The KNN performs better in some regions but worse than the random forest in other regions (JS). The random forest model performs better than the KNN model in all regions, especially in accurately classifying labels with fewer samples. The KNN model is prone to confusion between some similar regions (such as [similar regions 1] and [similar regions 2]), indicating that its discrimination ability in the feature space is weak (GS, HN). This shows that the random forest model is more suitable for the origin traceability prediction of this dataset.

[0087] The performance indicators of the random forest and KNN models are shown in Tables 5 and 6, including Sensitivity, Specificity, Positive Predictive Value (PPV), Negative Predictive Value (NPV), Detection Rate, Detection Prevalence, and Balanced Accuracy.

[0088] Table 5 Random Forest Indicator Table

[0089] HN YN SD GS JX JS SX HLJ Sensitivity 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Specificity 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Positive Predictive Value 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Negative Predictive Value 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Detection Rate 0.14 0.10 0.14 0.14 0.10 0.19 0.05 0.14 Detection Prevalence 0.14 0.10 0.14 0.14 0.10 0.19 0.05 0.14 Balanced Accuracy 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00

[0090] Table 6 KNN Indicator Table

[0091]

[0092]

[0093] Sensitivity represents the ability of the model to identify samples in a specific region. High sensitivity means that most samples in this region are correctly classified. Specificity represents the ability of the model to exclude this region, that is, it can accurately identify samples that are not in this region. High specificity means that fewer samples are misclassified as this region. If the sensitivity of a certain region is high but the specificity is low, it means that the model tends to classify samples into this region (easy to misjudge). On the contrary, if the sensitivity is low and the specificity is high, it means that the model does not fully recognize the samples in this region. PPV represents the proportion of samples predicted as this region that actually belong to this region. Regions with low PPV may have more misjudgments. NPV represents the proportion of samples predicted as not this region that actually do not belong to this region. High NPV indicates accurate prediction for non-this region. The PPV and NPV of each region can be observed. If the PPV of a certain region is low, more features may be needed to improve the discrimination of this region. Balanced accuracy combines sensitivity and specificity and is a comprehensive index suitable for imbalanced samples. High balanced accuracy indicates that the model has a good classification effect in this region. For regions with fewer samples, special attention can be paid to balanced accuracy to avoid bias caused by excessive focus on accuracy.

[0094] The random forest model demonstrated perfect performance across all metrics and regions, achieving 1.00 in sensitivity, specificity, PPV, NPV, and balanced accuracy. This consistency implies that the random forest is highly effective and reliable for classification tasks across different regions, accurately identifying both positive and negative cases without any errors. In contrast, the KNN model exhibited greater variability in performance across regions, with sensitivity, specificity, PPV, and balanced accuracy fluctuating significantly in some regions. For example, in regions such as SX and HLJ, the KNN model struggled with both sensitivity and balanced accuracy, indicating challenges in accurately identifying positive and negative cases. This variability suggests that the KNN may not generalize well across different regional datasets, likely due to its reliance on local data points, which may vary in distribution across different regions. Overall, these tables highlight the robustness of the random forest model for this application, suggesting it may be a better choice for achieving consistent classification performance across different regional datasets. The KNN model may still be useful in regions where high accuracy is achieved, but its lower performance in specific regions indicates that it may require additional tuning.

[0095] 4.3. By combining these two analyses of feature importance and relative concentration heatmaps (see Figure 11 ), a comprehensive understanding of the relationships, prediction accuracy, and geographical trends among sulfide compounds can be obtained. The feature importance ranking shows which compounds are most important for the model's prediction, while the relative concentration heatmap reveals the variations of these compounds across different regions.

[0096] The left bar chart shows the feature importance of various sulfide compounds in the random forest regression model. Each bar represents a sulfide compound, and the bar length indicates the importance level. In this case, DAS was determined to be the most important predictor, followed by DATTS. This ranking indicates that DAS plays a crucial role in the model's ability to predict the target outcome, suggesting that changes in DAS levels have the greatest impact on the prediction results. The importance of certain compounds relative to others implies that DAS and DATTS may be particularly influential in characterizing the quality, health benefits, or regional traits of garlic. The random forest model uses feature importance scores to show the extent to which each variable contributes to reducing prediction error. Therefore, compounds with high importance scores are crucial for the accuracy of the model and can serve as indicators of key garlic characteristics. In practical applications, understanding feature importance is valuable. For example, food quality controllers or agricultural producers can focus on more closely monitoring DAS and DATTS levels as they are more predictive of the target variable. If the target variable is related to quality assessment, these compounds may affect the sensory properties or health benefits of garlic. This insight is crucial for setting quality standards and can also support marketing efforts by highlighting specific compound concentrations as beneficial features of the product.

[0097] Figure 11 The relative concentration heatmap on the right outlines the relative concentrations of various sulfur compounds in different regions. Each row represents a sulfur compound, while each column corresponds to a specific region (e.g., GS, HN, YN). The color intensity represents the concentration level of each compound in a given region, with warmer colors (such as yellow and green) indicating higher concentrations and cooler colors (such as purple and blue) indicating lower concentrations. This heatmap supports comparative analysis across regions and enables the identification of trends in the distribution of sulfur compounds. For example, if a high concentration of DAS is observed in the GS region, this may indicate that the environmental factors or agricultural practices in GS promote the production of DAS. Such regional characteristics are valuable for identifying garlic varieties with unique chemical profiles. If certain compounds (such as DAS or DATTS) are consistently higher in specific regions, this may suggest that these regions are suitable for growing garlic with desirable chemical properties. The ability of the relative concentration heatmap is crucial for traceability and authentication in the garlic supply chain. By analyzing the sulfur compound profiles of each region, producers and consumers can trace the geographical origin of garlic. This information helps prevent fraud and ensure the authenticity of the product, especially for premium garlic varieties known for specific health benefits or culinary qualities. For example, if a particular region has a distinct concentration pattern, the product can be marketed based on this unique feature, adding value for consumers seeking certain health benefits related to sulfur compounds.

[0098] Based on the data of the random forest model and the training set, taking DAS as the feature variable and the place of origin as the target variable, a prediction model for predicting garlic variety and / or place of origin is constructed. The evaluation of the model is based on the test set. The confusion matrix diagram of the model is shown in Figure 12 .

[0099] Based on the data of the random forest model and the training set, taking DAS and AMDS as the feature variables and the place of origin as the target variable, a prediction model for predicting garlic variety and / or place of origin is constructed. The performance of the model is evaluated based on the test machine. The confusion matrix diagram of the model is shown in Figure 13 .

[0100] In summary, through the establishment of a prediction model based on sulfide compounds, the present application realizes the traceability of garlic variety and / or place of origin. This model has good stability and generalization ability, and can effectively distinguish garlic samples from different places of origin. In terms of model performance, the random forest model shows high accuracy in all regions, especially in regions with a small sample size, it can also maintain excellent recognition effects. In addition, through feature importance analysis and relative concentration heat maps, it is possible to clarify which sulfur compounds are most critical for model prediction (such as DAS, DATTS), and the distribution patterns of these compounds in different regions. This model not only provides a scientific basis for garlic traceability, but also has practical application value for garlic quality monitoring and product development in specific regions.

[0101] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. Use of a reagent for detecting markers in garlic in preparing a product for predicting garlic variety and / or origin, characterized in that: The markers include thioether compounds.

2. The use according to claim 1, characterized in that The sulfide compound includes any one or more of DAS, AMDS, MPDS, DMTS, DADS, DATTS, AMS and DATS.

3. The use according to claim 2, characterized in that The sulfide compounds include: DAS; Optionally, the sulfide compounds include: DAS and DATTS.

4. The use according to claim 3, characterized in that The sulfide compounds include: DAS, AMDS, MPDS, DMTS, DADS, DATTS, AMS and DATS.

5. A method for training a prediction model for garlic varieties and / or origins, characterized in that: It includes: Obtain the detection results of the marker content in the training samples and their corresponding annotation results; The marker is the marker according to any one of claims 1 to 4, and the labeling result is a label representing the garlic variety and / or origin; Input the test results of the training samples into the pre-built machine learning model to obtain the prediction results; The parameters of the machine learning model are updated based on the prediction results and the annotation results to obtain a prediction model.

6. The training method according to claim 5, characterized in that The machine learning model includes any one of logistic regression, support vector machine, random forest, gradient boosting classifier, K nearest neighbor, decision tree and naive Bayes.

7. A method for tracing garlic varieties and / or origins, characterized in that: It includes: Obtaining the test results of the marker content in the sample to be tested; The marker is the marker described in any one of claims 1 to 4; The detection result of the sample to be tested is input into the prediction model trained by the training method described in claim 5 or 6 to obtain the prediction result of the sample to be tested.

8. A device for predicting garlic varieties and / or origins, characterized in that: It includes: An acquisition module is used to obtain the detection result of the marker content in the sample to be tested; The marker is the marker described in any one of claims 1 to 4; The prediction module is used to input the detection result of the sample to be tested into the prediction model trained by the training method described in claim 5 or 6 to obtain the prediction result of the sample to be tested.

9. An electronic device, characterized in that: It includes a processor and a memory, the memory is used to store a program, and when the program is executed by the processor, the processor implements the method for tracing garlic varieties and / or origins according to claim 7.

10. A computer-readable medium, characterized in that The computer-readable medium stores a computer program, and when the computer program is executed by a processor, the method for tracing the garlic variety and / or origin according to claim 7 is implemented.