Chinese yam decoction piece producing area tracing method based on GC-MS and hyperspectral imaging fusion and application of Chinese yam decoction piece producing area tracing method

By integrating hyperspectral imaging and GC-MS technology, a method for tracing the origin of yam slices was constructed, which solved the problems of low detection efficiency and insufficient accuracy in existing technologies, and achieved rapid and accurate origin identification and visual traceability.

CN121978230APending Publication Date: 2026-05-05CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2025-12-14
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for tracing the origin of yam slices suffer from low detection efficiency, insufficient accuracy, and unreliable results. Traditional sensory evaluation is highly subjective, chromatographic techniques are complex and time-consuming, and spectroscopic techniques lack chemical mechanism support.

Method used

By combining hyperspectral imaging and gas chromatography-mass spectrometry (GC-MS) technologies, and integrating the rapid and non-destructive detection of hyperspectral imaging with the precise chemical analysis of GC-MS, a region of origin discrimination model is constructed through multi-source data feature analysis, feature screening and data fusion, forming a joint fingerprint spectrum of "chemical index-spectral index".

Benefits of technology

It enables rapid, accurate, and interpretable identification of the origin of yam slices, improves detection efficiency and accuracy, reduces costs, enhances the scientific validity and persuasiveness of the results, and provides a unique visual identifier.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121978230A_ABST
    Figure CN121978230A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese yam decoction piece origin traceability method based on GC-MS and hyperspectral imaging fusion and application. The method comprises the following steps: S1, preparing a sample; s2, data acquisition; the method specifically comprises the following steps: S21, acquiring hyperspectral data; s22, GC-MS data acquisition; S3, multi-source data feature analysis; s4, feature screening and data fusion; s5, model construction and performance verification; and S6, constructing a fusion fingerprint spectrum. The data fusion model overcomes the problems of easy overfitting and poor generalization ability of a single technology, and the traceability result is more reliable. The classification accuracy of the random forest model based on fusion data reaches 92.3%, which is significantly higher than that of a single data source method. A unique visual identifier is provided for each producing area by fusing the fingerprints, and technical transformation and industrial popularization are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of traditional Chinese medicine quality testing and traceability technology, and in particular to a system and method for tracing the origin of yam slices by combining hyperspectral imaging and gas chromatography-mass spectrometry (GC-MS) technology. Background Technology

[0002] Yam is a major medicinal herb used in both food and medicine, and its quality is closely related to its place of origin, possessing a "genuine" quality. Genuine medicinal herbs refer to those produced under specific natural conditions and ecological environments. Their production is relatively concentrated, and their cultivation, harvesting, and processing have specific requirements, resulting in superior quality and efficacy compared to the same medicinal herbs produced in other regions. Currently, the market is rife with confusion regarding the origin of yam slices, leading to problems such as the substitution of inferior products for superior ones, which seriously affects clinical medication safety and consumer rights.

[0003] Existing methods for identifying the origin of processed yam slices mainly include traditional sensory evaluation and modern instrumental analysis. Traditional sensory evaluation relies primarily on experienced pharmacists who use visual inspection, touch, smell, and taste to identify the ingredients. This method is highly subjective, has poor reproducibility, and is difficult to quantify. Modern instrumental analysis methods mainly include chromatographic techniques, such as gas chromatography-mass spectrometry (GC-MS) and high-performance liquid chromatography (HPLC), which can accurately identify the chemical components in traditional Chinese medicine and have the advantages of high sensitivity and high resolution. However, chromatographic techniques involve complex pretreatment, are time-consuming, costly, and destructive, making them unsuitable for large-scale rapid screening. Spectroscopic techniques, such as near-infrared spectroscopy and hyperspectral imaging, offer advantages such as speed, non-destructive testing, and on-site detection. Hyperspectral imaging can simultaneously acquire spatial and spectral information of samples, showing promising application prospects in the identification of traditional Chinese medicine. However, the identification results of spectroscopic techniques lack clear chemical substance indicators, the interpretability of the models is poor, and their authority is often questioned.

[0004] Single-technology traceability has significant limitations: while chromatography can provide precise chemical composition information, its processing procedures are complex and unsuitable for rapid screening; while spectroscopy, though fast and non-destructive, lacks support from chemical mechanisms, limiting the reliability of its results. Therefore, there is an urgent need in this field for a technology that integrates detection efficiency, accuracy, and reliability for tracing the origin of Chinese medicinal materials. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for tracing the origin of yam slices based on the fusion of GC-MS and hyperspectral imaging. By combining the rapid and non-destructive testing capabilities of hyperspectral imaging with the precise chemical analysis capabilities of GC-MS, the origin of yam slices can be identified quickly, accurately, and interpretably.

[0006] The steps of the method include: S1 Sample Preparation; S2 data acquisition; specifically including S21 hyperspectral data acquisition; S22 GC-MS data acquisition: S3 Multi-Source Data Feature Analysis; S4 Feature Filtering and Data Fusion; S5 model construction and performance verification; S6 fusion fingerprint mapping.

[0007] Furthermore, it also includes S7 mechanism analysis and statistical verification, and / or S8 classification performance verification.

[0008] Preferably, S1 includes taking fresh small white-mouthed yam from different origins, washing, peeling and slicing it, and further drying it. For example, it is dried at low temperature (50°C) for 5 hours using an open drying oven, then cooled and placed in a desiccator for later use. Several samples are prepared from each origin for hyperspectral imaging analysis; several parallel samples are also taken from each origin for GC-MS analysis.

[0009] Preferably, the hyperspectral data acquisition includes: acquiring hyperspectral images of yam slices using a hyperspectral imaging system (spectral range 400-1000nm). After acquisition, the original images are corrected using a black and white card, and the region of interest (ROI) for each sample is extracted to obtain the average reflectance spectral curve of each sample under different wavelength channels.

[0010] Preferably, the GC-MS data acquisition includes: accurately weighing yam slices into a test tube, adding acetonitrile for vortexing and extraction, followed by ultrasonic extraction. After standing, the supernatant is collected, and a derivatization reagent (e.g., BSTFA containing 1% TMCS) is added, vortexed, and then subjected to ultrasonic derivatization. After the reaction, the mixture is cooled to room temperature and GC-MS analysis is performed. By comparing with a standard mass spectrometry library, compounds, including sugars, organic acids, amino acids, and fatty acids, are identified.

[0011] Preferably, the multi-source data feature analysis includes: (A) hyperspectral feature curves showing significant differences in the characteristic wavelengths (818nm, 756nm, 946nm, 938nm, 742nm) of samples from different origins; (B) GC-MS chemical composition analysis showing significant differences in the content distribution of five key chemical markers (citric acid, mannitol, inositol, palmitic acid, and fructose) among different origins; (C) box plots of characteristic wavelength distribution further verifying the spectral differences among origins; and (D) heatmaps of chemical marker variability showing the stability of chemical components within each origin.

[0012] Preferably, the feature screening includes (1) Hyperspectral feature wavelength screening: using a competitive adaptive reweighted sampling algorithm to screen several feature wavelengths from the entire hyperspectral band, for example, 5 feature wavelengths: 818nm, 756nm, 946nm, 938nm, 742nm. (2) GC-MS key differential chemical marker screening: using partial least squares discriminant analysis (PLS-DA) combined with variable importance projection (VIP>1.0) and analysis of variance (p<0.05) to screen several key differential chemical markers from GC-MS data: citric acid (retention time 20.4min), mannitol (30.6min), inositol (29.3min), palmitic acid (28.2min), and fructose (25.2min).

[0013] Preferably, the data fusion includes performing Z-score normalization on the absorbance data of 5 characteristic wavelengths and the peak area data of 5 chemical markers respectively, and then concatenating them into a 10-dimensional fusion feature vector to construct a fusion data matrix.

[0014] Preferably, the model construction and performance verification include dividing the fused feature vectors and their origin labels of a large sample into training and test sets according to a certain ratio. Origin discrimination models are constructed using Support Vector Machine (SVM) and Random Forest algorithms, respectively. The ratio of the training set to the test set is (7-9.5):(3-0.5), for example, 8:2, 7:3, or 9:1.

[0015] Preferably, the construction of the fusion fingerprint spectrum includes constructing a "chemical index-spectral index" joint fingerprint spectrum for each place of origin: (A) the fingerprint spectrum in the form of a radar chart intuitively displays the unique distribution pattern of yam slices from different places of origin in 10 characteristic dimensions; (B) the fingerprint spectrum in the form of a heat map quantitatively characterizes the standardized distribution of characteristic values ​​of each place of origin, providing a quantifiable "identity authentication" label for each place of origin.

[0016] Preferably, mechanistic analysis and statistical validation included demonstrating the distribution characteristics of the fused data in origin classification using principal component analysis (PCA). The PCA distribution of the fused data showed good class separation, with samples from different origins forming independent clusters, validating the superiority of the fusion method in feature representation. Furthermore, cross-modal feature correlation analysis revealed the intrinsic correlation mechanism between hyperspectral features and chemical components: mannitol showed a correlation coefficient of 0.61 with 756 nm wavelength, and inositol showed a correlation coefficient of -0.71 with 742 nm wavelength. Statistical significance analysis validated the significance of the selected features. ANOVA analysis showed that all 10 features had extremely significant differences among different origins (p<0.001), confirming that the selected features were statistically significant for origin identification.

[0017] Preferably, the classification performance verification includes visually demonstrating the improvement in classification performance through a confusion matrix.

[0018] This invention also provides an application of a method for tracing the origin of yam slices based on the fusion of GC-MS and hyperspectral imaging in identifying the origin of yam slices.

[0019] The beneficial effects of this invention are: high precision and high robustness: the data fusion model overcomes the problems of overfitting and poor generalization ability of single technologies, resulting in more reliable traceability results. The random forest model based on fused data achieves a classification accuracy of 92.3%, significantly higher than methods using single data sources. In practical applications, hyperspectral imaging can be used for rapid initial screening, followed by GC-MS confirmation of suspected samples, greatly improving detection efficiency and reducing costs. The combined use of fingerprinting and mechanistic correlation analysis enhances the scientific validity and persuasiveness of the method. The fused fingerprinting provides a unique visual identifier for each production area. The selected characteristic wavelengths provide clear technical parameters for developing low-cost, portable dedicated detection instruments, facilitating technology transfer and industrial promotion.

[0020] This application, for the first time, fuses data from hyperspectral characteristic wavelengths (such as 742nm, 756nm, 938nm, etc.) with key differential chemical markers from GC-MS (such as citric acid, inositol, mannitol, etc.) to construct a fused data matrix. Machine learning models (such as SVM and random forest) are then used for origin identification, significantly improving the model's accuracy and robustness. Based on the fusion analysis results, a novel "chemical index-spectral index" joint fingerprint spectrum is constructed. This spectrum visualizes the relative content of chemical markers and the spectral response values ​​of characteristic wavelengths within the same framework, providing a unique and interpretable "identity verification" for each origin. Attached Figure Description

[0021] Figure 1 This invention provides a roadmap for multi-source information fusion technology. Figure 2 This is a feature analysis diagram of multi-source data for yam slices.

[0022] Figure 3 A graph comparing the performance of classification models from different data sources.

[0023] Figure 4 This is a fingerprint spectrum of yam slices based on multi-source information fusion.

[0024] Figure 5 Principal component analysis distribution plot for fused data.

[0025] Figure 6 This is a graph showing the correlation analysis of cross-modal features.

[0026] Figure 7 The graph shows the results of the feature statistical significance analysis.

[0027] Figure 8 A comparison chart of classification confusion matrices for different methods. Detailed Implementation

[0028] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments. The content mentioned in the embodiments is not intended to limit the present invention. Example

[0029] Taking the identification of "Xiaobaizui yam" slices from seven production areas in Hebei (Anguo, Wuji, Wenren, Maoshanwei, Anping, Boye, and Xinle) as an example.

[0030] 1. Sample Preparation Fresh white-mouthed yam was collected from seven production areas. After washing and peeling, it was sliced ​​into 2-4 mm thin slices using a rotary slicer. The slices were then dried in an open drying oven at low temperature (50℃) for 5 hours, cooled, and placed in a desiccator for later use. 100 samples were prepared from each production area, for a total of 700 samples, for hyperspectral imaging analysis; three additional parallel samples from each production area were used for GC-MS analysis.

[0031] 2. Data Acquisition (1) Hyperspectral data acquisition: Hyperspectral images of 700 yam slices were acquired using a hyperspectral imaging system (spectral range 400-1000 nm). The system was preheated for 30 minutes, the object distance was set to 35 cm, the exposure time to 8 ms, and the conveyor belt speed to 1.5 mm / s. After acquisition, the raw images were corrected using a black and white card, and the region of interest (ROI) for each sample was extracted to obtain the average reflectance spectrum curve of each sample across 301 wavelength channels.

[0032] (2) GC-MS data acquisition: Accurately weigh 0.3g of yam powder into a test tube, add 1.5mL of acetonitrile and soak for 24h, vortex for 2 minutes, and sonicate at 60℃ for 60 minutes. After standing, take 500μL of the supernatant, add 200μL of derivatization reagent (BSTFA containing 1% TMCS), vortex to mix, and sonicate for derivatization for 1.5h. After the reaction, cool to room temperature and perform GC-MS analysis. By comparison with the NIST standard mass spectrometry library, a total of 37 compounds were identified, including sugars, organic acids, amino acids, and fatty acids.

[0033] 3. Multi-source data feature analysis like Figure 2 As shown, a systematic analysis of the multi-source data characteristics of yam slices from different origins was conducted: (A) The hyperspectral characteristic curves show that there are significant differences between samples from different origins at characteristic wavelengths (818nm, 756nm, 946nm, 938nm, 742nm); (B) GC-MS chemical composition analysis showed that the content distribution of five key chemical markers (citric acid, mannitol, inositol, palmitic acid, and fructose) varied significantly among different production areas; (C) The box plot of characteristic wavelength distribution further verifies the spectral differences between production areas; (D) The heat map of chemical marker variability shows the stability of chemical components within each production area.

[0034] 4. Feature Filtering and Data Fusion (1) Screening of hyperspectral characteristic wavelengths: Five characteristic wavelengths were selected from the entire hyperspectral band using the Competitive Adaptive Reweighted Sampling (CARS) algorithm: 818nm, 756nm, 946nm, 938nm, and 742nm.

[0035] (2) Screening of key differential chemical markers by GC-MS: Partial least squares discriminant analysis (PLS-DA) combined with variable importance projection (VIP>1.0) and analysis of variance (p<0.05) was used to screen out five key differential chemical markers from GC-MS data: citric acid (retention time 20.4 min), mannitol (30.6 min), inositol (29.3 min), palmitic acid (28.2 min), and fructose (25.2 min).

[0036] (3) Data fusion: The absorbance data of the five characteristic wavelengths and the peak area data of the five chemical markers were Z-score normalized and then concatenated into a 10-dimensional fusion feature vector to construct a fusion data matrix.

[0037] 5. Model Building and Performance Validation The fused feature vectors and their origin labels of 700 samples were divided into training and test sets in a 7:3 ratio. Origin discrimination models were constructed using Support Vector Machine (SVM) and Random Forest algorithms, respectively.

[0038] like Figure 3 As shown, the model performance comparison results indicate that: SVM model: 84.3% accuracy when used alone (hyperspectral density), 61.2% accuracy when used alone (GC-MS), and 89.5% accuracy when using fused data. Random Forest model: Hyperspectral data used alone 81.2%, GC-MS used alone 14.3%, and fused data used 92.3%. The fusion method improves accuracy by more than 10% compared to the best single method, verifying the effectiveness of multi-source information fusion.

[0039] 6. Construction of fused fingerprint maps like Figure 4 As shown, a combined "chemical index-spectral index" fingerprint was constructed for each origin: (A) The fingerprint map in the form of a radar chart intuitively shows the unique distribution patterns of yam slices from different origins across 10 characteristic dimensions; (B) The fingerprint spectrum in the form of heat map quantitatively characterizes the standardized distribution of characteristic values ​​of each place of origin, providing a quantifiable "identity authentication" identifier for each place of origin.

[0040] 7. Mechanism Analysis and Statistical Validation like Figure 5 As shown, principal component analysis (PCA) reveals the distribution characteristics of the fused data in origin classification. The PCA distribution of the fused data exhibits good class separation, with samples from different origins forming independent clusters, validating the superiority of the fusion method in feature representation.

[0041] like Figure 6 As shown, cross-modal feature correlation analysis reveals the intrinsic correlation mechanism between hyperspectral features and chemical composition: - Mannitol showed a correlation coefficient of 0.61 with a wavelength of 756 nm. - The correlation coefficient between inositol and the 742nm wavelength is -0.71. These strong correlations provide chemical mechanism support for rapid hyperspectral detection.

[0042] like Figure 7 As shown, statistical significance analysis verified the significance of the selected features. ANOVA analysis showed that all 10 features had extremely significant differences among different origins (p<0.001), confirming that the selected features are statistically significant for origin identification.

[0043] 8. Classification performance verification like Figure 8 As shown, the improvement in classification performance is visually demonstrated through the confusion matrix: (A) Hyperspectral imaging alone has a high rate of misjudgment of origin, with an accuracy rate of 84.3%; (B) The fusion method significantly reduced misclassification, achieving an accuracy of 92.3%, while retaining only a small number of overlaps between adjacent origins, thus verifying the effectiveness and robustness of the present invention in complex origin identification tasks.

[0044] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0045] It should be understood that the above detailed description of the technical solutions of the present invention with reference to preferred embodiments is illustrative and not restrictive. Those skilled in the art can modify the technical solutions described in the embodiments or make equivalent substitutions for some of the technical features based on reading this specification; however, these modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for tracing the origin of yam slices based on GC-MS and hyperspectral imaging fusion, characterized by the following steps: include: S1 Sample Preparation; S2 data acquisition; details This includes S21 hyperspectral data acquisition; S22GC-MS Data Acquisition: S3 Multi-Source Data Feature Analysis; S4 Feature Filtering and Data Fusion; S5 model construction and performance verification; S6 fusion fingerprint mapping.

2. The method according to claim 1, characterized in that, It also includes S7 mechanism analysis and statistical verification, and / or S8 classification performance verification.

3. The method according to claim 1, characterized in that, S1 includes fresh small white-mouthed yams from different production areas, which are washed, peeled, sliced, and dried.

4. The method according to claim 1, characterized in that, The hyperspectral data acquisition includes: acquiring hyperspectral images of yam slices using a hyperspectral imaging system. After acquisition, the original images are corrected using a black and white card, the region of interest for each sample is extracted, and the average reflectance spectral curves of each sample under different wavelength channels are obtained.

5. The method according to claim 1, characterized in that, The GC-MS data acquisition process included: accurately weighing yam slices into a test tube, adding acetonitrile, vortexing, and ultrasonic extraction. After standing, the supernatant was collected, a derivatization reagent was added, vortexed, and then ultrasonically derivatized. After the reaction, the mixture was cooled to room temperature and GC-MS analysis was performed; compounds were identified by comparison with a standard mass spectrometry library.

6. The method according to claim 1, characterized in that, The multi-source data feature analysis includes: (A) hyperspectral feature curves showing significant differences in the characteristic wavelengths of samples from different origins; (B) GC-MS chemical composition analysis showing significant differences in the content distribution of key chemical markers among different origins; (C) box plots of characteristic wavelength distributions further verifying the spectral differences among origins; and (D) heatmaps of chemical marker variability showing the stability of chemical components within each origin.

7. The method according to claim 6, characterized in that, The feature screening includes (1) Hyperspectral feature wavelength screening: a competitive adaptive reweighted sampling algorithm is used to screen several feature wavelengths from the entire hyperspectral band, with 5 feature wavelengths: 818nm, 756nm, 946nm, 938nm, and 742nm. (2) GC-MS key differential chemical marker screening: partial least squares discriminant analysis combined with variable importance projection and analysis of variance is used to screen several key differential chemical markers from GC-MS data: citric acid, mannitol, inositol, palmitic acid, and fructose.

8. The method according to claim 7, characterized in that, The data fusion process includes performing Z-score normalization on the absorbance data of five characteristic wavelengths and the peak area data of five chemical markers, and then concatenating them into a 10-dimensional fusion feature vector to construct a fusion data matrix. The model construction and performance verification include dividing the fused feature vectors and their origin labels of a large sample into training and test sets according to a certain ratio. Origin discrimination models are then constructed using support vector machines and random forest algorithms, respectively.

9. The method according to claim 2, characterized in that, The construction of the fusion fingerprint spectrum includes the construction of a "chemical index-spectral index" joint fingerprint spectrum for each place of origin: (A) the fingerprint spectrum in the form of a radar chart intuitively shows the unique distribution pattern of yam slices from different places of origin in 10 characteristic dimensions; (B) the fingerprint spectrum in the form of a heat map quantitatively characterizes the standardized distribution of characteristic values ​​of each place of origin, providing a quantifiable "identity authentication" label for each place of origin. and / or Mechanism analysis and statistical verification, including principal component analysis, demonstrated the distribution characteristics of fused data in place of origin classification.

10. Application of a method for tracing the origin of yam slices based on GC-MS and hyperspectral imaging fusion in identifying the origin of yam slices.