A method and system for tracing the origin of radix paeoniae based on infrared spectrum fusion and machine learning
By using spectral fusion and machine learning methods, a model for identifying the origin of Taoerqi was constructed, which solved the problems of accuracy and cost in the existing technology of origin tracing, and achieved high-precision and low-cost origin identification and tracing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWEST INST OF PLATEAU BIOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2026-05-18
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies are insufficient to achieve high-precision and highly generalizable intelligent traceability of peach production areas, and traditional methods are costly and inefficient, making them unsuitable for large-scale sample screening and grassroots promotion.
By collecting near-infrared and mid-infrared spectral data, spectral fusion was performed, and machine learning algorithms were used to construct an origin discrimination model. The proportion of spectral preprocessing and modeling set was optimized to construct origin discrimination models under single and fused spectral conditions.
It enables high-precision traceability of the origin of Taoerqi, significantly improves the accuracy and generalization performance of origin identification, reduces testing costs, is suitable for large-scale rapid screening, and is applicable to the acquisition of Chinese medicinal materials and grassroots quality inspection.
Smart Images

Figure CN122490435A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Chinese medicinal material identification technology, and in particular to a method and system for tracing the origin of *Taoerqi* based on infrared spectral fusion and machine learning. Background Technology
[0002] Origin is a core determinant of the quality, efficacy stability, and clinical safety of Chinese medicinal materials. Significant differences exist in the ecological conditions, such as climate, soil, and altitude, among different origins, directly leading to marked variations in the quality of Chinese medicinal materials. Traditional methods for determining the origin of Chinese medicinal materials rely heavily on empirical morphological identification, which suffers from inherent flaws such as low accuracy, strong subjectivity, and poor repeatability, making it difficult to meet the precision requirements of modern Chinese medicine quality control.
[0003] *Sinopodophyllum hexandrum* (Royle), a perennial herbaceous plant belonging to the genus *Sinopodophyllum* in the family Berberidaceae, is a rare medicinal plant in my country. It is distributed in Qinghai, Yunnan, Sichuan, and Tibet, with scattered wild populations, mainly growing in flat valleys, understory forests, forest edges, and thickets at high altitudes of 2700–4500 m. The roots, rhizomes, leaves, and fruits are all used medicinally; the fruit is called "Xiaoye Lian," and the underground parts are called "Tao'erqi" or "Guijiu," all possessing significant medicinal value. However, *Sinopodophyllum hexandrum* currently faces prominent problems such as uneven quality across different production areas and outdated methods for identifying its origin, severely restricting its standardized quality control, rational resource development, and standardized clinical application.
[0004] With the deepening research on the quality control of traditional Chinese medicine (TCM), high-performance liquid chromatography (HPLC), gas chromatography (GC), mass spectrometry (MS), and their coupled techniques have become the mainstream methods for component analysis and origin identification of TCM materials. However, these methods rely on expensive and sophisticated instruments, involve cumbersome and complex sample pretreatment procedures, require highly skilled operators, and are costly and inefficient, making them difficult to apply to large-scale sample screening and grassroots promotion. Therefore, establishing a low-cost, easy-to-operate, rapid, efficient, and widely applicable technology for the origin analysis and detection of *Prunus persica* (a type of medicinal herb) has become a critical issue that urgently needs to be addressed.
[0005] Infrared spectroscopy, particularly near-infrared (NIR) and mid-infrared (MIR) spectroscopy, has experienced rapid development in recent years due to its unique advantages such as ease of operation, rapid detection, simple sample pretreatment, low analysis cost, and non-destructive nature. It has been widely applied in agriculture, food, petrochemicals, and pharmaceuticals, achieving large-scale industrial application. However, the information provided by a single infrared spectroscopy technique is limited, making it difficult to comprehensively capture the differences in the complex chemical components of traditional Chinese medicine. Currently, there is no reliable technical solution that can integrate multi-source spectral information and utilize advanced machine learning algorithms for high-precision, intelligent part identification. This paper proposes a method to achieve high-precision, high-generalization intelligent traceability of *Taoerqi* (a type of medicinal herb) from different origins. Summary of the Invention
[0006] The purpose of this invention is to overcome the technical problems existing in the prior art and provide a method and system for tracing the origin of *Prunus persica* based on infrared spectral fusion and machine learning. By optimizing key parameters such as modeling methods, spectral preprocessing methods, and model set proportions, origin discrimination models for *Prunus persica* under single infrared spectral (NIR, MIR) and fused spectral conditions are constructed respectively. The optimal origin discrimination model is selected through comparative analysis, aiming to provide a scientific basis and technical support for tracing the origin of *Prunus persica* medicinal materials.
[0007] The objective of this invention is achieved through the following technical solution:
[0008] In a first aspect, the present invention provides a method for tracing the origin of Taoerqi (a type of peach) based on infrared spectral fusion and machine learning, comprising the following steps:
[0009] S1. Collect multiple single infrared spectral data from different parts of peach samples to be traced from different origins, including near-infrared spectral data and mid-infrared spectral data;
[0010] S2. The near-infrared spectral data and mid-infrared spectral data are concatenated to obtain primary fused data; the primary fused data is then subjected to feature extraction and fusion to obtain intermediate fused spectral data;
[0011] S3. Based on the single infrared spectral data, primary fusion data, and intermediate fusion data at the same location, use the Python platform and TQ analyst software to construct origin discrimination models respectively, and obtain different levels of origin discrimination models corresponding to different locations;
[0012] S4. Use the origin discrimination model constructed in step S3 to determine the origin of the tested peach sample and output the origin traceability results.
[0013] In some embodiments, step S3 further includes:
[0014] By using the Python platform, different origin discrimination models are fused at the decision level to form an advanced fused origin discrimination model.
[0015] In some embodiments, the samples of *Prunus cerasifera* to be tested were collected from populations at 15 different collection sites in Qinghai Province, Yunnan Province, and Tibet Autonomous Region. The populations at each collection site were spaced more than 20 km apart. At least 20 healthy plants were collected from each population at the same collection site, and each plant was spaced 10 m apart. Each *Prunus cerasifera* sample was dried, pulverized, and passed through an 80-mesh sieve at different parts before being placed in a desiccator for analysis.
[0016] In some embodiments, the acquisition of the near-infrared spectral data includes:
[0017] Take appropriate amounts of powder from different parts of the peach kernel and place them on filter paper. Use a Fourier transform infrared spectrometer (NIR fiber module) to analyze the sample at 10000-4000 cm⁻¹. -1 Spectral data were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 6 cm⁻¹. -1 The scan was performed 64 times, with air as a reference. Each peach sample was collected 3 times, and the average spectrum was taken for analysis.
[0018] The acquisition of the mid-infrared spectral data includes:
[0019] Appropriate amounts of *Prunus persica* powder from different parts were placed on the attenuated total reflectance infrared probe of a Fourier transform infrared spectrometer, at 4000-400 cm⁻¹. -1 MIR spectra were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 4 cm⁻¹. -1 The sample was scanned 32 times, with air as a reference. Each sample was collected 3 times, and the average spectrum was taken for analysis.
[0020] In some embodiments, step S1 further includes:
[0021] The near-infrared and mid-infrared spectral data are preprocessed, including one or more combinations of scattering correction, derivative processing, and spectral smoothing. The scattering correction method is multivariate scattering correction or standard normal transformation. The derivative processing method is first derivative or second derivative. The spectral smoothing method is Savitzky-Golay smoothing or Norris smoothing.
[0022] In some embodiments, when constructing the origin discrimination model using the Python platform or TQ analyst software in step S3, multiple origin discrimination models are constructed using different machine learning methods, preprocessing methods, model set ratios, and optical path types.
[0023] In some embodiments, the machine learning methods used on the TQ analyst software include distance matching and discriminant analysis, while the machine learning methods used on the Python platform include support vector machines, decision trees, random forests, and limit trees.
[0024] In some embodiments, outlier removal is performed on single infrared spectral data, primary fusion data, and intermediate fusion data before constructing the origin discrimination model. The outlier removal methods include Mahalanobis distance and principal component analysis.
[0025] In some embodiments, it also includes:
[0026] In step S4, the optimal origin discrimination model is selected, and the origin of the tested peach sample is determined using the optimal origin discrimination model; the selection of the optimal origin discrimination model includes:
[0027] For samples from the same location, compare the discrimination effects of different level origin discrimination models and select the best-performing origin discrimination model as the optimal origin discrimination model for that location.
[0028] A second aspect of the present invention provides a method for tracing the origin of Taoerqi (a type of peach) based on infrared spectral fusion and machine learning, comprising:
[0029] The spectral acquisition module is used to acquire multiple single infrared spectral data from different parts of peach samples from different origins to be traced. The multiple single infrared spectral data include near-infrared spectral data and mid-infrared spectral data.
[0030] The spectral fusion module is used to concatenate the near-infrared spectral data and the mid-infrared spectral data to obtain primary fused data; and to perform feature extraction and fusion on the primary fused data to obtain intermediate fused spectral data.
[0031] The origin identification model construction module is used to construct origin identification models based on the single infrared spectral data, primary fusion data, and intermediate fusion data in the same part, respectively using the Python platform and TQ analyst software, to obtain different levels of origin identification models corresponding to different parts;
[0032] The identification output module is used to identify the origin of the peach 7 sample to be traced using the origin identification model constructed in the origin identification model construction module, and output the origin traceability results.
[0033] It should be further noted that the technical features corresponding to the above-mentioned options and embodiments can be combined or substituted with each other to form new technical solutions without conflict.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] 1. This invention effectively solves the technical problem of insufficient information dimension of a single spectral spectrum and inability to comprehensively reflect differences in chemical composition by simultaneously acquiring near-infrared (NIR) and mid-infrared (MIR) spectra. NIR spectroscopy mainly reflects the overtone and combination frequency absorption of hydrogen-containing groups (such as CH, NH, OH), and is sensitive to the overall information of organic components; MIR spectroscopy mainly reflects the fundamental frequency vibration of molecules, and can provide more refined information on functional group structure. The fusion of the two achieves complementary advantages of spectral information. Combined with the powerful nonlinear modeling capabilities of machine learning algorithms, it can more accurately capture the subtle but crucial chemical differences between samples of *Prunus persica* from different origins, significantly improving the accuracy of origin tracing. Experimental data (see examples) show that the intermediate fusion model of this invention can achieve a 100% prediction rate for origin discrimination of *Prunus persica* roots, with an external validation prediction rate of 93.75%, far exceeding that of a single spectral model (such as the MIR single model with an external validation rate of only 77.15%).
[0036] 2. This invention systematically optimizes spectral preprocessing methods (such as MSC, SNV, 1D, 2D, SG smoothing, etc.), effectively eliminating interference from light scattering, baseline drift, noise, etc., and extracting robust spectral features. Simultaneously, by optimizing the ratio of the modeling set to the validation set (2:1 to 5:1), overfitting or underfitting of the model is avoided. More importantly, by constructing three different levels of fusion models—primary, intermediate, and advanced—especially the intermediate fusion which uses logistic regression to select high-contribution features, and the advanced fusion which integrates the advantages of multiple models through decision-level voting, the generalization performance of the model on different datasets and unknown samples is significantly improved. Experiments demonstrate (see examples) that this invention can select the optimal fusion strategy and model combination for different parts of *Prunus persica* (root, rhizome, stem, leaf, fruit) (e.g., primary fusion is optimal for rhizome, and NIR single model is optimal for stem), exhibiting extremely strong adaptability.
[0037] 3. This invention employs infrared spectroscopy, which, compared to traditional chemical analysis methods such as HPLC and GC-MS, eliminates the need for complex sample pretreatment, consumes no organic solvents, does not damage the sample, and requires only a few minutes to detect a single sample, significantly reducing detection costs. Combined with a pre-trained intelligent classifier, operators do not need extensive knowledge of spectral analysis and chemometrics to quickly obtain origin identification results, making it ideal for large-scale rapid screening in the acquisition, warehousing inspection, and grassroots quality inspection departments of Chinese medicinal materials.
[0038] 4. This invention not only targets a single medicinal part, but comprehensively covers different medicinal parts of *Prunus persica*, including the root, rhizome, stem, leaves, and fruit, constructing a differentiated origin traceability model. This provides a scientific basis and technical tools for the refined, full-chain quality control of *Prunus persica*, which is characterized by "multiple parts from one plant, each with different effects," and is conducive to promoting its rational resource development and standardized clinical application. Attached Figure Description
[0039] Figure 1 This is a flowchart illustrating a method for tracing the origin of Taoerqi (a type of peach) based on infrared spectral fusion and machine learning, as shown in an embodiment of the present invention.
[0040] Figure 2 The following are NIR spectra of different parts of a *Prunus persica* sample as shown in an embodiment of the present invention;
[0041] Figure 3 The images shown are MIR spectra of different parts of a *Prunus persica* sample as illustrated in this embodiment of the invention.
[0042] Figure 4 This is a root sample anomaly spectrum discrimination diagram shown in an embodiment of the present invention;
[0043] Figure 5 This is an abnormal spectral discrimination diagram of a rhizome sample shown in an embodiment of the present invention;
[0044] Figure 6 This is a spectrum discrimination diagram of an abnormal stem sample shown in an embodiment of the present invention;
[0045] Figure 7 This is an abnormal spectral discrimination diagram of a leaf sample shown in an embodiment of the present invention;
[0046] Figure 8 This is an abnormal spectral discrimination diagram of a fruit sample shown in an embodiment of the present invention;
[0047] Figure 9 This is a schematic diagram illustrating the results of a root origin discrimination model based on a single infrared spectrum, as shown in an embodiment of the present invention.
[0048] Figure 10 This is a schematic diagram illustrating the modeling results of the root origin discrimination model based on fused spectra, as shown in an embodiment of the present invention.
[0049] Figure 11 This is a schematic diagram illustrating the results of a root and stem origin discrimination model based on a single infrared spectrum, as shown in an embodiment of the present invention.
[0050] Figure 12 This is a schematic diagram illustrating the modeling results of a root and stem origin discrimination model based on fused spectra, as shown in an embodiment of the present invention.
[0051] Figure 13This is a schematic diagram illustrating the results of a stem origin discrimination model based on a single infrared spectrum, as shown in an embodiment of the present invention.
[0052] Figure 14 This is a schematic diagram illustrating the modeling results of the stem origin discrimination model based on fused spectra, as shown in an embodiment of the present invention.
[0053] Figure 15 This is a schematic diagram illustrating the results of a leaf origin discrimination model based on a single infrared spectrum, as shown in an embodiment of the present invention.
[0054] Figure 16 This is a schematic diagram illustrating the modeling results of a leaf origin discrimination model based on fused spectra, as shown in an embodiment of the present invention.
[0055] Figure 17 This is a schematic diagram illustrating the results of a fruit origin discrimination model based on a single infrared spectrum, as shown in an embodiment of the present invention.
[0056] Figure 18 This is a schematic diagram illustrating the modeling results of a fruit origin discrimination model based on fused spectra, as shown in an embodiment of the present invention. Detailed Implementation
[0057] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] It should be noted that the defects in the solutions in the prior art are all the results of the inventors' practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of this application in the following text should be the inventors' contributions to this application in the process of invention and creation, and should not be understood as technical content known to those skilled in the art.
[0059] In view of the technical problems pointed out in the background art, the present invention provides the following embodiments:
[0060] In one exemplary embodiment, a method for tracing the origin of Taoerqi (a type of peach) based on infrared spectral fusion and machine learning is provided, such as... Figure 1 As shown, it includes the following steps:
[0061] S1. Collect multiple single infrared spectral data from different parts of peach samples to be traced from different origins, including near-infrared spectral data and mid-infrared spectral data;
[0062] S2. The near-infrared spectral data and mid-infrared spectral data are concatenated to obtain primary fused data; the primary fused data is then subjected to feature extraction and fusion to obtain intermediate fused spectral data;
[0063] S3. Based on the single infrared spectral data, primary fusion data, and intermediate fusion data at the same location, use the Python platform and TQ analyst software to construct origin discrimination models respectively, and obtain different levels of origin discrimination models corresponding to different locations;
[0064] S4. Use the origin discrimination model constructed in step S3 to determine the origin of the tested peach sample and output the origin traceability results.
[0065] Based on the above steps, this embodiment provides a specific experimental method, mainly including:
[0066] 1. Instruments
[0067] Fourier transform infrared spectrometer (iS 50, Thermo Nicolet, USA) (equipped with fiber optic cable and attenuated total reflection accessories), oven (Shanghai Yiheng Scientific Instruments Co., Ltd., China), and pulverizer (Tianjin Tester Co., Ltd., China).
[0068] 2. Sample Source
[0069] During the fruiting period of *Sinopodophyllum hexandrum* in July and August 2024, samples were collected from 15 different production areas in Qinghai Province, Yunnan Province, and Tibet Autonomous Region. All samples were wild *Sinopodophyllum hexandrum* (National Key Protected Wild Plant Collection Certificate, No.: 0043349), and the original plant specimens were identified as *Sinopodophyllum hexandrum* (Specimen No.: 2024-002), belonging to the genus *Sinopodophyllum* of the family Berberidaceae. The populations at each collection site were spaced at least 20 km apart, with at least 20 healthy plants collected from each population, spaced 10 m apart. The samples were divided into five different parts (root, stem, leaf, rhizome, and fruit), dried, pulverized, and passed through an 80-mesh sieve before being placed in a desiccator for analysis. Sample information is shown in Table 1.
[0070] Table 1. Source information of peach samples from different origins
[0071]
[0072] 3. Experimental Methods
[0073] 3.1 Infrared Spectral Acquisition and Spectral Fusion
[0074] 3.1.1. NIR Spectral Acquisition
[0075] Take appropriate amounts of sample powder from different parts of the sample and place them on filter paper. Use a Fourier transform infrared spectrometer (NIR) fiber optic module to measure the sample at 10000-4000 cm⁻¹. -1 Spectral data were acquired within the specified range. Background interference such as CO2 and water was subtracted in real-time during acquisition, with a scanning resolution of 6 cm⁻¹. -1 The number of scans was 64. Using air as a reference, each sample was collected 3 times, and the average spectrum was used for analysis.
[0076] 3.1.2. MIR Spectral Acquisition
[0077] Appropriate amounts of sample powder from different locations were placed on the attenuated total reflectance infrared probe of a Fourier transform infrared spectrometer, and the sample was analyzed at 4000-400 cm⁻¹. -1 MIR spectra were acquired within the specified range, with background interference from CO2 and water subtracted in real time during acquisition. The scanning resolution was 4 cm⁻¹. -1 The sample was scanned 32 times, with air as a reference. Each sample was collected 3 times, and the average spectrum was used for analysis.
[0078] 3.1.3. Spectral Fusion
[0079] Data fusion was performed using existing spectral fusion methods. The NIR and MIR data of the samples were concatenated using a Python platform to obtain primary fused data. Logistic regression was then used on the primary fused spectral data to calculate the contribution of each feature point, and features with high contributions were extracted for further fusion to obtain intermediate fused data. Advanced fusion, or decision-level fusion, used Python to fuse models from different modeling methods at the decision-level, forming the advanced fused model. The weighting coefficients were based on the performance of each model.
[0080] 3.2. Construction of a model for identifying medicinal parts and origins based on infrared spectroscopy technology
[0081] 3.2.1. Outlier Removal
[0082] Using infrared spectra of different parts of *Prunus persica* collected under section "3.1", an external validation set was randomly divided at a ratio of 5:1, with the remaining samples serving as the modeling set. Outliers were removed from the modeling set samples using Marginal Distance (MD) and Principal Component Analysis (PCA), specifically for NIR, MIR, primary fusion, and intermediate fusion spectra.
[0083] 3.2.2. Sample Set Partitioning
[0084] In the modeling set, the calibration set and the validation set are divided into ratios of 2:1, 3:1, 4:1, and 5:1, respectively. The optimized spectral preprocessing method and modeling method are used to establish a discriminant model, and the performance of the models built under different ratios is compared.
[0085] 3.2.3. Optimization of Spectral Preprocessing and Modeling Methods
[0086] 3.2.3.1. TQ Analyst Software Modeling
[0087] Spectral data were input into TQ Analyst software to model single infrared spectra of MIR and NIR, as well as primary and intermediate fusion data of the two. Preprocessing methods were selected, including scattering correction methods (Multiplicative Scatter Correction, MSC) and Standard Normal Variate Transformation (SNV); spectral derivative processing methods (First Derivative, 1D) and Second Derivative, 2D); and spectral smoothing methods (Savitzky-Golay smoothing, SG smoothing, and Norris smoothing). A three-factor, three-level orthogonal experiment was designed using these methods (see Table 2). Distance Match (DM) and Discriminant Analysis (DA) were used to build models and compare the optimal preprocessing methods. The Norris smoothing used 5 significant bits and a significant bit interval of 5.
[0088] Table 2. Level of Factors for Discriminant Modeling Conditions
[0089]
[0090] 3.2.3.2. Python Platform Modeling
[0091] There are four commonly used and effective modeling methods for discriminant analysis in the Python platform: Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), and Extra Tree (ET). This study utilizes Python to model single infrared spectra of MIR and NIR, as well as primary and intermediate fusion data of the two. During modeling, the spectra were optimized using both preprocessing methods (Norris smoothing, MSC, D1) and non-preprocessing methods. Since the TQ Analyst software lacks advanced fusion capabilities, advanced fusion was performed on the Python platform.
[0092] 3.2.4. Model Evaluation
[0093] Record the number of misclassifications in the calibration set and prediction set of the model under different modeling combinations, calculate the recognition rate and prediction rate using formula (1) and formula (2), use these as indicators to judge the model effect, and optimize the modeling conditions using single-factor experiments.
[0094] Recognition rate = (Total number of calibration sets - Number of misclassified calibration sets) / Total number of calibration sets × 100% (1)
[0095] Prediction rate = (Total number of prediction sets - Number of misclassifications in prediction sets) / Total number of prediction sets × 100% (2)
[0096] 3.2.5. Model Validation
[0097] Substitute the spectrum corresponding to the external validation sample into each optimized model to obtain the model's prediction result for the sample. The accuracy of the model's prediction result for the external validation sample is judged by calculating the model's external validation prediction rate. The calculation formula is shown in (3).
[0098] External verification (3)
[0099] 4. Results and Discussion
[0100] 4.1. Spectral Characteristic Analysis
[0101] The average NIR spectra of samples from different parts are shown below. Figure 2 It can be seen that the root system is at 8300 cm. -1 6804 cm -1 5623 cm -1 5172 cm -1 4763 cm -1 4378 cm -1 4297 cm -1Characteristic absorption peaks are observed at the location, such as Figure 2 (A); Rhizome at 8299cm -1 6804 cm -1 5623 cm -1 4758 cm -1 4383 cm -1 4297 cm -1 4250 cm -1 Characteristic absorption peaks are observed at the location, such as Figure 2 (B); stem at 8250 cm -1 6804 cm -1 5622 cm -1 5164 cm -1 4763 cm -1 4378 cm -1 4289cm -1 The area exhibits clear characteristic absorption peaks, such as Figure 2 (C); Leaf at 8301 cm -1 6803 cm -1 5624 cm -1 5164 cm -1 4763 cm -1 4378 cm -1 Characteristic absorption peaks are observed at the location, such as Figure 2 (D); Fruit at 8291 cm -1 6807 cm -1 5635 cm -1 5172 cm -1 4762 cm -1 4378 cm -1 It exhibits unique characteristic absorption peaks, such as Figure 2 (E). The main differences in overall spectral characteristics are reflected in the number and position of peaks. Regarding the number of peaks, roots, rhizomes, and stems each exhibit 7 characteristic peaks, while leaves and fruits each have 6. This is one less aromatic CH-related vibration peak compared to the former three, reflecting differences in chemical composition characterization between leaves / fruits and underground parts / stems. Regarding peak position, there are slight shifts in the positions of vibrational peaks of the same type of functional group in different parts, and these shifts are directly related to differences in composition among different parts. The NIR spectral baselines of different parts are relatively stable, with no obvious abnormal vibrational signals, and the peak shapes are generally regular without severe overlap.
[0102] There are 5 common peaks in the NIR spectra of samples from different parts: 8331 cm⁻¹ -1 6804 cm -1 5623 cm -1 4763 cm-1 4378 cm -1 Among them, 8331 cm -1 The vicinity shows CH overtone vibrations, indicating an aliphatic chain or aromatic ring structure; 6804 cm⁻¹ -1 The nearby absorption peak is related to the -NH symmetric stretching vibration, indicating the presence of nitrogen-containing organic components; 5623 cm⁻¹ -1 The nearby absorption peaks, representing overtone / combination vibrations of CH3, indicate the presence of aliphatic compounds; 4763 cm⁻¹ -1 The vicinity likely exhibits acid-OH combination frequency vibrations, indicating the presence of organic acid components; 4378 cm⁻¹ -1 This is related to the overtone / combination frequency vibrations of CH in aromatic hydrocarbons.
[0103] One-way ANOVA was performed on the absorbance values corresponding to the five common peaks in the NIR spectra of samples from different parts of the peach, and the results are shown in Table 3. The results showed that the absorbance differences corresponding to the common peaks among different parts were extremely significant (p < 0.01), reflecting the specificity of the internal chemical components in terms of type and content, providing a basis for the subsequent establishment of a discrimination model for the seven parts of the peach. (Zhang Yaya et al.) Statistical analysis of the near-infrared diffuse reflectance spectra of five parts of Angelica sinensis showed that there were significant differences in the spectral characteristics of different medicinal parts of Angelica sinensis, indicating that there are certain differences in the chemical composition of different medicinal parts, which is consistent with the results of the NIR spectral difference analysis of different parts of Angelica sinensis in the above study.
[0104] Table 3. Results of one-way ANOVA of common peak absorbance values in NIR spectra of samples from different locations.
[0105]
[0106] In Table 3, * indicates a significant difference (p < 0.05); ** indicates an extremely significant difference (p < 0.01).
[0107] The average MIR spectra of samples from different locations are shown below. Figure 3 The number of MIR spectral peaks varied significantly among samples from different locations. The root sample exhibited 17 characteristic peaks, located at 3728 cm⁻¹. -1 3285 cm -1 2918 cm -1 2049 cm -1 1735 cm -1 Near the same wave number, such as Figure 3 (A); The rhizome has 16 characteristic segments, located at 3728 cm. -1 3285 cm -1 2918 cm -12049 cm -1 1735 cm -1 Near the same wave number, such as Figure 3 (B); The stem has 13 absorption peaks, located at 3288 cm⁻¹. -1 2918 cm -1 1735 cm -1 1592 cm -1 1370 cm -1 1316 cm -1 1243 cm -1 Near the same wave number, such as Figure 3 (C); The leaf has 12 absorption peaks, located at 3278 cm⁻¹. -1 2918 cm -1 2850 cm -1 1730 cm -1 1594 cm -1 Near the same wave number, such as Figure 3 (D); The fruit has 15 characteristic points, located at 3278 cm. -1 2923 cm -1 2853 cm -1 1743 cm -1 Near the same wave number, such as Figure 3 (E). The number of characteristic peaks in the MIR spectra of roots and rhizomes is greater than that of stems, leaves and fruits, indicating that the chemical composition of underground parts is more complex.
[0108] There are 8 common peaks in the MIR spectra of samples from different parts: 3285 cm⁻¹ -1 2922 cm -1 1735 cm -1 1591 cm -1 1419 cm -1 1237 cm -1 1074 cm -1 572 cm -1 Among them, 3285 cm -1 The stretching vibration is -OH or -NH, reflecting polar components containing hydroxyl or amino groups; 2922 cm⁻¹ -1 The CH stretching vibration of CH3 / CH2 indicates the presence of an aliphatic compound; 1735 cm⁻¹ -1 This is a C=O stretching vibration, corresponding to esters, carboxylic acids, etc.; 1591 cm⁻¹ -1 The presence of C=C skeletal vibrations in the aromatic ring indicates the presence of aromatic compounds (such as flavonoids and lignans); 1419 cm -1For CH bending or OH stretching vibration; 1237 cm -1 The vibration is CO stretching; 1074 cm⁻¹ is COC or COH vibration, matching the ether bond and hydroxyl structure of lignans; 572 cm⁻¹ -1 The skeletal bending or low-frequency heterocyclic vibrations reflect the characteristics of complex heterocyclic structures.
[0109] One-way ANOVA was performed on the absorbance values corresponding to the eight common peaks in the MIR spectra of samples from different parts of the sample. The results are shown in Table 4. It can be found that each common peak showed extremely significant differences among different parts (p < 0.01), indicating that the composition and content of compounds in different medicinal parts of *Prunus persica* are different, and this difference can be reflected in the MIR spectra.
[0110] Table 4. Results of one-way ANOVA of absorbance values of common peaks in MIR spectra of samples from different locations.
[0111]
[0112] In Table 4, * indicates a significant difference (p < 0.05); ** indicates an extremely significant difference (p < 0.01).
[0113] 4.2. Abnormal Spectral Removal
[0114] Abnormal spectral removal was performed on the NIR, MIR, primary fusion, and intermediate fusion spectra of the root samples. Mahalanobis distance and principal component score plots are shown below. Figure 4 Of the 241 MIR spectra used for modeling, 21 anomalous spectra were removed, leaving 220 for modeling. Figure 4 (B); Of the 241 NIR spectra used for modeling, 14 anomalous spectra were removed, leaving 227 for modeling, such as... Figure 4 (A); The initial fusion yielded 241 modeling spectra. 26 anomalous spectra were removed, leaving 215 spectra for modeling, such as... Figure 4 (C); The initial fusion process removed 27 anomalous spectra, leaving 214 for modeling, such as... Figure 4 (D).
[0115] Abnormal spectral removal was performed on the NIR, MIR, primary fusion, and intermediate fusion spectra of the rhizomes. Mahalanobis distance and principal component score plots are shown below. Figure 5 After removing anomalous spectra in NIR, 208 spectra were retained for modeling, such as... Figure 5 (A); After removing anomalous spectra, 239 spectra remained in the MIR model, such as... Figure 5 (B); After removing anomalous spectra from the primary fusion spectra, 212 spectra remained for modeling, such as... Figure 5(C); After intermediate fusion removed anomalous spectra, 217 spectra remained for modeling, such as... Figure 5 (D).
[0116] Outlier removal was performed on the NIR, MIR, primary fusion, and intermediate fusion modeled spectra of the stem samples. Mahalanobis distance and principal component score plots of the spectral data are shown below. Figure 6 After removing anomalous spectra in NIR, 225 spectra were retained for modeling, such as... Figure 6 (A); After removing anomalous spectra, 220 spectra remained in the MIR model, such as... Figure 6 (B); After removing anomalous spectra from the primary fusion spectra, 224 spectra remained for modeling, such as... Figure 6 (C); After intermediate fusion removed anomalous spectra, 220 spectra remained for modeling, such as... Figure 6 (D).
[0117] Abnormal spectra were removed from the NIR, MIR, primary fusion, and intermediate fusion spectra of the leaf samples. The Mahalanobis distance and principal component scores are shown in Figure 7. After removing abnormal NIR spectra, 209 spectra were retained for modeling. Figure 7 (A); After removing anomalous spectra, 211 spectra remained in the MIR model, such as... Figure 7 (B); After removing anomalous spectra from the primary fusion spectra, 212 spectra remained for modeling, such as... Figure 7 (C); After intermediate fusion removed anomalous spectra, 233 spectra remained for modeling, such as... Figure 7 (D).
[0118] Abnormal spectra were removed from the MIR, NIR, primary fusion, and intermediate fusion spectra of the fruit samples. The Mahalanobis distance and principal component score plots of the spectral data are shown below. Figure 8 After removing anomalous spectra in NIR, 74 spectra were retained for modeling, such as... Figure 8 (A); After removing anomalous spectra, the remaining 82 spectra from the MIR model were used for modeling, such as Figure 8 (B); After removing anomalous spectra from the primary fusion spectra, the remaining 75 spectra were used for modeling, such as... Figure 8 (C); After intermediate fusion removes anomalous spectra, 75 spectra remain for modeling, such as... Figure 8 (D).
[0119] 4.3. Construction of Root Origin Discrimination Model
[0120] 4.3.1. Construction of a root origin discrimination model based on a single infrared spectrum
[0121] Table 5 shows the results of the origin discrimination models for peaches with different proportions established using TQ Analyst software. The model established using NIR spectroscopy shows that when the modeling set is divided into 2:1 proportions, and the DM modeling method is combined with 1D derivatives, SG smoothing, and no-scattering correction, the model performance is optimal. Figure 9 (A) The recognition rate was 100% and the prediction rate was 98.67%, indicating that the model has reliable quantitative prediction capabilities and can meet the prediction needs of the seven origins of peaches. The model established under MIR spectroscopy results show that, at a 5:1 partitioning ratio, the optimal effect is achieved when using DM combined with 1D, SG smoothing, and MSC. (See...) Figure 9 (B) The recognition rate is 100%, and the prediction rate is 75.00%. Comparing the root origin discrimination models established by the two spectra under TQ analyst software, it can be found that the MIR model has a lower prediction rate compared with the NIR model.
[0122] The results of the root origin discrimination model based on the Python platform and single infrared spectroscopy are shown in Table 12. The model established under NIR spectroscopy shows that when the modeling set is divided into 5:1 parts, the model performance is optimal when using the ET modeling method combined with spectral preprocessing. Figure 9 (C) shows a recognition rate of 100% and a prediction rate of 94.87%. MIR model optimization results indicate that, at a 2:1 partition ratio, the model performs best when using SVM combined with spectral preprocessing. (See...) Figure 9 (D) The recognition rate was 100% and the prediction rate was 84.93%. The prediction rate of the NIR spectral model was higher than that of the MIR spectral model, indicating that the NIR spectral model is more suitable for identifying the origin of peach seven roots, while the discrimination model established by the MIR spectral model has poor accuracy.
[0123] Table 5 Performance Comparison of Root Origin Discrimination Models Based on Single Infrared Spectroscopy
[0124]
[0125] Substituting the spectra of the external validation samples into the optimal model, the external validation prediction rates were calculated, and the results are shown in Table 6. Among the single infrared spectral models, the external validation prediction rates of the NIR models were all above 85%, demonstrating overall superior predictive ability compared to the MIR models. Specifically, the NIR model built on the Python platform achieved a prediction rate of 87.50%, while the NIR model built on the TQ analyst platform achieved 86.34%, both meeting the requirements for root origin determination. In contrast, the external validation prediction rates of the MIR models were generally low. The MIR model built on the TQ analyst software achieved an external validation prediction rate of 77.15%, while the MIR model built on the Python platform achieved only 62.50%, making them suitable only for rough determination and preliminary reference.
[0126] Table 6 External validation results of the single infrared spectral root site origin discrimination model
[0127]
[0128] 4.3.2. Construction of a root origin discrimination model based on fused spectra
[0129] The results of the fusion model built using TQ Analyst software are shown in Table 7. For the initial fusion model at a ratio of 4:1, the DM algorithm combined with 2D, Norris smoothing, and MSC preprocessing yielded the best results. Figure 10 (A) The recognition rate is 100%, and the prediction rate is 97.67%. For the intermediate fusion model, at a ratio of 5:1, the optimal model is achieved using the DM algorithm combined with 2D + SGsmoothing + MSC preprocessing. (See...) Figure 10 (B) Recognition rate 100%, prediction rate 100%. In TQ analyst software, model performance improves with increasing fusion level, with the intermediate fusion model achieving a prediction rate of 100% under optimal parameters, demonstrating the best overall performance. The external validation prediction rates of the primary and intermediate fusion spectral best models are both above 90%, exhibiting good prediction performance.
[0130] Results from the root origin discrimination model based on the Python platform and fused spectra show that the primary fused spectra perform best under the conditions of a 5:1 ratio and preprocessing in the ET modeling method. (See...) Figure 10 (C) The recognition rate is 100%, and the prediction rate is 91.89%. The intermediate fusion spectrum performs best under the conditions of a 2:1 ratio and preprocessing using the RF method. See [link to relevant documentation]. Figure 10 (D) The recognition rate is 100%, and the prediction rate is 81.69%. The model obtained by high-level fusion under the NIR spectrum ratio of 5:1 is better; see [reference needed]. Figure 10(E) The model's recognition rate is 100%, and its prediction rate is 97.44%. The primary and intermediate fusion models on the Python platform do not show a significant advantage; only the advanced fusion model achieves a high prediction rate, indicating that its fusion effect depends on the algorithm and feature processing method. Comparing the two modeling systems, the intermediate fusion model built with TQ analyst software outperforms the model built on the Python platform in overall prediction accuracy, meeting the practical needs for accurate identification of the origin of peaches.
[0131] Table 7 Performance Comparison of Root Origin Discrimination Model Based on TQ Analyst Software and Fusion Spectrum
[0132]
[0133] External validation was performed on the optimal discrimination models established for each fused spectrum, and the results are shown in Table 8. The optimal fused spectral models all demonstrated superior discrimination ability compared to the single infrared spectral models. The primary fusion model established using TQ Analyst software achieved the highest external validation prediction rate of 95.83%, demonstrating good predictive performance. The advanced fusion model established using the Python platform also achieved a prediction rate of 91.67%, similarly demonstrating good application value. A comprehensive comparison of the performance of models established using single infrared spectroscopy and fused spectroscopy shows that the fused spectral model effectively improves discrimination accuracy and reliability through the complementarity of multi-source information, and is generally superior to the single infrared spectral model. The advanced fusion model exhibits good predictive performance in the identification of the origin of Chinese medicinal materials; for example, Pei Yifei et al. By combining MIR and NIR with chemometric methods, spectral data of *Paris polyphylla* from central, western, northwestern, southeastern, and southwestern Yunnan were fused and analyzed. After primary, intermediate, and advanced fusion, the performance of all three fusion models was superior to that of the single NIR model. This indicates that information fusion technology helps improve the model's discrimination performance, and that data fusion strategies combined with infrared spectroscopy can be used for rapid identification of production areas.
[0134] Table 8 External validation results of the fusion spectral root region origin discrimination model
[0135]
[0136] 4.4. Construction of a model for identifying the origin of rhizomes
[0137] 4.4.1. Construction of a root and stem origin discrimination model based on single infrared spectrum
[0138] The results of the rhizome origin discrimination model established based on TQ analyst software and single infrared spectroscopy are shown in Table 9. The model established using NIR spectroscopy exhibits the best performance when the modeling set is divided at a ratio of 4:1, and when the DM algorithm is combined with 1D, SG smoothing, and SNV preprocessing. Figure 11 (A) The recognition rate was 100%, and the prediction rate was 92.68%. The model built using MIR spectroscopy performed worse than that built using NIR spectroscopy, with prediction rates all below 50%. At a 3:1 ratio, the model performed best when using DM combined with 2D, Norrissmoothing, and MSC preprocessing. (See...) Figure 11 (B) The recognition rate is 100% and the prediction rate is 43.64%.
[0139] The NIR spectroscopy model built using the Python platform exhibits optimal performance when the modeling set is divided into 4:1 ratios and the ET modeling method is combined with spectral preprocessing. Figure 11 As shown in (C), the recognition rate is 100% and the prediction rate is 88.10%. The model built using MIR spectroscopy performs worse than that built using NIR spectroscopy. At a 4:1 ratio, the model performs relatively better when using SVM combined with spectral preprocessing, such as... Figure 11 As shown in (D), the recognition rate was 96.59% and the prediction rate was 66.22%. The results of the root and stem origin discrimination model show that, among the single infrared spectral models, NIR spectroscopy still shows certain advantages, with prediction rates at a high level; while the MIR spectral model has a low prediction rate and cannot meet the requirements for origin discrimination.
[0140] Table 9. Performance comparison of root and tuber origin identification models based on TQ analyst software and single infrared spectroscopy.
[0141]
[0142] Substituting the spectra of the external validation samples into the optimal models, the external validation prediction rates of each model were calculated, and the results are shown in Table 10. Under TQ Analyst software, the external validation prediction rate of the NIR model was 83.33%, while that of the MIR model was 33.33%. Under the Python platform, the external validation prediction rate of the optimal NIR model was 75.56%, and that of the optimal MIR model was 44.44%. Among the single infrared spectral models, the external validation performance of the NIR spectral model was superior to that of the MIR spectral model. Specifically, the NIR model constructed using TQ Analyst software achieved an external validation prediction rate of 83.33%, enabling accurate prediction of the origin of *Prunus persica* rootstock samples; while the MIR models under TQ Analyst software and the Python platform achieved 33.33% and 44.44% respectively, failing to meet the requirements for accurate discrimination.
[0143] Table 10 External validation results of the single infrared spectrum root and rhizome part origin discrimination model
[0144]
[0145] 4.4.2. Construction of a Rhizome Origin Discrimination Model Based on Fuded Spectra
[0146] The results of the rhizome origin discrimination model established based on TQ Analyst software and fused spectroscopy are shown in Table 11. The predictive performance of the primary fusion model is close to that of the NIR model. When the modeling ratio is 2:1, the model with the DM algorithm combined with 2D + SGsmoothing preprocessing achieves the best results. Figure 12 (A) ), with a recognition rate of 100% and a prediction rate of 92.06%. The intermediate fusion model, at a ratio of 5:1, uses the DM algorithm combined with 2D + SG smoothing preprocessing, and the model is optimal. Figure 12 (B) The recognition rate was 100%, and the prediction rate was 78.79%.
[0147] Among the rhizome origin discrimination models based on the Python platform and fused spectra, the high-level fusion model performed best, followed by the basic fusion model, while the intermediate fusion model performed the worst. The basic fusion spectra model performed best under the conditions of a 5:1 ratio and preprocessing using the ET modeling method. (See...) Figure 12 (C) The recognition rate is 100%, and the prediction rate is 84.38%. The intermediate fusion spectrum performs best under the conditions of a 3:1 ratio and preprocessing using the RF method. See [link to relevant documentation]. Figure 12 (D), with a recognition rate of 100% and a prediction rate of 64.71%. The model obtained through advanced fusion under the NIR spectral ratio of 4:1 is better; see [link to relevant documentation]. Figure 12(E) The model has a recognition rate of 100% and a prediction rate of 90.48%.
[0148] Table 11 Performance Comparison of Root and Rhizome Origin Discrimination Models Based on Fuded Spectra
[0149]
[0150] The external validation results of the model are shown in Table 12. The generalization ability of the fusion model varies depending on the modeling method. In TQ Analyst software, the prediction rate of the basic fusion model is 85.42%, indicating high accuracy in distinguishing unknown samples. On the Python platform, the prediction rates of the advanced fusion model and the NIR model are similar at 75.56%, while the prediction rate of the intermediate fusion model drops to 42.22%, suggesting that unreasonable fusion strategies can easily introduce spectral information redundancy and noise, thereby reducing the model's generalization performance. In conclusion, TQ analyst's NIR single model has the highest prediction rate and is the optimal choice for identifying the origin of Taoerqi rhizomes.
[0151] Table 12 External validation results of the fusion spectral rhizome origin discrimination model
[0152]
[0153] 4.5. Construction of Stem Origin Discrimination Model
[0154] 4.5.1. Construction of a stem origin discrimination model based on a single infrared spectrum
[0155] Table 13 shows the modeling results of the stem origin discrimination model based on TQ Analyst software and single infrared spectroscopy. The NIR model showed the best performance when the modeling set was divided into 2:1 parts and the DM modeling method was used in combination with 2D, SG smoothing, and no-scattering correction preprocessing. Figure 13 (A) The recognition rate is 100%, the prediction rate is 100%, and the model performance is good. The model established under MIR spectroscopy results show that, at a 4:1 division ratio, the modeling parameters of DM combined with 2D, Norris smoothing, and MSC are the best model obtained by MIR optimization, but the overall prediction ability is relatively low. Figure 13 (B)).
[0156] The results of the single infrared spectral stem origin discrimination model built under the Python platform show that when the NIR spectrum is divided into 3:1 parts in the modeling set, and the ET modeling method is used in combination with spectral preprocessing, the model performance is the best. (See...) Figure 13(C) The recognition rate is 100%, and the prediction rate is 96.43%. For the MIR model with a 5:1 partition ratio, the optimal performance is achieved using SVM combined with spectral preprocessing. (See...) Figure 13 (D) The recognition rate was 86.34%, and the prediction rate was 83.78%. Among the models built by TQ Analyst software and the Python platform, the NIR single infrared spectral model showed excellent discrimination ability. The NIR model of TQanalyst software achieved a recognition rate and prediction rate of 100% after optimization, with the best performance. The NIR model of the Python platform also maintained high accuracy. Due to limitations such as information stability, the overall performance of the MIR model was always inferior to that of the NIR model.
[0157] Table 13 Performance Comparison of Stem Origin Discrimination Models Based on Single Infrared Spectroscopy
[0158]
[0159] Substituting the spectra corresponding to the external validation samples into the optimal models, the external validation prediction rates were calculated, and the results are shown in Table 14. Under the TQ analyst software, the external validation prediction rate of the optimal NIR model was 97.92%, exhibiting the strongest predictive ability among all models; the external validation prediction rate of MIR was lower, at only 35.42%. Under the Python platform, the external validation prediction rate of the optimal NIR model was 91.49%, while the prediction rate of the MIR model was lower at 55.32%. In summary, in the construction of stem origin discrimination models based on single infrared spectra, the NIR spectral-based stem origin discrimination models outperformed the MIR models under both the TQ analyst software and Python platform. The external validation performance of the optimal NIR model under the TQ analyst software was particularly outstanding, indicating that NIR spectra contain richer feature information related to stem origin, enabling more effective and accurate identification of stem origin. The low external validation prediction rates of MIR spectra under both algorithms suggest that its reliability as a sole criterion for stem origin discrimination is relatively insufficient.
[0160] Table 14 External validation results of the single infrared spectroscopy stem part origin discrimination model
[0161]
[0162] 4.5.2. Construction of Stem Origin Discrimination Model Based on Fuded Spectra
[0163] Table 15 shows the results of the stem origin discrimination model based on TQ Analyst software and fused spectra. The primary fused spectra, under the 3:1 scale DM modeling method and 2D+SG smoothing preprocessing method, show good model performance. Figure 14 As shown in (A), the recognition rate is 100.00% and the prediction rate is 98.21%. The intermediate fusion modeling method, using a 4:1 ratio DM modeling approach with 2D+SG smoothing+SNV preprocessing, achieves good modeling results, as shown in [example data]. Figure 14 As shown in (B), the recognition rate is 100.00% and the prediction rate is 86.36%.
[0164] The results of the stem origin discrimination model built under the Python platform show that the primary fusion spectrum performs best under the conditions of a 5:1 ratio and preprocessing in the ET modeling method. (See...) Figure 14 (C) The recognition rate is 100%, and the prediction rate is 91.89%. The intermediate fusion spectrum performs best under the conditions of a 4:1 ratio and preprocessing using the SVM method. See [link to relevant documentation]. Figure 14 (D), with a recognition rate of 100% and a prediction rate of 97.37%. The model obtained by performing advanced fusion under the condition of a 4:1 ratio of preprocessed primary fusion spectra is better. See [link to relevant documentation]. Figure 14 (E) The model has a recognition rate of 100% and a prediction rate of 97.78%.
[0165] Table 15 Performance Comparison of Stem Origin Discrimination Models Based on Fuded Spectra
[0166]
[0167] Table 16 shows the external validation results of the fusion spectral models. The external validation prediction rate of the primary fusion spectral model under TQ Analyst software reached 91.67%, compared to 64.58% for the intermediate fusion model, indicating better predictive ability. Under the Python platform, the advanced fusion model performed relatively well with a prediction rate of 78.72%, followed by the primary fusion model with 76.60%. The intermediate fusion model performed poorly, with an external validation prediction rate of only 59.57%. Comparing the modeling results, the performance of the primary fusion model under TQ Analyst software is close to that of the single NIR model, while the performance of the intermediate fusion model declines. The optimized fusion models under the Python platform, especially the advanced fusion model, exhibit high accuracy, similar to the NIR model. In summary, the performance of the peach seven-stem origin discrimination model is similar to that of the rhizome, mainly depending on the spectral type, fusion strategy, and algorithm adaptability. The single NIR spectral model or the optimized advanced fusion model can achieve accurate discrimination.
[0168] Table 16 External validation results of the fusion spectral stem origin discrimination model
[0169]
[0170] 4.6. Construction of Leaf Origin Discrimination Model
[0171] 4.6.1. Construction of a Leaf Origin Discrimination Model Based on a Single Infrared Spectrum
[0172] Table 17 shows the results of the leaf origin discrimination model based on TQ analyst software and single infrared spectroscopy. The model established under NIR spectroscopy indicates that when the modeling set is divided at a ratio of 3:1, and DM combined with 2D and SG smoothing is used, the model performance is optimal. Figure 15 (A)), with a recognition rate of 100% and a prediction rate of 98.08%. The prediction rate of the model built using MIR single infrared spectroscopy is worse than that of NIR spectroscopy, with prediction rates all below 80%. At a 4:1 ratio, the best results are achieved when using DM combined with 1D, Norris smoothing, and SNV preprocessing. Figure 15 (B) The recognition rate was 100%, and the prediction rate was 78.57%.
[0173] Results of a leaf origin discrimination model based on Python platform and single infrared spectroscopy show that the NIR model performs best when the modeling set is divided into 5:1 groups and the ET modeling method is combined with spectral preprocessing. (See...) Figure 15 (C) The recognition rate was 95.40%, and the prediction rate was 91.43%. Model results established under MIR spectroscopy show that, at a 5:1 partition ratio, the model performance is optimal when using SVM combined with spectral preprocessing. (See...) Figure 15 (D), with a recognition rate of 97.71% and a prediction rate of 83.33%. NIR spectroscopy can better reflect the chemical information related to the origin of the leaves, and the overall discrimination accuracy of the established model is high. However, MIR spectroscopy is easily affected by the surface condition of the sample, and the prediction effect is significantly lower, only achieving preliminary discrimination.
[0174] Table 17 Performance Comparison of Leaf Origin Discrimination Models Based on Single Infrared Spectroscopy
[0175]
[0176] The spectra corresponding to the external validation samples were substituted into the optimal model, and the external validation prediction rates were calculated. The results are shown in Table 18. The NIR spectral models exhibited strong generalization performance, with high external validation prediction rates. Among them, the optimal NIR model built with TQ Analyst software achieved an external validation prediction rate of 91.30%, demonstrating the most outstanding generalization ability and enabling accurate identification of the origin of *Prunus armeniaca* leaf samples. The external validation prediction rates of MIR spectral models were generally low, failing to reach 70% in both TQ Analyst software and the Python platform, indicating limited generalization ability and difficulty in meeting the requirements for high-precision identification.
[0177] Table 18 External validation results of the single infrared spectroscopy leaf origin discrimination model
[0178]
[0179] 4.6.2. Construction of Leaf Origin Discrimination Model Based on Fuded Spectra
[0180] The results of the fusion model built using TQ Analyst software are shown in Table 19. The initial fusion model, with a modeling ratio of 5:1, achieves the best results using the DM algorithm combined with 1D + SG smoothing preprocessing. Figure 16 As shown in (A), the recognition rate is 100% and the prediction rate is 100%. The intermediate fusion model, with a ratio of 2:1, achieves optimal performance using the DM algorithm combined with 2D and SG smoothing preprocessing. Figure 16 As shown in (B), the recognition rate is 94.12% and the prediction rate is 84.78%. The performance of the fusion models shows significant differences. Primary fusion can effectively integrate multi-source spectral information and can achieve extremely high discrimination accuracy under appropriate algorithm and preprocessing conditions.
[0181] The results of the fusion model built using the Python platform are shown in Table 33. The primary fusion spectrum performs best under the conditions of a 4:1 ratio and preprocessing using the ET modeling method. Figure 16 (C) The recognition rate is 100%, and the prediction rate is 93.02%. The intermediate fusion spectrum performs best under the conditions of a 3:1 ratio and preprocessing using the ET method. See [link to relevant documentation]. Figure 16 (D), with a recognition rate of 100% and a prediction rate of 92.59%. The model obtained by performing advanced fusion under the condition of a 4:1 ratio of preprocessed primary fusion spectra is better. See [link to relevant documentation]. Figure 16 (E) The model has a recognition rate of 100% and a prediction rate of 93.02%.
[0182] Table 19 Performance Comparison of Leaf Origin Discrimination Models Based on Fuded Spectra
[0183]
[0184] Table 20 shows the external validation results of the fusion spectral models. The generalization abilities of the fusion models vary, and a reasonable fusion strategy can effectively improve the model's generalization performance. Under TQ Analyst software, the external validation prediction rate of the primary fusion model reached 89.13%, which is better than the intermediate fusion model. Under the Python platform, the external validation prediction rates of both the primary and advanced fusion models were 84.78%, while the intermediate fusion model showed weaker generalization ability. In summary, the external validation effect of the *Prunus armeniaca* origin discrimination model mainly depends on the applicability of the spectral type and the rationality of the fusion strategy. The models built using NIR spectra and optimized fusion spectra can effectively improve the model's generalization ability.
[0185] Table 20 External validation results of the leaf origin discrimination model based on fused spectra
[0186]
[0187] 4.7. Construction of Fruit Origin Discrimination Model
[0188] 4.7.1. Construction of a Fruit Origin Discrimination Model Based on Single Infrared Spectroscopy
[0189] The results of the single infrared spectral fruit origin discrimination model established using TQ Analyst software are shown in Table 21. The NIR model showed the best performance when the modeling set was divided into 2:1 groups, and when using the DM modeling method combined with 1D and SG smoothing preprocessing. Figure 17 As shown in (A), the recognition rate is 100%, and the prediction rate is 87.05%. The prediction rate of the MIR model is worse than that of the NIR model, with prediction rates all below 70%. At a 2:1 ratio, the model performs best when using DM combined with 1D, SG smoothing, and MSC preprocessing. Figure 17 As shown in (B), the recognition rate is 100% and the prediction rate is 62.96%.
[0190] Results of a leaf origin discrimination model based on Python platform and single infrared spectroscopy show that the NIR model performs best when the modeling set is divided into 2:1 groups and the ET modeling method is combined with spectral preprocessing. (See...) Figure 17 (C) The recognition rate is 100%, and the prediction rate is 77.78%. At a 5:1 ratio, the MIR model performs best when using SVM combined with spectral preprocessing. (See...) Figure 17 (D), with a recognition rate of 98.55% and a prediction rate of 80.00%.
[0191] Table 21 Performance Comparison of Fruit Origin Discrimination Models Based on Single Infrared Spectroscopy
[0192]
[0193] Substituting the spectra corresponding to the external validation samples into the optimal model, the external validation prediction rates of the models were calculated, and the results are shown in Table 22. Under TQ Analyst software, the external validation prediction rate of the optimal NIR model was 73.33%; the external validation prediction rate of MIR was lower, at only 33.33%. Under the Python platform, MIR had the highest external validation prediction rate at 81.25%, followed by the NIR model with an external validation prediction rate of 68.75%.
[0194] Table 22 External validation results of the fruit origin discrimination model based on single infrared spectrum
[0195]
[0196] 4.7.2. Construction of a Fruit Origin Discrimination Model Based on Fuded Spectra
[0197] The fusion model results built using TQ Analyst software are shown in Table 23. The initial fusion model, with a modeling ratio of 3:1, achieved the best results using the DM algorithm combined with 2D + SG smoothing preprocessing. Figure 18 (A) The recognition rate is 100%, and the prediction rate is 77.78%. For the intermediate fusion model at a ratio of 4:1, the optimal model is achieved using the DM algorithm combined with 1D, Norrissmoothing, and MSC preprocessing. (See...) Figure 18 (B) The recognition rate is 100% and the prediction rate is 80.00%.
[0198] The results of the fusion model built on the Python platform show that the model performs best under the conditions of a 5:1 scale and preprocessing using the ET modeling method. The modeling results are shown in the figure below. Figure 18 As shown in (C), the recognition rate is 100% and the prediction rate is 84.62%. The intermediate-level fusion spectrum performs best under the conditions of a 3:1 ratio and preprocessing using the SVM method. The modeling results are shown in the figure below. Figure 18 As shown in (D), the recognition rate is 100% and the prediction rate is 84.21%. The model obtained by performing advanced fusion under the condition of a 4:1 ratio of preprocessed primary fusion spectra is better. The modeling results are shown in the figure below. Figure 18 As shown in (E), the model has a recognition rate of 100% and a prediction rate of 84.62%.
[0199] Table 23 Performance Comparison of Fruit Origin Discrimination Models Based on Fuded Spectra
[0200]
[0201] External validation of the models established using fused spectra was performed, and the results are shown in Table 24. The external validation prediction rates for each model did not exceed 70%, indicating poor predictive ability for unknown samples. Under TQ Analyst software, the external validation prediction rate for the basic fused spectra model was 58.82%, and for the intermediate fused model, it was 52.94%. Under the Python platform, the external validation prediction rate for the advanced fused model was 68.75%. The prediction rates for the basic and intermediate fused models were relatively low, at 62.50% and 56.25%, respectively.
[0202] Analysis of the external validation results of the peach seven-fruit origin discrimination model reveals significant differences in generalization ability among different spectral and modeling methods. Under TQ Analyst software, the NIR model outperforms the MIR model and various fusion models in external validation, indicating that fusion strategies failed to improve the model's generalization performance. On the Python platform, the opposite trend is observed: the MIR model achieves the highest external validation prediction rate, outperforming the NIR model and fusion models. Overall, this suggests that the tissue structure and component distribution of fruit parts are complex, and the information advantage of different spectra depends on the algorithm's adaptability. A reasonable combination of spectra and algorithms can improve the model's discrimination ability.
[0203] Table 24 External validation results of the fruit origin discrimination model based on fused spectra
[0204]
[0205] 5. Conclusion
[0206] Using 291 samples of *Prunus persica* from 15 production areas as the subjects, we integrated their NIR and MIR spectral data, and constructed a model for identifying medicinal parts and production areas based on single infrared spectroscopy and fused spectroscopy by optimizing modeling parameters, spectral preprocessing methods and sample set division ratios.
[0207] 1. The results of the origin discrimination model for the seven parts of peach (root, stem, leaf, and fruit) show that the spectral type, fusion strategy, and algorithm adaptability all have a significant impact on the discrimination effect of the model, and the optimal discrimination model differs for different parts.
[0208] 2. For root parts, the intermediate fusion model in TQ Analyst software was the best, with a prediction rate of 100% and an external validation prediction rate of 93.75%, which was superior to the single infrared spectroscopy model. For root and stem parts, the primary fusion model in TQ Analyst software was the best, with a prediction rate of 92.06% and an external validation prediction rate of 85.42%. For stem parts, the origin discrimination model established by NIR single infrared spectroscopy in TQ Analyst software was the best, with a prediction rate of 100% and an external validation prediction rate of 97.92%. For leaves, the primary fusion model in TQ Analyst software was the best, with a prediction rate of 100% and an external validation prediction rate of 89.13%. For fruits, the MIR model under the Python platform was the best, with a prediction rate of 80.00% and an external validation prediction rate of 81.25%, which was the best modeling scheme for this part.
[0209] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a method for tracing the origin of Taoerqi (a type of peach) based on infrared spectral fusion and machine learning is provided, including:
[0210] The spectral acquisition module is used to acquire multiple single infrared spectral data from different parts of peach samples from different origins to be traced. The multiple single infrared spectral data include near-infrared spectral data and mid-infrared spectral data.
[0211] The spectral fusion module is used to concatenate the near-infrared spectral data and the mid-infrared spectral data to obtain primary fused data; and to perform feature extraction and fusion on the primary fused data to obtain intermediate fused spectral data.
[0212] The origin identification model construction module is used to construct origin identification models based on the single infrared spectral data, primary fusion data, and intermediate fusion data in the same part, respectively using the Python platform and TQ analyst software, to obtain different levels of origin identification models corresponding to different parts;
[0213] The identification output module is used to identify the origin of the peach 7 sample to be traced using the origin identification model constructed in the origin identification model construction module, and output the origin traceability results.
[0214] The above detailed embodiments are a description of the present invention. It should not be considered that the specific embodiments of the present invention are limited to these descriptions. For those skilled in the art, several simple deductions and substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the protection scope of the present invention.
Claims
1. A method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning, characterized in that, Includes the following steps: S1. Collect multiple single infrared spectral data from different parts of peach samples to be traced from different origins, including near-infrared spectral data and mid-infrared spectral data; S2. The near-infrared spectral data and mid-infrared spectral data are concatenated to obtain primary fused data; the primary fused data is then subjected to feature extraction and fusion to obtain intermediate fused spectral data; S3. Based on the single infrared spectral data, primary fusion data, and intermediate fusion data at the same location, use the Python platform and TQ analyst software to construct origin discrimination models respectively, and obtain different levels of origin discrimination models corresponding to different locations; S4. Use the origin discrimination model constructed in step S3 to determine the origin of the tested peach sample and output the origin traceability results.
2. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Step S3 also includes: By using the Python platform, different origin discrimination models are fused at the decision level to form an advanced fused origin discrimination model.
3. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, The samples of *Prunus cerasifera* to be tested were collected from populations at 15 different collection sites in Qinghai Province, Yunnan Province, and Tibet Autonomous Region. The populations at each collection site were more than 20 km apart. At least 20 healthy plants were collected from each population at the same collection site, and each plant was spaced 10 m apart. Each *Prunus cerasifera* sample was dried, pulverized, and passed through an 80-mesh sieve at different parts before being placed in a desiccator for analysis.
4. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, The acquisition of the near-infrared spectral data includes: Take appropriate amounts of powder from different parts of the peach kernel and place them on filter paper. Use a Fourier transform infrared spectrometer (NIR fiber module) to analyze the sample at 10000-4000 cm⁻¹. -1 Spectral data were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 6 cm⁻¹. -1 The scan was performed 64 times, with air as a reference. Each peach sample was collected 3 times, and the average spectrum was taken for analysis. The acquisition of the mid-infrared spectral data includes: Appropriate amounts of *Prunus persica* powder from different parts were placed on the attenuated total reflectance infrared probe of a Fourier transform infrared spectrometer, at 4000-400 cm⁻¹. -1 MIR spectra were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 4 cm⁻¹. -1 The sample was scanned 32 times, with air as a reference. Each sample was collected 3 times, and the average spectrum was taken for analysis.
5. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Step S1 also includes: The near-infrared and mid-infrared spectral data are preprocessed, including one or more combinations of scattering correction, derivative processing, and spectral smoothing. The scattering correction method is multivariate scattering correction or standard normal transformation. The derivative processing method is first derivative or second derivative. The spectral smoothing method is Savitzky-Golay smoothing or Norris smoothing.
6. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, In step S3, when constructing the origin discrimination model using the Python platform or TQ analyst software, multiple origin discrimination models are constructed using different machine learning methods, preprocessing methods, model set ratios, and optical path types.
7. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 6, characterized in that, The machine learning methods used in the TQ analyst software include distance matching and discriminant analysis, while the machine learning methods used on the Python platform include support vector machines, decision trees, random forests, and limit trees.
8. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Before constructing the origin discrimination model, outlier removal is performed on single infrared spectral data, primary fusion data, and intermediate fusion data. The outlier removal methods include Mahalanobis distance method and principal component analysis method.
9. The method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Also includes: In step S4, the optimal origin discrimination model is selected, and the origin of the test peach sample is determined using the optimal origin discrimination model; The optimal origin selection model includes: For samples from the same location, compare the discrimination effects of different level origin discrimination models and select the best-performing origin discrimination model as the optimal origin discrimination model for that location.
10. A method for tracing the origin of Taoerqi peaches based on infrared spectral fusion and machine learning, characterized in that, include: The spectral acquisition module is used to acquire multiple single infrared spectral data from different parts of peach samples from different origins to be traced. The multiple single infrared spectral data include near-infrared spectral data and mid-infrared spectral data. The spectral fusion module is used to concatenate the near-infrared spectral data and the mid-infrared spectral data to obtain primary fused data; and to perform feature extraction and fusion on the primary fused data to obtain intermediate fused spectral data. The origin discrimination model construction module is used to construct origin discrimination models based on the single infrared spectral data, primary fusion data, and intermediate fusion data in the same part, respectively using the Python platform and TQ analyst software, to obtain different levels of origin discrimination models corresponding to different parts; The identification output module is used to identify the origin of the peach 7 sample to be traced using the origin identification model constructed in the origin identification model construction module, and output the origin traceability results.