A method and system for identifying parts of radix sinopodophyli based on infrared spectrum fusion and machine learning

CN122591602APending Publication Date: 2026-08-18NORTHWEST INST OF PLATEAU BIOLOGY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610684375.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

然而,现有技术通常仅采用单一红外光谱(仅NIR或仅MIR)结合传统化学计量学方法进行定性判别,光谱信息维度单一,未能充分挖掘不同波段下药物化学成分的特征互补信息,导致模型的判别准确率、稳定性和泛化能力仍有较大提升空间

Benefits of technology

[0034]1. This invention overcomes the limitations of single-dimensional infrared spectral information by effectively fusing near-infrared spectroscopy (reflecting the overtone and combination frequencies of hydrogen-containing groups) with mid-infrared spectroscopy (reflecting the fundamental vibrational frequencies of molecular groups) (including primary, intermediate, and advanced fusion), achieving complementarity and enhancement of chemical characteristic information across different bands. Experimental data shows that the overall performance of the fused spectral models is superior to that of the single infrared spectral models. The intermediate and advanced fusion models on the Python platform achieve 100% recognition and prediction rates, with the advanced fusion model also achieving 100% external validation prediction rate. All fused spectral models achieve external validation prediction rates above 80%, indicating that the established seven-part discrimination model for peaches can accurately distinguish different parts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122591602A_ABST
    Figure CN122591602A_ABST
Patent Text Reader

Abstract

The application discloses a peachwood root part discrimination method and system based on infrared spectrum fusion and machine learning, and belongs to the technical field of traditional Chinese medicinal material identification. The method comprises the following steps: collecting near-infrared spectrum and mid-infrared spectrum of a peachwood root sample; pre-processing the two kinds of spectrum; performing data fusion on the pre-processed spectrum to obtain fused spectrum data; inputting the fused spectrum data into a pre-trained machine learning classification model to output a medicinal part discrimination result. The application integrates the complementary chemical information of near-infrared spectrum and mid-infrared spectrum, and combines advanced machine learning algorithms such as support vector machines to construct a multi-level spectrum fusion model. Experimental results show that the discrimination accuracy of the intermediate fusion or high-level fusion model for different medicinal parts of the peachwood root, such as roots, rhizomes, stems, leaves and fruits, can reach 100%. The method is rapid, non-destructive, accurate and low in cost, and provides an efficient technical means for the quality control of the rare traditional Chinese medicinal material peachwood root and the clinical safe use of the peachwood root.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Chinese medicinal material identification technology, and in particular to a method and system for identifying seven parts of peach based on infrared spectral fusion and machine learning. Background Technology

[0002] The medicinal parts of Chinese medicinal herbs are one of the core factors determining their quality, efficacy stability, and safety in clinical use. Different parts of the same medicinal plant (roots, stems, leaves, fruits, etc.) exhibit variations in the types and contents of their active ingredients during growth and development, directly affecting the medicinal value and clinical efficacy of the herb.

[0003] *Sinopodophyllum hexandrum* (Royle) is a perennial herbaceous plant belonging to the genus *Sinopodophyllum* in the family Berberidaceae. It is a rare and endangered medicinal plant in my country, distributed in Qinghai, Yunnan, Sichuan, and Tibet. Wild populations mainly grow in flat valleys, understory forests, forest edges, and thickets at high altitudes of 2700–4500 m. The medicinal parts are diverse; the roots and rhizomes, leaves, and fruits can all be used medicinally. The fruit is called "Xiaoye Lian," while the underground parts (roots and rhizomes) are called "Tao'erqi" or "Guijiu." The medicinal effects and active ingredients of different parts vary significantly, each with its own emphasis on medicinal value.

[0004] In practical applications, *Prunus persica* also faces challenges such as difficulty in identifying medicinal parts and insufficiently standardized identification methods, which hinder its quality evaluation and safe use. Traditional Chinese medicine methods for identifying medicinal parts rely heavily on empirical morphological identification, which suffers from inherent defects such as low accuracy, strong subjectivity, and poor repeatability, making it difficult to meet the precision requirements of modern Chinese medicine quality control.

[0005] Currently, the identification of different medicinal parts of *Taoerqi* still relies on traditional experience, lacking rapid and accurate scientific testing methods. This easily leads to confusion and misuse of parts, seriously restricting its quality standardization control, rational resource development, and standardized clinical application.

[0006] With the deepening research on the quality control of traditional Chinese medicine, high-performance liquid chromatography (HPLC), gas chromatography (GC), mass spectrometry (MS), and their coupled techniques have become the mainstream methods for component analysis and medicinal part identification of Chinese medicinal materials. However, these methods rely on expensive and sophisticated instruments, have cumbersome and complex sample pretreatment procedures, require high levels of professional skills from operators, and are costly and inefficient, making them difficult to apply to large-scale sample screening and grassroots promotion. Therefore, establishing a low-cost, easy-to-operate, rapid, efficient, and widely applicable analytical technique for different medicinal parts of *Prunus persica* has become a key issue that urgently needs to be addressed.

[0007] With the advantages of simple operation, rapid detection, simple sample pretreatment, and low analysis cost, infrared spectroscopy technology has been widely used in research aspects such as the origin identification, species identification, part identification, content determination, and quality control in the production process of medicinal materials. Especially near-infrared spectroscopy and mid-infrared spectroscopy, with the advantages of rapidity, non-destructiveness, simple sample pretreatment, and low cost, have been preliminarily applied to the origin, species, and quality evaluation of Chinese herbal medicines. However, existing technologies usually only use a single infrared spectrum (only NIR or only MIR) combined with traditional chemometric methods for qualitative discrimination. The spectral information dimension is single, and the characteristic complementary information of drug chemical components in different bands fails to be fully exploited, resulting in a large room for improvement in the discrimination accuracy, stability, and generalization ability of the model. Especially for Sinopodophyllum hexandrum, a rare medicinal material with multiple medicinal parts and significant differences in components, there is currently no reliable technical solution that can fuse multi-source spectral information and use advanced machine learning algorithms for high-precision and intelligent part discrimination. Summary of the Invention

[0008] The purpose of the present invention is to overcome the technical problems existing in the prior art, and provide a method and system for discriminating the parts of Sinopodophyllum hexandrum based on infrared spectrum fusion and machine learning, optimize key conditions such as spectral pretreatment, modeling method, and sample division ratio, respectively construct medicinal part discrimination models for single spectrum and fusion spectrum, and determine the best modeling scheme through comparative analysis, so as to provide a scientific basis and technical support for the accurate identification of different medicinal parts of Sinopodophyllum hexandrum.

[0009] The purpose of the present invention is achieved through the following technical solutions:

[0010] In the first aspect of the present invention, a method for discriminating the parts of Sinopodophyllum hexandrum based on infrared spectrum fusion and machine learning is provided, including the following steps:

[0011] S1. Collect a variety of single infrared spectrum data of different parts of the待测 Sinopodophyllum hexandrum samples, and the variety of single infrared spectrum data includes near-infrared spectrum data and mid-infrared spectrum data;

[0012] S2. Concatenate the near-infrared spectrum data and mid-infrared spectrum data of different parts to obtain primary fusion data; perform feature extraction and fusion on the primary data fusion data to obtain intermediate fusion spectrum data;

[0013] S3. Based on the single infrared spectrum data, primary fusion data, and intermediate fusion data, respectively construct part discrimination models using the Python platform and TQ analyst software;

[0014] S4. Use the part discrimination models constructed in step S3 to discriminate different parts of the待测 Sinopodophyllum hexandrum samples, and output the discrimination results of the medicinal parts.

[0015] In some embodiments, step S3 further includes:

[0016] By using the Python platform, decision-level fusion of different part discrimination models is performed to form an advanced fused part discrimination model.

[0017] In some embodiments, the samples of *Prunus cerasifera* to be tested were collected from populations at 15 different collection sites in Qinghai Province, Yunnan Province, and Tibet Autonomous Region. The populations at each collection site were spaced more than 20 km apart. At least 20 healthy plants were collected from each population at the same collection site, and each plant was spaced 10 m apart. Each *Prunus cerasifera* sample was dried, pulverized, and passed through an 80-mesh sieve at different parts before being placed in a desiccator for analysis.

[0018] In some embodiments, the acquisition of the near-infrared spectral data includes:

[0019] Take appropriate amounts of powder from different parts of the peach kernel and place them on filter paper. Use a Fourier transform infrared spectrometer (NIR fiber module) to analyze the sample at 10000-4000 cm⁻¹. -1 Spectral data were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 6 cm⁻¹. -1 The scan was performed 64 times, with air as a reference. Each peach sample was collected 3 times, and the average spectrum was taken for analysis.

[0020] The acquisition of the mid-infrared spectral data includes:

[0021] Appropriate amounts of *Prunus persica* powder from different parts were placed on the attenuated total reflectance infrared probe of a Fourier transform infrared spectrometer, at 4000-400 cm⁻¹. -1 MIR spectra were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 4 cm⁻¹. -1 The sample was scanned 32 times, with air as a reference. Each sample was collected 3 times, and the average spectrum was taken for analysis.

[0022] In some embodiments, step S1 further includes:

[0023] The near-infrared and mid-infrared spectral data are preprocessed, including one or more combinations of scattering correction, derivative processing, and spectral smoothing. The scattering correction method is multivariate scattering correction or standard normal transformation. The derivative processing method is first derivative or second derivative. The spectral smoothing method is Savitzky-Golay smoothing or Norris smoothing.

[0024] In some embodiments, when constructing the part discrimination model using the Python platform or TQ analyst software in step S3, multiple part discrimination models are constructed using different machine learning methods, preprocessing methods, model set proportions, and optical path types.

[0025] In some embodiments, the machine learning methods used on the TQ analyst software include distance matching and discriminant analysis, while the machine learning methods used on the Python platform include support vector machines, decision trees, random forests, and limit trees.

[0026] In some embodiments, outlier removal is performed on single infrared spectral data, primary fusion data, and intermediate fusion data before constructing the location discrimination model. The outlier removal methods include Mahalanobis distance and principal component analysis.

[0027] A second aspect of the present invention provides a peach seven-part identification system based on infrared spectral fusion and machine learning, comprising:

[0028] The spectral acquisition module is used to acquire multiple single infrared spectral data from different parts of the sample to be tested, including near-infrared spectral data and mid-infrared spectral data.

[0029] The spectral fusion module is used to concatenate near-infrared and mid-infrared spectral data from different regions to obtain primary fused data; feature extraction and fusion are performed on the primary fused data to obtain intermediate fused spectral data.

[0030] The part discrimination model construction module is used to construct part discrimination models based on the single infrared spectral data, primary fusion data, and intermediate fusion data, respectively, using the Python platform and TQ analyst software.

[0031] The identification output module is used to identify different parts of the tested *Prunus persica* sample using the part identification model constructed in the part identification model construction module, and output the medicinal part identification results.

[0032] It should be further noted that the technical features corresponding to the above-mentioned options and embodiments can be combined or substituted with each other to form new technical solutions without conflict.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] 1. This invention overcomes the limitations of single-dimensional infrared spectral information by effectively fusing near-infrared spectroscopy (reflecting the overtone and combination frequencies of hydrogen-containing groups) with mid-infrared spectroscopy (reflecting the fundamental vibrational frequencies of molecular groups) (including primary, intermediate, and advanced fusion), achieving complementarity and enhancement of chemical characteristic information across different bands. Experimental data shows that the overall performance of the fused spectral models is superior to that of the single infrared spectral models. The intermediate and advanced fusion models on the Python platform achieve 100% recognition and prediction rates, with the advanced fusion model also achieving 100% external validation prediction rate. All fused spectral models achieve external validation prediction rates above 80%, indicating that the established seven-part discrimination model for peaches can accurately distinguish different parts.

[0035] 2. This invention, through orthogonal experimental optimization of preprocessing methods such as multivariate scattering correction, standard normal transformation, derivatives, and smoothing on the original spectrum, can effectively eliminate baseline drift, light scattering noise, and random noise, extracting purer and significantly different spectral features. Furthermore, by combining machine learning algorithms with strong generalization capabilities, such as support vector machines, and scientifically dividing the calibration and validation sets, the constructed discriminant model exhibits stable and excellent predictive performance on both the internal validation set and the unknown external validation set, with a low risk of overfitting.

[0036] 3. This invention only requires placing the medicinal powder on the spectrometer accessory to complete spectral acquisition within seconds to minutes, eliminating the need for complex sample pretreatment processes (such as tedious extraction, separation, and purification) required by techniques like chromatography and mass spectrometry. The infrared spectrometer used is significantly less expensive than large, precision instruments like mass spectrometers, and it also reduces the professional skill requirements for operators, making it easy to promote and use in grassroots units such as medicinal material purchasing stations, traditional Chinese medicine processing plants, and drug testing institutions, greatly improving detection efficiency. Attached Figure Description

[0037] Figure 1 This is a flowchart illustrating a method for identifying seven parts of a peach based on infrared spectral fusion and machine learning, as shown in an embodiment of the present invention.

[0038] Figure 2 The following are NIR spectra of different parts of a *Prunus persica* sample as shown in an embodiment of the present invention;

[0039] Figure 3 The images shown are MIR spectra of different parts of a *Prunus persica* sample as illustrated in this embodiment of the invention.

[0040] Figure 4 This is a root sample anomaly spectrum discrimination diagram shown in an embodiment of the present invention;

[0041] Figure 5 This is an abnormal spectral discrimination diagram of a rhizome sample shown in an embodiment of the present invention;

[0042] Figure 6 This is a spectrum discrimination diagram of an abnormal stem sample shown in an embodiment of the present invention;

[0043] Figure 7 This is an abnormal spectral discrimination diagram of a leaf sample shown in an embodiment of the present invention;

[0044] Figure 8 This is an abnormal spectral discrimination diagram of a fruit sample shown in an embodiment of the present invention;

[0045] Figure 9 This is a 3D diagram of a site discrimination model based on TQ analyst software and two single infrared spectra, as shown in an embodiment of the present invention.

[0046] Figure 10 The results of the site discrimination model based on the Python platform and two single infrared spectra are shown in the embodiments of the present invention.

[0047] Figure 11 This is a 3D diagram of a site discrimination model based on TQ analyst software and fused spectra, as shown in an embodiment of the present invention.

[0048] Figure 12 This is a schematic diagram illustrating the optimal result of the site discrimination model based on the Python platform and fused spectra in an embodiment of the present invention. Detailed Implementation

[0049] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] It should be noted that the defects in the solutions in the prior art are all the results of the inventors' practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of this application in the following text should be the inventors' contributions to this application in the process of invention and creation, and should not be understood as technical content known to those skilled in the art.

[0051] In one exemplary embodiment, a method for identifying seven parts of a peach based on infrared spectral fusion and machine learning is provided, such as... Figure 1 As shown, it includes the following steps:

[0052] S1. Collect multiple single infrared spectral data from different parts of the sample to be tested, including near-infrared spectral data and mid-infrared spectral data;

[0053] S2. Concatenate near-infrared and mid-infrared spectral data from different locations to obtain primary fused data; extract and fuse features from the primary fused data to obtain intermediate fused spectral data;

[0054] S3. Based on the single infrared spectral data, primary fusion data, and intermediate fusion data, construct part discrimination models using the Python platform and TQ analyst software, respectively;

[0055] S4. Use the part discrimination model constructed in step S3 to identify different parts of the sample of *Prunus persica* to be tested, and output the medicinal part discrimination results.

[0056] Based on the above steps, this embodiment provides a specific experimental method, mainly including:

[0057] 1. Instruments

[0058] Fourier transform infrared spectrometer (iS 50, Thermo Nicolet, USA) (equipped with fiber optic cable and attenuated total reflection accessories), oven (Shanghai Yiheng Scientific Instruments Co., Ltd., China), and pulverizer (Tianjin Tester Co., Ltd., China).

[0059] 2. Sample Source

[0060] During the fruiting period of *Sinopodophyllum hexandrum* in July and August 2024, samples were collected from 15 different production areas in Qinghai Province, Yunnan Province, and Tibet Autonomous Region. All samples were wild *Sinopodophyllum hexandrum* (National Key Protected Wild Plant Collection Certificate, No.: 0043349), and the original plant specimens were identified as *Sinopodophyllum hexandrum* (Specimen No.: 2024-002), belonging to the genus *Sinopodophyllum* of the family Berberidaceae. The populations at each collection site were spaced at least 20 km apart, with at least 20 healthy plants collected from each population, spaced 10 m apart. The samples were divided into five different parts (root, stem, leaf, rhizome, and fruit), dried, pulverized, and passed through an 80-mesh sieve before being placed in a desiccator for analysis. Sample information is shown in Table 1.

[0061] Table 1. Source information of peach samples from different origins

[0062]

[0063] 3. Experimental Methods

[0064] 3.1 Infrared Spectral Acquisition and Spectral Fusion

[0065] 3.1.1. NIR Spectral Acquisition

[0066] Take appropriate amounts of sample powder from different parts of the sample and place them on filter paper. Use a Fourier transform infrared spectrometer (NIR) fiber optic module to measure the sample at 10000-4000 cm⁻¹. -1 Spectral data were acquired within the specified range. Background interference such as CO2 and water was subtracted in real-time during acquisition, with a scanning resolution of 6 cm⁻¹. -1 The number of scans was 64. Using air as a reference, each sample was collected 3 times, and the average spectrum was used for analysis.

[0067] 3.1.2. MIR Spectral Acquisition

[0068] Appropriate amounts of sample powder from different locations were placed on the attenuated total reflectance infrared probe of a Fourier transform infrared spectrometer, and the sample was analyzed at 4000-400 cm⁻¹. -1 MIR spectra were acquired within the specified range, with background interference from CO2 and water subtracted in real time during acquisition. The scanning resolution was 4 cm⁻¹. -1 The sample was scanned 32 times, with air as a reference. Each sample was collected 3 times, and the average spectrum was used for analysis.

[0069] 3.1.3. Spectral Fusion

[0070] Data fusion was performed using existing spectral fusion methods. The NIR and MIR data of the samples were concatenated using a Python platform to obtain primary fused data. Logistic regression was then used on the primary fused spectral data to calculate the contribution of each feature point, and features with high contributions were extracted for further fusion to obtain intermediate fused data. Advanced fusion, or decision-level fusion, used Python to fuse models from different modeling methods at the decision-level, forming the advanced fused model. The weighting coefficients were based on the performance of each model.

[0071] 3.2. Construction of a model for identifying medicinal parts and origins based on infrared spectroscopy technology

[0072] 3.2.1. Outlier Removal

[0073] Using infrared spectra of different parts of *Prunus persica* collected under section "3.1", an external validation set was randomly divided at a ratio of 5:1, with the remaining samples serving as the modeling set. Outliers were removed from the modeling set samples using Marginal Distance (MD) and Principal Component Analysis (PCA), specifically for NIR, MIR, primary fusion, and intermediate fusion spectra.

[0074] 3.2.2. Sample Set Partitioning

[0075] In the modeling set, the calibration set and the validation set are divided into ratios of 2:1, 3:1, 4:1, and 5:1, respectively. The optimized spectral preprocessing method and modeling method are used to establish a discriminant model, and the performance of the models built under different ratios is compared.

[0076] 3.2.3. Optimization of Spectral Preprocessing and Modeling Methods

[0077] 3.2.3.1. TQ Analyst Software Modeling

[0078] Spectral data were input into TQ Analyst software to model single infrared spectra of MIR and NIR, as well as primary and intermediate fusion data of the two. Preprocessing methods were selected, including scattering correction methods (Multiplicative Scatter Correction, MSC) and Standard Normal Variate Transformation (SNV); spectral derivative processing methods (First Derivative, 1D) and Second Derivative, 2D); and spectral smoothing methods (Savitzky-Golay smoothing, SG smoothing, and Norris smoothing). A three-factor, three-level orthogonal experiment was designed using these methods (see Table 2). Distance Match (DM) and Discriminant Analysis (DA) were used to build models and compare the optimal preprocessing methods. The Norris smoothing used 5 significant bits and a significant bit interval of 5.

[0079] Table 2. Level of Factors for Discriminant Modeling Conditions

[0080]

[0081] 3.2.3.2. Python Platform Modeling

[0082] There are four commonly used and effective modeling methods for discriminant analysis in the Python platform: Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), and Extra Tree (ET). This study utilizes Python to model single infrared spectra of MIR and NIR, as well as primary and intermediate fusion data of the two. During modeling, the spectra were optimized using both preprocessing methods (Norris smoothing, MSC, D1) and non-preprocessing methods. Since the TQ Analyst software lacks advanced fusion capabilities, advanced fusion was performed on the Python platform.

[0083] 3.2.4. Model Evaluation

[0084] Record the number of misclassifications in the calibration set and prediction set of the model under different modeling combinations, calculate the recognition rate and prediction rate using formula (1) and formula (2), use these as indicators to judge the model effect, and optimize the modeling conditions using single-factor experiments.

[0085] Recognition rate = (Total number of calibration sets - Number of misclassified calibration sets) / Total number of calibration sets × 100% (1)

[0086] Prediction rate = (Total number of prediction sets - Number of misclassifications in prediction sets) / Total number of prediction sets × 100% (2)

[0087] 3.2.5. Model Validation

[0088] Substitute the spectrum corresponding to the external validation sample into each optimized model to obtain the model's prediction result for the sample. The accuracy of the model's prediction result for the external validation sample is judged by calculating the model's external validation prediction rate. The calculation formula is shown in (3).

[0089]

[0090] 4. Results and Discussion

[0091] 4.1. Spectral Characteristic Analysis

[0092] The average NIR spectra of samples from different parts are shown below. Figure 2 It can be seen that the root system is at 8300 cm. -1 6804 cm -1 5623 cm -1 5172 cm -1 4763 cm -1 4378 cm -1 4297 cm -1 Characteristic absorption peaks are observed at the location, such as Figure 2 (A); Rhizome at 8299cm -1 6804 cm -1 5623 cm -1 4758 cm -1 4383 cm -1 4297 cm -1 4250 cm -1 Characteristic absorption peaks are observed at the location, such as Figure 2 (B); stem at 8250 cm -1 6804 cm -1 5622 cm -1 5164 cm -1 4763 cm -1 4378 cm -1 4289cm -1 The area exhibits clear characteristic absorption peaks, such as Figure 2 (C); Leaf at 8301 cm -1 6803 cm -1 5624 cm -1 5164 cm -1 4763 cm -1 4378 cm -1 Characteristic absorption peaks are observed at the location, such as Figure 2 (D); Fruit at 8291 cm -1 6807 cm -1 5635 cm -1 5172 cm -1 4762 cm -1 4378 cm -1 It exhibits unique characteristic absorption peaks, such as Figure 2 (E). The main differences in overall spectral characteristics are reflected in the number and position of peaks. Regarding the number of peaks, roots, rhizomes, and stems each exhibit 7 characteristic peaks, while leaves and fruits each have 6. This is one less aromatic CH-related vibration peak compared to the former three, reflecting differences in chemical composition characterization between leaves / fruits and underground parts / stems. Regarding peak position, there are slight shifts in the positions of vibrational peaks of the same type of functional group in different parts, and these shifts are directly related to differences in composition among different parts. The NIR spectral baselines of different parts are relatively stable, with no obvious abnormal vibrational signals, and the peak shapes are generally regular without severe overlap.

[0093] There are 5 common peaks in the NIR spectra of samples from different parts: 8331 cm⁻¹ -1 6804 cm -1 5623 cm -1 4763 cm -14378 cm -1 Among them, 8331 cm -1 The vicinity shows CH overtone vibrations, indicating an aliphatic chain or aromatic ring structure; 6804 cm⁻¹ -1 The nearby absorption peak is related to the -NH symmetric stretching vibration, indicating the presence of nitrogen-containing organic components; 5623 cm⁻¹ -1 The nearby absorption peaks, representing overtone / combination vibrations of CH3, indicate the presence of aliphatic compounds; 4763 cm⁻¹ -1 The vicinity likely exhibits acid-OH combination frequency vibrations, indicating the presence of organic acid components; 4378 cm⁻¹ -1 This is related to the overtone / combination frequency vibrations of CH in aromatic hydrocarbons.

[0094] One-way ANOVA was performed on the absorbance values ​​corresponding to the five common peaks in the NIR spectra of samples from different parts of the peach, and the results are shown in Table 3. The results showed that the absorbance differences corresponding to the common peaks among different parts were extremely significant (p < 0.01), reflecting the specificity of the internal chemical components in terms of type and content, providing a basis for the subsequent establishment of a discrimination model for the seven parts of the peach. (Zhang Yaya et al.) Statistical analysis of the near-infrared diffuse reflectance spectra of five parts of Angelica sinensis showed that there were significant differences in the spectral characteristics of different medicinal parts of Angelica sinensis, indicating that there are certain differences in the chemical composition of different medicinal parts, which is consistent with the results of the NIR spectral difference analysis of different parts of Angelica sinensis in the above study.

[0095] Table 3. Results of one-way ANOVA of common peak absorbance values ​​in NIR spectra of samples from different locations.

[0096]

[0097] In Table 3, * indicates a significant difference (p < 0.05); ** indicates an extremely significant difference (p < 0.01).

[0098] The average MIR spectra of samples from different locations are shown below. Figure 3 The number of MIR spectral peaks varied significantly among samples from different locations. The root sample exhibited 17 characteristic peaks, located at 3728 cm⁻¹. -1 3285 cm -1 2918 cm -1 2049 cm -1 1735 cm -1 Near the same wave number, such as Figure 3 (A); The rhizome has 16 characteristic segments, located at 3728 cm. -1 3285 cm -1 2918 cm -1 2049 cm-1 1735 cm -1 Near the same wave number, such as Figure 3 (B); The stem has 13 absorption peaks, located at 3288 cm⁻¹. -1 2918 cm -1 1735 cm -1 1592 cm -1 1370 cm -1 1316 cm -1 1243 cm -1 Near the same wave number, such as Figure 3 (C); The leaf has 12 absorption peaks, located at 3278 cm⁻¹. -1 2918 cm -1 2850 cm -1 1730 cm -1 1594 cm -1 Near the same wave number, such as Figure 3 (D); The fruit has 15 characteristic points, located at 3278 cm. -1 2923 cm -1 2853 cm -1 1743 cm -1 Near the same wave number, such as Figure 3 (E). The number of characteristic peaks in the MIR spectra of roots and rhizomes is greater than that of stems, leaves and fruits, indicating that the chemical composition of underground parts is more complex.

[0099] There are 8 common peaks in the MIR spectra of samples from different parts: 3285 cm⁻¹ -1 2922 cm -1 1735 cm -1 1591 cm -1 1419 cm -1 1237 cm -1 1074 cm -1 572 cm -1 Among them, 3285 cm -1 The stretching vibration is -OH or -NH, reflecting polar components containing hydroxyl or amino groups; 2922 cm⁻¹ -1 The CH stretching vibration of CH3 / CH2 indicates the presence of an aliphatic compound; 1735 cm⁻¹ -1 This is a C=O stretching vibration, corresponding to esters, carboxylic acids, etc.; 1591 cm⁻¹ -1 The presence of C=C skeletal vibrations in the aromatic ring indicates the presence of aromatic compounds (such as flavonoids and lignans); 1419 cm -1 For CH bending or OH stretching vibration; 1237 cm-1 For CO stretching vibration; 1074 cm⁻ 1 The vibration is COC or COH, matching the ether bond and hydroxyl structure of lignans; 572 cm -1 The skeletal bending or low-frequency heterocyclic vibrations reflect the characteristics of complex heterocyclic structures.

[0100] One-way ANOVA was performed on the absorbance values ​​corresponding to the eight common peaks in the MIR spectra of samples from different parts of the sample. The results are shown in Table 4. It can be found that each common peak showed extremely significant differences among different parts (p < 0.01), indicating that the composition and content of compounds in different medicinal parts of *Prunus persica* are different, and this difference can be reflected in the MIR spectra.

[0101] Table 4. Results of one-way ANOVA of absorbance values ​​of common peaks in MIR spectra of samples from different locations.

[0102]

[0103] In Table 4, * indicates a significant difference (p < 0.05); ** indicates an extremely significant difference (p < 0.01).

[0104] 4.2. Abnormal Spectral Removal

[0105] Abnormal spectral removal was performed on the NIR, MIR, primary fusion, and intermediate fusion spectra of the root samples. Mahalanobis distance and principal component score plots are shown below. Figure 4 Of the 241 MIR spectra used for modeling, 21 anomalous spectra were removed, leaving 220 for modeling. Figure 4 (B); Of the 241 NIR spectra used for modeling, 14 anomalous spectra were removed, leaving 227 for modeling, such as... Figure 4 (A); The initial fusion yielded 241 modeling spectra. 26 anomalous spectra were removed, leaving 215 spectra for modeling, such as... Figure 4 (C); The initial fusion process removed 27 anomalous spectra, leaving 214 for modeling, such as... Figure 4 (D).

[0106] Abnormal spectral removal was performed on the NIR, MIR, primary fusion, and intermediate fusion spectra of the rhizomes. Mahalanobis distance and principal component score plots are shown below. Figure 5 After removing anomalous spectra in NIR, 208 spectra were retained for modeling, such as... Figure 5 (A); After removing anomalous spectra, 239 spectra remained in the MIR model, such as... Figure 5 (B); After removing anomalous spectra from the primary fusion spectra, 212 spectra remained for modeling, such as... Figure 5 (C); After intermediate fusion removed anomalous spectra, 217 spectra remained for modeling, such as... Figure 5 (D).

[0107] Outlier removal was performed on the NIR, MIR, primary fusion, and intermediate fusion modeled spectra of the stem samples. Mahalanobis distance and principal component score plots of the spectral data are shown below. Figure 6 After removing anomalous spectra in NIR, 225 spectra were retained for modeling, such as... Figure 6 (A); After removing anomalous spectra, 220 spectra remained in the MIR model, such as... Figure 6 (B); After removing anomalous spectra from the primary fusion spectra, 224 spectra remained for modeling, such as... Figure 6 (C); After intermediate fusion removed anomalous spectra, 220 spectra remained for modeling, such as... Figure 6 (D).

[0108] Abnormal spectra were removed from the NIR, MIR, primary fusion, and intermediate fusion spectra of the leaf samples. The Mahalanobis distance and principal component scores are shown in Figure 7. After removing abnormal NIR spectra, 209 spectra were retained for modeling. Figure 7 (A); After removing anomalous spectra, 211 spectra remained in the MIR model, such as... Figure 7 (B); After removing anomalous spectra from the primary fusion spectra, 212 spectra remained for modeling, such as... Figure 7 (C); After intermediate fusion removed anomalous spectra, 233 spectra remained for modeling, such as... Figure 7 (D).

[0109] Abnormal spectra were removed from the MIR, NIR, primary fusion, and intermediate fusion spectra of the fruit samples. The Mahalanobis distance and principal component score plots of the spectral data are shown below. Figure 8 After removing anomalous spectra in NIR, 74 spectra were retained for modeling, such as... Figure 8 (A); After removing anomalous spectra, the remaining 82 spectra from the MIR model were used for modeling, such as Figure 8 (B); After removing anomalous spectra from the primary fusion spectra, the remaining 75 spectra were used for modeling, such as... Figure 8 (C); After intermediate fusion removes anomalous spectra, 75 spectra remain for modeling, such as... Figure 8 (D).

[0110] 4.3. Construction of the Part Discrimination Model

[0111] 4.3.1. Construction of a Partial Discrimination Model Based on a Single Infrared Spectrum

[0112] Table 5 shows the modeling results of the seven-part discrimination model of peach based on TQ Analyst software and two single infrared spectra. The model established under NIR spectroscopy shows that when the modeling set division ratio is 2:1, the model performance is optimal when using the DM modeling method combined with 2D, SG smoothing, and no-scattering correction. Figure 9(A) The recognition rate was 96.06% and the prediction rate was 95.27%, indicating that the model has reliable quantitative prediction capabilities and can meet the requirements for distinguishing the seven parts of a peach. The model established under MIR spectroscopy results show that, at a 2:1 division ratio, the model performs best when using DM combined with 2D, Norris smoothing, and no spectral preprocessing, achieving a recognition rate of 91.52% and a prediction rate of 91.82%. (See [reference needed]). Figure 9 (B). Overall, comparing the models built using the two infrared spectra, the predictive performance of the MIR single infrared spectra model is lower than that of the NIR single infrared spectra model.

[0113] Table 5 Comparison of different part discrimination models based on TQ analyst software and two single infrared spectra

[0114]

[0115] Table 6 shows the modeling results based on the Python platform and two single infrared spectral models for site discrimination. The model established under NIR spectroscopy indicates that when the modeling set partition ratio is 4:1 or 5:1, the model performance is optimal when using the SVM modeling method combined with spectral preprocessing strategies. (See table below for details.) Figure 10 (A) The recognition rate was 99.83%, and the prediction rate was 100.00%. The model established under MIR spectroscopy showed that, at a 5:1 partition ratio, the model performance was optimal when using SVM combined with spectral preprocessing. Figure 10 (B) The recognition rate was 98.78%, and the prediction rate was 98.48%.

[0116] Comprehensive analysis shows that both NIR and MIR spectral models can achieve good site discrimination. In the models built using TQanalyst software, the NIR spectral model has a higher recognition rate and prediction rate than the MIR spectral model, indicating that NIR spectroscopy better reflects the differences in overall chemical composition between different sites and has stronger discrimination reliability. In the models built on the Python platform, both spectral models achieve extremely high discrimination accuracy, with the NIR model achieving a prediction rate of 100.00%, while the MIR model's prediction rate is slightly lower. (Sampaio et al.) Analysis revealed that the classification model developed using SVM exhibits strong robustness compared to other models, enabling high-confidence classification of sample origins using spectral data. In a single infrared spectral region discrimination model built on the Python platform, the SVM method demonstrated excellent predictive performance and good adaptability to both types of spectral data, effectively extracting spectral feature information.

[0117] Table 6 Comparison of site discrimination models based on Python platform and two single infrared spectra

[0118]

[0119] Substituting the MIR and NIR spectra corresponding to the external validation samples into the optimal model, the external validation prediction rates of the models were calculated, and the results are shown in Table 7. It can be seen that the external validation prediction rate of the optimal NIR model (2:1 scale, DM+2D+SG smoothing) based on TQ Analyst software is 93.59%, while the external validation prediction rate of the optimal MIR model (2:1 scale, DM+2D+Norrissmoothing) is 86.89%, with the NIR model showing the best performance. The optimal NIR model based on the Python platform (5:1 scale, SVM, with preprocessing) has an external validation prediction rate of 96.73%, exhibiting the best predictive ability among single infrared spectral models; while the optimal MIR model (5:1 scale, SVM, with preprocessing) has an external validation prediction rate of only 62.50%. From the external validation results, regardless of whether based on TQ Analyst software or the Python platform, the external validation prediction rates of the established single NIR spectral models are all above 90%, effectively enabling accurate differentiation of different parts of the peach.

[0120] Table 7 External validation results of the optimal model using a single infrared spectrum

[0121]

[0122] 4.3.2. Construction of a Partial Discrimination Model Based on Fuded Spectra

[0123] Although a single infrared spectral model has shown some potential in predicting different parts of the pyrrhotid gland, its performance is limited by the single-dimensional spectral information, and there is still room for improvement in prediction accuracy. To overcome this limitation, a multi-source spectral fusion model was further constructed. Referring to the optimization approach of the single model, fused spectral modeling was performed using TQ Analyst software and the Python platform, respectively.

[0124] The modeling results from the TQ Analyst software are shown in Table 8. When the ratio of the calibration set to the validation set of the primary fusion model is 5:1, the DM algorithm combined with 2D + Norris smoothing spectral preprocessing yields a model with better recognition performance, achieving a recognition rate of 96.27% and a prediction rate of 96.00%. The discrimination results are shown in Table 8. Figure 11(A). When the ratio of the intermediate fusion model is 4:1, using the DM algorithm combined with 1D + SG smoothing + SNV preprocessing, the model's recognition rate is 99.87% and its prediction rate is 100%. However, at a 5:1 ratio, the model achieves the same performance without requiring spectral correction. This is because the spectral preprocessing steps are fewer. Therefore, this model is the optimal model established under intermediate fusion spectral conditions. The discrimination results are shown in […]. Figure 11 (B). Based on the location discrimination models established using TQ Analyst software under single and fused spectra, the model performance improved as the fusion level increased, indicating that fusing infrared spectral data from different sources is one of the effective measures to improve model performance.

[0125] Table 8 Results of the site discrimination model based on TQ analyst software and fused spectra

[0126]

[0127] The results of the fusion model built under the Python platform are shown in Table 9. The primary fusion spectrum model achieves the best results under the conditions of a 2:1 ratio and preprocessing using the SVM modeling method. (See Table 9 for the results.) Figure 12 (A) The recognition rate was 98.91%, and the prediction rate was 99.64%. The intermediate-level fusion spectrum performed best under the conditions of a 2:1 ratio and SVM preprocessing. See the results below. Figure 12 (B) shows a 100% recognition rate and a 100% prediction rate. The advanced fusion model performs well with a 2:1 modeling set ratio, as shown in the results below. Figure 12 (C) The model has a recognition rate of 100% and a prediction rate of 100%.

[0128] A systematic analysis of the results of the seven-part discrimination fusion model for peaches reveals that spectral fusion can effectively integrate multi-source information and significantly improve the model's part discrimination accuracy. The model performance is closely related to the fusion level, algorithm selection, and preprocessing method. In TQ Analyst software, the fusion model outperforms the single infrared spectroscopy model in discrimination, with the intermediate fusion model achieving higher discrimination accuracy and a prediction rate of 100% under suitable parameter conditions. On the Python platform, all levels of fusion models exhibit excellent discrimination capabilities, with the intermediate and advanced fusion models achieving 100% recognition and prediction rates.

[0129] Table 9 Results of the site discrimination model based on Python platform and fused spectrum

[0130]

[0131] To further verify the predictive ability of the fusion spectral models, external validation experiments were conducted simultaneously. The results are shown in Table 10. It can be seen that the external validation prediction rates of different fusion spectral models are all above 80%. Among them, the intermediate fusion spectral model built based on TQ analyst software has the highest external validation prediction rate, reaching 100.00%, which can accurately determine the origin of the *Prunus persica* root samples. In the fusion models built on the Python platform, the predictive ability of the models gradually improves with the increase of the fusion level, and the advanced fusion model achieves a 100% external validation prediction rate.

[0132] Table 10 External validation results of the optimal fusion spectral model

[0133]

[0134] 5. Conclusion

[0135] Using samples of *Prunus persica* from 15 different production areas as examples, we integrated their NIR and MIR spectral data. By optimizing modeling parameters, spectral preprocessing methods, and sample set division ratios, we constructed a medicinal part discrimination model based on single infrared spectroscopy and fused spectroscopy. The results are as follows:

[0136] In the single infrared spectral model, the NIR model built with TQ Analyst software generally outperformed the MIR model. The model built using Python and SVM significantly outperformed the model built with TQ Analyst software. Specifically, the best NIR model on the Python platform achieved a 100% prediction rate and an external validation prediction rate of 96.73%, meeting the requirements for the discrimination and detection of *Prunus persica* parts. The MIR model performed slightly lower than the NIR model, but it could still be used for the discrimination of *Prunus persica* parts.

[0137] The overall performance of the fusion spectral model is better than that of the single infrared spectral model. The recognition rate and prediction rate of the intermediate and advanced fusion models under the Python platform are both 100%, and the external validation prediction rate of the advanced fusion model is also 100%. The external validation prediction rate of all fusion spectral models is above 80%, indicating that the established peach seven-part discrimination model can achieve accurate discrimination of different parts.

[0138] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a peach seven-part identification system based on infrared spectral fusion and machine learning is provided, including:

[0139] The spectral acquisition module is used to acquire multiple single infrared spectral data from different parts of the sample to be tested, including near-infrared spectral data and mid-infrared spectral data.

[0140] The spectral fusion module is used to concatenate near-infrared and mid-infrared spectral data from different regions to obtain primary fused data; feature extraction and fusion are performed on the primary fused data to obtain intermediate fused spectral data.

[0141] The part discrimination model construction module is used to construct part discrimination models based on the single infrared spectral data, primary fusion data, and intermediate fusion data, respectively, using the Python platform and TQ analyst software.

[0142] The identification output module is used to identify different parts of the tested *Prunus persica* sample using the part identification model constructed in the part identification model construction module, and output the medicinal part identification results.

[0143] The above detailed embodiments are a description of the present invention. It should not be considered that the specific embodiments of the present invention are limited to these descriptions. For those skilled in the art, several simple deductions and substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the protection scope of the present invention.

Claims

1. A method for identifying seven parts of a peach based on infrared spectral fusion and machine learning, characterized in that, Includes the following steps: S1. Collect multiple single infrared spectral data from different parts of the sample to be tested, including near-infrared spectral data and mid-infrared spectral data; S2. Concatenate near-infrared and mid-infrared spectral data from different locations to obtain primary fused data; extract and fuse features from the primary fused data to obtain intermediate fused spectral data; S3. Based on the single infrared spectral data, primary fusion data, and intermediate fusion data, construct part discrimination models using the Python platform and TQ analyst software, respectively; S4. Use the part discrimination model constructed in step S3 to identify different parts of the sample of *Prunus persica* to be tested, and output the medicinal part discrimination results.

2. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Step S3 also includes: By using the Python platform, decision-level fusion of different part discrimination models is performed to form an advanced fused part discrimination model.

3. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 1, characterized in that, The samples of *Prunus cerasifera* to be tested were collected from populations at 15 different collection sites in Qinghai Province, Yunnan Province, and Tibet Autonomous Region. The populations at each collection site were more than 20 km apart. At least 20 healthy plants were collected from each population at the same collection site, and each plant was spaced 10 m apart. Each *Prunus cerasifera* sample was dried, pulverized, and passed through an 80-mesh sieve at different parts before being placed in a desiccator for analysis.

4. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 1, characterized in that, The acquisition of the near-infrared spectral data includes: Take appropriate amounts of powder from different parts of the peach kernel and place them on filter paper. Use a Fourier transform infrared spectrometer (NIR fiber module) to analyze the sample at 10000-4000 cm⁻¹. -1 Spectral data were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 6 cm⁻¹. -1 The scan was performed 64 times, with air as a reference. Each peach sample was collected 3 times, and the average spectrum was taken for analysis. The acquisition of the mid-infrared spectral data includes: Appropriate amounts of *Prunus persica* powder from different parts were placed on the attenuated total reflectance infrared probe of a Fourier transform infrared spectrometer, at 4000-400 cm⁻¹. -1 MIR spectra were acquired within the specified range, with background interference subtracted in real time during acquisition. The scanning resolution was 4 cm⁻¹. -1 The sample was scanned 32 times, with air as a reference. Each sample was collected 3 times, and the average spectrum was taken for analysis.

5. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Step S1 also includes: The near-infrared and mid-infrared spectral data are preprocessed, including one or more combinations of scattering correction, derivative processing, and spectral smoothing. The scattering correction method is multivariate scattering correction or standard normal transformation. The derivative processing method is first derivative or second derivative. The spectral smoothing method is Savitzky-Golay smoothing or Norris smoothing.

6. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 5, characterized in that, In step S3, when constructing the part discrimination model using the Python platform or TQ analyst software, multiple part discrimination models are constructed using different machine learning methods, preprocessing methods, model set proportions, and optical path types.

7. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 6, characterized in that, The machine learning methods used in the TQ analyst software include distance matching and discriminant analysis, while the machine learning methods used on the Python platform include support vector machines, decision trees, random forests, and limit trees.

8. The method for identifying seven parts of a peach based on infrared spectral fusion and machine learning according to claim 1, characterized in that, Before constructing the location discrimination model, outlier removal is performed on the single infrared spectral data, primary fusion data, and intermediate fusion data. The outlier removal methods include Mahalanobis distance method and principal component analysis method.

9. A system for identifying seven parts of a peach based on infrared spectral fusion and machine learning, characterized in that, include: The spectral acquisition module is used to acquire multiple single infrared spectral data from different parts of the sample to be tested, including near-infrared spectral data and mid-infrared spectral data. The spectral fusion module is used to concatenate near-infrared and mid-infrared spectral data from different regions to obtain primary fused data; feature extraction and fusion are performed on the primary fused data to obtain intermediate fused spectral data. The part discrimination model construction module is used to construct part discrimination models based on the single infrared spectral data, primary fusion data, and intermediate fusion data, respectively, using the Python platform and TQ analyst software. The identification output module is used to identify different parts of the tested *Prunus persica* sample using the part identification model constructed in the part identification model construction module, and output the medicinal part identification results.