Infrared spectroscopy method, system and program for rapidly and nondestructively identifying semen cuscutae and processed product thereof
By combining infrared spectroscopy with cluster analysis and principal component analysis, the problem of rapid and non-destructive identification of processed Cuscuta chinensis (a traditional Chinese medicine) products has been solved. This has enabled efficient, low-cost, and environmentally friendly overall chemical composition analysis, ensuring the safety and effectiveness of quality control for traditional Chinese medicine.
Patent Information
- Application Number
- CN202511045951.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for identifying Cuscuta chinensis (a traditional Chinese medicine) are insufficient to quickly, non-destructively, and comprehensively reflect the changes in the overall chemical composition of the herb before and after processing. Furthermore, they are complex to operate, costly, and pose risks of environmental pollution.
Infrared spectroscopy combined with cluster analysis and principal component analysis was used to perform non-destructive testing on dodder and its processed products. Spectral information was collected and preprocessed, and rapid identification was achieved using cluster analysis and principal component analysis.
This method enables rapid and non-destructive identification of Cuscuta chinensis and its processed products, reducing testing costs, minimizing environmental pollution risks, improving the accuracy and efficiency of identification, and ensuring the safety and effectiveness of traditional Chinese medicine quality control.
Smart Images

Figure CN120948397A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traditional Chinese medicine identification and analysis technology, and in particular to a rapid and non-destructive infrared spectroscopy method, system and procedure for identifying Cuscuta chinensis and its processed products. Background Technology
[0002] Cuscuta ( Cuscutae Semen Cuscuta chinensis is a commonly used traditional Chinese medicine in clinical practice, possessing the effects of tonifying the kidneys and strengthening essence, benefiting the liver and improving eyesight. It is widely used for various diseases caused by liver and kidney deficiency. Modern research shows that the efficacy of Cuscuta chinensis is closely related to its processing method. Different processing methods (such as stir-frying, wine-frying, and salt-frying) can not only change its properties and efficacy but also affect the content and activity of its chemical components, thus impacting clinical efficacy. For example, raw Cuscuta chinensis is warm in nature and is often used for blurred vision; stir-fried Cuscuta chinensis has similar effects to raw Cuscuta chinensis, but stir-frying can improve the decoction effect; salt-fried Cuscuta chinensis is neither warm nor cold and is often used to gently tonify the liver and kidneys; wine-fried Cuscuta chinensis tends to warm and tonify the spleen and kidneys. Therefore, for treating impotence and seminal emission, if the deficiency is due to both yin and yang deficiency of the kidneys, salt-processed Cuscuta chinensis can be used; if the deficiency is mainly due to yang deficiency of the kidneys, wine-processed Cuscuta chinensis can be used. Some studies have also shown that in terms of liver protection, wine-fried and salt-fried Cuscuta chinensis have better therapeutic effects than stir-fried and raw Cuscuta chinensis. Therefore, the correct use of dodder seed and its different processed products helps to maximize its clinical efficacy. Thus, the identification and differentiation of different processed dodder seed products is of significant practical importance in ensuring the effectiveness and safety of its clinical application.
[0003] Currently, the commonly used identification methods for dodder seeds (including raw and processed products) include thin-layer chromatography, HPLC fingerprinting, and high-performance capillary electrophoresis.
[0004] 1. Thin-layer chromatography (TLC): Thin-layer chromatography is typically used to separate and detect certain characteristic or indicator components in medicinal materials. This method is simple to operate and has relatively low detection costs. However, due to the complexity of the components of Chinese medicinal materials, the differences in characteristic components between different processed products may not be significant, and the detection targets are mostly one or a few characteristic compounds, making it difficult to comprehensively reflect the changes in the overall chemical composition of Chinese medicinal materials, thus having certain limitations.
[0005] 2. High-Performance Liquid Chromatography (HPLC) and Fingerprint Analysis: HPLC fingerprint analysis can reflect the overall chemical composition characteristics of traditional Chinese medicine to a certain extent and is widely used in the quality control of various traditional Chinese medicines and their preparations. However, this method usually requires the use of organic solvents, the pretreatment steps are cumbersome, the instrument operation and maintenance costs are high, and the testing cycle is long, which can easily cause potential pollution or health risks to the environment and laboratory personnel. For rapid detection and on-site identification of large batches of samples, its operational efficiency is still insufficient.
[0006] 3. Other separation and analysis techniques such as high-performance capillary electrophoresis: In addition to HPLC, some studies have also used high-performance capillary electrophoresis (HPCE) and gas chromatography (GC) to detect Cuscuta chinensis and its processed products. However, these techniques, like HPLC, also face common problems such as limited detection targets, complex pretreatment, and large solvent consumption. Furthermore, they do not yet have advantages in rapid resolution and overall identification in practical applications.
[0007] From the perspective of traditional Chinese medicine theory and the holistic mechanism of action of Chinese herbal medicines, the pharmacological efficacy of Chinese herbal medicines is often the result of the synergistic effect of multiple active ingredients. However, existing analytical methods mostly focus on the detection of single or small amounts of components, making it difficult to comprehensively reflect the changes in the overall composition of the medicinal material before and after processing. Furthermore, to avoid harm to laboratory personnel and the environment, and to reduce the use of chemical reagents and the discharge of organic waste, there is an urgent need for a simple, rapid, non-destructive method that can characterize the overall information for the identification and quality control of Cuscuta chinensis and its different processed products. Summary of the Invention
[0008] To address the aforementioned technical problems, the present invention aims to provide a rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its processed products. This method directly detects the sample using infrared spectroscopy, acquiring spectral information, and simultaneously performs qualitative analysis of the sample using chemometric methods such as cluster analysis and principal component analysis. This achieves rapid and non-destructive identification of Cuscuta chinensis and its different processed products, better ensuring its safety and efficacy in clinical applications, improving the quality control level of traditional Chinese medicine, and meeting the needs of the modern traditional Chinese medicine industry for rapid detection and accurate identification.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: An infrared spectroscopy method for rapid and non-destructive identification of dodder seed and its processed products, the method comprising the following steps: 1) Sample preparation: Place dodder seed and / or its processed products in an oven to dry, cool and then pulverize. The powder is sieved and sealed for storage. 2) Infrared spectral acquisition: The powder sample obtained in step 1) was analyzed using an infrared spectrometer with an integrating sphere diffuse reflectance method. The spectral acquisition range was 4000 cm⁻¹. -1 ~500cm -1 The resolution is 0.1-10 nm, the number of scans is 20-50, and the infrared spectrum is recorded; 3) Spectral preprocessing: Perform one or more preprocessing operations on the infrared spectrum. The preprocessing operations are selected from automatic baseline calibration, smoothing, first derivative, and vector normalization. The infrared spectral wavelength range is selected from 4000 cm⁻¹. -1 ~2500cm -1 2000cm-1 ~500cm -1 ; 4) Pattern recognition analysis: The infrared spectral data preprocessed in step 3) is imported into the cluster analysis and principal component analysis models. By observing and judging the cluster analysis dendrogram and principal component score map, raw dodder and / or its processed products can be quickly distinguished and identified.
[0010] Preferably, a cubic polynomial is used in step 3). p(n)=a 0 +a 1 n+a 2 n 2 +a 3 n 3 As the initial baseline model, where ν is the wavelength sequence and a0, a1, a2, a3 are the coefficients to be fitted; the least squares method is used iteratively in conjunction with the following adaptive elimination rule: (a) After the initial fitting, absorption peaks above a certain threshold are identified and temporarily removed; (b) Fit the remaining data twice or multiple times to make the polynomial baseline fit the non-absorbing region as closely as possible; (c) After the iteration is completed, subtract the final baseline function from all the original data to obtain the baseline-corrected spectrum.
[0011] Preferably, in step 3), Savitzky-Golay filtering is used to perform vector normalization on the first derivative spectra of segment I and segment II, respectively:
[0012] in: n i The i-th wavenumber point is the abscissa of the infrared spectrum; D 1 (n i ) Spectral band I, at wavenumber ν i The first derivative spectral value at the location; D 2 (n i ) In spectral band II, at wavenumber ν i The first derivative spectral value at the location; ∑ j [D1 (n j )] 2 The summation of the squares of the first derivative values of all wavenumber points j within spectral region I represents the squared magnitude of the first derivative spectral vector of region I. ∑ j [D2(ν j )] 2 The summation of the squares of the first derivative values of all wavenumber points j within spectral region II represents the squared magnitude of the first derivative spectral vector of region II. D 1,norm (ν i ): D1(ν i The value after vector normalization according to the modulus of the entire first derivative spectrum; D 2,norm (ν i ): D2(ν i The value after vector normalization according to the modulus of the entire first derivative spectrum; At the segment splicing point, the data at both ends are normalized and then spliced using a transition weighting strategy to obtain the overall first derivative spectrum. D norm (n) .
[0013] Preferably, the cluster analysis in step 4) includes the following steps: 4.1.1) Nonlinear eigenmap: Input the preprocessed spectral matrix X into the radial basis function kernel: , Where x i ,x j For the first i , No. j The feature vectors of each sample are used to construct the centered kernel matrix K by taking the median of the pairwise Euclidean distances between the samples. 4.1.2) Kernel Principal Component Reduction: Perform eigenvalue decomposition on the kernel matrix K, and select the top p eigenvectors with a cumulative contribution rate of not less than 95% to obtain the low-dimensional score matrix Z=[z1,z2,…,zn]. T ∈R n×p ; 4.1.3) Local density estimation: For each sample i in matrix Z, calculate the number of samples ρ within its k nearest neighbor radius. i , as local density; 4.1.4) Adaptive distance construction: Define a new distance between samples i and j:
[0014] Among them, zi , z j Let ρ be the low-dimensional score vector obtained by kernel PCA for samples i and j. i , ρ j This provides the local density estimates for samples i and j in a low-dimensional space, aiming to compress the distance between high-density regions and amplify the distance between low-density regions. 4.1.5) Density-weighted Ward aggregation: using the adaptive distance d * ij The Ward chain criterion is used to perform clustering on the input, and the local density is included as a weight in the calculation of the intra-class squared error increment. 4.1.6) Determining the optimal number of clusters: During the aggregation process, the average profile coefficient s(k) and GapStatistic are calculated in real time. When s(k) reaches a local maximum and the Gap value no longer increases significantly for the first time, the optimal number of clusters k̂ is automatically determined. 4.1.7) Output Results: Based on the cluster identifiers obtained in step 6), perform rapid and non-destructive classification of raw dodder seed, stir-fried dodder seed, salt-fried dodder seed, and wine-fried dodder seed.
[0015] Preferably, step 4.1.6) includes the following steps: In the bottom-up agglomerative hierarchical clustering process, when the number of clusters decreases from k+1 to k, calculate for the current cluster set C(k)={C1,…,Ck}: ① Average profile coefficient
[0016] Where a i Let b be the average distance between sample i and other members in the cluster. i Let i be the average distance between sample i and its nearest neighbor cluster; ②Gap Statistic
[0017] Among them W K *b Let B be the intra-class dispersion of the b-th reference distribution when it is distributed across k clusters, and B be the number of reference data groups. Record the curve of (k, s(k), Gap(k)) changing with aggregation, and detect it in real time: When s(k) reaches a local maximum and satisfies: Gap(k) ≥ Gap(k+1)-s k+1 ;where s k+1 The standard error of Gap(k+1) is considered the first time the Gap value no longer increases significantly; the current cluster number k that satisfies the above condition is output as the optimal cluster number k. ∧ It serves as the final pruning level in the hierarchical tree diagram.
[0018] As a preferred embodiment, the principal component analysis method in step 4) includes the following steps: 4.2.1) Multi-scale wavelet packet decomposition: For each normalized spectrum s(ν), a wavelet packet decomposition of no more than 5 levels is performed using the Daubechies-4 mother wavelet to obtain the lowest-level approximation coefficients A. L and detail coefficients of each layer ; 4.2.2) Energy Threshold Denoising: Calculate the energy of each detail subband:
[0019] Where u is a discrete point, and the threshold τ is set to 0.8 × all The median, for satisfying The subband coefficients of <τ are all set to zero; 4.2.3) Coefficient vector concatenation and matrix construction: Concatenate A according to the fixed sub-band order. L and those not set to zero The wavelet coefficient vectors are sequentially concatenated into a wavelet coefficient vector w. i Then, by stacking these matrices, we obtain the coefficient matrix W = [w1;w2;…;wn] ∈ R. n×q ; 4.2.4) Z-score standardization: Perform mean centering and standard deviation normalization on each column of matrix W to obtain the standardized matrix Ŵ; 4.2.5) Principal Component Decomposition of Functions: Constructing the Covariance Operator We perform eigenvalue decomposition on the system and select the top two eigenvectors q1 and q2 that contribute at least 85% of the cumulative variance. 4.2.6) Score Calculation and Visualization: [This section likely refers to a specific step or function, but without Projecting the matrix onto the q1 and q2 directions yields the score matrix Z. [q1,q2]∈R n×2 The PC1-PC2 score map was plotted using the two columns of Z as the horizontal and vertical axes, respectively, and different labels were used to distinguish between raw dodder seed, stir-fried dodder seed, salt-fried dodder seed, and wine-fried dodder seed, thereby achieving rapid and non-destructive separation of the four types of processed products and traceable interpretation of key vibration bands.
[0020] Furthermore, this invention also discloses an infrared spectroscopy system for rapid and non-destructive identification of Cuscuta chinensis and its different processed products. This system is used to implement the method, including: The infrared spectroscopy detection device is configured to acquire infrared spectra of prepared dodder seed or its processed powder samples using an integrating sphere diffuse reflectance method, with an acquisition range of 4000 cm⁻¹. -1 ~500cm -1 It has a resolution of 0.1-10nm and supports multiple scans; The data processing module is communicatively connected to the infrared spectral detection device and is used to receive the infrared spectral data output by the detection device and perform one or more of the following: automatic baseline calibration, smoothing, first derivative, and vector normalization of the spectrum. The analysis and determination module, connected to the data processing module, is used to perform cluster analysis and / or principal component analysis to determine the processing types of raw dodder seed and its stir-fried, wine-fried, and salt-fried dodder seed based on the obtained spectra and the results of the chemometric model.
[0021] Furthermore, the present invention also discloses the application of the method in the quality control of Cuscuta chinensis slices, which enables accurate determination of the type and quality of processed products by rapid and non-destructive testing of raw Cuscuta chinensis or processed products such as stir-fried, stir-fried with wine, or stir-fried with salt.
[0022] Furthermore, the present invention also discloses a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements steps 3)-4) of the method.
[0023] Furthermore, the present invention also discloses a computer program product, including a computer program or instructions, which, when executed by a processor, implement steps 3)-4) of the method.
[0024] This invention, employing the aforementioned technical solution and combining infrared spectroscopy with cluster analysis and principal component analysis, achieves rapid identification of Cuscuta chinensis and its three processed products. This method is not limited to individual components; it can complete the overall analysis of a sample within tens of seconds using only a small amount of sample without damaging the sample. It features low cost, ease of operation, and no chemical reagent contamination, providing a new reference for the identification and analysis of different processed products of the same traditional Chinese medicine, and offering new technical assurance for the safety and efficacy of processed products in clinical use. It has high practical application value. Attached Figure Description
[0025] Figure 1 Infrared spectrum overlay of raw dodder seeds; Figure 2 The infrared spectrum overlay of salt-fried dodder seeds; Figure 3 Infrared spectrum overlay of stir-fried dodder seeds with wine; Figure 4 The infrared spectrum overlay of stir-fried dodder seeds; Figure 5 The cluster analysis dendrogram of raw, salt-fried, wine-fried, and plain-fried dodder seeds in Example 1; Figure 6 The main components of Cuscuta chinensis in Example 1 are shown in the following diagrams: raw, salt-fried, wine-fried, and plain-fried. Figure 7 Two-dimensional score diagrams of the main components of raw, salt-fried, wine-fried, and plain-fried dodder seeds in Example 1. Detailed Implementation
[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0027] 1. Experimental Materials experimental drugs Authentic raw dodder seed slices: produced in Inner Mongolia, totaling 4 batches, numbered raw product 1 to 4.
[0028] Salt-fried dodder seeds: Take clean dodder seed slices and fry them according to the salt-frying method (General Rule 0213) until they are slightly puffed up. They are numbered as salt-fried 1 to 5.
[0029] Stir-fried dodder seeds: Place them in a preheated wok and stir-fry over low heat until the surface is slightly yellow and fragrant. Remove them and let them cool to obtain stir-fried dodder seeds, numbered 1-5.
[0030] Stir-fried dodder seeds with wine: Mix with rice wine, moisten slightly, and wait for the rice wine to be absorbed. Stir-fry over low heat until the surface is slightly yellow and slightly cracked. Remove and let cool to obtain wine-stirred dodder seeds (100g dodder seeds with 20g rice wine), numbered wine-stirred 1 to 4.
[0031] Experimental instruments and equipment Spectrum TwoFT-IR near-infrared spectrometer (PerkinElmer, USA); 1000C multi-functional pulverizer (Yongkang Hongtaiyang Electromechanical Co., Ltd.); BAS124S electronic analytical balance (Sartorius Scientific Instruments Beijing Co., Ltd.); DHG-9246A electric heating constant temperature drying oven (Shanghai Jinghong Experimental Equipment Co., Ltd.); 60-mesh standard test sieve (Shaoxing Shangyu Huafeng Hardware Instrument Co., Ltd.).
[0032] 2 Experimental Methods and Results 2.1 Acquisition of Infrared Spectra The samples to be tested were placed in a 60℃ oven and dried for 6 hours. After drying, they were removed, cooled, pulverized, and the powder was passed through a 60-mesh sieve and then sealed for storage. 1.0 g of each sample was weighed and placed in a quartz bottle. Infrared spectra were collected using the integrating sphere diffuse reflectance method, with a spectral acquisition range of 4000 cm⁻¹. -1 ~500cm -1 The resolution was 1 nm, the number of scans was 32, and each spectral line consisted of 1869 data points. Each batch of samples was measured in triplicate. The infrared spectra of raw, salt-fried, wine-fried, and plain-fried dodder seeds are shown below. Figure 1-4 .
[0033] The results showed that the infrared spectra of the four samples of raw, salt-fried, wine-fried, and plain-fried dodder seeds were extremely similar. Among them, the four samples had little difference in chemical composition, but there were significant differences in content.
[0034] 2.2 Infrared Spectrum Data Processing 2.2.1 Multi-segment iterative automatic baseline calibration 1) Band segmentation: The infrared spectrum is divided into two main segments: Section I: 4000cm -1 ~2500cm -1 Section II: 2000cm -1 ~500cm -1 ; 2) Iterative baseline fitting Automatic baseline calibration is performed for each segment separately. Taking segment I as an example, a cubic polynomial is selected. p(n)=a 0 +a 1 n+ a 2 n 2 +a 3 n 3 As the initial baseline model, iterative processing is performed using the least squares method combined with the following "adaptive elimination" rule: (a) After the initial fitting, identify absorption peaks above a certain threshold (such as points exceeding twice the standard deviation of the fitted baseline) and temporarily remove them; (b) Fit the remaining data twice or multiple times to make the polynomial baseline fit the non-absorbing region as closely as possible; (c) After the iteration is completed, subtract the final baseline function from all the original data to obtain the baseline-corrected spectrum.
[0035] Repeating this process for section II yields the following results: S b1 (n) (Segment I corrected spectrum) and S b2 (n) (Spectrum after correction for section II).
[0036] 3) Splicing and transition At 2500cm -1 ~2000cm -1To ensure a smooth transition between the spliced spectra, the overlapping portions of the two segments can be fused using a weighted average or smooth transition algorithm to obtain a complete baseline-corrected spectrum. S b (n) .
[0037] This invention employs multi-segment iterative baseline calibration, performing iterative baseline fitting and eliminating abnormal absorption peaks in two key segments, thus avoiding the accumulation of errors in the overall spectrum caused by simple single-segment baseline calibration.
[0038] 2.2.2 Piecewise Adaptive Smoothing Processing 1) Select Savitzky-Golay filter In section I (4000~2500cm) -1 An SG filter window with 5 smoothing points can be used; in section II (2000~500cm) -1 An SG filter window with 7 smoothing points can be used, and the order of the polynomial fitting can be selected as 2 or 3, respectively.
[0039] 2) Mathematical Model Taking segment I as an example, after smoothing, a certain point n i spectral values at S s1 (n i ) Represented as
[0040] Where 2m+1=5, c k The coefficients, pre-calculated based on the Savitzky-Golay principle, depend on the window length 2m+1 and the order of the local polynomial fitting (order 2 in this case), and the coefficients satisfy... This ensures that the amplitude remains consistent after filtering.
[0041] For section II, a window of 2m+1=7 can be used for smoothing to obtain... S s2 (n) .
[0042] 3) Multi-segment fusion Similar to baseline calibration, at 2500cm -1 ~2000cm -1 The transition regions are weighted or smoothed together to obtain a smooth overall spectrum. S s (n) .
[0043] The above method flexibly sets the number of smoothing points to 5 (segment I) and 7 (segment II) for different noise levels and spectral peak densities, and uses smooth splicing of transition segments to preserve the effective peak shape to the greatest extent.
[0044] 2.2.3 First-order derivative transformation The first derivative can highlight variations in absorption peaks, eliminate some baseline or background influences, and further enhance subtle differences related to processing methods. The calculation method uses the central difference approximation.
[0045] Where Δν=ν i+1 -ν i ≈4 cm -1 (This can be set according to the instrument resolution or interpolation settings), thus obtaining the first derivative spectrum. D(n) For further smoothing, a low-order SG filter with a small window can be applied again after the derivative.
[0046] 2.2.4 Multi-segment fusion vector normalization To better balance the differences in absorption intensity between the two segments, the first derivative spectra of segments I and II can be vector normalized separately: ' in: n i The i-th wavenumber point (usually in cm) -1 (), is the abscissa of the infrared spectrum; D 1 (n i ) Spectral range I (e.g., 4000–2500 cm⁻¹) -1 In the wave number ν i The first derivative spectral value at the location; D 2 (n i ) Spectral band II (e.g., 2000–500 cm⁻¹) -1 In the wave number ν i The first derivative spectral value at the location; ∑ j [D 1 (n j )] 2The summation of the squares of the first derivative values of all wavenumber points j within spectral region I represents the squared magnitude of the first derivative spectral vector of region I. ∑ j [D2(ν j )] 2 The summation of the squares of the first derivative values of all wavenumber points j within spectral region II represents the squared magnitude of the first derivative spectral vector of region II. D 1,norm (ν i ): D1(ν i The value after vector normalization according to the modulus of the entire first derivative spectrum; D 2,norm (ν i ): D2(ν i The value after vector normalization based on the modulus of the entire first derivative spectrum.
[0047] At the section splicing point (2500cm) -1 ~2000cm -1 The data at both ends are normalized and then concatenated using a transition weighting strategy to obtain the overall first derivative spectrum. D norm (n) .
[0048] For a full spectrum comparison, you can obtain the spliced spectrum. D norm (n) Based on this, a global vector normalization is performed to ensure consistency in cross-segment comparisons.
[0049] The above process significantly improves the sensitivity and accuracy of spectral analysis in identifying differences in processed products (as shown in Table 1), providing higher-quality preprocessing input for subsequent cluster analysis, principal component analysis, or discriminant analysis. This combined strategy of multi-segment iterative calibration, segmented smoothing, and segmented normalization is more innovative than the traditional "single-segment calibration + fixed smoothing" approach, maximizing the utilization of the infrared band in the 4000–2500 cm⁻¹ range. -1 2000~500cm -1 The feature information within the range provides reliable support for the qualitative and quantitative identification of processed Chinese medicine products.
[0050] Table 1. Infrared Spectrum Data Processing Results
[0051] Conclusion: The method of the present invention significantly reduces drift and noise while preserving the fine shoulder peak characteristics.
[0052] 2.3 Cluster Analysis Cluster analysis was performed on the spectra that underwent automatic baseline calibration, smoothing, first derivative and vector normalization preprocessing to achieve accurate classification of Cuscuta chinensis and its three processed products.
[0053] 2.3.1 Nonlinear Eigenmap Using the preprocessed vector-normalized spectral matrix X as input, X is processed through a radial basis function (RBF) kernel: , Construct the centralized kernel matrix K. Where k(x) i ,x j ) represents the similarity value between the radial basis function (RBF) kernel or the Gaussian kernel, ranging from 0 to 1. i ,x j For the first i , No. j The feature vector of each sample. The kernel width σ adopts the "median heuristic", that is, the median of the pairwise Euclidean distances of all samples is taken as σ, which avoids manual parameter tuning and ensures adaptability to different batches of data.
[0054] 2.3.2 Core PCA Dimensionality Reduction Perform eigenvalue decomposition on K, and select the top p eigenvectors with a cumulative contribution rate ≥ 95% to obtain the low-dimensional score matrix: Z=[z1,z2,…,zn] T ∈R n×p .
[0055] Nonlinear mapping preserves subtle differences in the infrared spectrum that are latent but not easily linearly separated, and is key to distinguishing the three types of processed products.
[0056] 2.3.3 Local Density Estimation In the Z-space, for each sample i, calculate the number ρ of samples within its k-nearest neighbor radius. i (Default k = 5). High ρ indicates that the sample is located in a dense spectral region; low ρ may indicate that it is located on the edge of a class or in an outlier region. It is usually defined as its... k- Number of samples within the nearest neighbor sphere: , Where z r The score vector obtained by dimensionality reduction of the r-th sample in the dataset, r k It is the distance from sample i to its k-th nearest neighbor.
[0057] 2.3.4 Adaptive Distance Construction Define a new between-sample metric: , The original Euclidean distances between samples in high-density areas are compressed, while the distances in low-density areas are amplified, thus strengthening the boundaries between different processing categories. Among them, z i , z j Let ρ be the low-dimensional score vector obtained by kernel PCA for samples i and j. i , ρ j This represents the local density estimates of samples i and j in the low-dimensional space.
[0058] 2.3.5 Weighted Ward Hierarchical Aggregation With d * As input, Ward's chain rule aggregation is performed: each merge minimizes the intra-class squared error increment, while local density information is included as a weight in the error calculation. Therefore, it maintains the sensitivity of the Ward method to the overall structure and avoids the mis-aggregation of high-density regions.
[0059] 2.3.6 Automatic Determination of Optimal Cutting During the agglomeration process, the average profile coefficient s(k) and gap statistic of the clusters after shearing are calculated in real time. When s(k) reaches a local maximum and the gap value no longer increases significantly for the first time, the optimal number of clusters is automatically determined. (Stable obtained in practice) = 4, i.e., raw, stir-fried, salt-fried, and wine-fried. No human experience threshold is required.
[0060] 2.3.6.1 After each "merge", a set of cluster partitions is obtained. Hierarchical clustering starts with k=n (each sample is independent) and merges samples once from k to k-1. For the current cluster set C(k)={C1,…,Ck}, two metrics are calculated simultaneously: 2.3.6.2 Average profile coefficient s(k) 1) Single sample contour values For sample i∈C r : ai: Average distance to other samples in the same cluster ; bi: the average distance to the nearest neighbor cluster Ct (t≠r) ; Contour value 2) Average profile coefficient ; A new s(k) can be obtained by updating once with each aggregation step. 2.3.6.3 Gap Statistic 1) Intra-class dispersion W k , 2) Reference Distribution B groups of "uniform" dummy data are randomly generated within the original data bounding box (B is usually set to 100). The same hierarchical aggregation is performed on each group of dummy data to obtain W. k *b (b=1…B).
[0061] 3) Gap value and standard deviation , .
[0062] 2.3.6.4 Automatic Determination Rules 1) Scrolling recording curve While condensing, store (k, s(k), Gap(k)) into a vector.
[0063] 2) Detect convergence of "local maxima + gap" Find the local maximum point k=k of the current s(k). m That is, s(k) m )≥s(k m ±1). Calculate the gap difference: ΔGap=Gap(k m +1)-Gap(k m ) ; Stopping criterion: If ΔGap < (If the gap curve rises for the first time without crossing the threshold), then k... m Determined as the optimal number of clusters .
[0064] 2.3.7 Interpretation and Visualization of Results The final cluster identifiers are mapped back to Z-space to create 2D or 3D scatter plots, and a hierarchical tree diagram is also output. For further analysis, the original spectrum can be mapped back to examine the differences in the spectral shape of each cluster center to explain the correlation between changes in chemical composition and processing techniques. The analysis results are shown below. Figure 5The results showed that raw, stir-fried, and salt-fried dodder seeds were initially grouped into one large category, while wine-fried dodder seeds were grouped into another. Next, salt-fried and stir-fried dodder seeds were grouped into two separate categories along with the two types of raw dodder seeds. Further, salt-fried and stir-fried dodder seeds were each grouped into a separate category, with two types of raw dodder seeds grouped into one category along with both stir-fried and salt-fried dodder seeds. This indicates that the chemical composition of dodder seeds changes after processing, and the chemical composition varies among the three different processed products. The difference in chemical composition between wine-fried and raw dodder seeds is greater than that between stir-fried and salt-fried dodder seeds. Therefore, cluster analysis can be used for the identification and analysis of dodder seeds and their different processed products. As shown in Table 2, the method of this invention significantly outperforms commonly used linear spectral processing, HPLC fingerprinting, and TLC comparison schemes in terms of classification performance.
[0065] Table 2 Comparison of Cluster Analysis Performance
[0066] 2.4 Principal component analysis of infrared spectra Before performing 2D PC1-PC2 visualization, each infrared spectrum is first converted into a wavelet coefficient sequence. By leveraging the multi-scale characteristics of wavelets, global background changes and local subtle peak differences are simultaneously captured. Then, a functional PCA is performed on the coefficients. This strategy is particularly sensitive to nonlinearity, local displacement, or weak shoulder peaks, and can more clearly distinguish between the three types of roasted and raw products.
[0067] 2.4.1 Spectral Preprocessing and Matrix Preparation For each spectrum, baseline calibration, SG smoothing, first derivative calculation, and vector normalization are performed. The processed spectra are then arranged into a matrix X ∈ R. n × m n is the number of samples, and m is the number of wavenumber points.
[0068] 2.4.2 Wavelet Packet Decomposition (Multi-Scale Feature Extraction) Mother wavelet selection: Daubechies-4 (db4), which balances tight support and sensitivity to sharp peaks.
[0069] Maximum number of decomposition layers: L = 5 (ensuring minimum subband bandwidth ≈ 31 cm) -1 (The narrow shoulder peak can be distinguished).
[0070] Perform discrete wavelet packet decomposition on each sample spectrum s(ν); obtain the lowest level approximation coefficients A. L With multiple sets of detail coefficients ( = 1…L, b represents the sub-band index).
[0071] Coefficient splicing: Fix the sub-band order, and put A L and all Concatenate them sequentially into vectors of the same length w=[A L ,D5,∗,D4,∗,…,D1,*] Let q be the total dimension of the wavelet coefficients.
[0072] 2.4.3 Energy Threshold Denoising Calculate the energy of each detail subband: , Let the threshold τ = 0.8 × median (E); if If the value is less than τ, then the subband coefficient is set to zero to suppress high-frequency noise and instrument drift.
[0073] 2.4.4 Coefficient Matrix Construction and Standardization Stack the wavelet vectors w of all samples into W ∈ R n × q Perform Z-score on each column of W: .
[0074] Obtain the standardized matrix .
[0075] 2.4.5 Functional Principal Component Analysis (fPCA) 1) Functional perspective: Treat each column of wavelet coefficients as discrete values of a spectral function on a specific wavelet basis. That is, the set of function samples.
[0076] 2) Covariance operator .
[0077] 3) Eigenvalue decomposition Find the eigenpair (λ, q) of Σ; the eigenvector q. j That is, "principal component loadings of the function", eigenvalues λ j This indicates the explained variance.
[0078] 4) Selecting principal components: Take the first two digits. .
[0079] 5) Calculate the score Z= [q1,q2]∈R n×2 row vector z i =(PC1 i PC2 i ) represents the sample coordinates.
[0080] 2.4.6 PC1-PC2 Score Chart Plot the scatter plot: PC1 on the horizontal axis and PC2 on the vertical axis. Label the four classes (Raw, Pan, Salt, and Wine) with different symbols / colors. Indicate the contribution rate in the axis titles; if confidence limits are required, calculate the covariance for each class and plot a 95% Hotelling T² ellipse. See the analysis results below. Figure 6-7 The results showed that raw, salt-fried, wine-fried, and plain-fried dodder seeds were grouped into separate categories without overlap, indicating significant differences between the groups. The results obtained by principal component analysis were consistent with those obtained by cluster analysis; therefore, both pattern recognition methods can be used for the identification of dodder seeds and their different processed forms. Furthermore, as shown in Table 3, the method of this invention is significantly superior to the linear PCA comparison scheme.
[0081] Table 3 Comparison of Principal Component Visualization and Explanatory Power
[0082] The foregoing description of embodiments of the present invention, through which those skilled in the art are able to implement or use the present invention, will be readily apparent to those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novelty disclosed herein.
Claims
1. A rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its processed products, characterized in that, The method includes the following steps: 1) Sample preparation: Place dodder seed and / or its processed products in an oven to dry, cool and then pulverize. The powder is sieved and sealed for storage. 2) Infrared spectral acquisition: The powder sample obtained in step 1) was analyzed using an infrared spectrometer with an integrating sphere diffuse reflectance method. The spectral acquisition range was 4000 cm⁻¹. -1 ~500cm -1 The resolution is 0.1-10 nm, the number of scans is 20-50, and the infrared spectrum is recorded; 3) Spectral preprocessing: Perform one or more preprocessing operations on the infrared spectrum. The preprocessing operations are selected from automatic baseline calibration, smoothing, first derivative, and vector normalization. The infrared spectral wavelength range is selected from the following: segment I 4000 cm⁻¹ -1 ~2500cm -1 Section II 2000cm -1 ~500cm -1 ; 4) Pattern recognition analysis: The infrared spectral data preprocessed in step 3) is imported into the cluster analysis and principal component analysis models. By observing and judging the cluster analysis dendrogram and principal component score map, raw dodder and / or its processed products can be quickly distinguished and identified.
2. The rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its different processed products according to claim 1, characterized in that, In step 3), a cubic polynomial is selected. p(ν)=a 0 +a 1 ν+a 2 ν 2 +a 3 ν 3 As the initial baseline model, where ν is the wavelength sequence and a0, a1, a2, a3 are the coefficients to be fitted; the least squares method is used iteratively in conjunction with the following adaptive elimination rule: (a) After the initial fitting, absorption peaks above a certain threshold are identified and temporarily removed; (b) Fit the remaining data twice or multiple times to make the polynomial baseline fit the non-absorbing region as closely as possible; (c) After the iteration is completed, subtract the final baseline function from all the original data to obtain the baseline-corrected spectrum.
3. The rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its different processed products according to claim 1, characterized in that, In step 3), Savitzky-Golay filtering is used to perform vector normalization on the first derivative spectra of segments I and II, respectively. ’ in: ν i The i-th wavenumber point is the abscissa of the infrared spectrum; D 1 (ν i ) Spectral region I, at wavenumber ν i The first derivative spectral value at the location; D 2 (ν i ) In spectral band II, at wavenumber ν i The first derivative spectral value at the location; ∑ j [D 1 (ν j )] 2 The summation of the squares of the first derivative values of all wavenumber points j within spectral region I represents the squared magnitude of the first derivative spectral vector of region I. ∑ j [D2(ν j )] 2 The summation of the squares of the first derivative values of all wavenumber points j within spectral region II represents the squared magnitude of the first derivative spectral vector of region II. D 1,norm (ν i ): D1(ν i The value after vector normalization according to the modulus of the entire first derivative spectrum; D 2,norm (ν i ): D2(ν i The value after vector normalization according to the modulus of the entire first derivative spectrum; At the segment splicing point, the data at both ends are normalized and then spliced using a transition weighting strategy to obtain the overall first derivative spectrum. D norm (ν) .
4. The rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its different processed products according to claim 1, characterized in that, The cluster analysis in step 4) includes the following steps: 4.1.1) Nonlinear eigenmap: Input the preprocessed spectral matrix X into the radial basis function kernel , Where x i ,x j For the first i , No. j The feature vectors of each sample are used to construct the centered kernel matrix K by taking the median of the pairwise Euclidean distances between the samples. 4.1.2) Kernel Principal Component Reduction: Perform eigenvalue decomposition on the kernel matrix K, and select the top p eigenvectors with a cumulative contribution rate of not less than 95% to obtain the low-dimensional score matrix Z=[z1,z2,…,zn]. T ∈R n×p ; 4.1.3) Local density estimation: For each sample i in matrix Z, calculate the number of samples ρ within its k nearest neighbor radius. i , as local density; 4.1.4) Adaptive distance construction: Define a new distance between samples i and j: , Among them, z i , z j Let ρ be the low-dimensional score vector obtained by kernel PCA for samples i and j. i , ρ j This provides the local density estimates for samples i and j in a low-dimensional space, aiming to compress the distance between high-density regions and amplify the distance between low-density regions. 4.1.5) Density-weighted Ward aggregation: using the adaptive distance d * ij The Ward chain criterion is used to perform clustering on the input, and the local density is included as a weight in the calculation of the intra-class squared error increment. 4.1.6) Determining the optimal number of clusters: During the aggregation process, the average silhouette coefficient s(k) and GapStatistic are calculated in real time. When s(k) reaches a local maximum and the Gap value no longer increases significantly for the first time, the optimal number of clusters K is automatically determined. ∧ ; 4.1.7) Output Results: Based on the cluster identifiers obtained in step 6), perform rapid and non-destructive classification of raw dodder seed, stir-fried dodder seed, salt-fried dodder seed, and wine-fried dodder seed.
5. The rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its different processed products according to claim 4, characterized in that, Step 4.1.6) includes the following steps: In the bottom-up agglomerative hierarchical clustering process, when the number of clusters decreases from k+1 to k, calculate the following for the current cluster set C(k)={C1,…,Ck}: ① Average profile coefficient ; Where a i Let b be the average distance between sample i and other members in the cluster. i Let i be the average distance between sample i and its nearest neighbor cluster; ②Gap Statistic ; Among them W K *b Let B be the intra-class dispersion of the b-th reference distribution when it is distributed across k clusters, and B be the number of reference data groups. Record the curve of (k, s(k), Gap(k)) changing with aggregation, and detect it in real time: When s(k) reaches a local maximum and satisfies: Gap(k) ≥ Gap(k+1)-s k+1 ;where s k+1 The standard error of Gap(k+1) is considered as the first time that the Gap value no longer increases significantly; the current number of clusters k that satisfies the above conditions is output as the optimal number of clusters k^, and is used as the final pruning level of the hierarchical tree diagram.
6. The rapid and non-destructive infrared spectroscopy method for identifying Cuscuta chinensis and its different processed products according to any one of claims 1 to 5, characterized in that, The lower principal component analysis method in step 4) includes the following steps: 4.2.1) Multi-scale wavelet packet decomposition: For each normalized spectrum s(ν), a wavelet packet decomposition of no more than 5 levels is performed using the Daubechies-4 mother wavelet to obtain the lowest-level approximation coefficients A. L and detail coefficients of each layer ; 4.2.2) Energy Threshold Denoising: Calculate the energy of each detail subband: , Where u is a discrete point, and the threshold τ is set to 0.8 × all The median, for satisfying The subband coefficients of <τ are all set to zero; 4.2.3) Coefficient vector concatenation and matrix construction: Concatenate A according to the fixed sub-band order. L and those not set to zero The wavelet coefficient vectors are sequentially concatenated into a wavelet coefficient vector w. i Then, by stacking these matrices, we obtain the coefficient matrix W = [w1;w2;…;wn] ∈ R. n×q ; 4.2.4) Z-score standardization: Perform mean centering and standard deviation normalization on each column of matrix W to obtain the standardized matrix. ; 4.2.5) Principal Component Decomposition of Functions: Constructing the Covariance Operator , Perform eigenvalue decomposition on it and select the top two eigenvectors q1 and q2 that make the cumulative variance contribution rate no less than 85%; 4.2.6) Score Calculation and Visualization: [This section likely refers to a specific step or function, but without The score matrix is obtained by projecting the matrix onto the q1 and q2 directions. Z= [q1,q2]∈R n×2 , Using the two columns of Z as the horizontal and vertical axes respectively, a PC1-PC2 score diagram was plotted, and different labels were used to distinguish between raw dodder seed, stir-fried dodder seed, salt-fried dodder seed, and wine-fried dodder seed, thereby achieving rapid and non-destructive separation of the four types of processed products and traceable interpretation of key vibration bands.
7. An infrared spectroscopy system for rapid and non-destructive identification of Cuscuta chinensis and its different processed products, characterized in that, The system is used to implement the method according to any one of claims 1 to 6, comprising: The infrared spectroscopy detection device is configured to acquire infrared spectra of prepared dodder seed or its processed powder samples using an integrating sphere diffuse reflectance method, with an acquisition range of 4000 cm⁻¹. -1 ~500cm -1 It has a resolution of 0.1-10nm and supports multiple scans; The data processing module is communicatively connected to the infrared spectral detection device and is used to receive the infrared spectral data output by the detection device and perform one or more of the following: automatic baseline calibration, smoothing, first derivative, and vector normalization of the spectrum. The analysis and determination module, connected to the data processing module, is used to perform cluster analysis and / or principal component analysis to determine the processing types of raw dodder seed and its stir-fried, wine-fried, and salt-fried dodder seed based on the obtained spectra and the results of the chemometric model.
8. The application of the method according to any one of claims 1 to 6 in the quality control of Cuscuta chinensis slices, characterized in that, By conducting rapid, non-destructive testing on raw dodder seeds or processed products such as stir-fried, stir-fried with wine, or stir-fried with salt, the type and quality of processed products can be accurately determined.
9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by the processor, they implement steps 3)-4) of the method according to any one of claims 1-6).
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement steps 3)-4) of the method according to any one of claims 1-6).