Pericarpium citri reticulatae production place identification method based on graph regularization sparse principal component analysis and support vector machine

By combining sparse principal component analysis with graph regularization and support vector machine algorithm, and using terahertz spectroscopy technology to process tangerine peel samples, the problems of tangerine peel origin identification methods in the existing technology, such as strong subjectivity, cumbersome operation and poor recognition effect, are solved, and fast and accurate tangerine peel origin identification is achieved.

CN120594441APending Publication Date: 2025-09-05GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510734010.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing methods for identifying the origin of tangerine peel are highly subjective, cumbersome, time-consuming, and not sensitive enough. They cannot effectively utilize the sample spatial correlation information implicit in the spectral data, resulting in poor recognition results.

Method used

Sparse principal component analysis combined with graph regularization and support vector machine algorithm was used. Tangerine peel samples were processed using terahertz spectroscopy technology, and the particle swarm optimization algorithm was used to find the optimal model parameters to construct a fast and accurate tangerine peel origin identification model.

Benefits of technology

It can realize the rapid and accurate identification of dried tangerine peel from different production areas, simplify the operation process, improve the sensitivity and accuracy of detection, and solve the problem of counterfeit and inferior products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120594441A_ABST
    Figure CN120594441A_ABST
Patent Text Reader

Abstract

The invention relates to a terahertz spectrum detection technology, in particular to a pericarpium citri reticulatae producing area identification method based on graph regularization sparse principal component analysis and a support vector machine algorithm, and belongs to the field of pericarpium citri reticulatae producing area quality detection. The method comprises the following steps: grinding and tabletting a dried orange peel sample to be detected, firstly detecting the dried orange peel sample in a nitrogen environment by adopting a terahertz time-domain spectroscopy system in a transmission mode to obtain a terahertz time-domain spectroscopy signal of the sample, and performing Fourier transform on the time-domain spectroscopy signal to obtain a frequency-domain spectrum of the sample; and obtaining a corresponding terahertz absorption spectrum according to the frequency domain spectrum. Savitzky-Golay smoothing preprocessing is carried out on the obtained terahertz absorption spectrum, the terahertz absorption spectrum is divided into a training set and a test set, and then feature extraction is carried out on data by utilizing sparse principal component analysis in combination with graph regularization. And taking the processed data as the input of a classification model. And finally, a combined parameter of an optimal regularization coefficient c and a kernel function parameter g of the support vector machine is obtained through a particle swarm optimization algorithm, so that an optimal pericarpium citri reticulatae producing area classification model is established. The method provided by the invention is convenient in sample preparation and simple in operation, can effectively realize rapid and accurate identification of dried orange peel from different producing areas, and provides a new method for identification of dried orange peel in high-value producing areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of terahertz spectrum detection technology, and specifically relates to a method for identifying the origin of tangerine peel based on graph regularization sparse principal component analysis and support vector machine algorithm. Background Art

[0002] Tangerine peel is the peel of the ripe fruit of the Rutaceae plant citrus and its cultivated varieties, which is obtained by sun-drying or low-temperature drying. The 2020 edition of the "Pharmacopoeia of the People's Republic of China" clearly states that tangerine peel has the effects of regulating qi and strengthening the spleen, drying dampness and resolving phlegm, and can effectively relieve symptoms such as abdominal distension, loss of appetite, vomiting and diarrhea, and coughing and sputum. At the same time, tangerine peel can also be made into tea. Different from its medicinal effects, it has the effects of refreshing the mind, removing phlegm and regulating qi when used in tea. Many studies have shown that hesperidin can lower blood lipids, prevent diabetes, reduce cholesterol, and also has antibacterial and analgesic effects. The value and price of tangerine peel can vary greatly depending on its origin. There are three main types of tangerine peels: the first is tangerine peels other than Guangdong tangerine peels, which are made from the dried and mature peels of plants such as Dahongpao and Fuju, and are mainly produced in Yunnan, Hubei, Sichuan, Fujian, and other places; the second is Guangdong tangerine peels, which are derived from tea-branch tangerine peels, with Guangdong being its main production area; and the third is mixed tangerine peels, which are mixed with the peels of other citrus and orange varieties and used as tangerine peels for medicinal purposes. Since the value of tangerine peels from different production areas varies greatly, counterfeit and inferior products are often found, so it is particularly important to detect and identify the origin of tangerine peels.

[0003] At present, the conventional method of analyzing the quality of dried tangerine peel is the analytical method of sensory evaluation or high performance liquid chromatography, wherein the sensory evaluation method is too subjective, and the result is not stable enough, and conventional liquid chromatography has cumbersome operating steps, consumes a large amount of chemical reagents, and detection is time-consuming and longer, and there are problems such as environmental protection simultaneously. Therefore, the terahertz spectroscopy technology with higher safety and sensitivity obtains more and more attention. However, the processing for terahertz spectral data at present often ignores the sample space correlation information (as the spectral similarity of dried tangerine peel in the place of production) implicit in the spectral data, causes the discriminant power of feature extraction to be insufficient, and recognition effect is not good. Therefore, the present invention proposes the sparse principal component analysis in conjunction with graph regularization, and then filters out the key features most relevant to dried tangerine peel quality to the greatest extent, strengthens the model to the capture ability of the implicit category structure in the spectral data, and by introducing particle swarm optimization algorithm (Particle Swarm Optimization, PSO) support vector machine (SVM) model is carried out parameter optimization, seeks model optimal parameter combination. In recent years, feature extraction technology is applied to food quality analysis in conjunction with chemometric method and obtains good effect, and has laid the foundation for the application of this technology in dried tangerine peel detection and analysis.

[0004] Therefore, it is urgent to design a method for identifying the origin of tangerine peel based on graph regularization sparse principal component analysis and support vector machine algorithm to solve the above technical problems. Summary of the Invention

[0005] The present invention aims to provide a method for identifying tangerine peel origin using sparse principal component analysis combined with graph regularization and a support vector machine algorithm. This method can accurately and quickly identify tangerine peel from multiple different producing areas, addressing the problem of counterfeiting and passing off inferior tangerine peel from different producing areas as genuine products. The detection process is simple and convenient, allowing for identification and analysis of tangerine peel origin without complex operations.

[0006] The technical problem to be solved by the present invention is to quickly, accurately and conveniently identify the origin of dried tangerine peel. To solve the above technical problem, the present invention is implemented through the following technical solutions.

[0007] A method for identifying the origin of tangerine peel by combining sparse principal component analysis with graph regularization and support vector machine algorithm comprises the following steps:

[0008] S1: Crush, sieve, and slice dried tangerine peels from different origins to obtain round pieces of tangerine peel with smooth surfaces;

[0009] S2: Use the terahertz time-domain spectroscopy system to detect tangerine peel samples from different origins, and obtain terahertz time-domain spectroscopy datasets of tangerine peel from different origins;

[0010] S3: The terahertz time-domain spectrum obtained in S2 is converted into a frequency-domain signal by fast Fourier transform (FFT), and the frequency-domain signal is then calculated to obtain the absorption spectrum;

[0011] S4: Perform Savitzky-Golay smoothing preprocessing on the absorption spectrum obtained in S3 to suppress noise and smooth the signal;

[0012] S5: Combine the data obtained in S4 with its category information and the hybrid nearest neighbor strategy to construct a weighted graph and obtain the Laplace matrix;

[0013] S6: Based on the original L1 and L2 regularization of sparse principal component analysis, graph regularization is introduced to complete data dimensionality reduction and feature extraction;

[0014] S7: Use support vector machine (SVM) to build a qualitative identification model for dried tangerine peel from different origins, and use the characteristic data of the spectrum as the input data set;

[0015] S8: Use particle swarm optimization (PSO) algorithm to improve the support vector machine model and obtain the optimal parameter combination of the support vector machine;

[0016] S9: Input the terahertz spectrum dataset of tangerine peel from different origins to be identified into the optimized qualitative identification model and output the identification results;

[0017] Optionally, the thickness of the sample is controlled by a weight method during the sample production process, and the mass of each sample is weighed to be 200 mg, and the same mold is used for tableting.

[0018] Optionally, the improvement of the support vector machine is to search for the optimal value of the regularization parameter c of the support vector machine and the parameter g of the Gaussian radial basis kernel function through the particle swarm optimization algorithm (PSO), and finally obtain the combination of parameters c and g that best suits the model.

[0019] Optionally, the input data set is divided into a training set and a test set in a ratio of 3:2, the training set is used for model learning, and the test set is used for model judgment and identification.

[0020] Optionally, the output result of the qualitative identification model uses accuracy as an evaluation indicator of model performance.

[0021] A method for identifying dried tangerine peel origin using sparse principal component analysis combined with graph regularization and a support vector machine algorithm has been developed. Based on terahertz time-domain spectroscopy, this method combines sparse principal component analysis with graph regularization to qualitatively analyze the more expensive Xinhui dried tangerine peel and tangerine peel from other origins with significantly different prices, enabling identification of dried tangerine peel samples from different origins. Feature extraction from terahertz spectral data often overlooks the implicit spatial correlation information within the spectral data. Sparse principal component analysis combined with graph regularization maximizes the selection of key features most relevant to dried tangerine peel quality, enhancing the model's ability to capture the implicit category structure within the spectral data. This addresses the issue of insufficient spectral feature information and low model recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein

[0023] Figure 1 The present invention is a flow chart of a method for identifying the origin of tangerine peel based on graph regularization sparse principal component analysis and support vector machine algorithm.

[0024] Figure 2 It is the original time domain spectrum of the tangerine peel samples from different origins in the range of 0-50 ps.

[0025] Figure 3 It is the original absorption spectrum of the tangerine peel samples from different origins in the characteristic band of 0.2-1.6 THz.

[0026] Figure 4 It is the weight distribution histogram of the weighted adjacency matrix W(i,j) described in the present invention.

[0027] Figure 5This is the PSO optimization SVM parameter process when the input is terahertz spectral data as described in the present invention. DETAILED DESCRIPTION

[0028] In order to make the objects and advantages of the present invention more clearly understood, the specific implementation of the present invention is further described in detail below with reference to the accompanying drawings.

[0029] The implementation flow chart of the present invention is as follows: Figure 1 As shown, the specific implementation steps are as follows:

[0030] S1: Crush, sieve, and slice tangerine peels from different origins to produce round tangerine peel samples with smooth surfaces;

[0031] S2: Use the terahertz time-domain spectroscopy system to detect tangerine peel samples from different origins, and obtain terahertz time-domain spectroscopy datasets of tangerine peel from different origins;

[0032] S3: The terahertz time-domain spectrum obtained in S2 is converted into a frequency-domain signal by fast Fourier transform (FFT), and the frequency-domain signal is then calculated to obtain the absorption spectrum;

[0033] S4: Perform Savitzky-Golay smoothing preprocessing on the absorption spectrum obtained in S3 to suppress noise and smooth the signal;

[0034] S5: Combine the data obtained in S4 with its category information and the hybrid nearest neighbor strategy to construct a weighted graph and obtain the Laplace matrix;

[0035] S6: Based on the original L1 and L2 regularization of sparse principal component analysis, graph regularization is introduced to complete data dimensionality reduction and feature extraction;

[0036] S7: Use support vector machine (SVM) to build a qualitative identification model for tangerine peel from different origins, and use the characteristic data of the spectrum as the input data set;

[0037] S8: Use particle swarm optimization (PSO) algorithm to improve the support vector machine model and obtain the optimal parameter combination of the support vector machine;

[0038] S9: Input the terahertz spectrum dataset of tangerine peel from different origins to be identified into the optimized qualitative identification model and output the identification results;

[0039] Specifically, in the step (1), sample pieces are prepared by placing dried tangerine peels from different origins into a vacuum dryer to dry and remove moisture brought during storage, each type of sample is placed into a grinder for pulverization, sieved to remove coarse particles, and then re-ground and mixed evenly with an appropriate amount of polyethylene powder in a grinding dish, wherein the mixing ratio of the sample to the polyethylene powder is 3:1, and 200 mg of the mixed powder is weighed out, placed into a mold, and pressed for 30 seconds under a tableting pressure of 12 MPa to obtain dry and smooth experimental samples.

[0040] Specifically, in step (2), the sample spectrum is collected, and the terahertz time-domain spectrum collection is carried out with the help of a spectrum collection system composed of a Huaxun Ark CCT-1800 terahertz spectrometer (Huaxun Ark Technology, China) and a control computer. The detection range of the system is 0.1-4THz, the resolution can reach 20GHz, and the signal-to-noise ratio in the low frequency band can be as high as 75dB. The experiment is carried out at room temperature. First, the sample chamber is filled with nitrogen. After the real-time spectrum signal is stable, the reference signal is collected. Then, the sample is placed on the rack, and the terahertz spectrum signal of the sample is collected after the signal is stable. After completing the parameter setting, the terahertz spectrum data is collected, and the spectrum after three averages is selected. A total of 240 sets of data are obtained in this collection, and there are 48 sets of spectrum data for each origin.

[0041] Specifically, in step (3), the terahertz absorption spectrum and preprocessing method are calculated, the terahertz time domain spectrum is obtained by fast Fourier transform (FFT) and the absorbance formula is used to calculate the absorbance A. sam The calculation formula is as follows:

[0042]

[0043] Where, E sam (ω) is the spectral signal of the sample, E ref (ω) is the reference signal.

[0044] Specifically, in step (4), the terahertz absorption spectrum is preprocessed using the Savitzky-Golay preprocessing method. The preprocessed absorption spectrum eliminates interference noise, greatly reducing the impact on the sample analysis results and making the data information more effective.

[0045] Specifically, in step (5), for the terahertz spectrum dataset and its corresponding labels, a weighted adjacency matrix W(i, j) is constructed. This graph structure is constructed based on the cosine similarity between samples, where the connection weights between similar samples are larger, while the weights between heterogeneous samples are smaller. The calculation formula of W(i, j) is as follows.

[0046]

[0047] Where, sim(xi ,x j ) is the cosine similarity between samples, α is the weight coefficient, and σ is the Gaussian kernel bandwidth parameter. The final graph Laplacian matrix can be obtained through this adjacency matrix.

[0048] Specifically, in step (6), sparse principal component analysis is combined with graph regularization to perform data dimensionality reduction and feature extraction. Sparse principal component analysis (SPCA) is a dimensionality reduction method that combines principal component analysis (PCA) with sparse constraints, aiming to extract interpretable key features from high-dimensional data. Its core idea is to retain the main variance information of the data while forcing most elements in the principal component load vector to be zero by introducing a sparse regularization term (L1 norm), thereby screening out features that contribute significantly to the principal component and achieving the dual goals of "feature selection" and "dimensionality reduction". Graph regularization is a technology that integrates the topological structure of data (such as the similarity relationship between samples) into the dimensionality reduction process. The core goal of introducing graph regularization in SPCA is to explicitly retain the local geometric relationship between samples while retaining the variance structure of the data, thereby improving the discriminative ability of features after dimensionality reduction. After combination, its objective function is as follows.

[0049] min B,A Tr(X T XX T AB T X)+λ1||B||1+λ2Tr(B T X T LXB) (3)

[0050] Where L is the Laplace matrix constructed in step (5), X is the data matrix, A is the projection matrix, and B is the sparse loading matrix. λ1 and λ2 are the regularization strengths. The goal is to minimize information loss and retain the main variance information of the data, measuring the degree to which the local structure of the sample in the low-dimensional space is preserved after dimensionality reduction. It is hoped that during the dimensionality reduction process, the original similarity relationship between samples will be maintained, that is, similar samples will still maintain a close distance in the low-dimensional space, without destroying the local geometric structure of the data, making the features after dimensionality reduction more conducive to subsequent classification tasks.

[0051] Specifically, a support vector machine model is established in the step (7). The support vector machine (SVM) is a supervised classifier based on statistical learning theory, which can map input data to a higher-dimensional space, making the data linearly separable in the space, and avoiding complex multi-dimensional calculations in the feature space by introducing the kernel function. At the same time, regularization processing is adopted to reduce the risk of overfitting of the model and make it have a strong generalization ability. In practical applications, the regularization constant c and the kernel function parameter g are important parameters that affect the performance of the model, so it is necessary to select the best combination parameters of c and g. The present invention adopts the Gaussian radial basis kernel function k(x i ,xj ) is used as the kernel function of the model, and its mathematical expression is

[0052]

[0053] where σ rate is the arrival rate of the Gaussian radial basis kernel function. The smaller the value, the narrower the kernel function.

[0054] Specifically, in step (8), a particle swarm optimization algorithm (PSO) is used to optimize parameters, and the best combination parameters of c and g are selected in SVM to achieve the best classification accuracy.

[0055] Specifically, the evaluation index of the model output result in step (9) uses the ratio of overall accuracy to evaluate the performance of the classification model. The higher the accuracy rate obtained, the better the corresponding model classification effect.

[0056] See also Figures 2 to 5 , the present invention provides a specific embodiment:

[0057] (1) Selection of research subjects and preparation of test samples

[0058] Tangerine peel from different production areas is the specific implementation object of the present invention. The production areas are respectively Guangdong Xinhui, Guangxi Pubei, Sichuan Chengdu, Zhejiang Quzhou and Hunan Zhangjiajie, and are purchased from local farmers. All samples are well preserved and pulverized to obtain clean and dry tangerine peel powder after sample drying. Finally, 200mg of powder is pressed into sheets under a pressure of 12MPa, with a sample diameter of 13mm and a thickness of 1mm. 48 sheets of sample sheets are made for each year, making a total of 240 sheets. The sample data set is divided into a training set and a test set with a ratio of 3:2, with 24 groups of data used as training sets and 16 groups of data used as test sets for each year. The total training set is 144 groups of data, and the test set is 96 groups.

[0059] (2) Terahertz spectrum acquisition

[0060] All samples were collected using a terahertz time-domain spectroscopy system. The time-domain spectra are shown in Figure 2 , first perform fast Fourier transform (FFT) to convert it into frequency domain spectrum, and then calculate the frequency domain spectrum into absorption spectrum through absorbance formula. Figure 3 .

[0061] (3) Spectral preprocessing

[0062] The terahertz absorption spectrum is preprocessed, and the preprocessed spectral data removes most of the interference signals such as noise.

[0063] (4) Constructing a weighted adjacency matrix

[0064] The dataset is constructed based on its corresponding labels and the cosine similarity between samples to form a graph structure. The two-dimensional graph structure is shown in the figure. The connection weight between similar samples is larger, while the weight between heterogeneous samples is smaller. The sample weight histogram is as follows: Figure 4 .

[0065] (5) Constructing PSO-improved SVM model

[0066] The particle swarm optimization algorithm (PSO) is used to find the parameters of the support vector machine. The absorbance data is used as the input of the model. The results of the particle swarm optimization algorithm for the support vector machine parameter optimization process are shown in Figure 5 During the iteration process, the best fitness obtained by PSO-SVM reached 97.52%, and the best fitness was also achieved in the later stages of the iteration.

[0067] (6) Evaluation and analysis of qualitative identification models

[0068] Table 1 shows that the model trained using PCA feature extraction and PSO-SVM achieved 97.93% classification accuracy on the training set, but only 92.63% recognition accuracy on the test set. This compares to the model trained using SPCA feature extraction and PSO-SVM, which achieved 100% classification accuracy on the training set and 93.68% recognition accuracy on the test set. Finally, the model trained using Graph Regularized Sparse Principal Component Analysis (GF-SPCA) feature extraction and PSO-SVM achieved 100% classification accuracy on the training set and 96.84% recognition accuracy on the test set. Experimental results demonstrate that GF-SPCA significantly improves model accuracy compared to simple PCA or SPCA feature extraction algorithms.

[0069] Table 1 Comparison of the performance of tangerine peel origin identification models

[0070]

[0071] The above is only a representative embodiment of the present invention and this example cannot be used to limit the scope of the rights of the present invention. Those skilled in the art should be aware that merely modifying or equivalently replacing the technical solution of the present invention without departing from the purpose and scope of the technical solution of the present invention still falls within the scope covered by the present invention.

Claims

1. A method for identifying the origin of dried tangerine peel based on a graph regularization sparse principal component analysis algorithm and a support vector machine algorithm, characterized in that: The following steps are involved: S1: Crush, sieve, and slice dried tangerine peels from different origins to obtain round pieces of tangerine peel with smooth surfaces; S2: Use the terahertz time-domain spectroscopy system to detect tangerine peel samples from different origins, and obtain terahertz time-domain spectroscopy datasets of tangerine peel from different origins; S3: The terahertz time-domain spectrum obtained in S2 is converted into a frequency-domain signal by fast Fourier transform (FFT), and the frequency-domain signal is then calculated to obtain the absorption spectrum; S4: Perform Savitzky-Golay smoothing preprocessing on the absorption spectrum obtained in S3 to suppress noise and smooth the signal; S5: Combine the data obtained in S4 with its category information and the hybrid nearest neighbor strategy to construct a weighted graph and obtain the Laplace matrix; S6: Based on the original L1 and L2 regularization of sparse principal component analysis, graph regularization is introduced to complete data dimensionality reduction and feature extraction; S7: Use support vector machine (SVM) to build a qualitative identification model for tangerine peel from different origins, and use the characteristic data of the spectrum as the input data set; S8: Use particle swarm optimization (PSO) algorithm to improve the support vector machine model and obtain the optimal parameter combination of the support vector machine; S9: Input the terahertz spectrum data set of tangerine peels from different origins that need to be identified into the optimized SVM qualitative identification model and output the identification results.

2. The method for identifying the origin of dried tangerine peel based on graph regularization sparse principal component analysis algorithm and support vector machine algorithm according to claim 1, characterized in that: The circular sample obtained by pressing is detected by a terahertz time-domain spectroscopy system to obtain terahertz time-domain spectroscopy data.

3. The Laplacian matrix constructed according to claim 1, characterized in that By combining the category information of samples with a hybrid nearest neighbor strategy to construct a weighted graph, a semi-supervised graph learning framework was established that can simultaneously capture the local manifold structure and category discrimination information of the data.

4. The sparse principal component analysis method combined with graph regularization according to claim 1, characterized in that: By adding graph regularization in the objective function optimization stage of sparse principal component analysis to modify the objective function of traditional SPCA, the graph structure information is incorporated into the feature extraction process. Improve the performance of the algorithm in nonlinear scenarios and establish graph regularized sparse principal component analysis (GR-SPCA).

5. The improved support vector machine model according to claim 1, characterized in that The PSO optimization algorithm is used to optimize the regularization coefficient c and kernel function parameter g of the support vector machine to accelerate the convergence speed and establish a POS-SVM classification model.

Citation Information

Cited By

  • Method for identifying production place and aging year of pericarpium citri reticulatae

    CN121540743A