Method for identifying fructus aurantii and fructus aurantii immaturus categories and origins

By combining ATR-FTIR spectroscopy and wavelet transform with SVM classifier, the accuracy problem of identifying the categories and origins of Citrus aurantium and Citrus aurantium was solved, and efficient and accurate identification of Citrus aurantium and Citrus aurantium was achieved, improving the accuracy and speed of medicinal material identification.

CN119804375BActive Publication Date: 2025-10-21HUNAN UNIV OF CHINESE MEDICINE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411859711.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-10-21
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing technology makes it difficult to accurately distinguish the categories and origins of Citrus aurantium and Citrus aurantium fructus, which affects the efficacy and safety of medication.

Method used

ATR-FTIR spectroscopy technology combined with wavelet transform and support vector machine (SVM) classifier was used to process and extract features from the ATR-FTIR spectra of Citrus aurantium and Citrus aurantium, and the SVM classifier was trained and optimized to identify the type and origin of Citrus aurantium and Citrus aurantium.

Benefits of technology

It achieved high-accuracy identification of the types and origins of Citrus aurantium and Citrus aurantium fructus, with an accuracy rate of 96.7%, which is much higher than traditional methods. It simplified the sample preparation process and improved the identification speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119804375B_ABST
    Figure CN119804375B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of traditional Chinese medicine material identification, and particularly relates to a method for identifying the category and origin of Fructus Aurantii Immaturus and Fructus Aurantii, which comprises the following steps: collecting ATR-FTIR spectra of training samples for multiple times, obtaining multiple ATR-FTIR spectra, averaging the ATR-FTIR spectra, and obtaining average ATR-FTIR spectra; decomposing the average ATR-FTIR spectra, obtaining scale signals; stretching and translating the scale signals, obtaining wavelet signals; reconstructing the wavelet signals, retaining large trend signals, obtaining smooth average ATR-FTIR spectra for training, extracting characteristic peaks of the smooth average ATR-FTIR spectra for training, and obtaining characteristic vectors; inputting the characteristic vectors into an SVM classifier, training an optimized SVM classifier; inputting characteristic vectors of smooth average ATR-FTIR spectra of samples to be tested into the optimized SVM classifier, and obtaining the category and origin of the samples to be tested. The application can identify the category and origin of Fructus Aurantii Immaturus and Fructus Aurantii, and the accuracy of identification is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of identification of traditional Chinese medicines, and particularly relates to a method for identifying the types and origins of Citrus aurantium immaturus and Citrus aurantium fructus. Background Art

[0002] Zhishi (Fructus Aurantii) and Zhipi (Fructus Aurantii) are commonly used Qi-regulating herbs. They share a common origin. Zhishi (Fructus Aurantii) is the dried young fruit of Citrus aurantium L. and its cultivars, or Citrus sinensis Osbeck (Sweet Orange). The fruit is collected from May to June each year after it falls, cut in half across the middle, and then sun-dried or low-temperature dried. Smaller fruit are directly sun-dried or low-temperature dried. Zhipi (Fructus Aurantium) is the dried immature fruit of Citrus aurantium L. and its cultivars, also in the Rutaceae family. The fruit is harvested in July while the peel is still green, cut in half across the middle, and then sun-dried or low-temperature dried.

[0003] The different harvesting times of Citrus aurantium and Citrus aurantium fructus result in different medicinal effects. Citrus aurantium is known to dissipate qi and resolve stagnation, resolve phlegm, and relieve abdominal distension; Citrus aurantium is known to regulate qi and relieve fullness, dissipate stagnation, and alleviate distension. Citrus aurantium is highly potent and is often used to dissipate stagnation; Citrus aurantium fructus is known to be milder and is often used to regulate qi and relieve fullness, dissipate bloating, and alleviate fullness. Citrus aurantium and Citrus aurantium need to be used differently in clinical practice, and therefore are included as two distinct medicinal herbs in the Chinese Pharmacopoeia.

[0004] The types of Citrus aurantium and Citrus aurantium fructus are determined according to the harvesting time of the fruit. The quality is affected by many factors, such as the growth environment (origin) of the plant and storage conditions. If they are not distinguished, it will directly affect the doctor's medication safety and efficacy. Therefore, it is necessary to identify the types and origins of Citrus aurantium and Citrus aurantium fructus.

[0005] In the 2020 edition of the Chinese Pharmacopoeia, synephrine is measured in Citrus aurantium immaturus, and naringin and neohesperidin are measured in Citrus aurantium fructus. Although these components are active ingredients, they are not exclusive to the genus Citrus, and cannot distinguish between related Chinese herbs from different growth stages of the same plant. Selecting only a few common indicator components cannot achieve effective identification, and it cannot distinguish between origins.

[0006] A thin-layer chromatography-based study on the identification of Citrus aurantium and its counterfeits, by Yuan Hanwen et al., Journal of Hunan University of Chinese Medicine, Vol. 41, No. 10, October 2021, published the following information: A comparative thin-layer chromatography study of Citrus aurantium and its easily confused counterparts, Citrus aurantium fructus, Citrus aurantium citrifolia, and Citrus reticulata, in the Chinese Pharmacopoeia, found that while all four herbs could be analyzed, their chemical compositions were highly similar and could not be specifically distinguished. Citrus aurantium and Citrus aurantium fructus share the same origin, but are harvested at different times. Experimental results indicate that their chemical compositions are highly similar, making them difficult to distinguish from a chemical perspective.

[0007] CN201710555332.X discloses a method for identifying Citrus aurantium, Citrus aurantium and Citrus aurantium based on chemical classification and UPLCT of MS. It distinguishes the varieties and growth stages of samples by measuring and analyzing the primary and secondary metabolites in the samples, but it cannot effectively distinguish the origins of Citrus aurantium and Citrus aurantium. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a method for identifying the category and origin of Citrus aurantium and Citrus aurantium, which can identify the category and origin of Citrus aurantium and Citrus aurantium with high accuracy.

[0009] The embodiment of the present invention provides a method for identifying the type and origin of Citrus aurantium and Citrus aurantium, comprising the following steps:

[0010] ATR-FTIR spectra of training samples were collected multiple times to obtain multiple ATR-FTIR spectra, and the training samples were Fructus Aurantii Immaturus or Fructus Aurantii Immaturus from different origins;

[0011] Averaging multiple ATR-FTIR spectra to obtain an average ATR-FTIR spectrum;

[0012] The average ATR-FTIR spectrum was decomposed to obtain the scale signal;

[0013] The scale signal is stretched and translated to obtain a wavelet signal;

[0014] The wavelet signal is reconstructed, the large trend signal is retained, and the smoothed average ATR-FTIR spectrum for training is obtained;

[0015] Extracting characteristic peaks of the training smoothed average ATR-FTIR spectrum to obtain a characteristic vector of the training smoothed average ATR-FTIR spectrum;

[0016] The feature vector of the smoothed average ATR-FTIR spectrum used for training is input into the SVM classifier to obtain an optimized SVM classifier through training;

[0017] The feature vector of the smoothed average ATR-FTIR spectrum of the sample to be tested is input into the optimized SVM classifier to obtain the category and origin of the sample to be tested;

[0018] Reconstruct the wavelet signal, retain the large trend signal, and obtain the smoothed average ATR-FTIR spectrum for training, including:

[0019] Design the wavelet signal as an orthogonal signal of the scale signal;

[0020] Obtain the spanned signal space of the wavelet signal;

[0021] The square integrable space of wavelet signal is obtained according to the spanned signal space of wavelet signal;

[0022] According to the scale signal, wavelet signal and square integrable space, the time domain signal is obtained;

[0023] The time domain signal is decomposed and reconstructed to obtain a decomposition signal set, and a large trend signal is selected from the decomposition signal set to obtain a smoothed average ATR-FTIR spectrum for training.

[0024] Optionally, the step of designing the wavelet signal as an orthogonal signal of a scale signal includes:

[0025] The wavelet signal is expressed as:

[0026] ψ j,l (t) = 2 j / 2 ψ(2t-l)

[0027] The wavelet signal ψ j,l (t) is designed as a scale signal The orthogonal signal is expressed as:

[0028]

[0029] Optionally, the training sample is prepared by removing impurities from the Fructus Aurantii Immaturus and Fructus Aurantii Immaturus, splitting them, drying them, crushing and sieving them to obtain the training sample.

[0030] Optionally, in the ATR-FTIR spectrum, the absorbance is collected in the range of 4000 to 600 cm -1 , the resolution of ATR-FTIR spectrum is 4 cm -1 .

[0031] Optionally, the average ATR-FTIR spectrum is decomposed to obtain scale signals including:

[0032] Get the preset wavelet basis;

[0033] According to the preset wavelet basis, the corresponding scaling function is obtained;

[0034] The average ATR-FTIR spectrum is decomposed according to the scaling function to obtain the scaling signal.

[0035] Optionally, scaling and translating the scaled signal to obtain a wavelet signal includes:

[0036] According to the recursive relationship of the scaling function, the wavelet function is obtained;

[0037] The scale signal is stretched and retracted by the scale function, and the scale signal is translated by the wavelet function to obtain scale signals at different scales.

[0038] The scale signals at different scales are expressed linearly to obtain a signal space composed of linear expression signals;

[0039] According to the signal space, the spanned signal space is obtained;

[0040] According to the spanned signal space, the recursive signal is obtained;

[0041] The recursive signal is scaled and translated through the wavelet function to obtain the wavelet signal.

[0042] Optionally, the time domain signal is decomposed and reconstructed to obtain a decomposed signal set, and a large trend signal is selected from the decomposed signal set to obtain a smoothed average ATR-FTIR spectrum for training, including:

[0043] The wavelet signal is linearly represented by the scaling function to obtain a linear wavelet signal;

[0044] The linear wavelet signal is deformed by recursive equation through scaling function and wavelet function to obtain the decomposed signal set;

[0045] The large trend signal is selected from the decomposed signal set to obtain the smoothed average ATR-FTIR spectrum for training.

[0046] Optionally, extracting characteristic peaks of the training smoothed average ATR-FTIR spectrum to obtain a characteristic vector of the training smoothed average ATR-FTIR spectrum includes:

[0047] The first-order derivative of the training smoothed average ATR-FTIR spectrum is taken to obtain all zero-point positions, and the peak values ​​of all zero-point positions are obtained to obtain the characteristic vector of the training smoothed average ATR-FTIR spectrum.

[0048] Optionally, the feature vector of the smoothed average ATR-FTIR spectrum used for training is input into an SVM classifier, and the optimized SVM classifier obtained by training includes:

[0049] According to the SVM classifier, obtain the preset kernel function and decision function;

[0050] Establish the objective function based on the kernel function;

[0051] The SVM classifier is trained according to the characteristic vector of the smoothed average ATR-FTIR spectrum used for training and the objective function to obtain the first parameter, the second parameter and the third parameter of the objective function;

[0052] Obtaining an adjustment decision function according to the first parameter, the second parameter, the third parameter, and the decision function;

[0053] The decision function in the SVM classifier is replaced by the adjusted decision function to obtain the optimized SVM classifier.

[0054] Optionally, the SVM classifier is trained according to the feature vector of the training smoothed average ATR-FTIR spectrum and the objective function, and the first parameter, the second parameter, and the third parameter of the objective function are obtained, including:

[0055] According to the characteristic vector of the smoothed average ATR-FTIR spectrum used for training and the objective function, the constraint function of the objective function is obtained;

[0056] Convert the constraint function into a Lagrangian function;

[0057] Set the gradient of the first and second parameters in the Lagrangian function to 0, and calculate the third parameter;

[0058] Substitute the third parameter into the constraint function and calculate the first and second parameters through the KKT condition.

[0059] The beneficial effects of the present invention are:

[0060] 1. By averaging the ATR-FTIR spectra of the training samples, an average ATR-FTIR spectrum is obtained, and then the average ATR-FTIR spectrum is decomposed to obtain a scale signal, the scale signal is stretched and translated to obtain a wavelet signal, and a large trend signal is selected from the wavelet signal to obtain a smooth average ATR-FTIR spectrum for training, and the characteristic peaks of the smooth average ATR-FTIR spectrum are extracted to form a feature vector, and then the feature vector is input into the SVM classifier for training to obtain an optimized SVM classifier. At this time, the SVM classifier training is completed, and the feature vector of the smooth average ATR-FTIR spectrum of the sample to be tested is input into the optimized SVM classifier to obtain the category and origin of the sample to be tested. Compared to the chemical classification method, the present application can not only identify the category of Citrus aurantium and Citrus aurantium, but also identify the origin of Citrus aurantium and Citrus aurantium, and the accuracy of identification is higher, reaching 96.7%, which is much higher than 80.2% of the orthogonal partial least squares-discriminant analysis method.

[0061] 2. When the ATR-FTIR spectrum is processed, the wavelet signal obtained is first designed as an orthogonal signal, and then other processes are carried out. Compared with the traditional processing mode, it is designed as an orthogonal signal to ensure that the signals do not interfere with each other. Each signal can independently represent the information in the signal space, avoiding the inaccurate signal extracted due to interference. Secondly, the use of an orthogonal basis makes the signal simpler during wavelet transformation. The signal can be represented as a linear combination of orthogonal bases, thereby the composition of each basis in the signal can be calculated by inner product. Using the processing method of the present invention, the processing effect of the ATR-FTIR spectrum is the best, the smoothness is the highest, which is 96.64%, the similarity is 98.02%, and the average value is 97.83%. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 The present invention is a flow chart of a method for identifying the types and origins of Citrus aurantium and Citrus aurantium fructus.

[0063] Figure 2 This is the ATR-FTIR spectrum of the three-region Citrus aurantium.

[0064] Figure 3 Multiple ATR-FTIR spectra were obtained for multiple measurements.

[0065] Figure 4 The smoothed average ATR-FTIR spectra obtained by different treatment methods are shown in Figure 2. (A) is the smoothed average ATR-FTIR spectrum obtained by the MSC treatment method, (B) is the smoothed average ATR-FTIR spectrum obtained by the SG treatment method, (C) is the smoothed average ATR-FTIR spectrum obtained by the SNV treatment method, and (D) is the smoothed average ATR-FTIR spectrum obtained by the treatment method of the present invention. DETAILED DESCRIPTION

[0066] A method for identifying the types and origins of Citrus aurantium and Citrus aurantium, as shown in the attached Figure 1 As shown, the present invention includes:

[0067] S1. Collecting ATR-FTIR spectra of training samples multiple times to obtain multiple ATR-FTIR spectra, where the training samples are Fructus Aurantii Immaturus or Fructus Aurantii Immaturus from different origins;

[0068] The training sample preparation method comprises removing impurities from Citrus aurantium immaturus and Citrus aurantium fructus, splitting, drying, crushing, and sieving to obtain the training sample. The training sample of the present invention is Citrus aurantium immaturus or Citrus aurantium fructus of known origin and species. Different Citrus aurantium immaturus or Citrus aurantium fructus species and origins have different feature vectors. The feature vectors are input into an SVM classifier to obtain an optimized SVM classifier, thereby distinguishing Citrus aurantium immaturus or Citrus aurantium fructus of unknown weight and origin.

[0069] The training sample of the present invention is a powder and does not need to be prepared into a solution. The sample preparation is simpler, which can greatly save the time of sample preparation and improve the identification speed.

[0070] In the ATR-FTIR spectrum, the absorbance is collected in the range of 4000-600 cm -1 , the resolution of ATR-FTIR spectrum is 4 cm -1 ATR-FTIR spectroscopy has higher sensitivity and resolution, which can collect better detailed information about the sample and improve the accuracy of identification.

[0071] Specifically, the sample preparation method for differentiation was as follows: the samples used were from Hunan Province, a major production area for Citrus aurantium and Citrus aurantium fructus. Citrus aurantium and Citrus aurantium fructus samples were collected from three locations in Hunan Province (Yiyang, Huaihua, and Changde). The Citrus aurantium and Citrus aurantium fructus were collected in June and July, for a total of 60 samples. The fresh Citrus aurantium and Citrus aurantium fructus were cleaned of impurities, cut in half along the middle, sun-dried, and crushed. The samples were placed in sealed bags, labeled with the place of origin and variety, and stored in a desiccator until further use. All samples were kept in powder form for subsequent analysis.

[0072] Before collecting sample spectra, background information was collected and subtracted from the sample spectra before further analysis. Each sample spectrum was collected three times, and the spectral data were averaged for further analysis. After each sample was collected, the ATR accessory was cleaned with an alcohol pad until it was free of contamination before collecting the next sample.

[0073] ATR-FTIR spectra of each Fructus Aurantii Immaturus and Fructus Aurantii were obtained using an FTIR spectrometer equipped with an ATR accessory. All spectra were recorded using OMNIC software and saved as absorbance values. The absorbance acquisition range of each spectrum was 4000–600 cm -1 , spectral resolution is 4cm -1 The number of scans was 32 to obtain the average analysis results and improve the signal-to-noise ratio. The spectrum consists of absorbance and wavenumber, with absorbance on the vertical axis and wavenumber on the horizontal axis.

[0074] For example, the ATR-FTIR spectra of Fructus Aurantii Immaturus samples from three different places (Yiyang, Huaihua, and Changde) were measured. The ATR-FTIR spectra obtained were as follows: Figure 2 shown.

[0075] S2. Averaging multiple ATR-FTIR spectra to obtain an average ATR-FTIR spectrum.

[0076] Specifically, multiple scans are averaged to avoid inaccurate average ATR-FTIR spectra due to errors or large errors in the data from a single scan. The average absorbance values ​​at the same wavelength are averaged to obtain the average ATR-FTIR spectrum.

[0077] S3. Decompose the average ATR-FTIR spectrum to obtain a scale signal.

[0078] The average ATR-FTIR spectrum is decomposed to obtain scale signals including:

[0079] S31, obtaining a preset wavelet basis;

[0080] S32. Obtaining a corresponding scaling function according to a preset wavelet basis;

[0081] S33. Decompose the average ATR-FTIR spectrum according to the scaling function to obtain a scaling signal.

[0082] Specifically, before processing the averaged ATR-FTIR spectrum, the signal must first be preprocessed, such as denoising and smoothing, to ensure accurate analysis. The next step is to select an appropriate wavelet basis, including Haar wavelets, Daubechies wavelets, and Symlets. Each wavelet has its own specific scaling and waveform functions. Using the wavelet transform, the signal is decomposed to obtain representations at different scales.

[0083] The scaling function reflects the low-frequency components of the signal at that scale.

[0084] According to the selected wavelet basis, the corresponding scaling filter and initial function are determined. Usually, the unit pulse function can be selected, and the scaling function can be obtained by continuous iterative calculation.

[0085] The standard signal representation is:

[0086]

[0087] S4. Scale and translate the scale signal to obtain a wavelet signal.

[0088] The scale signal is scaled and translated to obtain the wavelet signal including:

[0089] S41. According to the recursive relationship of the scaling function, a wavelet function is obtained;

[0090] For the selected wavelet basis, a low-pass filter and a high-pass filter are selected, and the recursive relationship is used to start from the basic scaling function (such as the unit impulse function) to calculate the scaling functions of different scales. The recursive relationship of the scaling function is used in combination with the high-pass filter to generate the corresponding wavelet function.

[0091] S42. Perform scaling processing on the scale signal using a scaling function, and perform translation processing on the scale signal using a wavelet function to obtain scale signals at different scales.

[0092] Scaling function After scaling and translation transformation, we can get scale signals at different scales j Defined as:

[0093]

[0094] S43, expressing scale signals at different scales linearly to obtain a signal space composed of linearly expressed signals;

[0095] Specifically, the scale signal The signal space composed of linearly expressed signals is V j :

[0096]

[0097] S44. Obtaining a spanned signal space according to the signal space;

[0098] Specifically, the scaling function at high resolution can contain more information, that is, the signal space spanned by the high-resolution scaling signal contains the signal space spanned by the low-resolution scaling signal, which can be expressed as:

[0099]

[0100] S45. Obtain a recursive signal according to the spanned signal space;

[0101] Depend on Get the recursive signal:

[0102]

[0103] Where h0[k] is the scaling function coefficient.

[0104] S46. Perform scaling and translation transformation on the recursive signal through a wavelet function to obtain a wavelet signal.

[0105] Specifically, according to the square integrable space L 2 (R), by the scaling function The corresponding wavelet function ψ(t) can be defined. The wavelet function is defined in accordance with the scaling function, and the space W spanned by the wavelet function is j-1 Vj and V j-1 The difference of .

[0106] Wavelet signal ψ j,l (t) is expressed as:

[0107] ψ j,l (t) = 2 j / 2 ψ(2t-l).

[0108] Reconstruct the wavelet signal, retain the large trend signal, and obtain the smoothed average ATR-FTIR spectrum for training, including:

[0109] Design the wavelet signal as an orthogonal signal of the scale signal;

[0110] Obtain the spanned signal space of the wavelet signal;

[0111] The square integrable space of wavelet signal is obtained according to the spanned signal space of wavelet signal;

[0112] According to the scale signal, wavelet signal and square integrable space, the time domain signal is obtained;

[0113] The time domain signal is decomposed and reconstructed to obtain a decomposition signal set, and a large trend signal is selected from the decomposition signal set to obtain a smoothed average ATR-FTIR spectrum for training.

[0114] Specifically, the wavelet signal ψ j,l (t) The spanned signal space is denoted as W j ,but:

[0115]

[0116] Specifically, the square integrable space of wavelet signals is:

[0117] The time domain signal x(t) is a square integrable space L 2 (R), so the signal x(t) can be represented by the scaled signal and the wavelet signal ψ j,l (t) Expression:

[0118]

[0119] scale signal Expressing rough information, wavelet signal ψ j,l (t) Express detailed information.

[0120] S5, reconstructing the wavelet signal, retaining the large trend signal, and obtaining the smoothed average ATR-FTIR spectrum for training;

[0121] The time domain signal is decomposed and reconstructed to obtain a decomposition signal set. A large trend signal is selected from the decomposition signal set to obtain a smoothed average ATR-FTIR spectrum for training, including:

[0122] The wavelet signal is linearly represented by the scaling function to obtain a linear wavelet signal;

[0123] The linear wavelet signal is deformed by recursive equation through scaling function and wavelet function to obtain the decomposed signal set;

[0124] The large trend signal is selected from the decomposed signal set to obtain the smoothed average ATR-FTIR spectrum for training.

[0125] To obtain and d j,k , the recursive equation of the scaling function and wavelet function is transformed:

[0126]

[0127] So:

[0128]

[0129] Right now:

[0130]

[0131] d j,l =∑ n h1[2n-l]·d j+1,n .

[0132] The decomposed signal set is:

[0133]

[0134] Considering that the purpose of analysis is to smooth the spectrum, only the signal containing the general trend should be retained, that is, the reconstruction of the signal only retains item.

[0135] The orthogonal wavelet function family includes sym, coif and db, and the decomposition level is limited to 1-12. The signal x(t) is decomposed and reconstructed to retain only Item, about to The item is output as a smoothing result for subsequent analysis and processing.

[0136] Multiple ATR-FTIR spectra of Fructus Aurantii Immaturus were obtained by multiple measurements. Figure 3 As shown. Take the average of multiple ATR-FTIR spectra to obtain the average ATR-FTIR spectrum, and use different processing methods to process the average ATR-FTIR spectrum to obtain the smooth average ATR-FTIR spectrum for training. The results are shown in Figure 4 shown.

[0137] Different processing methods yielded different results. The smoothing effect evaluation indicators for different processing methods are shown in Table 1. In Table 1, the smoothing degree and similarity of each region were equal under the same smoothing treatment, indicating that factors such as environmental noise had a consistent impact on each sample during infrared spectrum scanning, and no abnormalities occurred during the scanning process.

[0138] Table 1 Smoothing effects of four smoothing methods

[0139]

[0140] From Table 1 and Figure 4It can be seen that the treatment method of the present invention produces the best smoothed average ATR-FTIR spectrum, with the highest smoothness of 96.64%, a similarity of 98.02%, and an average of 97.83%. No bands with large fluctuations appear in the spectrum, resulting in the best overall smoothing effect. The smoothing degree for the MSC and SNV treatments is 0, while the smoothing degree for the SG treatment is 0.48%, indicating that these three methods do not effectively mitigate the impact of regional factors, and some absorption peaks with large fluctuations still exist in the smoothed spectrum.

[0141] S6, extracting characteristic peaks of the training smoothed average ATR-FTIR spectrum to obtain a characteristic vector of the training smoothed average ATR-FTIR spectrum; specifically comprising the following steps:

[0142] The first-order derivative of the training smoothed average ATR-FTIR spectrum is taken to obtain all zero-point positions, and the peak values ​​of all zero-point positions are obtained to obtain the characteristic vector of the training smoothed average ATR-FTIR spectrum.

[0143] Specifically, the first derivative M1 of the smoothed average ATR-FTIR spectrum is calculated, and all zero-point positions a0 of M1 are obtained, where a0 is also the peak position of the average smoothed spectrum M. The peak value x of each smoothed spectrum at a0 is obtained, where x is the characteristic vector of each spectrum.

[0144] Set the category label y2, Fructus Aurantii corresponds to 1, Fructus Aurantii Immaturus corresponds to 2, and the region label y1, Changde corresponds to 1, Huaihua corresponds to 2, and Yiyang corresponds to 3. Sample set {x i ,y i}, i=1,2,......,30, where x i is the eigenvector, y i For the region label.

[0145] S7, inputting the feature vector of the smoothed average ATR-FTIR spectrum used for training into the SVM classifier, and training to obtain an optimized SVM classifier; specifically comprising the following steps:

[0146] S71. Obtain a preset kernel function and decision function according to the SVM classifier;

[0147] S72, establishing an objective function according to the kernel function;

[0148] S73, training an SVM classifier according to the characteristic vector of the training smoothed average ATR-FTIR spectrum and the objective function to obtain a first parameter, a second parameter, and a third parameter of the objective function;

[0149] S74, obtaining an adjustment decision function according to the first parameter, the second parameter, the third parameter, and the decision function;

[0150] S75. Replace the decision function in the SVM classifier with the adjusted decision function to obtain an optimized SVM classifier.

[0151] Specifically, the kernel function used in this application is

[0152] The decision function is: f(x) = sgn(w T K(x i ,x)+b.

[0153] The objective function and constraints are sty i (w T K(x i ,x)+b)≥1-ξ i ,ξ i ≥0.

[0154] The SVM classifier is trained according to the characteristic vector of the smoothed average ATR-FTIR spectrum used for training and the objective function, and the first parameter, second parameter and third parameter of the objective function are obtained as follows:

[0155] According to the characteristic vector of the smoothed average ATR-FTIR spectrum used for training and the objective function, the constraint function of the objective function is obtained;

[0156] Convert the constraint function into a Lagrangian function;

[0157] Set the gradient of the first and second parameters in the Lagrangian function to 0, and calculate the third parameter;

[0158] Substitute the third parameter into the constraint function and calculate the first and second parameters through the KKT condition.

[0159] Specifically, the constraint function of the objective function is converted into a Lagrangian function:

[0160]

[0161] where λ i ,r i ,ξ i ≥0,λ i =Cr i .

[0162] The objective function and constraints are equivalent to:

[0163]

[0164] Specifically, the first parameter is w, the second parameter is b, and the third parameter is λ. Set the gradient of L at w and b to 0, and substitute the simplified formula to get:

[0165]

[0166] Where C is the box constraint. Solve the quadratic programming problem and use the SMO algorithm to get the optimal solution

[0167] Then according to the KKT condition, we can get:

[0168] Solved by the KKT condition:

[0169]

[0170] Specifically, the calculated parameters are substituted into the decision function to obtain the adjusted decision function:

[0171]

[0172] The decision function in the SVM classifier is replaced with an adjusted decision function to obtain an optimized SVM classifier. When the training samples of the Citrus aurantium and Citrus aurantium category are input into the adjusted decision function, the adjusted decision function outputs 1 and 2, corresponding to Citrus aurantium and Citrus aurantium, respectively. When the training sample input is the Citrus aurantium and Citrus aurantium origin training sample, the parameters of the decision function are adjusted to be different from the parameters of the adjusted decision function of the category training sample. At this time, the adjusted decision function outputs the corresponding value according to the input value, for example, 1 for the Changde region, 2 for the Huaihua region, and 3 for the Yiyang region. The output of the adjusted decision function is 1, 2, and 3.

[0173] Optimization parameters for feature extraction include the wavelet function and the number of decomposition levels; SVM optimal point estimation model parameters include the corresponding mode (Coding), box constraint (BoxConstraint), scaling parameter (KernelScale), kernel function (KernelFunction), polynomial order (PolynomialOrder), and whether to standardize the data (Standardize). Different wavelet functions and decomposition levels produce different smoothing results and waveform differences. Smoothed signals that retain more detail have more peaks and correspondingly more extracted eigenvalues, indicating a higher dimensionality. When the number of features in the SVM classifier is too large, linear and Gaussian kernels are suitable. When the number of samples is greater than the number of features, a polynomial kernel is suitable.

[0174] S8, inputting the feature vector of the smoothed average ATR-FTIR spectrum of the sample to be tested into the optimized SVM classifier to obtain the category and origin of the sample to be tested;

[0175] Specifically, after the model training is completed, you only need to input the feature vector of the smoothed average ATR-FTIR spectrum of the sample to be tested into the optimized SVM classifier to automatically output the result, that is, the category and origin of the sample to be tested. Depending on the different features input during training, the results output by the SVM classifier are different. For example, if only the regional features are trained, only the regional results will be output. If the region and category are trained at the same time, the category and origin of the sample to be tested will be output at the same time.

[0176] The optimized SVM classifier of the present invention can process high-dimensional data and achieve classification. Using the optimized SVM classifier of the present invention to classify and identify the origin of Fructus Aurantii Immaturus, the identification accuracy reached 96.7%, significantly higher than the 80.2% achieved by the orthogonal partial least squares-discriminant analysis method. The identification accuracy was calculated as the number of correctly identified samples divided by the total number of identified samples.

[0177] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of the present application is limited to these examples. In line with the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as above, which are not provided in detail for the sake of simplicity.

[0178] The one or more embodiments of this application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this application should be included in the scope of protection of this application.

Claims

1. A method for identifying the type and origin of Citrus aurantium and Citrus aurantium fructus, characterized by: The following steps are included: ATR-FTIR spectra of training samples are collected multiple times to obtain multiple ATR-FTIR spectra, wherein the training samples are Fructus Aurantii Immaturus or Fructus Aurantii Immaturus from different origins; Averaging multiple ATR-FTIR spectra to obtain an average ATR-FTIR spectrum; Decomposing the average ATR-FTIR spectrum to obtain a scale signal; Scaling and translating the scale signal to obtain a wavelet signal; Reconstructing the wavelet signal, retaining the large trend signal, and obtaining a smoothed average ATR-FTIR spectrum for training; Extracting characteristic peaks of the training smoothed average ATR-FTIR spectrum to obtain a characteristic vector of the training smoothed average ATR-FTIR spectrum; Inputting the feature vector of the training smoothed average ATR-FTIR spectrum into the SVM classifier, and training to obtain an optimized SVM classifier; The feature vector of the smoothed average ATR-FTIR spectrum of the sample to be tested is input into the optimized SVM classifier to obtain the category and origin of the sample to be tested; The reconstructing the wavelet signal, retaining the large trend signal, and obtaining a smoothed average ATR-FTIR spectrum for training comprises: Designing the wavelet signal as an orthogonal signal of the scale signal; Obtain the spanned signal space of the wavelet signal; Obtaining a square integrable space of the wavelet signal according to the spanned signal space of the wavelet signal; Obtaining a time domain signal according to the scale signal, the wavelet signal, and the square integrable space; Decomposing and reconstructing the time domain signal to obtain a decomposition signal set, selecting a large trend signal from the decomposition signal set, and obtaining a smoothed average ATR-FTIR spectrum for training; Decomposing the average ATR-FTIR spectrum to obtain a scale signal includes: Get the preset wavelet basis; According to the preset wavelet basis, a corresponding scaling function is obtained; Decomposing the average ATR-FTIR spectrum according to the scaling function to obtain a scaling signal; Decomposing and reconstructing the time domain signal to obtain a decomposed signal set, and selecting a large trend signal from the decomposed signal set to obtain a smoothed average ATR-FTIR spectrum for training includes: Linearly representing the wavelet signal using the scaling function to obtain a linear wavelet signal; Performing recursive equation deformation on the linear wavelet signal by using the scaling function and the wavelet function to obtain a decomposed signal set; Selecting a large trend signal from the decomposition signal set to obtain a smoothed average ATR-FTIR spectrum for training; The inputting the feature vector of the training smoothed average ATR-FTIR spectrum into the SVM classifier, and training the optimized SVM classifier comprises: According to the SVM classifier, a preset kernel function and a decision function are obtained; Establishing an objective function according to the kernel function; Training an SVM classifier according to the characteristic vector of the training smoothed average ATR-FTIR spectrum and the objective function to obtain a first parameter, a second parameter, and a third parameter of the objective function; Obtaining an adjustment decision function according to the first parameter, the second parameter, the third parameter and the decision function; The decision function in the SVM classifier is replaced by the adjusted decision function to obtain the optimized SVM classifier; The SVM classifier is trained according to the characteristic vector of the training smoothed average ATR-FTIR spectrum and the objective function to obtain the first parameter, the second parameter and the third parameter of the objective function, which include: Obtaining a constraint function of the objective function according to the characteristic vector of the training smoothed average ATR-FTIR spectrum and the objective function; Converting the constraint function into a Lagrangian function; Setting the gradients of the first and second parameters in the Lagrangian function to 0, and calculating a third parameter; The third parameter is substituted into the constraint function, and the first parameter and the second parameter are obtained by KKT condition calculation.

2. The identification method according to claim 1, wherein: The training sample is prepared by removing impurities from the immature bitter orange and the fructus aurantii, splitting them, drying them, crushing them and screening them to obtain the training sample.

3. The identification method according to claim 1, wherein: In the ATR-FTIR spectrum, the absorbance acquisition range is 4000~600 cm -1 The resolution of ATR-FTIR spectrum is 4 cm -1 .

4. The identification method according to claim 1, wherein: The scaling and translating the scaled signal to obtain a wavelet signal includes: According to the recursive relationship of the scaling function, the wavelet function is obtained; Performing scaling processing on the scale signal through a scaling function and performing translation processing on the scale signal through a wavelet function to obtain scale signals at different scales; The scale signals at different scales are expressed linearly to obtain a signal space composed of linear expression signals; According to the signal space, a spanned signal space is obtained; Obtaining a recursive signal according to the spanned signal space; The recursive signal is subjected to scaling and translation transformation through a wavelet function to obtain a wavelet signal.

5. The identification method according to any one of claims 1 to 4, characterized in that: The step of extracting characteristic peaks of the training smoothed average ATR-FTIR spectrum to obtain a characteristic vector of the training smoothed average ATR-FTIR spectrum includes: Performing a first-order derivative of the training smoothed average ATR-FTIR spectrum to obtain all zero-point positions, obtaining peak values ​​at all zero-point positions, and obtaining a characteristic vector of the training smoothed average ATR-FTIR spectrum.

Citation Information

Patent Citations

  • A method for identifying green tangerine peel, dried tangerine peel, immature bitter orange, and immature bitter orange peel based on chemical classification and UPLC-Tof-MS.

    CN107389813B

  • Spectral analysis type high-precision on-line detector for gas and liquid detection

    CN104062264A

  • General rapid method for rapidly identifying whether natural products are produced in specific places or not

    CN117874609A