Terahertz spectrum-based sparse representation classification method, system and equipment for saccharide analysis and medium

The dimensionality reduction of the terahertz spectral data is solved by sparse representation classification method, which solves the problem of large amount of calculation and long time in terahertz time domain spectroscopy analysis, and achieves fast and accurate carbohydrate detection.

CN120404650APending Publication Date: 2025-08-01HENAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510634488.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-07-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing terahertz time domain spectroscopy requires a lot of calculations for qualitative analysis of sugars, and the model classification time is long and complex, resulting in large and long calculations.

Method used

The sparse representation classification method is adopted, including fast Fourier transform filtering and sparse representation algorithm, the sparse coefficient matrix is solved through orthogonal matching tracking OMP algorithm, and the dictionary matrix is updated using the K-SVD algorithm to reduce the dimensionality of terahertz spectral data.

Benefits of technology

It realizes rapid and accurate analysis of terahertz spectral data, reduces the computing time of the model, improves detection accuracy and data simplicity, and enriches terahertz detection theory and methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120404650A_ABST
    Figure CN120404650A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse representation classification method, system, equipment and medium for saccharide analysis based on terahertz spectrum, and belongs to the technical field of terahertz spectrum. Comprising the following steps: selecting three saccharides including sucrose, fructose and lactose as samples, and obtaining terahertz THz absorption spectrum characteristics of the saccharide samples; carrying out filtering pretreatment on the THz absorption spectrum data of the saccharides; and carrying out sparse representation on the preprocessed THz absorption spectrum data of the saccharides, namely constructing a sparse representation algorithm model, solving a sparse coefficient matrix of the THz absorption spectrum by adopting an orthogonal matching pursuit (OMP) algorithm, updating a dictionary matrix by adopting a K-SVD algorithm, and carrying out cyclic calculation to obtain sparse coefficient matrixes of the absorption spectrums of the three samples when iteration conditions are met. According to the method, saccharides are taken as research objects, the sparse coefficient matrix is taken as input of the classification model to obtain time for processing data by the model, and compared with other dimensionality reduction modes, the operation time is greatly shortened, the effectiveness of the model efficiency is improved, and sparse high-quality classification of saccharides spectrums is realized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The original application has the invention title of "Sparse Representation Classification Method for Carbohydrate Analysis Based on Terahertz Spectroscopy", with the application number 202110868596.7 and the filing date of July 29, 2021. Technical Field

[0002] The present invention relates to a method for detecting carbohydrates using terahertz spectroscopy, specifically to a sparse representation classification method for carbohydrate analysis based on terahertz spectroscopy, belonging to the technical field of terahertz spectroscopy. Background Art

[0003] Terahertz wave (THz) is an electromagnetic wave with a frequency between 0.1 THz and 10 THz and a wavelength between 3 mm and 0.03 mm. The position of THz wave is unique, having some characteristics of both microwave and infrared. It has good penetrability, safety, and spectral resolution, and has molecular fingerprint characteristics, low energy, and strong penetrability. THz technology has been applied in many fields, such as aerospace, biomedicine, security detection, material property analysis, etc.; currently, its applications in the fields of agricultural products and food safety are also increasing, such as the detection of water content in agricultural products and food, the detection of internal quality of food, and seed identification.

[0004] Carbohydrates are important components in grains. The carbohydrates contained in grains mainly include glucose, sucrose, lactose, etc. Changes in carbohydrates during grain storage will cause changes in grain quality. And carbohydrates have obvious absorption peaks in the terahertz spectrum. Using sucrose, fructose, and lactose as carbohydrate samples for qualitative analysis of carbohydrates can lay a certain foundation for the detection of stored grain quality.

[0005] Regarding the qualitative analysis of carbohydrates by terahertz time-domain spectroscopy, the literature "Study on the Qualitative and Quantitative Analysis of D-Glucose Anhydrous Based on Terahertz Spectroscopy Technology, Spectroscopy and Spectral Analysis, 2017, 37(07): 2165 - 2170" uses methods of multiple linear regression and partial least squares regression for qualitative analysis. Since THz spectral data not only contains linear relationships but also non-linear relationships, has a high spectral dimension, and there is a large amount of redundant information, directly using it as the input of the qualitative analysis model will affect the prediction accuracy of the model; at the same time, a large amount of data will increase the computational complexity of the model, resulting in large computational volume and long time consumption. When the existing algorithms process THz spectral data, they often require a large amount of calculation, and the model classification time is long and the complexity is high. Therefore, spectral dimensionality reduction has become an indispensable part of spectral processing. Summary of the Invention

[0006] The object of the present invention is to solve the problems that the existing algorithms for qualitative analysis of sugars by terahertz time-domain spectroscopy require a large amount of calculation when processing data, and the model classification time is long and the complexity is high. A sparse representation classification method for sugar analysis based on terahertz spectroscopy is provided, introducing sparse representation into the terahertz spectroscopy field to perform dimensionality reduction processing on data, making the terahertz spectroscopy more rapid and accurate in modeling analysis, and representing the data as concisely as possible.

[0007] To achieve the above object, the present invention adopts the following technical solutions: A sparse representation classification method for sugar analysis based on terahertz spectroscopy, comprising the following steps:

[0008] S1. Obtain the THz absorption spectral characteristics of sugar samples: Select three sugars, namely sucrose, fructose and lactose, as samples. First, use a terahertz time-domain spectroscopy system THz-TDS-spectrometer to measure a time-domain spectrum as a reference spectrum, then measure each type of sample multiple times to calculate the average time-domain spectrum, and use the time-domain spectrum to calculate the frequency-domain spectrum. Finally, calculate the absorption spectrum of each type of sample by comparing with the reference spectrum;

[0009] The reference spectrum is obtained from the time-domain spectral signal directly measured by the THz-TDS system, and the defined formula is f r (t), where t is the frequency and r is the subscript of f(t);

[0010] The average time-domain spectrum is obtained by taking the average of the time-domain spectral signals of each type of sample measured repeatedly, and the defined formula is f s (t), where t is the frequency and s is the subscript of f(t);

[0011] The frequency-domain spectrum is obtained by performing Fourier transform on the average time-domain spectrum, and the transformation calculation formula is: where F(ω) is the frequency-domain spectral signal and f(t) is the time-domain spectral signal;

[0012] The absorption spectrum is obtained by comparing the frequency-domain spectrum of each type of sample with the frequency-domain spectrum of the reference spectrum, and the calculation formula is: where α is the absorption spectrum, d is the thickness of the sample tablet, d is the amplitude of the frequency-domain spectrum of the reference spectrum, and A s is the amplitude of the frequency-domain spectrum of the sugar sample.

[0013] S2. Preprocess the THz absorption spectral data of sugars: Select fast Fourier transform FFT filtering to preprocess the THz absorption spectral data of each type of sample calculated in S1, remove the noise signals with frequencies higher than the cut-off frequency in the THz spectral data, obtain a data set of the filtered THz absorption spectral data, and obtain a smoother waveform of the THz absorption spectral data of each type of sample;

[0014] The specific steps of filtering using the fast Fourier transform (FFT) are as follows:

[0015] S21: Calculate the average values of the first 1% and the last 1% of the data points for the THz absorption spectrum data of each type of sample calculated in S1;

[0016] S22: Construct a baseline through the two average value points calculated in S21, and subtract all the THz absorption spectrum data of each type of sample with this baseline;

[0017] S23: Perform FFT transformation on the data difference obtained in S22;

[0018] S24: Perform low-pass filtering on the THz absorption spectrum data obtained after the transformation in S23 to remove the Fourier components with frequencies higher than the cut-off frequency. The cut-off frequency is: where n is the specified number of window points, n is set to 15, and t is the time interval between two adjacent data points;

[0019] S25: Perform inverse fast Fourier transform (IFFT) calculation on the filtered spectrum, and obtain a data set of the THz absorption spectrum data of each type of sample after sorting;

[0020] S26: Add the baseline constructed in S22 to the data set obtained in S25.

[0021] S3. Perform sparse representation on the preprocessed THz absorption spectrum data of sugars. The specific steps are as follows:

[0022] S31. Use the data set of the preprocessed THz absorption spectrum data of sugars to construct a sparse representation algorithm model: Assume that the data set Y of the preprocessed THz absorption spectrum data of a sugar sample is represented as the original sample matrix Y ∈ R m×n , where the original sugar sample absorption spectrum signal y is the THz absorption spectrum data of one type of sugar sample in the data set Y. Then the sparse representation model is:

[0023]

[0024] where D is an over-complete dictionary matrix, d i is an atom, D = {d1, d2,... d K}, x is the spectral sparse representation coefficient and has only a finite number of non-zero elements, x = {x1, x2,... x n};

[0025] Solve to obtain: where t is the upper limit of the number of non-zero components in the spectral sparse representation coefficient;

[0026] S32. Use the orthogonal matching pursuit (OMP) algorithm to solve the spectral sparse coefficient matrix:

[0027] Set the sparsity k to obtain the number of iterations. By inputting the dictionary matrix D, the absorption spectrum signal y of the original sugar sample, and the sparsity k, the sparse coefficient x is solved. The specific algorithm steps are as follows:

[0028] Initialization: Initialize the dictionary matrix D and the sparsity k. The residual R0 = y, the index set Λ = O, and t = 1;

[0029] a1: Find the atom in the residual R that best matches the dictionary matrix D, that is, find a certain column d with the largest inner product in the dictionary matrix D i , whose corresponding subscript is λ, λ t = argmax j=1,...N |<R t-1 , d j >|;

[0030] a2: Update the index set Λ t = Λ t-1 ∪{λ t} and the atom in the dictionary matrix D

[0031] a3: Use the least squares method to substitute into the solution formula in step S31 to calculate and obtain the sparse coefficient x of the absorption spectrum signal of the sugar sample t = argmax||y - D t x t ||;

[0032] a4: Update the residual R of the absorption spectrum signal of the sugar sample t = y - D t x t , t = t + 1;

[0033] a5: Judge whether t is greater than k. If t > k, the algorithm stops and the sparse coefficient matrix of the absorption spectrum of the sugar sample is obtained; if t ≤ t + 1, then jump back to step a1 to execute the algorithm again.

[0034] S33: Update the dictionary matrix using the K - SVD algorithm:

[0035] By inputting the dictionary matrix D, the absorption spectrum signal y of the original sugar sample, and the sparse coefficient matrix calculated in step S32, the updated dictionary matrix D is obtained. The specific algorithm steps are as follows:

[0036] Initialization: Randomly select K column vectors from the absorption spectrum signal y of the original sugar sample or take the first K vectors from the left singular matrix of the original sample matrix Y as the atoms of the initial dictionary. Initialize the dictionary D0 ∈ R m×K , j = 0, to obtain the initial dictionary D j ;

[0037] b1: Use the obtained dictionary D j to perform spectral sparse coding and obtain the sparse coefficient matrix X of the absorption spectra of sugar samples j ∈R K×n ;

[0038] b2: Update the dictionary matrix D column by column j , for the column d of the dictionary k ∈{d1, d2,..., d K}, calculate the error matrix

[0039] b3: Take the k-th non-zero row vector of the sparse matrix X of the absorption spectra of sugar samples j and its set of indices ;

[0040] b4: Take out the columns corresponding to the non-zero indices ω k from E k to obtain E' k ;

[0041] b5: Perform singular value decomposition on E' k E k = UΔV T , take the first column of U to update the k-th column of the dictionary, d k = U(·, 1), and let Update to the original position, j = j + 1;

[0042] b6: Determine whether the specified number of iteration steps and tolerance error are reached. If so, stop the algorithm and obtain the sparse coefficient matrix X of the absorption spectrum signals of three sugar samples; if not, continue to enter the OMP algorithm for loop calculation until the iteration condition is satisfied.

[0043] The beneficial effects of the present invention are as follows:

[0044] 1) The method of the present invention takes sugars as the research object. Aiming at problems such as large amount of THz spectral data, high dimensionality, complex structure, and long model calculation time, sparse representation is used to reduce the dimensionality of the preprocessed THz spectral data; making the THz spectrum more rapid and accurate in modeling analysis, and representing the data as concise as possible.

[0045] 2) In the method of the present invention, the OMP algorithm is selected to calculate the sparse coefficients, and K-SVD is used to update the dictionary to achieve the sparse representation of terahertz spectral data. Then, it is compared with traditional dimensionality reduction methods such as principal component analysis, linear discriminant analysis, locally linear embedding, and isometric mapping, and an SVM classification model is established. Compared with other dimensionality reduction methods, the operation time of the constructed terahertz sparse representation model is greatly reduced. The experimental results verify the effectiveness of the sparse representation non-linear dimensionality reduction method in improving the model efficiency, enhance the quality of terahertz spectral data and the detection accuracy of the model, and enrich and develop the terahertz detection theory and methods. Description of the Drawings

[0046] Figure 1 It is the THz absorption spectrogram of three kinds of sugar samples;

[0047] Figure 2 It is the smoothed effect diagram after FFT filtering processing of the THz absorption spectra of three kinds of sugar samples;

[0048] Figure 3 It is the algorithm diagram of sparse representation;

[0049] Figure 4 It is the algorithm flowchart of the sparse representation of THz absorption spectra;

[0050] Figure 5 It is the comparison result table of sparse representation and common dimensionality reduction methods;

[0051] Figure 6 It is the bar chart of the operation time of the THz spectral sparse representation model and the common dimensionality reduction method models. Detailed Embodiments

[0052] The present invention will be further explained below with reference to the drawings and specific embodiments.

[0053] Embodiment: As Figure 1-6 shown, the sparse representation classification method based on terahertz spectroscopy for sugar analysis of the present invention includes the following steps:

[0054] S1. Obtain the THz absorption spectral characteristics of sugar samples:

[0055] Select three kinds of sugars, namely sucrose, fructose, and lactose, as samples. First, use a terahertz time-domain spectroscopy system THz-TDS-spectrometer to measure a time-domain spectrum as a reference spectrum, then measure each type of sample multiple times to calculate the average time-domain spectrum, and use the time-domain spectrum to calculate the frequency-domain spectrum. Finally, calculate the absorption spectrum of each type of sample by comparing with the reference spectrum. As Figure 1 shows the absorption spectrograms of different sugar samples.

[0056] When using a time-domain spectrometer to measure time-domain spectral data, the same sample is continuously measured multiple times and then averaged to obtain an average time-domain spectrum, which can reduce systematic errors during the experiment. In actual experiments, it is generally measured 10 times and averaged.

[0057] The reference spectrum is obtained from the time-domain spectral signal directly measured by the THz-TDS system, and the defined formula is f r (t), where t is the frequency and r is the subscript of f(t);

[0058] The average time-domain spectrum is obtained by averaging the time-domain spectral signals of each type of sample measured repeatedly, and the defined formula is f s (t), where t is the frequency and s is the subscript of f(t);

[0059] The frequency-domain spectrum is obtained by performing a Fourier transform on the average time-domain spectrum, and the calculation formula for the transform is: where F(ω) is the frequency-domain spectral signal and f(t) is the time-domain spectral signal;

[0060] The absorption spectrum is obtained by comparing the frequency-domain spectra of each type of sample with the frequency-domain spectrum of the reference spectrum, and the calculation formula is: where α is the absorption spectrum, d is the thickness of the sample tablet, A r is the amplitude of the frequency-domain spectrum of the reference spectrum, and A s is the amplitude of the frequency-domain spectrum of the sugar sample.

[0061] S2. Preprocess the THz absorption spectral data of sugars:

[0062] Select fast Fourier transform (FFT) filtering to preprocess the THz absorption spectral data of each type of sample calculated in S1, remove the noise signals with frequencies higher than the cut-off frequency in the THz spectral data, and obtain a data set of the filtered THz absorption spectral data, making the waveform of the THz absorption spectral data of each type of sample smoother; the processing results are as Figure 2 shown.

[0063] The specific steps of using FFT filtering are as follows:

[0064] S21: For the THz absorption spectral data of each type of sample calculated in S1, calculate the average values of the first 1% and the last 1% of the data points;

[0065] S22: Construct a baseline through the two average value points calculated in S21, and subtract this baseline from all the THz absorption spectral data of each type of sample;

[0066] S23: Perform an FFT transform on the data difference obtained in S22;

[0067] S24: Perform low-pass filtering on the THz absorption spectrum data obtained after transformation in S23 to remove Fourier components with frequencies higher than the cut-off frequency, where the cut-off frequency is: where n is the specified number of window points, set to 15, and t is the time interval between two adjacent data points;

[0068] S25: Perform inverse fast Fourier transform IFFT calculation on the filtered spectrum, and obtain a dataset of THz absorption spectrum data for each type of sample after arrangement;

[0069] S26: Add the baseline constructed in S22 to the dataset obtained in S25.

[0070] S3. Perform sparse representation on the preprocessed THz absorption spectrum data of sugars, and the flowchart is as Figure 4 shown. The specific steps are as follows:

[0071] S31. Construct a sparse representation algorithm model using the dataset of preprocessed THz absorption spectrum data of sugars:

[0072] Assume that the dataset Y of preprocessed THz absorption spectrum data of a sugar sample is represented as the original sample matrix Y ∈ R m×n , where the original absorption spectrum signal y of the sugar sample is the THz absorption spectrum data of one type of sugar sample in the dataset Y. Then the sparse representation model is:

[0073]

[0074] where D is an over-complete dictionary matrix, d i is an atom, D = {d1, d2,... d K} and x is the spectral sparse representation coefficient with only a finite number of non-zero elements, x = {x1, x2,... x n};

[0075] Solve to obtain: where t is the upper limit of the number of non-zero components in the spectral sparse representation coefficient; the algorithm diagram of the sparse representation is as Figure 3 shown.

[0076] S32. Use the orthogonal matching pursuit OMP algorithm to solve the spectral sparse coefficient matrix:

[0077] Set the sparsity k to obtain the number of iterations. By inputting the dictionary matrix D, the original absorption spectrum signal y of the sugar sample, and the sparsity k, solve to obtain the sparse coefficient x. The specific algorithm steps are as follows:

[0078] Initialization: Initialize the dictionary matrix D and the sparsity k, the residual R0 = y, the index set Λ = O, and t = 1;

[0079] a1: Find the atom in the dictionary matrix D that best matches the residual R, i.e., find the column d in the dictionary matrix D with the largest inner product. i , and its corresponding subscript is λ, λ t = argmax j=1,...N |<R t-1 , d j >|;

[0080] a2: Update the index set Λ t = Λ t-1 ∪{λ t} and the atom in the dictionary matrix D

[0081] a3: Use the least squares method to substitute into the solution formula in step S31 for calculation to obtain the sparse coefficient x of the absorption spectrum signal of the sugar sample t = argmax||y - D t x t ||;

[0082] a4: Update the residual R of the absorption spectrum signal of the sugar sample t = y - D t x t , t = t + 1;

[0083] a5: Determine whether t is greater than k. If t > k, the algorithm stops and the sparse coefficient matrix of the absorption spectrum of the sugar sample is obtained; if t ≤ t + 1, then jump back to step a1 to execute the algorithm again.

[0084] S33: Update the dictionary matrix using the K - SVD algorithm:

[0085] By inputting the dictionary matrix D, the original absorption spectrum signal y of the sugar sample, and the sparse coefficient matrix calculated in step S32, an updated dictionary matrix D is obtained. The specific algorithm steps are as follows:

[0086] Initialization: Randomly select K column vectors from the original absorption spectrum signal y of the sugar sample or take the first K vectors from the left singular matrix of the original sample matrix Y as the atoms of the initial dictionary. Initialize the dictionary D0 ∈ R m×K , j = 0, to obtain the initial dictionary D j ;

[0087] b1: Use the obtained dictionary D j to perform spectral sparse coding to obtain the sparse coefficient matrix x of the absorption spectrum of the sugar sample j ∈ R K×n ;

[0088] b2: Update the dictionary matrix D column by column j , and the column d of the dictionary k ∈ {d1, d2,..., dK}, calculate the error matrix

[0089] b3: Take the sparse matrix X of the absorption spectrum of the saccharide sample j The k-th non-zero row vector The set of indices

[0090] b4: Take out the corresponding indices ω k from E k to obtain the non-zero columns, getting E' k ;

[0091] b5: Perform singular value decomposition on E' k to get E k = UΔV T , take the first column of U to update the k-th column of the dictionary, d k = U(·, 1), let Substitute into the original position, j = j + 1;

[0092] b6: Determine whether the specified number of iteration steps and tolerance error are reached. If so, stop the algorithm to obtain the sparse coefficient matrix X of the absorption spectrum signals of the three saccharide samples; if not, continue to enter the OMP algorithm for loop calculation until the iteration condition is satisfied.

[0093] Substitute the sparse coefficient matrix calculated by the OMP algorithm into the K-SVD algorithm, use the sparse coefficient matrix to update the dictionary matrix, and substitute the updated dictionary matrix D back into the orthogonal matching pursuit OMP algorithm to calculate the sparse coefficient matrix until the set number of iterations is reached and the calculation stops, finally obtaining the sparse coefficient matrix X calculated in the last time.

[0094] Using the absorption spectrum data of the three samples of sucrose, fructose, and lactose after pretreatment as the original matrix samples, when initializing the dictionary, use part of the terahertz spectrum data in the original matrix samples as the initial dictionary, and then set the number of iterations, where the maximum number of iterations is 300 times. Use the OMP algorithm to solve the sparse coefficient matrix of the THz absorption spectrum, use the K-SVD algorithm to update the dictionary matrix, and calculate the error matrix; when the set iteration condition is reached, the training ends, and at this time, the sparse coefficient matrix of the absorption spectrum signals of the three saccharide samples can be obtained.

[0095] To verify the efficiency of sparse representation in THz spectral data processing, it was compared with traditional dimensionality reduction methods such as principal component analysis (PCA), linear discriminant analysis (LDA), local linear embedding (LLE), and isometric mapping (ISOMAP). The absorption spectra of sucrose, fructose, and lactose samples after pretreatment were selected as the research object, and the sparse coefficient matrix calculated after their sparse representation was used as the input of the model to construct a support vector machine (SVM) classification model. The comparison results are as Figure 5 shown.

[0096] From Figure 6 it can be seen that the operation time of the model with the sparse representation of THz spectra as the input of the SVM classification model is greatly reduced compared with that of the other three dimensionality reduction methods; therefore, selecting sparse representation as the dimensionality reduction method for THz spectral data can effectively reduce the operation time of the model.

[0097] The method of the present invention aims at problems such as large amount of THz spectral data, high dimensionality, complex structure, and long model calculation time, and uses sparse representation to reduce the dimensionality of the preprocessed THz spectral data; making the THz spectra more rapid and accurate in modeling analysis, and representing the data as concise as possible.

[0098] The OMP algorithm was selected to calculate the sparse coefficients, and K-SVD was used to update the dictionary to realize the sparse representation of terahertz absorption spectral data. It was compared with four traditional dimensionality reduction methods of principal component analysis, linear discriminant analysis, local linear embedding, and isometric mapping, and an SVM classification model was established. Compared with other dimensionality reduction methods, the operation time of the constructed terahertz sparse representation model was reduced to 0.132 s. The experimental results verified the effectiveness of the sparse representation non-linear dimensionality reduction method in improving the model efficiency, improved the quality of terahertz spectral data and the detection accuracy of the model, and enriched and developed the terahertz detection theory and methods.

[0099] The above is only used to illustrate the technical solution of the present invention and not to limit it. Other modifications or equivalent replacements made by those of ordinary skill in the art to the technical solution of the present invention shall be covered within the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.

Claims

1. A sparse representation classification method for sugar analysis based on terahertz spectroscopy, characterized in that: It includes the following steps: S1. Obtain the terahertz (THz) absorption spectral characteristics of the saccharide sample: Select three saccharides, namely sucrose, fructose, and lactose, as samples. First, use a terahertz time-domain spectroscopy system (THz-TDS-spectrometer) to measure a time-domain spectrum as a reference spectrum. Then, measure each type of sample multiple times to calculate the average time-domain spectrum, and use the time-domain spectrum to calculate the frequency-domain spectrum. Finally, calculate the absorption spectrum of each type of sample by comparing with the reference spectrum; S2. Preprocess the THz absorption spectral data of the saccharides: Select fast Fourier transform (FFT) filtering to preprocess the THz absorption spectral data of each type of sample calculated in S1, remove the noise signals with frequencies higher than the cut-off frequency in the THz spectral data, obtain a dataset of the filtered THz absorption spectral data, and get a smoother waveform of the THz absorption spectral data of each type of sample; In the step S2, the specific steps of using fast Fourier transform (FFT) filtering are as follows: S21: Calculate the average values of the first 1% data points and the last 1% data points for the THz absorption spectral data of each type of sample calculated in S1; S22: Construct a baseline through the two average value points calculated in S21, and subtract this baseline from all the THz absorption spectral data of each type of sample; S23: Perform FFT transformation on the data difference obtained in S22; S24: Perform low-pass filtering on the THz absorption spectral data obtained after transformation in S23 to remove Fourier components with frequencies higher than the cut-off frequency, where the cut-off frequency is: where n is the specified number of window points, n is set to 15, and t is the time interval between two adjacent data points; S25: Perform inverse fast Fourier transform (IFFT) calculation on the filtered spectrum, and after sorting, obtain a dataset of the THz absorption spectral data of each type of sample; S26: Add the baseline constructed in S22 to the dataset obtained in S25; S3. Perform sparse representation on the preprocessed THz absorption spectral data of the saccharides. The specific steps are as follows: S31: Use the dataset of the preprocessed THz absorption spectral data of the saccharides to construct a sparse representation algorithm model; Suppose the dataset Y after preprocessing the THz absorption spectral data of a sugar sample is represented as the original sample matrix Y ∈ R m ×n , where the original sugar sample absorption spectral signal y is the THz absorption spectral data of one type of sugar sample in the dataset Y. Then the sparse representation model is as follows: where D is an over-complete dictionary matrix, and d i is an atom, D = {d1, d2,..., d K}, x is the spectral sparse representation coefficient and has only a finite number of non-zero elements, x = {x1, x2,... x n}; The solution is obtained as follows: where t is the upper limit of the number of non-zero components in the spectral sparse representation coefficients; S32: Use the orthogonal matching pursuit (OMP) algorithm to solve the spectral sparse coefficient matrix: Set the sparsity k to obtain the number of iterations. By inputting the dictionary matrix D, the original absorption spectral signal y of the saccharide sample, and the sparsity k, solve to obtain the sparse coefficient x. The specific algorithm steps are as follows: Initialization: Initialize the dictionary matrix D and the sparsity k, the residual R0 = y, the index set Λ = 0, and t = 1; a1: Find the atom in the dictionary matrix D that best matches the residual R, i.e., find a certain column d in the dictionary matrix D with the largest inner product i , whose corresponding subscript is λ, λ t = argmax j=1,...N |<R t-1 , d j >| a2: Update the index set Λ t = Λ t-1 ∪ {λ t} and the atoms in the dictionary matrix D a3: Calculate using the least squares method by substituting into the solution formula in step S31 to obtain the sparse coefficient x of the absorption spectrum signal of the sugar sample t = argmax ||y - D t x t ||; a4: Update the residual R of the absorption spectral signal of the sugar sample t = y - D t x t , t = t + 1; a5: Judge whether t is greater than k. If t > k, the algorithm stops, and the sparse coefficient matrix of the absorption spectrum of the saccharide sample is obtained; if t ≤ t + 1, then jump back to step a1 to execute the algorithm again; S33: Use the K-SVD algorithm to update the dictionary matrix: By inputting the dictionary matrix D, the original absorption spectral signal y of the saccharide sample, and the sparse coefficient matrix calculated in step S32, obtain the updated dictionary matrix D. The specific algorithm steps are as follows: Initialization: Randomly select K column vectors from the absorption spectrum signal y of the original sugar sample or take the first K vectors from the left singular matrix of the original sample matrix Y as the atoms of the initialized dictionary, and initialize the dictionary D0 ∈ R m×K , j = 0, to obtain the initial dictionary D j ; b1: Using the obtained dictionary D j Perform spectral sparse coding to obtain the sparse coefficient matrix X of the absorption spectrum of the sugar sample j ∈R K×n ; b2: Update the dictionary matrix D column by column j , for each column d k ∈ {d1, d2,..., d K}, calculate the error matrix b3: Obtain the sparse matrix X of the absorption spectrum of the sugar sample j The k-th non-zero row vector The set of indices of b4: Extract the column corresponding to the index ω k from E k where the index ω is not zero, to obtain E′ k ; b5: For E' k Perform singular value decomposition on E k = UΔV T , take the first column of U to update the k-th column of the dictionary, d k = U(·, 1), let Put Update to the original position, j = j + 1; b6: Judge whether the specified number of iteration steps and the allowable error are reached. If so, stop the algorithm, and obtain the sparse coefficient matrix X of the absorption spectral signals of the three saccharide samples; if not, continue to enter the OMP algorithm for loop calculation until the iteration condition is met; 2. The method according to claim 1, wherein: In the step S1, The reference spectrum is obtained from the time-domain spectral signal directly measured by the THz-TDS system, and the defined formula is f r (t), where t is the frequency and r is the subscript of f(t); The average time-domain spectrum is obtained by averaging the time-domain spectral signals of each type of sample measured repeatedly. The defined formula is f s (t), where t is the frequency and s is the subscript of f(t); The frequency-domain spectrum is obtained by performing a Fourier transform on the average time-domain spectrum, and the calculation formula for the transform is as follows: where F(ω) is the frequency-domain spectrum signal and f(t) is the time-domain spectrum signal; The absorption spectrum is obtained by comparing the frequency-domain spectra of each type of sample with the frequency-domain spectrum of the reference spectrum. The calculation formula is as follows: where α is the absorption spectrum, d is the thickness of the sample tablet, A r is the amplitude of the frequency-domain spectrum of the reference spectrum, and A s is the amplitude of the frequency-domain spectrum of the saccharide sample.

3. A sparse representation classification system for sugar analysis based on terahertz spectroscopy, characterized in that, It includes: A spectral characteristic acquisition module is used to acquire the terahertz (THz) absorption spectral characteristics of sugar samples: Three types of sugars, namely sucrose, fructose, and lactose, are selected as samples. First, a terahertz time-domain spectroscopy system (THz-TDS-spectrometer) is used to measure a time-domain spectrum as a reference spectrum. Then, each type of sample is measured multiple times to calculate the average time-domain spectrum, and the frequency-domain spectrum is calculated using the time-domain spectrum. Finally, the absorption spectrum of each type of sample is calculated by comparing with the reference spectrum. A preprocessing module is used to preprocess the THz absorption spectral data of sugars: The fast Fourier transform (FFT) filtering is selected to preprocess the THz absorption spectral data of each type of sample calculated in S1, removing the noise signals with frequencies higher than the cut-off frequency in the THz spectral data, obtaining a dataset of the filtered THz absorption spectral data, and getting a smoother waveform of the THz absorption spectral data of each type of sample. A sparse representation module is used to perform sparse representation on the preprocessed THz absorption spectral data of sugars. The specific steps are as follows: Construct a sparse representation algorithm model using the dataset of the preprocessed THz absorption spectral data of sugars. Use the orthogonal matching pursuit (OMP) algorithm to solve the spectral sparse coefficient matrix. Specifically: Set the sparsity k to obtain the number of iterations. By inputting the dictionary matrix D, the original absorption spectral signal y of the sugar sample, and the sparsity k, the sparse coefficient x is solved. The specific algorithm steps are as follows: Initialization: Initialize the dictionary matrix D and the sparsity k. The residual R0 = y, the index set Λ = 0, and t = 1. Find the atom in the dictionary matrix \(D\) that best matches the residual \(R\), that is, find a certain column \(d\) in the dictionary matrix \(D\) with the largest inner product i , whose corresponding subscript is \(\lambda\), \(\lambda\) t =\arg\max j=1,...N |\lt R t-1 , d j \gt|; Update the index set Λ t = Λ t-1 ∪ {λ t} and the atoms in the dictionary matrix D Calculate using the least squares method by substituting into the solution formula in step S31 to obtain the sparse coefficient x of the absorption spectrum signal of the sugar sample t = argmax ||y - D t x t ||; Update the residual R of the absorption spectrum signal of the saccharide sample t = y - D t x t , t = t + 1; Judge whether t is greater than k. If t > k, the algorithm stops, and the sparse coefficient matrix of the absorption spectrum of the sugar sample is obtained; if t ≤ t + 1, then jump back to step a1 to execute the algorithm again. Use the K-SVD algorithm to update the dictionary matrix. Specifically: By inputting the dictionary matrix D, the original absorption spectral signal y of the sugar sample, and the sparse coefficient matrix calculated in step S32, the updated dictionary matrix D is obtained. The specific algorithm steps are as follows: Initialization: Randomly select K column vectors from the absorption spectrum signal y of the original sugar sample or take the first K vectors from the left singular matrix of the original sample matrix Y as the atoms of the initialized dictionary, and initialize the dictionary D0 ∈ R m×K , j = 0, to obtain the initial dictionary D j ; Using the obtained dictionary D j perform spectral sparse coding to obtain the sparse coefficient matrix X of the absorption spectrum of the sugar sample j ∈R K ×n ; Update the dictionary matrix D column by column j , for the column d of the dictionary k ∈ {d1, d2,..., d K}, calculate the error matrix Take the sparse matrix X of the absorption spectrum of the sugar sample j The k-th non-zero row vector The set of indices Take the corresponding index ω k from E k and obtain the non-zero columns to get E′ k ; For E′ k Perform singular value decomposition on E k = UΔV T , take the first column of U to update the k-th column of the dictionary, d k = U(:, 1), and let Replace with the updated one at the original position, j = j + 1; Judge whether the specified number of iteration steps and tolerance error are reached. If so, stop the algorithm, and the sparse coefficient matrix X of the absorption spectral signals of the three types of sugar samples is obtained; if not, continue to enter the OMP algorithm for loop calculation until the iteration condition is met.

4. An electronic device, characterized in that, It includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a sparse representation classification method for sugar analysis based on terahertz spectroscopy according to any one of claims 1 - 2.

5. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by the processor, it implements a sparse representation classification method for sugar analysis based on terahertz spectroscopy according to any one of claims 1 - 2.