Terahertz spectrum and machine learning combined trace organic matter type and concentration detection method
By combining terahertz spectroscopy with machine learning methods, using metasurface sensors to enhance signals and optimize models, the problems of slow detection speed and low sensitivity of trace organic matter were solved, and efficient and accurate detection of the types and concentrations of trace organic matter was achieved.
Patent Information
- Application Number
- CN202510701046.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-05
AI Technical Summary
Existing methods are slow and have low sensitivity in detecting trace organic matter. Traditional spectral analysis methods are complex and destructive to samples. Terahertz spectral signals are weak and easily overwhelmed by noise.
By combining terahertz spectroscopy with machine learning, and using metasurface sensors to enhance signals, an improved particle swarm optimization algorithm is used to optimize support vector classification and regression models, and a qualitative and quantitative dual-model collaborative analysis system is established to improve detection capabilities through preprocessing and dimensionality reduction.
It achieves sensitive and efficient detection of the types and concentrations of trace organic matter, shortens the detection time, and improves detection accuracy and robustness.
Smart Images

Figure CN120594440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of methods for detecting types and concentrations of organic matter, and specifically relates to a method for detecting types and concentrations of trace organic matter by combining terahertz spectroscopy with machine learning. Background Art
[0002] The detection of trace organic matter has important applications in environmental monitoring, food safety, and drug preparation. Traditional methods such as gas chromatography (GC) or liquid chromatography (HPLC) offer high sensitivity and accuracy in detecting organic matter, but they often require complex procedures, long detection times, and may damage the sample. In contrast, THz waves, located between microwaves and infrared, can sensitively identify the molecular vibrational characteristics of organic matter. Their inherent low photon energy effectively reduces the risk of sample damage and photoionization. THz spectroscopy analysis typically does not require complex sample pretreatment and can be performed rapidly and non-contact. Therefore, THz spectroscopy has enormous potential for real-time detection and analysis of organic matter.
[0003] In fields such as environmental monitoring, food safety, and pharmaceutical analysis, concentrations in the parts per million (ppm) range are generally considered trace amounts. When detecting trace amounts of organic matter, the relatively low energy and short wavelength of terahertz waves make the signal generated when they interact with organic matter weak and easily overwhelmed by noise. To more accurately measure trace samples, metasurfaces are used to enhance the interaction between the analyte and the terahertz wave, increasing signal strength and thus improving the ability to detect the target. Many enhanced sensing metasurfaces have been designed for collecting terahertz spectra of trace organic matter, but few have been used in practice.
[0004] Machine learning has long demonstrated significant advantages in the field of spectral analysis. By automatically learning and optimizing models, it can improve data processing efficiency and significantly enhance analytical accuracy and sensitivity. Integrating machine learning into predictive models for data analysis can automatically identify and differentiate the characteristics of different organic compounds, thereby improving analytical accuracy. Furthermore, machine learning enables THz spectroscopy to achieve rapid detection, meeting the demands for rapid response and high sensitivity. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting the types and concentrations of trace organic matter by combining terahertz spectroscopy with machine learning, which solves the problems of slow detection speed and low sensitivity of trace organic matter in existing methods.
[0006] The technical solution adopted by the present invention is a method for detecting the types and concentrations of trace organic matter using terahertz spectroscopy combined with machine learning, which is specifically implemented according to the following steps: Step 1: Preprocess the terahertz time-domain spectrum data of the sample to obtain the qualitative frequency-domain transmission spectrum data and quantitative frequency-domain transmission spectrum data of the sample; Step 2: Preprocess the data; Step 3: Divide the data into qualitative subsets, quantitative subsets, qualitative training set data, qualitative test set data, quantitative training set data, quantitative test set data, qualitative full set, and quantitative full set; Step 4: Use the qualitative training set data to train the qualitative model; use the quantitative training set data to train the quantitative model, and use the improved particle swarm optimization algorithm to optimize the model parameters; Step 5: Train the optimized qualitative model based on 5-fold cross-validation on the qualitative subset and the full qualitative dataset; train the optimized quantitative model based on 5-fold cross-validation on the quantitative subset and the full quantitative dataset, and evaluate the model performance; Step 6: Input the qualitative test set data into the trained qualitative model for classification and output the corresponding classification; input the quantitative training set data into the trained quantitative model for prediction and output the corresponding concentration.
[0007] The present invention is also characterized in that: In step 1, specifically: Step 1.1: Using a metasurface sensor on a terahertz time-domain spectroscopy system, a drop-drying method is used to obtain time-domain signals of organic compounds at different ppm concentration levels and time-domain signals of the same organic compound at different ppm concentration levels, generating qualitative time-domain signal data and quantitative time-domain signal data, respectively. Step 1.2: Perform Gaussian noise processing on the qualitative and quantitative time domain signal data of the sample, and after wavelet denoising, convert them into corresponding frequency domain signals through FFT, and then calculate the frequency domain transmission spectrum of the sample; the calculation formula is shown in formula (1); (1) Where ω is the frequency, is the reference frequency domain signal, ) is the sample frequency domain signal; is the frequency domain transmission spectrum of the sample.
[0008] In step 2, specifically, Savitzky-Golay filtering and normalization processing are performed on the qualitative frequency domain transmission spectrum data and the quantitative frequency domain transmission spectrum data of the sample obtained in step 1.
[0009] In step 3, specifically: The equal distribution method is used to select data from qualitative frequency domain transmission data samples and quantitative frequency domain transmission data samples as qualitative subsets and quantitative subsets; the qualitative subsets and quantitative subsets are divided into training sets and test sets to form qualitative training set data, qualitative test set data, quantitative training set data, and quantitative test set data; the training set and test set data are reduced in dimensionality through principal component analysis; all qualitative time domain signal data samples are called the qualitative full set; all quantitative time domain signal data samples are called the quantitative full set.
[0010] In step 4, specifically: Step 4.1: The qualitative model is the support vector classification model SVC, and the quantitative model is the support vector regression model SVR; Initialize the particle swarm. The position of each particle represents the logarithm of the support vector machine hyperparameter. The particle position encoding satisfies the following constraints:
[0011] Step 4.2: Define the fitness of SVC as maximizing accuracy and minimizing training time and test time; define the fitness of SVR as maximizing the coefficient of determination and minimizing the root mean square error and test time; Step 4.3: Dynamically adjust the search range, reducing the parameter space to 80% of the current optimal solution every 5 generations; Step 4.4: Introduce the Levy flight perturbation mechanism, update the particle velocity and position, and use dynamic inertia weights and boundary constraints.
[0012] In step 5, specifically: For parameters that meet the threshold conditions, the results are output; for parameters that do not meet the conditions, adjustments will be made using the IPSO optimized parameters as the baseline parameters until the optimized parameters are determined or all candidate parameters are evaluated; if all candidate parameters do not meet the requirements, the adjustment process will be terminated by reverting to the initial baseline parameter configuration; The threshold of the support vector classification model is defined as:
[0013]
[0014]
[0015] in, is the accuracy based on the qualitative full set, the average training time per fold, and the average testing time, is the accuracy based on the qualitative subset, the average training time per fold, and the average testing time, is the number of qualitative full set samples and the number of qualitative subset samples; The threshold of the support vector regression model is defined as:
[0016]
[0017]
[0018] in, is the R², RMSE, and average test time per fold based on the quantitative full set, is the R² based on the quantitative subset, RMSE, and average test time per fold, is the number of quantitative full set samples and the number of quantitative subset samples.
[0019] The beneficial effects of the present invention are as follows: the method of the present invention enhances the sensitivity of terahertz signals by optimizing the metasurface structure, combines the improved particle swarm algorithm to optimize the support vector classification model and the support vector regression model, establishes a qualitative and quantitative dual-model collaborative analysis system, and realizes sensitive and efficient detection of the types and concentrations of organic substances at the parts per million (ppm) level. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of the method for detecting the types and concentrations of trace organic matter using terahertz spectroscopy combined with machine learning. Figure 2 is a flow chart of the improved particle swarm optimization algorithm (IPSO) in the method of the present invention; Figure 3a This is a time domain spectrum data diagram of cysteine based on noise level 0.1; Figure 3b This is a time domain spectrum data diagram of cysteine based on noise level 0.2; Figure 3c This is a time domain spectrum data diagram of cysteine based on noise level 0.3; Figure 3d This is a time domain spectrum data diagram of cysteine based on noise level 0.4; Figure 4a These are the terahertz frequency domain transmission spectra of four organic compounds; Figure 4b This is the terahertz frequency domain transmission spectrum of 8 concentrations of sodium taurocholate; Figure 4c are the original transmission spectrum and the noisy transmission spectrum of bovine serum albumin; Figure 4d The original transmission spectrum and the noisy transmission spectrum of 5ppm sodium taurocholate are shown; Figure 5a This is the signal-to-noise ratio diagram of the terahertz frequency domain spectra of four organic compounds after preprocessing using different methods; Figure 5bThis is the signal-to-noise ratio diagram of the terahertz frequency domain spectra of sodium taurocholate at 8 concentrations after preprocessing using different methods; Figure 5c This is the signal-to-noise ratio diagram of the terahertz frequency domain spectra of four organic compounds after SG filtering under different conditions; Figure 5d This is the signal-to-noise ratio diagram of the terahertz frequency domain spectra of sodium taurocholate with 8 concentrations after SG filtering under different conditions. DETAILED DESCRIPTION
[0021] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Example 1 The present invention provides a method for detecting the types and concentrations of trace organic matter by combining terahertz spectroscopy with machine learning, such as Figure 1 As shown, please follow the steps below: Step 1: Preprocess the terahertz time-domain spectrum data of the sample to obtain the frequency-domain transmission spectrum data of the sample; specifically: Step 1.1: Using a metasurface sensor on a terahertz time-domain spectroscopy system, a drop-drying method was used to obtain time-domain signals of organic compounds at different ppm concentration levels and time-domain signals of the same organic compound at different ppm concentration levels. The volume of the analyte solution selected for the drop-drying method was 20 μL. Qualitative time-domain signal data and quantitative time-domain signal data were generated, respectively. Step 1.2: Perform Gaussian noise processing on the qualitative and quantitative time domain signal data of the sample, and after wavelet denoising, convert them into corresponding frequency domain signals through FFT, and then calculate the frequency domain transmission spectrum of the sample; the calculation formula is shown in formula (1); (1) Where ω is the frequency, is the reference frequency domain signal, ) is the sample frequency domain signal; is the frequency domain transmission spectrum of the sample; Step 2: Preprocess the frequency domain transmission spectrum of the sample obtained in step 1, including Savitzky-Golay filtering and normalization. The normalization is to divide the spectrum by the maximum value. Step 3: Use the equal distribution method to select data from the qualitative frequency domain transmission data samples and the quantitative frequency domain transmission data samples as the qualitative subset and the quantitative subset, with a ratio of 0.6; split the qualitative subset and the quantitative subset into training set and test set at a ratio of 6:4, respectively, to form qualitative training set data, qualitative test set data, quantitative training set data, and quantitative test set data; perform dimensionality reduction on the training set and test set data through principal component analysis (PCA), retaining a 0.9 contribution rate for the PCA dimensionality reduction of the qualitative training set data and the qualitative test set data, and retaining a 0.99 contribution rate for the PCA dimensionality reduction of the quantitative training set data and the quantitative test set data; refer to all qualitative time domain signal data samples as the qualitative full set; refer to all quantitative time domain signal data samples as the quantitative full set; Step 4: Use the reduced-dimensional qualitative training data to train the qualitative model; use the reduced-dimensional quantitative training data to train the quantitative model, define the fitness, and use the improved particle swarm optimization algorithm (IPSO) to optimize the model parameters based on the performance fitness evaluation of the model on the test set, such as Figure 2 As shown; The qualitative analysis model was the support vector classification model (SVC), and the quantitative analysis model was the support vector regression model (SVR); Specifically: Step 4.1: Initialize the particle swarm. The position of each particle represents the logarithm of the support vector machine hyperparameters (C and gamma). The particle position encoding satisfies the following constraints:
[0023] Step 4.2: Define the fitness function: Define the fitness of SVC as maximizing the accuracy (Acc) and minimizing the training time and test time; define the fitness of SVR as maximizing the coefficient of determination (R^2) and minimizing the root mean square error (RMSE) and test time.
[0024] Step 4.3: Dynamically adjust the search range, reducing the parameter space to 80% of the current optimal solution every 5 generations; Step 4.4: Introduce the Levy flight perturbation mechanism, update the particle velocity and position, and use dynamic inertia weights and boundary constraints; The update formula for particle velocity and position is as follows
[0025] in, It is a particle In the The speed of generation, is the dynamic inertia weight, and is the acceleration constant, which controls the individual experiences (particle The influence of historical optimality) and global experience (optimum of all particles), and are two numbers randomly generated in the interval [0,1]. is the optimal position of an individual, is the global optimal position.
[0026] Dynamic inertia weight, the update formula is as follows:
[0027] in, , are the initial weight and the final weight respectively; is the current iteration number, is the maximum number of iterations; The Levy flight perturbation probability is triggered, and the step size scaling factor decays linearly with the number of iterations. The step size generation follows the formula:
[0028]
[0029] in, It is obedience independent random variables, is located in The constant between is the current scale factor of the step size, , are the initial factor and the current factor respectively. The current iteration number , Maximum number of iterations.
[0030] Step 5: Train the optimized support vector classification model based on 5-fold cross-validation on the qualitative subset and the full qualitative data set; train the optimized support vector regression model based on 5-fold cross-validation on the quantitative subset and the full quantitative data set, and evaluate the model performance; For parameters that meet the threshold conditions, the results are output; for parameters that do not meet the conditions, fine-tuning will be performed using the IPSO optimized parameters as the baseline parameters until the optimized parameters are determined or all candidate parameters are evaluated. If all candidate parameters do not meet the requirements, the fine-tuning process will be terminated by reverting to the initial baseline parameter configuration; The threshold of the support vector classification model is defined as:
[0031]
[0032]
[0033] in, is the accuracy based on the qualitative full set, the average training time per fold, and the average testing time, is the accuracy based on the qualitative subset, the average training time per fold, and the average testing time, is the number of qualitative full set samples and the number of qualitative subset samples.
[0034] The threshold of the support vector regression model is defined as:
[0035]
[0036]
[0037] in, is the R², RMSE, and average test time per fold based on the quantitative full set, is the R² based on the quantitative subset, RMSE, and average test time per fold, is the number of quantitative full set samples and the number of quantitative subset samples.
[0038] Step 6: Input the qualitative test set data into the trained support vector classification model for classification, and output the corresponding classification; input the quantitative training set data into the trained support vector regression model for prediction, and output the corresponding concentration.
[0039] The advantages of the method of the present invention are: (1) Sensitive detection of trace organic matter can be achieved by using metasurface sensors; (2) By optimizing multiple targets using an improved particle swarm algorithm, multiple targets are included in the optimization range to achieve faster detection of large amounts of data and reduce testing time; (3) Using the subsampling method, we optimize on the subset and verify the transferability of the subset and the full set, and fine-tune the parameters that do not meet the conditions, thus reducing the optimization time; Example 2 By using the default parameters and IPSO optimization parameters to train the model, the 5-fold cross-validation results are shown in Tables 1 and 2. The training time and test time in the table are the average time of the 5-fold cross-validation.
[0040] From the prediction results of the qualitative analysis model SVC for the spectra of four organic compounds shown in Table 1, it can be found that metasurface-enhanced terahertz spectroscopy can effectively distinguish trace organic compounds. The optimization algorithm is used to optimize the training time and test time while maintaining high prediction accuracy. Compared with PSO, IPSO optimization reduces the training time by 7% and the test time by 12.5%.
[0041] Table 1 Prediction results of the qualitative analysis model SVC for the spectra of four organic compounds
[0042] Table 2 shows the prediction results of the quantitative analysis model SVR for eight different concentrations of sodium taurocholate. It shows that surface-enhanced terahertz spectroscopy can effectively distinguish different trace concentrations of the same organic compound. The optimization algorithm, while enhancing the coefficient of determination and reducing the root mean square error (RMSE), optimizes test time. Compared with the PSO and IPSO optimizations, the predicted labels are rounded to the nearest integer and used as the statistical number of correctly predicted samples.
[0043] Table 2 Quantitative analysis of the prediction results of the SVR model for the spectra of 8 concentrations of sodium taurocholate
[0044] Based on the optimized optimal parameters and the full set of preprocessed and dimensionality-reduced data, the final qualitative analysis model, quantitative analysis model, and corresponding PCA model are trained and saved.
[0045] Example 3 The spectral data of different organic compounds with noise levels of 0.2, 0.3, and 0.4 were input into the qualitative model respectively. The prediction results are shown in Table 3. The qualitative model can accurately distinguish the noisy spectra of the four different samples. The prediction time of IPSO-SVC for distinguishing different noise levels is better than that of SVC and PSO-SVC.
[0046] Table 3 Prediction results of the qualitative analysis model for the spectra of four organic compounds at different noise levels
[0047] Example 4 The quantitative model was used to input spectral data of sodium taurocholate at different concentrations, with noise levels of 0.2, 0.3, and 0.4. The prediction results are shown in Table 4. The quantitative model can effectively predict the sample at a noise level of 0.2. However, the prediction performance is poor under high noise conditions. Denoising the time domain data may improve the prediction accuracy. IPSO-SVC achieves better prediction times than SVC and PSO-SVC.
[0048] Table 4 Prediction results of the quantitative analysis model for sodium taurocholate spectra at different noise levels
[0049] In summary, metasurface-enhanced terahertz spectroscopy can effectively distinguish the types and concentrations of trace organic matter. Combined with machine learning, it can achieve accurate and efficient analysis of the types and concentrations of trace organic matter. The qualitative model has good robustness, and the quantitative model can effectively predict the sample concentration under low-noise conditions. By optimizing the algorithm, the time can be further shortened and the efficiency can be improved.
[0050] Example 5 Figure 3a This is a time domain spectrum data diagram of cysteine based on noise level 0.1; Figure 3b This is a time domain spectrum data diagram of cysteine based on noise level 0.2; Figure 3c This is a time domain spectrum data diagram of cysteine based on noise level 0.3; Figure 3d This is a time-domain spectrum data diagram of cysteine with noise added at a noise level of 0.4. As the noise level increases, the time-domain signal fluctuates significantly in the low-amplitude region. The generated noise signal can effectively expand the training dataset while retaining the basic spectral features, thereby enhancing the model's generalization ability and noise robustness.
[0051] Figure 4a These are terahertz frequency-domain transmission spectra of four organic compounds. Using a metasurface sensor, it is possible to characterize trace amounts of different organic compounds in the terahertz band. The differences in the characteristics and shapes of the spectral lines for different organic compounds improve the separability of the data, indicating that metasurface-enhanced terahertz spectroscopy can effectively distinguish trace amounts of different organic compounds. Figure 4b These are terahertz frequency-domain transmission spectra of sodium taurocholate at eight different concentrations. Using a metasurface sensor, we can characterize trace organic compounds at varying concentrations in the terahertz band. While the spectral characteristics and shapes of the organic compounds at different concentrations are similar, there are still differences. By extracting these spectral features through machine learning, we can predict the concentration of trace organic compounds of the same type.
[0052] Figure 4c These are the original transmission spectrum and noisy transmission spectrum of bovine serum albumin. As the noise level increases, the frequency domain spectrum data used for qualitative analysis fluctuates in amplitude in the high-frequency region, but the noise signal retains its basic spectral characteristics, which ensures the noise robustness of the model. Figure 4d These are the original transmission spectrum and noisy transmission spectrum of 5ppm sodium taurocholate. As the noise level increases, the frequency domain spectrum data used for quantitative analysis has significant fluctuations in amplitude in the high-frequency region, and the spectral characteristics of the noise signal change, which will affect the noise robustness of the model.
[0053] Example 6 Figure 5a This is the signal-to-noise ratio diagram of the terahertz frequency domain spectra of four organic compounds after preprocessing using different methods; Figure 5bThe figure shows the signal-to-noise ratio of the terahertz frequency domain spectra of sodium taurocholate at eight concentrations after preprocessing using different methods. For both the qualitative and quantitative data sets, SG filtering was selected as the preprocessing method, and the processed data are closest to the original data.
[0054] Figure 5c This is the signal-to-noise ratio diagram of the terahertz frequency domain spectra of four organic compounds after SG filtering under different conditions; Figure 5d Figure 2 shows the signal-to-noise ratio of the terahertz frequency domain spectra of eight sodium taurocholate concentrations after applying different SG filtering conditions. For both the qualitative and quantitative datasets, SG filtering with a window size of 7 and an order of 4 was selected as the preprocessing method, resulting in data that most closely resembled the original data.
Claims
1. A method for detecting the type and concentration of trace organic matter using terahertz spectroscopy combined with machine learning, characterized in that: Please follow the steps below to implement it: Step 1: Preprocess the terahertz time-domain spectrum data of the sample to obtain the qualitative frequency-domain transmission spectrum data and quantitative frequency-domain transmission spectrum data of the sample; Step 2: Preprocess the data; Step 3: Divide the data into qualitative subsets, quantitative subsets, qualitative training set data, qualitative test set data, quantitative training set data, quantitative test set data, qualitative full set, and quantitative full set; Step 4: Use the qualitative training set data to train the qualitative model; use the quantitative training set data to train the quantitative model, and use the improved particle swarm optimization algorithm to optimize the model parameters; Step 5: Train the optimized qualitative model based on 5-fold cross-validation on the qualitative subset and the full qualitative dataset; train the optimized quantitative model based on 5-fold cross-validation on the quantitative subset and the full quantitative dataset, and evaluate the model performance; Step 6: Input the qualitative test set data into the trained qualitative model for classification and output the corresponding classification; The quantitative training set data is input into the trained quantitative model for prediction and the corresponding concentration is output.
2. The method for detecting the type and concentration of trace organic matter by combining terahertz spectroscopy with machine learning as claimed in claim 1, characterized in that: In the step 1, specifically: Step 1.1: Using a metasurface sensor on a terahertz time-domain spectroscopy system, a drop-drying method is used to obtain time-domain signals of organic compounds at different ppm concentration levels and time-domain signals of the same organic compound at different ppm concentration levels, generating qualitative time-domain signal data and quantitative time-domain signal data, respectively. Step 1.2: Perform Gaussian noise processing on the qualitative and quantitative time domain signal data of the sample, perform wavelet denoising, and convert them into corresponding frequency domain signals through FFT, and then calculate the frequency domain transmission spectrum of the sample; The calculation formula is shown in formula (1); (1) Where ω is the frequency, is the reference frequency domain signal, ) is the sample frequency domain signal; is the frequency domain transmission spectrum of the sample.
3. The method for detecting the type and concentration of trace organic matter by combining terahertz spectroscopy with machine learning as claimed in claim 2, characterized in that: In the step 2, specifically, Savitzky-Golay filtering and normalization processing are performed on the qualitative frequency domain transmission spectrum data and the quantitative frequency domain transmission spectrum data of the sample obtained in the step 1.
4. The method for detecting the type and concentration of trace organic matter by combining terahertz spectroscopy with machine learning as claimed in claim 2, characterized in that: In the step 3, specifically: The equal distribution method is used to select data from qualitative frequency domain transmission data samples and quantitative frequency domain transmission data samples as qualitative subsets and quantitative subsets; the qualitative subsets and quantitative subsets are divided into training sets and test sets to form qualitative training set data, qualitative test set data, quantitative training set data, and quantitative test set data; the training set and test set data are reduced in dimensionality through principal component analysis; all qualitative time domain signal data samples are called the qualitative full set; all quantitative time domain signal data samples are called the quantitative full set.
5. The method for detecting the types and concentrations of trace organic matter by combining terahertz spectroscopy with machine learning as claimed in claim 4, characterized in that: In the step 4, specifically: Step 4.1: The qualitative model is the support vector classification model SVC, and the quantitative model is the support vector regression model SVR; Initialize the particle swarm. The position of each particle represents the logarithm of the support vector machine hyperparameter. The particle position encoding satisfies the following constraints: Step 4.2: Define the fitness of SVC as maximizing accuracy and minimizing training time and test time; define the fitness of SVR as maximizing the coefficient of determination and minimizing the root mean square error and test time; Step 4.3: Dynamically adjust the search range, reducing the parameter space to 80% of the current optimal solution every 5 generations; Step 4.4: Introduce the Levy flight perturbation mechanism, update the particle velocity and position, and use dynamic inertia weights and boundary constraints.
6. The method for detecting the type and concentration of trace organic matter by combining terahertz spectroscopy with machine learning according to claim 5, characterized in that: In step 4.4, the updating formulas for particle velocity and position are as follows: in, It is a particle In the The speed of generation, is the dynamic inertia weight, and are acceleration constants, controlling the effects of individual experience and global experience, and are two numbers randomly generated in the interval [0,1], is the optimal position of an individual, is the global optimal position.
7. The method for detecting the type and concentration of trace organic matter by combining terahertz spectroscopy with machine learning according to claim 6, wherein: In step 4.4, the dynamic inertia weight is updated using the following formula: in, , are the initial weight and the final weight respectively; is the current iteration number, is the maximum number of iterations; The Levy flight perturbation probability is triggered, and the step size scaling factor decays linearly with the number of iterations. The step size generation follows the formula: in, It is obedience independent random variables, is located in The constant between is the current scale factor of the step size, , are the initial factor and the current factor respectively. The current iteration number , Maximum number of iterations.
8. The method for detecting the type and concentration of trace organic matter by combining terahertz spectroscopy with machine learning according to claim 7, wherein: In the step 5, specifically: For parameters that meet the threshold conditions, the results are output; for parameters that do not meet the conditions, adjustments will be made using the IPSO optimized parameters as the baseline parameters until the optimized parameters are determined or all candidate parameters are evaluated; if all candidate parameters do not meet the requirements, the adjustment process will be terminated by reverting to the initial baseline parameter configuration; The threshold of the support vector classification model is defined as: in, is the accuracy based on the qualitative full set, the average training time per fold, and the average testing time, is the accuracy based on the qualitative subset, the average training time per fold, and the average testing time, is the number of qualitative full set samples and the number of qualitative subset samples; The threshold of the support vector regression model is defined as: in, is the R², RMSE, and average test time per fold based on the quantitative full set, is the R² based on the quantitative subset, RMSE, and average test time per fold, is the number of quantitative full set samples and the number of quantitative subset samples.
Citation Information
Cited By
Wood density detection method and system, computer equipment and storage medium
CN120948286A
Trace naphthylacetic acid detection method and system based on terahertz spectrum detection technology
CN121027029A