A rapid and nondestructive identification method for coffee bean varieties based on multi-modal terahertz spectroscopy and multi-scale adaptive weighted feature fusion
By using a multimodal terahertz spectroscopy and multi-scale adaptive weighted feature fusion method, combined with Bayesian optimization of SVM hyperparameters, the problems of subjectivity, low efficiency, destructiveness, and low information utilization in existing coffee bean variety identification methods are solved, achieving rapid, non-destructive, and accurate coffee bean variety identification.
Patent Information
- Application Number
- CN202610290256.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-11
- Publication Date
- 2026-06-12
AI Technical Summary
Existing methods for identifying coffee bean varieties are subjective, inefficient, destructive, and environmentally polluting. Terahertz spectroscopy identification relies on a single mode and has low information utilization. SVM hyperparameter optimization is difficult, making it hard to achieve rapid, accurate, and non-destructive identification of coffee bean varieties.
By employing multimodal terahertz spectroscopy (time domain, frequency domain, absorbance, first derivative) combined with multi-scale adaptive weighted feature fusion and Bayesian optimization to achieve SVM hyperparameter optimization, a support vector machine model is constructed to realize rapid and non-destructive identification of coffee bean varieties.
It achieves rapid, non-destructive, and accurate identification of coffee bean varieties, with an overall identification accuracy rate of 95.83%. The model has good robustness and generalization ability, and is suitable for food quality supervision and coffee industry production.
Smart Images

Figure CN122193144A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of non-destructive testing technology for food, specifically relating to a rapid non-destructive identification method for coffee bean varieties based on multimodal terahertz spectroscopy and multi-scale adaptive weighted feature fusion. Background Technology
[0002] Coffee, as the world's second most traded commodity after oil, boasts vast economic value and market prospects. Different varieties of coffee beans exhibit subtle differences in chemical composition and physical structure, yet their appearances are remarkably similar. This provides unscrupulous merchants with an opportunity to sell inferior or adulterated products, harming consumer rights and hindering the healthy development of the coffee industry. Therefore, achieving rapid, non-destructive, and accurate identification of coffee bean varieties is of great significance to both consumers and the coffee industry.
[0003] Traditional methods for identifying coffee bean varieties mainly include manual identification and chemical analysis. Manual identification relies on experience, is highly subjective, and inefficient. While chemical analysis methods such as high-performance liquid chromatography (HPLC), liquid chromatography-mass spectrometry (LC-MS), and gas chromatography (GC) offer high precision, they suffer from drawbacks such as cumbersome pretreatment, long processing times, sample destruction, and environmental pollution from the use of organic solvents. Although traditional spectroscopic techniques such as near-infrared, mid-infrared, and hyperspectral imaging can also be used for coffee bean variety identification, they generally suffer from peak overlap, susceptibility to environmental interference, insufficient stability, or limited depth detection capabilities, making it difficult to meet the requirements for rapid, stable, and non-destructive accurate identification.
[0004] Terahertz (THz) waves are electromagnetic waves with frequencies ranging from 0.1 to 10 THz. They possess unique advantages such as fingerprint spectral characteristics, penetrability, and low energy. The vibrational and rotational energy levels of many biomolecules fall within the terahertz band, making this technology highly promising for food testing. In recent years, some scholars have applied terahertz time-domain spectroscopy to coffee bean identification. However, most existing studies utilize only single-mode spectral features, failing to fully explore the diverse physicochemical information of coffee beans contained in multi-mode terahertz spectra. This results in low information utilization and the need to further improve the identification accuracy of classification models. Furthermore, the Support Vector Machine (SVM) commonly used in existing identification models suffers from difficulties in hyperparameter optimization. Traditional methods such as manual parameter tuning, grid search, and random search are either highly subjective and inefficient, or computationally expensive and prone to getting trapped in local optima, making it difficult to guarantee the model's classification performance and generalization ability.
[0005] To address the shortcomings of existing technologies, this invention proposes a rapid and non-destructive method for identifying coffee bean varieties based on multimodal terahertz spectroscopy and multi-scale adaptive weighted feature fusion. This method combines the complementary advantages of multimodal spectroscopy with a multi-scale adaptive weighted fusion strategy, and employs Bayesian optimization to achieve efficient optimization of SVM hyperparameters. This solves the problems of low information utilization, insufficient identification accuracy, and difficulty in hyperparameter optimization in existing methods, providing reliable technical support for coffee bean variety identification. Summary of the Invention
[0006] The purpose of this invention
[0007] The purpose of this invention is to overcome the shortcomings of existing coffee bean variety identification methods, such as strong subjectivity, low efficiency, destructiveness, and environmental pollution, as well as the shortcomings of terahertz spectroscopy identification, which only utilizes a single mode, has low information utilization, and is difficult to optimize SVM hyperparameters. This invention provides a rapid and non-destructive identification method for coffee bean varieties based on multi-modal terahertz spectroscopy and multi-scale adaptive weighted feature fusion, which can achieve rapid, non-destructive, and accurate identification of coffee bean varieties, and improve the identification accuracy and model robustness.
[0008] Technical solution
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] A rapid and non-destructive method for identifying coffee bean varieties based on multimodal terahertz spectroscopy and multi-scale adaptive weighted feature fusion includes the following steps:
[0011] Step 1: Sample preparation and spectral acquisition
[0012] Four types of coffee beans with significant differences in market price—Colombian, Catim, Blue Mountain, and Robusta—were selected as experimental samples. Among them, Colombian and Blue Mountain are Arabica varieties, Catim is a hybrid of Arabica and Robusta, and Robusta is an independent variety. The four samples showed significant differences in caffeine content, fat content, and flavor composition, which can effectively verify the classification and identification performance of the model.
[0013] Samples were prepared using a tableting method. The specific procedure was as follows: coffee beans were dried at a constant temperature of 50℃ for 2 hours to remove moisture (moisture has strong absorption in the terahertz band, which will seriously interfere with the spectral signal); after preliminary crushing and thorough grinding, the particle size was reduced to less than 50μm, and large particles were removed by passing the powder through a 300-mesh sieve; 200mg of coffee powder was weighed using a precision balance with an accuracy of 0.1mg; the powder was then pressed into circular tablets with a diameter of 13mm and a thickness of about 1mm in a tableting mold under a pressure of 5~8t. The tablets were required to have a smooth surface, no cracks, and no delamination. Unqualified samples were re-prepared; 120 tablets were made for each type of coffee bean, for a total of 480 tablets for the four types of coffee beans.
[0014] The CCT-1800 terahertz time-domain spectrometer manufactured by Shenzhen Huaxun Fangzhou Technology Co., Ltd. was used as the experimental data acquisition device. This spectrometer has a spectral range of 0.1-4.5 THz, a laser pulse repetition frequency of 80 MHz, a pulse center wavelength of 780 nm, and a pulse width of less than 100 fs. It mainly consists of a terahertz radiation generator, a terahertz radiation detector, a delay device, and a femtosecond pulse laser. Qualified samples were placed in a sample chamber equipped with a transmission terahertz experimental platform, maintaining a constant temperature (22±1℃), dry environment (relative humidity <3%), and continuously filled with nitrogen to eliminate interference from water vapor in the air. Each sample underwent three spectral acquisitions, and the average value was taken as the final spectral data for that sample to reduce random errors. A total of 480 raw time-domain spectral data points were obtained for four types of coffee beans.
[0015] Step 2: Spectral Preprocessing
[0016] The raw terahertz spectral signal may contain interference factors such as baseline drift, random noise, invalid time period signals, and outlier samples, which can affect the accuracy of subsequent feature extraction and model building. Therefore, the raw spectrum needs to be preprocessed, specifically including:
[0017] (2.1) Effective time period selection: Select the time domain spectrum in the range of 5-20ps with a high signal-to-noise ratio to reduce redundant interference;
[0018] (2.2) SG smoothing: The Savitzky-Golay (SG) smoothing method is used to reduce the noise of the spectrum. The smoothing window size is 11 points and the order of the fitting polynomial is 2. While effectively reducing noise, the feature details of the spectrum are preserved to the greatest extent and the feature loss is avoided due to excessive smoothing.
[0019] (2.3) ALS baseline correction: The adaptive iterative reweighted penalized least squares (ALS) method is used for baseline correction. The number of iterations is 5, the penalty coefficient is 1e5, and the smoothing factor is 0.01. The signal peaks and the baseline in the spectrum are distinguished, and a smooth baseline is gradually fitted. Then, the original spectrum is subtracted from the fitted baseline to obtain the spectrum after baseline correction, ensuring that the characteristic information of the spectrum itself is not destroyed.
[0020] (2.4) SNV normalization: The time-domain spectrum after baseline correction is normalized by standard normal variable (SNV). The mean of the spectral signal of each sample is subtracted and then divided by its standard deviation, so that the mean of the processed spectral signal is 0 and the standard deviation is 1, which effectively eliminates the interference caused by factors such as sample particle size and thickness.
[0021] (2.5) Outlier detection and median replacement: A robust strategy based on the intra-class median is used to identify and process outlier samples. The samples are grouped by coffee bean category, and the median of the spectrum of each category is calculated. Then, the Euclidean distance of each sample to the median spectrum of the category is calculated. The "distance mean + 3 times the standard deviation" is used as the outlier judgment threshold. If the sample distance exceeds the threshold, it is judged as an outlier sample. For outlier samples, the median spectrum of the category is used for replacement. At the same time, the corresponding outlier data in the time domain spectrum is replaced simultaneously to ensure data consistency and avoid the impact of dataset bias on model performance.
[0022] Step 3: Multimodal spectral acquisition
[0023] The preprocessed time-domain spectrum was subjected to a Fast Fourier Transform (FFT) to convert it into a frequency-domain spectrum. The effective frequency-domain signal within the range of 0.2–2.5 THz was extracted, and the amplitude of the frequency-domain signal was calculated. Based on the Lambert-Beer law and combining the frequency-domain amplitudes of the reference and sample signals, the absorbance spectrum was calculated using the following formula:
[0024]
[0025] in, The sample's frequency domain amplitude. For the reference signal frequency domain amplitude, The constant is a small constant used to avoid zero denominators, and the matrix dimension is adapted using the repmat function to ensure calculation accuracy. The first derivative of the absorbance spectrum is obtained by taking the first derivative spectrum, ultimately yielding four modal spectra: time domain, frequency domain, absorbance, and first derivative.
[0026] Step 4: Dataset Partitioning
[0027] To avoid overfitting or underfitting of the model due to the randomness of the dataset partitioning, the total dataset (4 classes × 120 records) was divided into a calibration set (4 classes × 84 records) and a test set (4 classes × 36 records) in a 7:3 ratio using the hierarchical Kennard-Stone (KS) algorithm. The calibration set was used for model training and parameter optimization, while the test set was used only for model performance verification and did not participate in model training, thus ensuring its independence and confidentiality. The hierarchical partitioning ensured that the sample ratio of the four types of coffee beans in the calibration set and the test set was consistent with that in the original dataset, avoiding an excessively high proportion of samples of a certain variety in the calibration set or the test set, which could lead to model bias.
[0028] Step 5: Feature Extraction
[0029] Based on the preprocessed time-domain, frequency-domain, absorbance, and first derivative modal spectra, single-scale and multi-scale features are extracted, laying the foundation for subsequent Bayesian optimization SVM modeling and feature fusion. Invalid data is processed simultaneously during the extraction process to ensure the effectiveness of the features.
[0030] (5.1) Single-scale feature extraction: Global feature extraction is performed for a single-mode spectrum. The representativeness of the features is improved by combining statistical features and PCA features. At the same time, the cumulative contribution rate of PCA features is calculated to provide a basis for subsequent weighted fusion. The core process is as follows:
[0031] ① Statistical characteristics: Seven statistical characteristics were extracted for each modal spectrum, namely mean, standard deviation, maximum value, minimum value, median, skewness, and kurtosis, to reflect the overall distribution characteristics of the spectrum;
[0032] ② PCA features: After Z-score normalization of the spectral data, principal component analysis (PCA) is performed. Principal components with a cumulative contribution rate of ≥99% are selected as PCA features. This reduces feature redundancy while retaining key information, and the cumulative contribution rate of each principal component is recorded simultaneously.
[0033] ③ Feature integration: Statistical features and PCA features are concatenated sequentially to obtain a single-scale feature vector for each modal spectrum;
[0034] ④ Anomaly handling: Replace invalid values such as non-numeric values and infinity in the feature vector with 0 to ensure the validity of the feature data.
[0035] Single-scale features were extracted from all four modal spectra using the method described above for single-modal modeling and subsequent feature fusion. Before modeling, the hierarchical KS method was used to divide the calibration set and the test set, and all models shared this division result to ensure modeling fairness.
[0036] (5.2) Multi-scale feature extraction: In order to balance global and local features and fully explore the detailed information in the spectrum, multi-scale feature extraction adopts a three-scale division strategy. Each scale follows the single-scale feature extraction method. The core process is as follows:
[0037] ① Scale division: Each modal spectrum is divided into three scales: the complete spectrum, the first half of the spectrum, and the second half of the spectrum;
[0038] ② Feature extraction: For spectral data at each scale, statistical features and PCA features are extracted separately, using the same extraction method as for single-scale feature extraction;
[0039] ③ Feature integration: The feature vectors of the three scales are concatenated in sequence, and invalid values are processed simultaneously to achieve effective fusion of global and local features, thereby further improving the feature representation capability.
[0040] Step 6: Multi-scale adaptive weighted feature fusion
[0041] First, based on the Bayesian Optimized Support Vector Machine (BO-SVM) model, the classification performance of four single-modal spectra was quantitatively evaluated, and the two modal spectra with the best classification effect (time-domain spectrum and absorbance spectrum) were selected. Then, based on the single-scale and multi-scale features of these two optimal modalities, three advanced feature fusion strategies were designed. The optimal strategy is multi-scale adaptive weighted feature fusion. The specific process is as follows: weights are assigned based on the cumulative contribution rate of PCA features at each scale. The feature vector of each scale is multiplied by the corresponding weight and then concatenated and fused. After Z-score normalization, it is used for modeling. This strategy can highlight the key information of high contribution rate scales and improve the relevance of features.
[0042] Step 7: Model Building and Optimization
[0043] Support Vector Machine (SVM) is used as the classification model. SVM is a supervised learning algorithm based on statistical learning theory. It classifies samples by finding the optimal hyperplane and has advantages such as strong generalization ability, resistance to overfitting, and suitability for small sample classification, making it suitable for small sample scenarios such as coffee bean variety identification. This study uses a nonlinear SVM model and selects the Radial Basis Function (RBF) as the kernel function. The RBF kernel function can map the low-dimensional feature space to the high-dimensional feature space, effectively handling nonlinear classification problems. Its expression is:
[0044]
[0045] in For the input feature vector, For support vectors, The kernel function parameters determine the complexity of the feature mapping.
[0046] The classification performance of the SVM model mainly depends on two hyperparameters: the kernel scale parameter γ and the penalty coefficient C. The penalty coefficient C is used to balance the model's classification accuracy and generalization ability. The larger the C value, the better the model fits the calibration set, making it prone to overfitting; the smaller the C value, the stronger the model's generalization ability, making it prone to underfitting. Therefore, Bayesian optimization (BO) algorithm is used to optimize the hyperparameters C and γ. Bayesian optimization is based on Bayes' theorem, models the objective function (model accuracy) through a probabilistic model, and adaptively selects the next parameter combination to be evaluated in combination with the acquisition function. It efficiently approaches the global optimum within a limited number of iterations, overcomes the shortcomings of traditional optimization methods, improves the accuracy and efficiency of parameter optimization, and obtains the optimal BO-SVM model.
[0047] Step 8: Identifying Coffee Bean Varieties
[0048] The spectral features of the test set are input into the optimal BO-SVM model, which outputs the identification results to achieve rapid and non-destructive identification of coffee bean varieties. Confusion matrix, classification accuracy, precision, recall, and F1 score are used as evaluation metrics for model performance. Five-fold cross-validation is employed to verify the model's robustness and generalization ability. Five-fold cross-validation involves randomly dividing the calibration set into five non-overlapping subsets of similar size. Four subsets are selected as the training set and one subset as the validation set in each iteration, repeated five times. The average accuracy of the five validations is calculated. If the five-fold cross-validation accuracy is similar to the test set classification accuracy, it indicates that the model has no significant overfitting and good robustness.
[0049] Beneficial effects
[0050] Compared with the prior art, the present invention has the following advantages:
[0051] 1. This invention employs multimodal terahertz spectroscopy (time domain, frequency domain, absorbance, first derivative) to fully explore the physicochemical information of coffee beans contained in terahertz spectroscopy, solving the problem that existing terahertz identification only utilizes a single mode and has low information utilization, thus providing rich feature support for accurate identification;
[0052] 2. A multi-scale adaptive weighted feature fusion strategy is proposed, which assigns weights based on the cumulative contribution rate of PCA features at each scale, highlights the key information of scales with high contribution rates, balances global features and local features, significantly improves feature representation ability, and thus improves discrimination accuracy;
[0053] 3. The Bayesian optimization algorithm is used to optimize the hyperparameters of SVM, overcoming the shortcomings of traditional methods such as manual parameter tuning, grid search, and random search. It can quickly focus on the optimal parameter range under small sample conditions, reduce computational costs, and improve the classification performance and generalization ability of the model.
[0054] 4. The entire method achieves rapid, non-destructive, and green identification of coffee bean varieties. It requires no complex pretreatment, does not damage the sample, and does not pollute the environment. It has high detection efficiency and high identification accuracy, with an overall identification accuracy of 95.83%. The identification accuracy of Catimor and Blue Mountain coffee beans both reached 100%. The deviation of the model's 5-fold cross-validation and test set accuracy is within ±1.04%, demonstrating good robustness and generalization ability.
[0055] 5. This method can effectively replace traditional sensory evaluation and cumbersome chemical testing methods, providing a reliable technical solution for the rapid identification of coffee bean varieties. It is applicable to scenarios such as food quality supervision and coffee industry production, and has important practical application value and promotion prospects. Attached Figure Description
[0056] Figure 1 : Schematic diagram of the experimental process of this invention;
[0057] Figure 2 Schematic diagram of CCT-1800 terahertz spectrometer;
[0058] Figure 3 Terahertz spectra of four coffee beans in four modes (a-time domain spectrum, b-frequency domain spectrum, c-absorbance spectrum, d-first derivative spectrum).
[0059] Figure 4 Confusion matrices of four single-modal spectral models (a-time domain spectrum, b-frequency domain spectrum, c-absorbance spectrum, d-first derivative spectrum).
[0060] Figure 5 Confusion matrices for three feature fusion models (a-direct feature fusion, b-multi-scale feature fusion, c-multi-scale adaptive weighted feature fusion). Detailed Implementation
[0061] The present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0062] Example 1: A rapid and non-destructive method for identifying coffee bean varieties based on multimodal terahertz spectroscopy and multi-scale adaptive weighted feature fusion
[0063] This embodiment uses four types of coffee beans—Colombia, Catimor, Blue Mountain, and Robusta—as research subjects to achieve rapid and non-destructive identification of their varieties. The specific steps are as follows:
[0064] 1. Sample preparation and spectral acquisition
[0065] Four types of coffee beans—Colombian, Catimor, Blue Mountain, and Robusta—were selected. 120 samples of each type were prepared, totaling 480 samples. The preparation process was as follows: drying at 50℃ for 2 hours → pulverizing and grinding to a particle size <50μm → passing through a 300-mesh sieve → weighing 200mg of powder → pressing into circular flakes with a diameter of 13mm and a thickness of approximately 1mm under pressure of 5-8t. Using a CCT-1800 terahertz time-domain spectrometer, under a constant temperature of 22±1℃, relative humidity <3%, and continuous nitrogen purging, the spectra of each sample were collected three times and averaged, yielding 480 raw time-domain spectral data points.
[0066] 2. Spectral preprocessing
[0067] The effective time-domain spectrum of 5-20 ps is extracted, and noise reduction is performed by SG smoothing (11-point window, second-order polynomial), ALS baseline correction (5 iterations, penalty coefficient 1e5, smoothing factor 0.01), SNV normalization, and outlier detection and replacement based on the intra-class median to obtain the preprocessed time-domain spectrum.
[0068] 3. Multimodal spectral acquisition
[0069] The preprocessed time-domain spectrum was subjected to FFT transformation to obtain the 0.2-2.5 THz frequency domain spectrum. The absorbance spectrum and the first derivative spectrum were calculated to obtain four modal spectra.
[0070] 4. Dataset Partitioning
[0071] The hierarchical KS algorithm was used to divide the calibration set (336 records) and the test set (144 records) in a 7:3 ratio to ensure that the proportion of the four types of coffee bean samples was consistent.
[0072] 5. Feature Extraction
[0073] Seven statistical features and PCA features with a cumulative contribution rate ≥99% were extracted for each modal spectrum and concatenated to obtain a single-scale feature vector. Each modal spectrum was divided into three scales, and features were extracted and concatenated for each scale to obtain a multi-scale feature vector. Invalid values were processed to ensure validity.
[0074] 6. Multi-scale adaptive weighted feature fusion
[0075] The classification performance of four single-mode spectra was evaluated, and the time-domain spectrum (accuracy 93.06%) and absorbance spectrum (accuracy 93.06%) were selected as the optimal modes. Based on the multi-scale features of the two optimal modes, weights were assigned according to the cumulative contribution rate of PCA features at each scale, and the weighted features were spliced and fused. The fused features were obtained by SNV normalization.
[0076] 7. Model Building and Optimization
[0077] An RBF kernel SVM model was constructed, and Bayesian optimization was used to optimize the hyperparameters C and γ to obtain the optimal BO-SVM model.
[0078] 8. Identification and Verification
[0079] The test set features were input into the optimal model, and the identification results were obtained: the overall accuracy was 95.83%, the accuracy of Katim and Blue Mountain was 100%, the accuracy of Robusta was 97.2%, and the accuracy of Columbia was 86.1%; the deviation between the accuracy of 5-fold cross-validation and the accuracy of the test set was 1.04%, indicating that the model has good robustness and generalization ability.
[0080] The results of the embodiments show that the method of the present invention can quickly, non-destructively, and accurately identify coffee bean varieties, effectively solve the shortcomings of existing methods, and has good practical application value.
Claims
1. A rapid and non-destructive method for identifying coffee bean varieties based on multimodal terahertz spectroscopy and multi-scale adaptive weighted feature fusion, characterized in that, Includes the following steps: (1) Sample preparation and spectral acquisition: Different varieties of coffee beans were selected as experimental samples. Samples were prepared by pressing. The original time-domain spectral data of the samples were acquired by a terahertz time-domain spectrometer. The constant temperature and dry environment was maintained and nitrogen was introduced to eliminate environmental interference. The samples were collected multiple times and the average value was taken. (2) Spectral preprocessing: The original time-domain spectrum is subjected to effective time segmentation, Savitzky-Golay (SG) smoothing, adaptive iterative reweighted penalized least squares (ALS) baseline correction, standard normal variable (SNV) normalization, outlier detection and median replacement in sequence to obtain the preprocessed time-domain spectrum. (3) Multimodal spectrum acquisition: The preprocessed time-domain spectrum is subjected to fast Fourier transform (FFT) to convert it into frequency-domain spectrum. The amplitude of the frequency-domain signal is calculated and the absorbance spectrum is obtained by combining Lambert-Beer law. The first derivative of the absorbance spectrum is obtained by taking the first derivative spectrum, thus obtaining four modal spectra: time-domain, frequency-domain, absorbance, and first derivative. (4) Data set partitioning: The hierarchical Kennard-Stone (KS) algorithm is used to partition the calibration set and the test set in a 7:3 ratio to ensure that the proportion of coffee bean samples of each variety in the calibration set and the test set is consistent with that in the original dataset; (5) Feature extraction: Single-scale features and multi-scale features were extracted for the four modal spectra respectively; single-scale features include statistical features and PCA features. The statistical features are mean, standard deviation, maximum value, minimum value, median, skewness and kurtosis. The PCA features are selected from principal components with a cumulative contribution rate ≥99%. The multi-scale features adopt a three-scale partitioning strategy, dividing each modal spectrum into three scales: the complete spectrum, the first half of the spectrum, and the second half of the spectrum. Statistical features and PCA features are extracted from each scale and integrated. (6) Multi-scale adaptive weighted feature fusion: Select the two modal spectra with the best classification effect, assign weights based on the cumulative contribution rate of PCA features at each scale, multiply the feature vector of each scale with the corresponding weight and then concatenate and fuse them, and obtain the fused feature vector after Z-score normalization. (7) Model construction and optimization: A support vector machine (SVM) model is constructed based on the fused feature vectors. The radial basis function (RBF) is selected as the kernel function. The Bayesian optimization (BO) algorithm is used to optimize the kernel function parameter γ and the penalty coefficient C of the SVM model to obtain the optimal BO-SVM model. (8) Coffee bean variety identification: Input the spectral features of the test set into the optimal BO-SVM model, output the identification results, and complete the rapid and non-destructive identification of coffee bean varieties.
2. The method according to claim 1, characterized in that, The specific process of the tableting method described in step (1) is as follows: <1> The coffee beans were dried at a constant temperature of 50℃ for 2 hours to remove moisture. <2> The dried coffee beans are initially crushed and thoroughly ground until the particle size is less than 50μm, and then passed through a 300-mesh sieve to remove large particles. <3> Weigh 200mg of coffee powder using a precision balance with an accuracy of 0.1mg, and press it into a round sheet with a diameter of 13mm and a thickness of about 1mm in a tableting mold with a pressure of 5~8t. The sheet should have a smooth surface, no cracks, and no delamination. <4> 120 sample pieces were made for each type of coffee bean, for a total of 480 sample pieces for the four types of coffee beans.
3. The method according to claim 1, characterized in that, In step (2): The effective time period is 5-20 ps; the window size for SG smoothing is 11 points, and the order of the fitted polynomial is 2; the number of iterations for ALS baseline correction is 5, the penalty coefficient is 1e5, and the smoothing factor is 0.01; outlier detection adopts a robust strategy based on the intra-class median, using "distance from the mean + 3 times the standard deviation" as the outlier judgment threshold, and outlier samples are replaced with the median spectrum of the corresponding category.
4. The method according to claim 1, characterized in that, The effective range of the frequency domain spectrum mentioned in step (3) is 0.2-2.5 THz, and the formula for calculating the absorbance spectrum is: in, The sample's frequency domain amplitude. For the reference signal frequency domain amplitude, It is a tiny constant used to avoid denominators of 0, and the matrix dimension is adapted by the repmat function.
5. The method according to claim 1, characterized in that, In the feature extraction process described in step (5), invalid values such as non-numeric values and infinity in the feature vector are replaced with 0 to ensure the validity of the feature data.
6. The method according to claim 1, characterized in that, The two modal spectra with the best classification effect in step (6) are time-domain spectra and absorbance spectra, and the single-scale feature identification accuracy of both modal spectra reaches 93.06%.
7. The method according to claim 1, characterized in that, The expression for the radial basis function (RBF) mentioned in step (7) is: in For the input feature vector, For support vectors, The kernel function parameter determines the complexity of the feature mapping.
8. The method according to claim 1, characterized in that, In step (8), confusion matrix, classification accuracy, precision, recall and F1 score are used as model performance evaluation indicators. At the same time, 5-fold cross-validation is used to verify the robustness and generalization ability of the model. The deviation between the 5-fold cross-validation accuracy and the test set accuracy is within ±1.04%.
9. The method according to claim 1, characterized in that, The different coffee bean varieties mentioned are Colombian, Catimor, Blue Mountain, and Robusta. The overall identification accuracy of the optimal BO-SVM model reached 95.83%, with the identification accuracy of Catimor and Blue Mountain coffee beans both reaching 100%.
10. The method according to claim 1, characterized in that, The terahertz time-domain spectrometer mentioned in step (1) is a CCT-1800 terahertz time-domain spectrometer with a spectral range of 0.1-4.5THz, a laser pulse repetition frequency of 80MHz, a pulse center wavelength of 780nm, and a pulse width of less than 100fs.