Chlorogenic acid component detection method fused with deep learning

By integrating deep learning and hyperspectral imaging technologies, and combining fuzzy C-means clustering and a high-dimensional multi-scale support vector machine regression model, the problem of noise interference in chlorogenic acid hyperspectral image processing by traditional clustering algorithms is solved, and accurate detection and stable prediction of chlorogenic acid concentration are achieved.

CN121476086APending Publication Date: 2026-02-06SHAANXI TIANXINGJIAN BIOCHEMICAL TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202610013407.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Traditional clustering algorithms cannot effectively isolate the interference of noise pixels on the calculation of cluster centers when processing chlorogenic acid hyperspectral images, resulting in blurred cluster boundaries and difficulty in stably reflecting the true spectral characteristics of the effective chlorogenic acid region, leading to inaccurate concentration prediction.

Method used

A deep learning-integrated approach was adopted to acquire hyperspectral images of chlorogenic acid using a hyperspectral imaging device. Clustering was performed using a fuzzy C-means clustering algorithm based on several sub-centers to extract chlorogenic acid spectral cluster feature vectors. A high-dimensional multi-scale support vector machine regression model was used for concentration prediction, and a sliding window method was used to detect concentration outliers.

Benefits of technology

It enables precise detection of chlorogenic acid concentration, improves the stability and prediction accuracy of clustering results, reduces the impact of noise on clustering results, and enhances the reliability and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121476086A_ABST
    Figure CN121476086A_ABST
Patent Text Reader

Abstract

The invention discloses a chlorogenic acid component detection method fused with deep learning, and relates to the technical field of chlorogenic acid. The method comprises the following steps: acquiring a chlorogenic acid hyperspectral image; clustering the chlorogenic acid hyperspectral image through a fuzzy C-means clustering algorithm based on a plurality of subclass centers to obtain a chlorogenic acid hyperspectral cluster, and calculating a chlorogenic acid spectral cluster feature vector; the method comprises the following steps: calculating mutual information of chlorogenic acid hyperspectral cluster feature vectors to obtain hyperspectral cluster feature weights; obtaining a chlorogenic acid hyperspectral cluster weighted vector by combining the chlorogenic acid hyperspectral cluster feature vector and the hyperspectral cluster feature weight; inputting the chlorogenic acid hyperspectral cluster weighted vector into a support vector machine regression model based on high-dimensional multiple scales, and outputting to obtain a chlorogenic acid concentration predicted value; and constructing a chlorogenic acid concentration diagram according to the chlorogenic acid concentration predicted value, calculating a concentration abnormal value through a sliding window method, and performing early warning if the concentration abnormal value is greater than a preset threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chlorogenic acid technology, specifically to a method for detecting chlorogenic acid components that integrates deep learning. Background Technology

[0002] Chlorogenic acid, an important bioactive substance, is widely found in various agricultural products and medicinal herbs such as coffee beans and honeysuckle. Its content is a key indicator for evaluating product quality and medicinal value. Therefore, developing rapid and accurate chlorogenic acid detection technologies is of great significance for quality control in agricultural production, food processing, and the pharmaceutical industry. Hyperspectral imaging technology combines the advantages of spectral analysis and image processing, enabling the simultaneous acquisition of spatial and spectral information of the sample, providing a powerful technical means for non-destructive and visualized detection of chlorogenic acid.

[0003] Traditional methods typically employ clustering algorithms to perform preliminary segmentation of the hyperspectral images of chlorogenic acid samples. By calculating the membership degree of each pixel to all cluster centers, the image is divided into several regions with similar spectral characteristics (i.e., clusters), and the spectral features of each region are extracted to predict the concentration of chlorogenic acid.

[0004] However, when processing chlorogenic acid hyperspectral images, traditional clustering algorithms rely directly on the spectral data of all pixels for their cluster centers. This makes it difficult to effectively isolate the interference of noisy pixels on the calculation of cluster centers, causing deviations in the cluster centers during the iteration process. As a result, the cluster centers cannot stably reflect the true spectral characteristics of the effective chlorogenic acid region, leading to blurred cluster boundaries and ultimately inaccurate prediction of chlorogenic acid concentration. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a chlorogenic acid component detection method that integrates deep learning, thereby resolving the problems existing in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for detecting chlorogenic acid using deep learning, comprising the following steps: Step S1: Take a picture of the chlorogenic acid sample using a hyperspectral imaging device to obtain a hyperspectral image of chlorogenic acid; Step S2: Cluster the chlorogenic acid hyperspectral image using a fuzzy C-means clustering algorithm based on several sub-centers to obtain chlorogenic acid hyperspectral clusters; Step S3: Extract features from the chlorogenic acid hyperspectral clusters to obtain the chlorogenic acid spectral cluster feature vectors; Step S4: Obtain the hyperspectral cluster feature weights by calculating the mutual information of the chlorogenic acid hyperspectral cluster feature vectors; obtain the chlorogenic acid hyperspectral cluster weighted vector by combining the chlorogenic acid hyperspectral cluster feature vectors and the hyperspectral cluster feature weights. Step S5: Input the weighted vector of the chlorogenic acid hyperspectral cluster into the support vector machine regression model based on high-dimensional multi-scale, and output the predicted value of chlorogenic acid concentration; construct a chlorogenic acid concentration map based on the predicted value of chlorogenic acid concentration, and calculate the concentration anomaly value by the sliding window method. If the concentration anomaly value is greater than the preset threshold, an early warning will be issued, thereby realizing the detection of chlorogenic acid components.

[0007] Preferably, the step of capturing a hyperspectral image of chlorogenic acid from the chlorogenic acid sample using a hyperspectral imaging device includes the following specific steps: Place the chlorogenic acid sample stably on the loading plane of the electric translation stage, ensuring that its main surface area faces the lens. Turn on the adjustable halogen lamp light source, with its initial incident angle at a 45-degree angle to the normal of the chlorogenic acid sample. Start the hyperspectral camera and place the diffuse reflection reference white plate next to the sample on the loading stage as a reference for illumination calibration. The effect of specular reflection was eliminated by adjusting the incident angle of the halogen lamp source and the parameters of the hyperspectral camera using an electrically controlled rotary stage; finally, the calibrated hyperspectral image of chlorogenic acid was calculated.

[0008] in, This represents a hyperspectral image of chlorogenic acid. This indicates that the chlorogenic acid sample is measured at pixels (u,v). The original spectral signal value corresponding to the wavelength, Indicates the diffuse reference whiteboard at The average value of the original signal of all spatial pixels at the specified wavelength. Indicates the diffuse reference whiteboard at Standard reflectance at wavelength; Finally, a hyperspectral image of chlorogenic acid was obtained. .

[0009] Preferably, the step of clustering chlorogenic acid hyperspectral images using a fuzzy C-means clustering algorithm based on several sub-centers to obtain chlorogenic acid hyperspectral clusters includes the following steps: The chlorogenic acid hyperspectral image obtained in step S1 Transform into a two-dimensional matrix X∈ Where: N is the total number of pixels in the chlorogenic acid hyperspectral image, D is the number of spectral bands, and each row of the two-dimensional matrix X represents a different element. It is a D-dimensional vector. The d-th element It represents the center wavelength of the d-th band. The reflectance value at that location; Initialize the first fuzzy partitioning matrix Y, the second fuzzy partitioning matrix Z, and the class center matrix; by fixing Y, Z, and center, calculate the vector of each subclass center. This yields the subclass center matrix, centersun. With centers fixed at Y and Z, calculate the center of each cluster. :

[0010] in, The vector represents the c-th cluster center, where Q is the number of sub-centers and q is the sub-center index; Given fixed centersun, center, and Z, calculate the membership degree of each pixel to the subclass center:

[0011] in, This represents the membership degree of the i-th pixel to the q-th subclass center. Indicates the first Vectors of subclass centers and ; With centersun, center, and Y fixed, calculate the membership degree of each subclass center to the cluster center using the following formula:

[0012] in, This represents the membership degree of the q-th sub-cluster center to the c-th cluster center. Indicates the first The vector of cluster centers and ; Calculate the objective function value:

[0013] Where J is the objective function value, Let be the weight of the i-th pixel. This is the adjustment coefficient; In each iteration, after sequentially updating the subclass center matrix centersun, the class center matrix center, the first fuzzy partition matrix Y, and the second fuzzy partition matrix Z, the objective function value of the DD-th iteration is calculated. Calculate the relative rate of change of the objective function value between two adjacent iterations. , For the objective function value of the DD-1th iteration, when If the number of iterations is less than the preset threshold, or the number of iterations exceeds the maximum number of iterations, When the threshold value is greater than or equal to the preset threshold, the fuzzy C-means clustering algorithm is determined to have converged and the iteration is terminated, finally yielding the chlorogenic acid hyperspectral cluster.

[0014] Preferably, obtaining the subclass center matrix includes the following steps: Initialize the cluster center matrix center by using the K-means++ algorithm to cluster the standardized two-dimensional matrix X, and use the resulting cluster centers as the initial values ​​of the cluster center matrix Center to avoid the instability of random initialization of cluster centers. With Y, Z, and center fixed, calculate the vector of the center of each subclass. The subclass center matrix centersun is obtained:

[0015] in, This vector represents the center of the q-th subclass, where i represents the pixel index. This represents the complete spectral curve of the i-th pixel. This represents the membership degree of the i-th pixel to the q-th subclass center. Here, r is the adjustment coefficient, and r is the fuzzy factor, which defaults to 2. This represents the membership degree of the q-th subclass center to the c-th cluster center. Let be the vector of the c-th cluster center, where C is the number of clusters; Calculate the vector of the center of each subclass one by one. All Arranged by rows, the final subclass center matrix is ​​obtained as centersun.

[0016] Preferably, the step of extracting features from the chlorogenic acid hyperspectral clusters to obtain chlorogenic acid spectral cluster feature vectors includes the following steps: For each cluster in the chlorogenic acid hyperspectral cluster, calculate its position within the chlorogenic acid characteristic band range. The pixel ratio, where DKEY is the number of keybands, and dkey is the index of the keyband, in each band If the reflectance values ​​are divided into B bins, then the c-th cluster in the band... The pixel ratio is:

[0017] in, Indicates that cluster c is in the band The reflectance of the pixels in the b-th box is the proportion of the pixel reflectance. This represents the number of pixels in cluster c. Represents pixels ( In the band absolute reflectivity, This represents the maximum reflectivity. () is an indicator function, when When it equals b, The value is 1 if it is 1, otherwise it is 0. Here, b is the floor function, and b is the histogram bin index. Calculate the statistical measure of the pixel proportion in the characteristic band for each chlorogenic acid hyperspectral cluster: mean. ,variance skewness ,energy ; Finally, the color histogram features are obtained. =[ , , , ]; The gray-level co-occurrence matrix of the chlorogenic acid hyperspectral cluster is calculated, thus obtaining the probability matrix; then, based on the normalized probability matrix, the probability matrix for each direction is calculated. The four core directional moments include the second-order moment, the moment of inertia, the correlation, and the entropy. Finally, the four core directional moments of each of the four directions are concatenated in sequence to form a texture feature with a dimension of 16. ; For each chlorogenic acid hyperspectral cluster, the average spectral curve was subjected to continuum removal processing: for local maxima of the average spectral curve, cubic spline interpolation was used to fit a continuum envelope, and the average spectral curve was divided by this continuum envelope to obtain the normalized spectrum after continuum removal. Based on the normalized spectrum, four key absorption parameters were extracted, including absorption depth, absorption width, absorption area, and absorption symmetry. Finally, absorption depth, absorption width, absorption area, and absorption symmetry were used as the spectral absorption characteristics of each chlorogenic acid hyperspectral cluster. ; Finally, the color histogram features, texture features, and spectral absorption features are combined to obtain the chlorogenic acid hyperspectral cluster feature vector. =[ , , ].

[0018] Preferably, the step of obtaining the hyperspectral cluster feature weights by calculating the mutual information of the chlorogenic acid hyperspectral cluster feature vectors includes the following steps: Collect chlorogenic acid samples and perform three key operations on each sample i in sequence: accurately measure the overall chlorogenic acid concentration using high-performance liquid chromatography. The chlorogenic acid sample was scanned using hyperspectral imaging technology to obtain hyperspectral images. Then, in the clustering and feature extraction stage, the chlorogenic acid hyperspectral images of each sample were spatially segmented using a fuzzy C-means clustering algorithm based on several sub-cluster centers to form several clusters. The chlorogenic acid hyperspectral cluster feature vector of each cluster was obtained, and all the chlorogenic acid hyperspectral cluster feature vectors were summed and averaged to finally obtain the hyperspectral feature vector of the sample to be tested. =[ , , ], This represents the color histogram features of the sample to be tested. This represents the texture features of the sample to be tested. Indicates the spectral absorption characteristics of the sample to be tested; For the hyperspectral feature vector of the sample to be tested Each feature and chlorogenic acid concentration Mutual information:

[0019] in, The hyperspectral feature vector of the sample to be tested The h-th feature and chlorogenic acid concentration The mutual information is given by K, which represents the number of bins for the h-th feature in the hyperspectral feature vector of the sample to be tested, and L, which is the number of bins for the chlorogenic acid concentration. This represents the joint probability of the k-th bin in the hyperspectral feature vector of the sample under test with the l-th bin of chlorogenic acid concentration. Let h be the marginal probability of the h-th feature in the hyperspectral feature vector of the sample to be tested in the k-th bin. This represents the marginal probability of chlorogenic acid concentration in the l-th bin; Finally, the color histogram feature weights, texture feature weights, and spectral absorption feature weights are obtained: ; ; ; in, For the feature weights of the color histogram, For texture feature weights, The weights are for spectral absorption characteristics.

[0020] Preferably, the weighted vector of the chlorogenic acid hyperspectral cluster is: After Z-score normalization of the chlorogenic acid hyperspectral cluster feature vector, a weighted vector of the chlorogenic acid hyperspectral cluster is obtained by combining the chlorogenic acid hyperspectral cluster feature vector and the hyperspectral cluster feature weights. =[ , , ].

[0021] Preferably, the step of inputting the chlorogenic acid hyperspectral cluster weighted vector into a high-dimensional multi-scale support vector machine regression model to output the predicted chlorogenic acid concentration includes the following specific steps: Based on the distribution characteristics of the weighted vector of chlorogenic acid hyperspectral clusters, a trapezoidal fuzzy membership function is introduced for the itrain training sample. Assign membership degree ; Construct a multi-scale kernel function to map different types of features in the weighted vector of chlorogenic acid hyperspectral clusters:

[0022] in, , Representing the first and the The comprehensive feature vector of each chlorogenic acid training sample. , The kernel function weights satisfy the following conditions: + =1, used to balance the contribution of local and global features; For Gaussian kernel scaling parameters, It is the dot product of vectors; This is the polynomial kernel offset, which defaults to 1. d is the polynomial kernel degree, which controls the fitting order of the global features. The objective function of the support vector machine model is:

[0023] in, Let b be the weight vector, and b be the offset. and As slack variables, For regularization parameters, For sample membership degree, The total number of samples in the training set is denoted by , and ittrain is the index of the training samples in the training set. The support vector machine model is trained using the training set to obtain a trained support vector machine model. For each weighted vector of the chlorogenic acid hyperspectral cluster to be tested, the weighted vector of the chlorogenic acid hyperspectral cluster to be tested is input into the trained high-dimensional multi-scale support vector machine regression model, and the predicted value of chlorogenic acid concentration is output.

[0024] in, This represents the predicted chlorogenic acid concentration for the c-th hyperspectral cluster. Let m be the number of support vectors, and m be the index of the support vector. and For Lagrange multipliers, The weighted vector of the chlorogenic acid hyperspectral cluster of the sample to be tested. This is the weighted vector for the chlorogenic acid hyperspectral cluster corresponding to the m-th support vector. () represents the multi-scale kernel function. This represents the offset of the support vector machine model.

[0025] Preferably, the membership degree is:

[0026] in, The Euclidean distance from the comprehensive feature vector of the i-th chlorogenic acid training sample to the centroid of the training set, where the centroid of the training set is the mean center of all comprehensive feature vectors in the training set, is used as an indicator to measure the importance of the sample. The 10th percentile of the distances from all composite feature vectors in the training set to the centroid of the training set. The 25th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set. The 75th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set. It is the 90th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set.

[0027] Preferably, the specific steps for calculating concentration outliers using the sliding window method are as follows: A chlorogenic acid concentration map is plotted based on the predicted chlorogenic acid concentration. The size and number of sliding windows are set. For each window w, the average concentration within the window is calculated based on the number of windows. and window concentration standard deviation Then, the overall average concentration of chlorogenic acid in the concentration map is calculated based on the number of windows. and overall concentration standard deviation Finally, the concentration outliers within the window are calculated:

[0028] in, For the concentration outlier in the w-th window, Let w be the average concentration of the w-th window. Let w be the standard deviation of the concentration in the w-th window. This represents the overall average concentration of chlorogenic acid in the concentration graph. The overall concentration standard deviation of the chlorogenic acid concentration plot. The first adjustment coefficient, This is the second adjustment coefficient.

[0029] This invention provides a method for detecting chlorogenic acid by integrating deep learning, involving machine learning and deep learning technologies, which has the following beneficial effects: (1) The fuzzy C-means clustering algorithm is the core foundation for processing chlorogenic acid hyperspectral data, and its significance lies in accurately realizing the spatial segmentation of hyperspectral images. This algorithm eliminates the influence of dimensions through Z-score normalization, and combined with the iteratively optimized fuzzy partitioning matrix and class center matrix, it can effectively capture the spectral differences between different pixels and divide the hyperspectral data into clusters with similar features. Its termination criterion based on the convergence of the objective function ensures the stability and reliability of the clustering results, providing a clear partitioning basis for subsequent targeted extraction of chlorogenic acid-related features and avoiding interference from irrelevant information.

[0030] (2) The design based on several sub-centers aims to improve the precision and anti-interference ability of clustering. By setting the number of sub-centers Q to be greater than the number of clusters C, the calculation of the sub-center matrix takes into account the relationship between pixels and sub-centers, and between sub-centers and cluster centers. Combined with the adjustment coefficient α to balance the fitting error and aggregation error, it effectively avoids the instability caused by the random initialization of cluster centers in traditional clustering. This design breaks through the single-level mapping mode of "pixel-cluster center" in the existing technology and designs a two-level architecture of "pixel-sub-center-cluster center". This multi-level clustering structure can more meticulously depict the inherent distribution characteristics of hyperspectral data, reduce the impact of noise on the clustering results, and retain richer effective information for subsequent feature extraction.

[0031] (3) The support vector machine regression model is used to predict chlorogenic acid concentration by establishing a reliable mapping relationship between hyperspectral features and chlorogenic acid concentration. This model combines feature weights based on mutual information to enhance the contribution of features strongly correlated with concentration. By optimizing hyperparameters through grid search and 5-fold cross-validation, the prediction accuracy is significantly improved. At the same time, the trapezoidal fuzzy membership function introduced by the model weakens the influence of outliers, and the dual-weight penalty mechanism adapts to the changes in data distribution of newly added samples, avoiding repeated training on the full dataset. This balances the accuracy and efficiency of prediction, providing strong support for quantitative concentration analysis.

[0032] (4) The high-dimensional multi-scale design endows the support vector machine regression model with stronger adaptability and practicality. Its core significance lies in comprehensively mining the multi-dimensional features of hyperspectral data and optimizing the model's adaptability. Combining Gaussian kernel (capturing local features) and polynomial kernel (capturing global features) enhances the model's adaptability to high-dimensional data and achieves more flexible fitting. The multi-scale kernel function, through the combination of Gaussian kernel and polynomial kernel, takes into account both the fine capture of local spectral features and the overall fitting of global features, thereby improving the flexibility of concentration prediction. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart of the steps of a chlorogenic acid component detection method integrating deep learning proposed in this invention; Figure 2 This is a step hierarchy diagram of obtaining the chlorogenic acid hyperspectral cluster weighted vector in a chlorogenic acid component detection method integrating deep learning proposed in this invention; Figure 3 This is a step hierarchy diagram of obtaining concentration anomalies in a chlorogenic acid component detection method that integrates deep learning proposed in this invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] Please see Figures 1-3 The present invention provides a technical solution: a method for detecting chlorogenic acid by integrating deep learning.

[0037] Step S1: Take a picture of the chlorogenic acid sample using a hyperspectral imaging device to obtain a hyperspectral image of chlorogenic acid.

[0038] After grinding the chlorogenic acid sample (such as coffee beans) and passing it through an 80-mesh sieve, take 2g of powder and press it into a uniform thin film with a diameter of 10mm and a thickness of 2mm using a tablet press (pressure 5MPa, holding pressure for 30s). For medicinal samples (such as honeysuckle): after crushing, spread it evenly on a quartz glass slide (thickness 1mm) and press it firmly with a coverslip (to avoid air gaps). Fix the pretreated sample on the plane of the electric translation stage, ensuring that the center of the sample is aligned with the center of the camera's field of view and that the sample covers ≥90% of the field of view.

[0039] Turn on the adjustable-power halogen lamp light source. Set the initial incident angle based on the sample surface roughness: 45° for rough samples (e.g., ground coffee beans) and 60° for smooth samples (e.g., compressed sheets), both using an angle calibrator (accuracy ±0.5°). Start the cooled hyperspectral camera and set its spectral scanning range to 300 nm to 1100 nm, covering the characteristic absorption peak of chlorogenic acid (327 nm) and the absorption range of the sample matrix (e.g., carbohydrates, cellulose), ensuring a direct correlation between spectral data and chlorogenic acid concentration. Set the initial exposure time based on the camera noise curve (provided by the manufacturer). In the 327 nm characteristic band, with a signal-to-noise ratio ≥50:1 as the standard, set the initial exposure time to 15 ms (for samples with reflectivity ≥30%) or 25 ms (for samples with reflectivity <30%). Place a diffuse reflection reference white board (e.g., Spectralon material) next to the sample on the stage as a reference for illumination calibration.

[0040] Subsequently, adaptive illumination equalization adjustment was performed. The incident angle of the halogen lamp source was fine-tuned from the initial incident angle using an electronically controlled rotary stage, in 5-degree increments, ranging from 30 to 60 degrees. After each angle adjustment, a hyperspectral camera acquired a preview image of a reference white board at the 550 nm band (a representative band). The overall grayscale mean and standard deviation of this preview image were calculated. Simultaneously, the camera's exposure time was finely adjusted in 1-millisecond increments within the range of 10 to 50 milliseconds. This adjustment process is an iterative optimization process: if the average grayscale value of the whiteboard image is lower than the target range, the exposure time is increased. For example, for a 12-bit camera, the target can be set to 2000-3000 counts. This range is determined based on the standard reflectance of the Spectralon whiteboard (400-1100nm and reflectance ≥95%) and the camera's linear influence range (0-3500 counts). If the image grayscale standard deviation is too large (e.g., higher than 1% of the average), the light source angle is first fine-tuned (in 5° increments). If adjusting the angle to 30° or 60° still does not meet the requirements, the light source power is fine-tuned (in 5% increments, power range 50%-100%, to avoid overheating and affecting the stability of the light source). Iterative optimization is performed a maximum of 10 times. If the mean and standard deviation indicators are not met within 10 iterations, the device's optical path (e.g., whether the lens is clean or the light source is aged) needs to be checked and reinitialized to avoid invalid iterations.

[0041] Next, polarization filtering was initiated to eliminate specular reflection. Maintaining the optimized lighting conditions, the linear polarizer mounted in front of the halogen lamp light source was rotated to the 0-degree position. Subsequently, the linear polarizer mounted in front of the hyperspectral camera lens was rotated in 10-degree increments from 0 to 180 degrees. At each analyzer angle, the camera acquired a single-band image. By comparing the suppression effect on highlight areas in the images at different angles, it was determined that specular reflection from the sample surface was maximally suppressed when the analyzer was rotated orthogonal to the polarizer (i.e., an angle difference of 90 degrees ± 5 degrees). The analyzer was then fixed at this optimal angle.

[0042] With optimized parameter settings, a full-band scan of the diffuse reflection reference whiteboard is first performed to obtain the original spectral signal value of the whiteboard, Cube_white_raw( , ), and calculate the signal average value of all pixels: Cube_white_raw_avg( The white reference plate was then used as the calibration benchmark. Subsequently, the white reference plate was removed, and while keeping all parameters strictly constant, the sample to be tested was placed at the center of the field of view and scanned across the entire spectral band to obtain the raw spectral signal value of the chlorogenic acid sample, Cube_sample_raw( , ).

[0043] The hyperspectral image of chlorogenic acid is as follows:

[0044] in, This represents a hyperspectral image of chlorogenic acid, where the value is the number of pixels in the spatial region ( ). ),wavelength The reflectance at a certain point (ranging from 0 to 1, without units) reflects the sample's ability to reflect light of that wavelength—and the content of chlorogenic acid directly affects the reflectance. This indicates that the chlorogenic acid sample is in ( ) pixels The original spectral signal value corresponding to the wavelength, Indicates the diffuse reference whiteboard at The average value of the original signal of all spatial pixels at the specified wavelength. Indicates the diffuse reference whiteboard at Standard reflectance at wavelength.

[0045] It should be noted that the diffuse reference whiteboard is in Standard reflectance at wavelength Use the standard reflectance curve provided by the diffuse reflectance reference plate manufacturer (the manufacturer's calibration certificate number SpectralonSRM-99-01 is required). If the curve is missing, it needs to be calibrated by yourself using a UV-Vis-NIR spectrophotometer (accuracy ±0.5%). The calibration wavelength interval should be consistent with that of the hyperspectral camera (e.g., 1 nm / step).

[0046] Finally, a hyperspectral image of chlorogenic acid was obtained. .

[0047] Step S2: Cluster the chlorogenic acid hyperspectral image using a fuzzy C-means clustering algorithm based on several sub-centers to obtain chlorogenic acid hyperspectral clusters.

[0048] The chlorogenic acid hyperspectral image obtained in step S1 Transform into a two-dimensional matrix X∈ Where: N is the total number of pixels in the hyperspectral image of chlorogenic acid, and D is the number of spectral bands (the number of spectral bands D needs to be filtered for "chlorogenic acid characteristic correlation bands": the correlation coefficient between each band d and the standard concentration of chlorogenic acid is calculated using the Pearson correlation coefficient method, and bands with an absolute value of correlation coefficient > 0.6 are retained. Example: after filtering in the 300-1100nm range, D=50, containing the chlorogenic acid characteristic peak at 327nm). Each row in the matrix It is a D-dimensional vector representing the reflectance value of a single pixel across all wavelengths, i.e., the complete spectral curve of that pixel. This vector... The d-th element It represents the center wavelength of the d-th band. The reflectance value at that location.

[0049] The two-dimensional matrix X is Z-score normalized to eliminate the influence of dimensions (first, 3 is applied to each band d). Criteria are used to filter out extreme values, and then the elements in the two-dimensional matrix X are processed. Z-score standardization: The calculation results are reassigned to This achieves standardization for each element. Let be the average reflectance of all pixels in the d-th band. Let be the standard deviation of pixel reflectance in the d-th band. Initialize the number of clusters C for the fuzzy C-means clustering algorithm (using the time-part rule: run the K-means clustering method, C traverses 2-8, calculate the sum of squares (WSS) within each cluster corresponding to C. When the rate of decrease of WSS with increasing C slows down significantly, i.e., the elbow point, C is the optimal value. Example: coffee bean sample C=3, honeysuckle sample C=4.), and the number of sub-centers Q, Q>C (Q can be set to C+2 to balance anti-interference ability and computational load, and Q ≤ C must be satisfied). To avoid excessive sub-centers leading to data sparsity; for example, Q=5 when C=3, and Q=6 when C=4. The fuzzy factor r (which iterates through [1.5, 2.5] with a step size of 0.2, calculating the silhouette coefficient of the clustering result for each r, with a value of [-1, 1], the closer to 1 the better the clustering effect, the larger the r coefficient is selected). The fuzzy C-means clustering algorithm optimizes the sub-center matrix, fuzzy partition matrix, and class center matrix through alternating bagging, ultimately achieving clustering.

[0050] Initialize the first fuzzy partitioning matrix Y∈ N represents the total number of hyperspectral imaging pixels. First, the standardized two-dimensional matrix X is clustered using the K-means algorithm, with the number of clusters set to the number of sub-centers Q, to obtain the initial positions of the Q sub-centers. Subsequently, based on the distance relationship between pixel spectral data and the initial subclass centers, the initial value of the first fuzzy partitioning matrix Y is calculated. Specifically, the initial membership degree of the i-th pixel to the center of the q-th subclass. The following formula ensures that membership degree is correlated with spectral similarity and satisfies the constraints of fuzzy partitioning:

[0051] in, This represents the membership degree of the i-th pixel to the q-th subclass center. The initial vector that identifies the center of the q-th subclass. Indicates the first The initial vector of each subclass center and .

[0052] Next, using the Q subclass centers obtained above... As input data, the K-means algorithm is used again for clustering, with the number of clusters set to the final number of clusters C, to obtain the initial positions of the C cluster centers. (That is, the initial value of the class center matrix). Similarly, based on the distance relationship between the subclass centers and the initial cluster centers, the initial value of the second fuzzy partitioning matrix Z is calculated. The initial membership degree of the q-th subclass center to the c-th cluster center. The following formula is used to calculate that it also satisfies the constraints of fuzzy partitioning. (≥0 and the sum of elements in each row is 1):

[0053] in, This represents the membership degree of the q-th sub-cluster center to the c-th cluster center. Indicates the first The initial vector of each cluster center This represents the initial vector for the c-th cluster center. and .

[0054] Initialize the class center matrix center∈ The K-means++ algorithm is used to cluster the standardized data matrix X (the number of clusters is C). In the first step, one pixel is randomly selected as the initial cluster center. In the subsequent steps, each cluster center is selected from the remaining pixels (the one farthest from the selected cluster center) to reduce randomness. At the same time, K-means++ is run 10 times to select the cluster center with the smallest WSS as the initial value of center, thus avoiding the instability of random initialization of cluster centers.

[0055] With Y, Z, and center fixed, calculate the vector of the center of each subclass. The subclass center matrix is ​​obtained as centersun∈ :

[0056] in, This vector represents the center of the q-th subclass, where i represents the pixel index. This represents the complete spectral curve of the i-th pixel. This represents the membership degree of the i-th pixel to the q-th subclass center. This is the adjustment coefficient, the same as the adjustment coefficient in the objective function value formula; r is the fuzzy factor, which defaults to 2. This represents the membership degree of the q-th subclass center to the c-th cluster center. Let be the vector of the c-th cluster center, where C is the number of clusters.

[0057] Calculate the vector of the center of each subclass one by one. All Arranged by row, the final subclass center matrix is ​​obtained as centersun∈ Q represents the number of subclass centers, and D represents the number of spectral bands.

[0058] With centers fixed at Y and Z, calculate the center of each cluster. :

[0059] in, Let Q be the vector representing the c-th cluster center, Q be the number of sub-centers, and q be the sub-center index.

[0060] Given fixed centersun, center, and Z, calculate the membership degree of each pixel to the subclass center:

[0061] in, This represents the membership degree of the i-th pixel to the q-th subclass center. Indicates the first Vectors of subclass centers and .

[0062] With centersun, center, and Y fixed, calculate the membership degree of each subclass center to the cluster center using the following formula:

[0063] in, This represents the membership degree of the q-th sub-cluster center to the c-th cluster center. Indicates the first The vector of cluster centers and .

[0064] It should be noted that the calculation and At that time, if =0 (or If the denominator term is 0, then let the denominator term be... (To avoid infinity, the minimum value is set to 1, and the remaining memberships are set to 0 (to ensure...) or ).

[0065] Calculate the objective function value:

[0066] Where J is the objective function value, Let be the weight of the i-th pixel. This is the adjustment coefficient.

[0067] It should be noted that the adjustment coefficient The clustering granularity and model complexity are controlled by balancing the fitting error from data points to subclass centers and the aggregation error from subclass centers to cluster centers in the objective function. The optimization needs to be determined empirically through grid search or cross-validation based on the specific characteristics of the dataset. Typically, the initial value is set in the range of 1-10, with a step size of 0.5. Calculate the Davies-Bouldin index (DBI) of the clustering results (the smaller the DBI, the higher the intra-cluster similarity and the greater the inter-cluster difference), and select the one with the smallest DBI. Example: Coffee bean sample =3.5 (DBI=1.2), honeysuckle sample core =4 (DBI=1.15). Fine-tuning was performed using clustering evaluation metrics (such as OA, AA, and Kappa coefficient) to avoid overfitting or underfitting, ultimately selecting the value that ensures stable convergence and optimal clustering performance.

[0068] It should be noted that the weight of the i-th pixel... Based on the reflectance of chlorogenic acid in the 327nm characteristic band, the importance weight of each effective pixel is calculated: =1- The higher the chlorogenic acid content, the lower the 327nm reflectance. The larger.

[0069] In the iterative optimization process of the fuzzy C-means clustering algorithm based on several sub-centers, the convergence of the objective function value is the basis for terminating the algorithm. The maximum number of iterations is set to 200. Specifically, after each complete iteration (updating the sub-center matrix centersun, the sub-center matrix center, the first fuzzy partition matrix Y, and the second fuzzy partition matrix Z in sequence), the current objective function value is calculated. (The objective function value of the DD-th iteration), which is composed of the weighted distance from the data point to the subclass center and the weighted distance from the subclass center to the cluster center, adjusted by α; then the relative rate of change of the objective function value between two adjacent iterations is calculated. ,when Less than a preset threshold (take 10 groups of chlorogenic acid samples, 5 groups of coffee beans, and 5 groups of honeysuckle, and record the change curve of J with the number of iterations: when the number of iterations is ≥150, the relative change rate of J for all samples is less than 150). When, the preset threshold can be set to When the algorithm reaches a certain threshold, it is determined that the algorithm has converged and the iteration is terminated. If the preset threshold is not met after 200 iterations, the iteration is terminated and the clustering is not fully converged (data quality or parameter settings need to be checked). Finally, the chlorogenic acid hyperspectral cluster is obtained.

[0070] It should be noted that step S2 employs an improved fuzzy C-means clustering method based on several sub-centers to achieve accurate, stable, and interference-resistant spatial segmentation of chlorogenic acid hyperspectral data. This breaks through the existing single-level mapping mode of "pixel-cluster center" and designs a two-level architecture of "pixel-sub-center-cluster center." A noise isolation layer is constructed through sub-centers to prevent abnormal pixels from directly interfering with cluster center calculation, thus solving the problem of easy cluster center shift in traditional methods. Initialization strategy optimization: Abandoning the blindness of existing techniques' "random initialization" or "traditional K-means initialization," the K-means++ algorithm is used to optimize cluster center initialization (distance maximization principle + multi-round iterative selection). Combined with initialization of the membership matrix related to spectral similarity, this avoids iteration getting trapped in local optima from the source, improving the reproducibility of results. Objective function reconstruction: Breaking the constraint logic of existing techniques' "minimizing a single fitting error," a dual-error objective function of "fitting error + aggregation error" is constructed. This is achieved by adjusting the coefficients... (Optimized by grid search + DBI index) Dynamically balances local pixel fitting accuracy with global subclass center aggregation robustness, forcing cluster centers to fall in the core area of ​​subclass center distribution, avoiding being pulled off course by individual abnormal pixels.

[0071] Step S3: Extract features from the chlorogenic acid hyperspectral clusters to obtain the chlorogenic acid spectral cluster feature vectors.

[0072] Feature extraction is performed on the chlorogenic acid hyperspectral clusters to obtain the chlorogenic acid spectral cluster feature vector, which includes color histogram features, texture features and spectral absorption features.

[0073] For each cluster in the chlorogenic acid hyperspectral cluster, calculate its position within the chlorogenic acid characteristic band range. The pixel ratio, where DKEY is the number of keybands, and dkey is the index of the keyband, in each band The reflectance values ​​are divided into B bins (the cluster c is calculated in the above). standard deviation of reflectance ,like If the reflectance is greater than 0.2, the reflectance fluctuates greatly; therefore, B = 64 is chosen. If the reflectance is less than 0.1... <0.2, take B=32; if <0.1, take B=16 to ensure pixels are uniformly distributed within the bin. ), then the c-th cluster is in band. The pixel ratio is:

[0074] in, Indicates that cluster c is in the band The reflectance of the pixels in the b-th box is the proportion of the pixel reflectance. This represents the number of pixels in cluster c. Represents pixels ( In the band absolute reflectivity, This represents the maximum reflectivity. () is an indicator function, when When it equals b, The value is 1 if it is 1, otherwise it is 0. is the floor function, and b is the histogram bin index.

[0075] It should be noted that the band index has been changed from the global band index d (representing all D bands in the full spectrum range of 400-1100nm) to the key band index dkey (selecting DKEY characteristic bands closely related to chlorogenic acid, calculating the Pearson correlation coefficient between the band and the concentration of chlorogenic acid, and retaining bands with a Pearson correlation coefficient greater than 0.6 with the concentration of chlorogenic acid, such as the band from 320nm to 350nm). This change stems from the stage difference of the algorithm's objectives: the clustering stage (step S2) needs to use full-band spectral information to capture the overall differences between pixels and ensure the accuracy of clustering; while the feature extraction stage (step S3) focuses on the characteristic absorption bands of chlorogenic acid, improving the feature discrimination power by eliminating irrelevant band noise, while reducing the computational complexity, thus retaining key information and enhancing practicality.

[0076] Calculate the statistical measure of the pixel proportion in the characteristic band for each chlorogenic acid hyperspectral cluster: mean. ,variance skewness ,energy .

[0077] Finally, the color histogram features are obtained. =[ , , , Among them, the mean Reflects cluster c in a specific band The mean value represents the average reflectance level (i.e., the overall brightness level). It indirectly indicates the chlorogenic acid content: a higher mean value indicates a higher average reflectance, potentially corresponding to areas with lower chlorogenic acid content (because chlorogenic acid absorbs light, leading to reduced reflectance); conversely, a low mean value may indicate areas with high chlorogenic acid content; variance... The variance reflects the dispersion or heterogeneity of reflectance distribution. A large variance indicates significant variations in reflectance values ​​within a cluster, potentially including mixed structures or noise, such as highlights or shadows; a small variance indicates uniform material and stable chlorogenic acid distribution. Skewness Reflecting the asymmetry of the distribution, positive skewness indicates a longer right tail with more high reflectance values, potentially corresponding to regions with high reflectance; negative skewness indicates a longer left tail with more low reflectance values, potentially corresponding to chlorogenic acid-rich regions with strong absorption. Skewness can capture subtle differences in distribution shape, enhancing feature discrimination; energy Reflecting the uniformity or concentration of the distribution, a high energy value indicates that the reflectivity is highly concentrated in a few boxes (with sharp distribution peaks); a low energy value indicates that the distribution is dispersed, and a high energy value may correspond to a pure region with a uniform chlorogenic acid content.

[0078] Grayscale images of clusters are generated by selecting representative bands, and texture orientation moments (reflecting the uniformity and directionality of the texture) are calculated using the gray-level co-occurrence matrix to form texture feature vectors. From the hyperspectral clusters obtained in step S2, for each cluster c, a representative band that balances brightness and detail is first selected (e.g., =327nm), extract the grayscale image of this cluster. The gray value corresponding to each pixel spatial coordinate (u,v) in the image. (u,v) represents the pixel coordinates in the representative band. absolute reflectance at (u,v, ), and only pixels within cluster c are retained; subsequently, the grayscale image is processed. Gray-level compression is performed, mapping the reflectance range of [0,1] to M gray levels (M is set according to the reflectance fluctuation range, when...). When ≥0.2, For cluster c in The standard deviation of reflectance at a given location is given. Reflectance fluctuates greatly; M can be taken as 32. When 0.1 ≤ ≤0.2, moderate fluctuation, M can be 16, when When the value is less than 0.1, the fluctuation is small (M can be 8). During mapping, a floor function is used, which is to multiply the reflectance value by (M-1), round it down, and then add 1. The resulting gray value g falls between 1 and M.

[0079] Where g is the compressed grayscale value, which is an integer ranging from 1 to M, and M is the grayscale level.

[0080] Set the pixel pitch d (d=1 pixel recommended) and 4 main directions. (0°, 45°, 90°, 135°) to cover textures in different directions, constructing a gray-level co-occurrence matrix. The matrix is ​​an M×M square matrix, where the elements are... The calculation formula is:

[0081] in, This indicates that cluster c has a pixel spacing of d and a direction Below, the grayscale value starts from... arrive symbiotic frequency, and These represent the gray levels of the center pixel and its adjacent pixels, respectively. and These are the horizontal offset vector and the vertical offset vector, respectively, determined by direction. Decide, For direction The number of effective pixel pairs, Height is the height of the hyperspectral image, and Width is the width of the hyperspectral image. This is the logical AND operator.

[0082] It should be noted that when When =0°, =d, =0; when At 45°, =d, =d; when When =90°, =0, =d; when At 135°, =-d, =d.

[0083] Dividing each element of the gray-level co-occurrence matrix by the sum of all elements in the matrix yields the normalized probability matrix. First, ensure that the sum of all elements in the matrix is ​​1; then, based on the normalized probability matrix, calculate the probability for each direction. The four core directional moments (second-order moment, moment of inertia, correlation, and entropy) yield a core directional moment with a dimension of 4×4=16. Finally, the four core directional moments for each of the four directions are concatenated sequentially, starting with the four core directional moments at 0°, followed by the four core directional moments at 45°, 90°, and 135°, ultimately forming a clustered texture feature with a dimension of 4×4=16. .

[0084] For each chlorogenic acid hyperspectral cluster c, calculate its position in the chlorogenic acid characteristic band range. Average spectral curve on:

[0085] in, Indicates that cluster c is in the band The average reflectance, This represents the number of pixels in cluster c. Represents pixels ( In the band The absolute reflectance.

[0086] For each chlorogenic acid hyperspectral cluster, the average spectral curve was subjected to continuum removal processing: for local maxima of the average spectral curve (requiring a maximum value within a 5-7 band neighborhood), cubic spline interpolation was used to fit a continuum envelope. The average spectral curve was then divided by this continuum envelope to obtain the normalized spectrum after continuum removal. Based on the normalized spectrum, four key absorption parameters were extracted: absorption depth, absorption width, absorption area, and absorption symmetry. Finally, absorption depth, absorption width, absorption area, and absorption symmetry were used as the spectral absorption characteristics of each chlorogenic acid hyperspectral cluster. .

[0087] It should be noted that the absorption depth is 1 minus the global minimum value in the normalized spectral curve; the absorption width is obtained by determining the left and right boundary points of the absorption peak at half the height of the absorption depth using linear interpolation, and then calculating the wavelength difference between the two points; the absorption area is obtained by integrating the normalized generalized area over the characteristic band interval using the trapezoidal numerical integration method to capture the total energy of the entire absorption peak; the absorption symmetry is quantified by first determining the absorption peak position and the centroid position of the absorption peak, and then evaluating the ratio of the absolute offset between the two to the absorption width.

[0088] Finally, the color histogram features, texture features, and spectral absorption features are combined to obtain the chlorogenic acid hyperspectral cluster feature vector. =[ , , ].

[0089] Step S3 extracts multi-dimensional features from the chlorogenic acid hyperspectral clusters output in step S2 to generate chlorogenic acid hyperspectral cluster feature vectors. The vector includes three types of features, including color histogram features. Based on the reflectance distribution of key bands, calculate statistics (mean, variance, skewness, and magnitude) to describe the macroscopic spectral characteristics of the cluster. Texture features. Based on grayscale images of representative bands, directional moments (such as second-order moments and moments of inertia) are calculated using the gray-level co-occurrence matrix, describing spatial structure information and spectral absorption characteristics. Based on the average spectral curve, absorption parameters (depth, width, area, symmetry) are extracted to describe the spectral absorption characteristics. Traditional methods typically use full-band optical data or simple band ratios for feature extraction, without spatial distribution for clusters, resulting in features containing irrelevant noise and high computational complexity. Texture features are often based on RGB images or single bands, lacking adaptive band selection and failing to effectively capture microstructural differences related to chlorogenic acid. This paper focuses on key bands: by screening feature bands closely related to chlorogenic acid and excluding irrelevant bands, the feature dimensionality is reduced, and computational efficiency is improved. Combining color, texture, and spectral absorption features, the characteristics of clusters are described from multiple perspectives, avoiding the limitations of single features.

[0090] Step S4: Obtain the hyperspectral cluster feature weights by calculating the mutual information of the chlorogenic acid hyperspectral cluster feature vectors; obtain the chlorogenic acid hyperspectral cluster weighted vector by combining the chlorogenic acid hyperspectral cluster feature vectors and the hyperspectral cluster feature weights.

[0091] First, prepare a sufficient number of chlorogenic acid samples to be tested (e.g., coffee beans, with a quantity greater than or equal to 30 to ensure statistical significance). For each sample i, perform three key operations in sequence: accurately measure its overall chlorogenic acid concentration using high-performance liquid chromatography. The chlorogenic acid samples were scanned using hyperspectral imaging technology to obtain hyperspectral images. Then, in the clustering and feature extraction stage, the hyperspectral images of each sample were spatially segmented using a fuzzy C-means clustering algorithm based on multiple sub-cluster centers to form several clusters. The chlorogenic acid hyperspectral cluster feature vector of each cluster was then obtained, and all chlorogenic acid hyperspectral cluster feature vectors were weighted and averaged according to the number of pixels. , Let c be the total number of pixels in the c-th cluster of the sample. Let C be the chlorogenic acid hyperspectral cluster feature vector of cluster c, where C is the number of clusters. This will ultimately yield the hyperspectral feature vector of the sample to be tested. =[ , , ].

[0092] For the hyperspectral feature vector of the sample to be tested Each feature and chlorogenic acid concentration Mutual information:

[0093] in, The hyperspectral feature vector of the sample to be tested The h-th feature and chlorogenic acid concentration The mutual information is given by K, which represents the number of bins for the h-th feature in the hyperspectral feature vector of the sample to be tested, and L, which is the number of bins for the chlorogenic acid concentration. This represents the joint probability of the k-th bin and the l-th bin of chlorogenic acid concentration in the hyperspectral feature vector of the sample to be tested. It is calculated as the ratio of the number of samples where the h-th feature belongs to the k-th bin and the chlorogenic acid concentration belongs to the l-th bin to the total number of samples, and satisfies the following conditions: , Let h be the marginal probability of the h-th feature in the k-th bin of the hyperspectral feature vector of the sample to be tested. It is calculated as the ratio of the number of samples in the k-th bin to which the h-th feature belongs, satisfying the following condition: , This represents the marginal probability of chlorogenic acid concentration in the l-th bin. It is calculated as the ratio of the number of samples with chlorogenic acid concentration belonging to the l-th bin to the total number of samples, and satisfies the following condition: .

[0094] It should be noted that the number of feature bins K needs to be determined based on the sample size. When the sample size is 30, the number of bins K is 10, and each bin contains... Sample size / 10 Each sample (≥30 samples, ≥3 samples per bin) meets the minimum sample size requirement for probability statistics (probability fluctuations are large when the number of samples <3). The number of bins L for chlorogenic acid concentration is usually 8. Chlorogenic acid concentration is usually concentrated in 0.2%-2.0% (coffee bean, honeysuckle samples), with a narrow distribution range. L=8 is sufficient to distinguish the gradient of "low concentration (0.2%-0.5%) to medium concentration (0.5%-1.2%) to high concentration (1.2%-2.0%)". If the number of samples in a bin is <3 due to a small sample size, such as 25 or uneven concentration distribution, adjacent bins need to be merged: starting from the bin with the fewest samples, merge adjacent bins to the left or right (prioritize merging bins with the same few samples) until the number of samples in all bins is ≥3.

[0095] It should be noted that, =[ , , In ], each subset (such as It does indeed contain multiple independent features (e.g., the mean, variance, skewness, energy, etc. of the color histogram). To calculate the overall mutual information of each subset and use it for weight normalization, the average redundancy is calculated for each subset. : ,in, This represents the number of combinations of features within a subset (e.g., 4 features). =6), This represents the Pearson correlation coefficient between the h1-th and h2-th features within the subset. The closer the value is to 1, the more severe the feature redundancy within the subset. The mutual information of all features within the subset is summed. The sum of mutual information of the subsets is = , then calculate Effective mutual information of subsets Similarly, The effective mutual information of the subset is , The effective mutual information of the subset is The weights of each subset are obtained through normalization.

[0096] Finally, the color histogram feature weights, texture feature weights, and spectral absorption feature weights are obtained: ; ; ;

[0097] in, For the feature weights of the color histogram, For texture feature weights, The weights are for spectral absorption characteristics.

[0098] It should be noted that after completing the feature weight calculation based on mutual information, the obtained weight values ​​( , , The weights are a holistic measure derived from global sample data (the global form of the chlorogenic acid test samples), reflecting the average correlation strength between different feature types (such as color histograms, texture orientation moments, and spectral absorption parameters) and chlorogenic acid concentration. The weights are calculated based on the statistical properties of all samples, rather than relying on local data from specific clusters (the clusters of chlorogenic acid test samples), thus possessing global consistency and transferability. Features of the same type within each cluster (e.g., color histogram features across all clusters) share the same weight value, ensuring consistent weight application. Therefore, the hyperspectral cluster feature weights include… , , .

[0099] By combining the chlorogenic acid hyperspectral cluster feature vector and the hyperspectral cluster feature weights, and performing Z-score normalization on each feature subset, a weighted vector of the chlorogenic acid hyperspectral cluster is obtained. =[ , , ].

[0100] It should be noted that traditional methods typically use fixed weights (e.g., based on experience or simple statistics) or principal component analysis (PCA) for feature dimensionality reduction, but these methods cannot dynamically adjust feature contributions and may lose information strongly correlated with concentration. Mutual information calculation often fails to consider feature redundancy and binning strategies, leading to inaccurate weights. This approach dynamically calculates weights based on mutual information, strengthening the contribution of features strongly correlated with concentration and weakening the influence of irrelevant features. The weights are calculated based on all training samples, ensuring transferability and guaranteeing that similar features from different clusters share the same weights, thus improving consistency. The mutual information calculation employs a data-driven binning strategy, reducing discretization bias and improving the accuracy of MI estimation.

[0101] Step S5: Input the weighted vector of the chlorogenic acid hyperspectral cluster into the support vector machine regression model based on high-dimensional multi-scale, and output the predicted value of chlorogenic acid concentration; construct a chlorogenic acid concentration map based on the predicted value of chlorogenic acid concentration, and calculate the concentration anomaly value by the sliding window method. If the concentration anomaly value is greater than the preset threshold, an early warning will be issued, thereby realizing the detection of chlorogenic acid components.

[0102] Based on the hyperspectral cluster weighted vector obtained in step S4 The core of constructing a high-dimensional multi-scale support vector machine regression model lies in the multi-scale feature mining and regression fitting of high-dimensional weighted vectors.

[0103] First, to train the support vector machine regression model, training data needs to be prepared. Prepare a set of training samples (≥30) with known chlorogenic acid concentrations (accurately measured by high-performance liquid chromatography). For each chlorogenic acid training sample, perform the following operations: execute steps S2-S4 on the hyperspectral image of the sample to obtain the weighted vectors of all clusters for that sample. The weighted average of the weighted vectors of all clusters of the same sample is calculated based on the number of pixels within each cluster. ,in, This is the comprehensive feature vector of the sample. Let c be the number of pixels. Let C be the weighted vector of the chlorogenic acid hyperspectral clusters of cluster c, where C is the number of clusters. This will ultimately yield the comprehensive feature vector of the sample. The comprehensive feature vector of the sample is used as the feature of the training sample, and the known chlorogenic acid concentration of the sample is used as the label. The comprehensive feature vectors and concentration labels of all chlorogenic acid training samples constitute the training set of the support vector machine regression model. Then, the comprehensive feature vector of the i-th chlorogenic acid training sample is... .

[0104] To improve model robustness, a sample importance metric is introduced. Based on the distribution characteristics of the chlorogenic acid hyperspectral cluster weighted vector, a trapezoidal fuzzy membership function is introduced for the itrain training sample. Assign membership degree This is to differentiate the importance of samples and mitigate the negative impact of outliers on model training.

[0105] in, The Euclidean distance from the comprehensive feature vector of the i-th chlorogenic acid training sample to the centroid of the training set, where the centroid of the training set is the mean center of all comprehensive feature vectors in the training set, is used as an indicator to measure the importance of the sample. The 10th percentile of the distances from all composite feature vectors in the training set to the centroid of the training set. The 25th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set. The 75th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set. It is the 90th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set.

[0106] A multi-scale kernel function is constructed using a Gaussian kernel (local kernel) and a polynomial kernel (global kernel) to map different types of features in the weighted vector of chlorogenic acid hyperspectral clusters.

[0107] in, , Representing the first and the The comprehensive feature vector of each chlorogenic acid training sample. , The kernel function weights satisfy the following conditions: + =1, used to balance the contribution of local and global features; For Gaussian kernel scaling parameters, It is the dot product of vectors; This is the polynomial kernel offset, which defaults to 1. d is the polynomial kernel degree, which controls the fitting order of the global features.

[0108] The objective function of the support vector machine model is:

[0109] in, Let b be the weight vector, and b be the offset. and As slack variables, For regularization parameters, For sample membership degree, This represents the total number of samples in the training set.

[0110] The root mean square error (RMSE) of the support vector machine regression model is minimized by traversing the predefined parameter space using grid search and 5-fold cross-validation, thereby improving the prediction accuracy. Specifically, the search range of key hyperparameters is first defined: regularization parameter... The search is performed within the interval [0.1, 200] with a step size of logarithmic scale. , , , When the characteristic noise of chlorogenic acid is high, Then it needs to be increased (e.g., 200), when the characteristic noise of chlorogenic acid is small. This needs to be reduced (e.g., to 0.1) to balance model complexity and error penalty strength; the multi-scale kernel weights μ1 and μ2 are adjusted within the range of [0.3, 0.7] to coordinate the contribution ratio of local and global kernels in feature mapping; Gaussian kernel scale... The model is optimized within the range [0.2, 0.8] to adapt to the local correlation strength of different spectral features. The polynomial kernel degree d is searched in the integer set {2, 3, 4, 5} to determine the optimal nonlinearity of the global feature mapping. Under the 5-fold cross-validation framework, the training set is randomly divided into 5 subsets. Four subsets are used for model training, and the remaining subset is used for validation. The mean RMSE is calculated as the evaluation metric. Grid search exhaustively searches all parameter combinations. When the RMSE change rate of 5 consecutive parameter sets is <1%, the search stops to avoid invalid traversal. The configuration with the smallest RMSE is selected as the optimal hyperparameter set. After determining the optimal hyperparameter combination of the support vector machine regression model through 5-fold cross-validation, the generalization ability of the model needs to be evaluated. Finally, the entire dataset of samples with known chlorogenic acid concentrations is divided into training, validation, and test subsets in a 7:1:2 ratio. The training subset is used for model parameter learning; the validation subset is used for hyperparameter optimization; and the test subset is used for generalization ability evaluation. The determination coefficient of the test set is required. If the result is >0.85 and RMSE≤0.1%, then repeat the grid search and 5-fold cross-validation until the generalization performance requirements are met.

[0111] For each chlorogenic acid hyperspectral cluster weighted vector to be measured, the weighted vector is input into a pre-trained support vector machine regression model based on high-dimensional multi-scale, and the predicted chlorogenic acid concentration is output.

[0112] in, Let be the predicted chlorogenic acid concentration for the c-th hyperspectral cluster, and m be the index of the support vector. The number of support vectors, and For Lagrange multipliers, satisfy , The weighted vector of the chlorogenic acid hyperspectral cluster of the sample to be tested. This is the weighted vector for the chlorogenic acid hyperspectral cluster corresponding to the m-th support vector. () represents the multi-scale kernel function. This represents the offset of the support vector machine model.

[0113] When constructing the chlorogenic acid concentration map, a direct mapping method based on clusters was adopted. For each pixel within each hyperspectral cluster c, the cosine similarity between its spectral vector and the average spectral vector within the cluster was calculated. Valid pixels with a cosine similarity ≥ 0.85 were retained (excluding impurity pixels). The resulting predicted chlorogenic acid concentration values ​​for each hyperspectral cluster c were then used. The valid pixels contained in the cluster are directly assigned a value, while invalid pixels are marked as undefined (in black). For a valid pixel (u, v) in the image, if it belongs to cluster c, then the value of that pixel in the density region map is... The final result is a graph showing the concentration of chlorogenic acid.

[0114] Optionally, the pixel values ​​of the entire chlorogenic acid concentration map can be linearly mapped to a continuous gradient color band from dark blue (representing the lowest concentration) to red (representing the highest concentration) to visualize the chlorogenic acid concentration map.

[0115] Concentration anomaly scanning was performed using the sliding window method, with the window size set to: Wsize=max(5,(⌊ ⌋,⌊ ⌋)),in , Given the image width and height, ensure the minimum window size is 5×5 pixels to avoid excessively small image windows, and the maximum window size does not exceed 1 / 10 of the image size to avoid obscuring local anomalies. The sliding step size is step=⌊ ⌋; Number of windows: Nwindows=⌊ +1⌋×⌊ +1⌋ If the percentage of valid pixels in the window is less than 70% (too many undefined pixels), skip the window to avoid invalid calculations.

[0116] For each window w, the average concentration within the window is calculated based on the window size. and window concentration standard deviation Then, the overall average concentration of chlorogenic acid in the concentration map is calculated based on the number of windows. and overall concentration standard deviation Finally, the concentration outliers within the window are calculated:

[0117] in, For the concentration outlier in the w-th window, Let w be the average concentration of the w-th window. Let w be the standard deviation of the concentration in the w-th window. This represents the overall average concentration of chlorogenic acid in the concentration graph. The overall concentration standard deviation of the chlorogenic acid concentration plot. The first adjustment coefficient, This is the second adjustment coefficient.

[0118] It should be noted that the first adjustment coefficient Second adjustment coefficient The optimization objective was to validate the calculated results on 10 confirmed anomalous samples (such as samples with localized mold growth or significantly uneven concentrations) using a grid search method. The statistical differences (such as KL divergence) between normal and abnormal regions are maximized, thereby ensuring a balance between the contributions of mean deviation and standard deviation fluctuations.

[0119] Concentration anomaly index obtained from sliding window analysis Set multi-level early warning thresholds: when A value ≤ 2.0 is within the normal range and no warning is needed; when 2.0 < A value ≤3.0 indicates a level of concern, triggering a yellow alert, suggesting a possible slight uneven distribution; when 3.0 < A value ≤4.0 indicates a warning level, triggering an orange alert, indicating a significant abnormal distribution; when A score of >4.0 indicates a severe level, triggering a red alert and indicating a significant quality anomaly.

[0120] It should be noted that traditional methods typically use simple regression models (such as linear regression) or single-kernel SVMs, failing to fully consider the nonlinear relationships of high-dimensional features, and anomaly detection based on fixed thresholds lacks adaptability. Concentration visualization often directly uses raw values ​​without considering spatial smoothing, resulting in unnatural boundaries. This approach employs multi-scale kernel functions, combining Gaussian kernels (capturing local features) and multinomial kernels (capturing global features), improving the model's adaptability to high-dimensional data and achieving more flexible fitting. Optimal hyperparameters are automatically selected through grid search and cross-validation, avoiding the dominance of manual parameter tuning. The sliding window method is then used to calculate concentration anomalies, identifying local anomalies and providing multi-level warnings, enhancing practicality.

[0121] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, the phrase "comprising an element defined as..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0122] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for detecting chlorogenic acid using deep learning, characterized in that: Includes the following steps: Step S1: Take a picture of the chlorogenic acid sample using a hyperspectral imaging device to obtain a hyperspectral image of chlorogenic acid; Step S2: Cluster the chlorogenic acid hyperspectral image using a fuzzy C-means clustering algorithm based on several sub-centers to obtain chlorogenic acid hyperspectral clusters; Step S3: Extract features from the chlorogenic acid hyperspectral clusters to obtain the chlorogenic acid spectral cluster feature vectors; Step S4: Obtain the hyperspectral cluster feature weights by calculating the mutual information of the chlorogenic acid hyperspectral cluster feature vectors; obtain the chlorogenic acid hyperspectral cluster weighted vector by combining the chlorogenic acid hyperspectral cluster feature vectors and the hyperspectral cluster feature weights. Step S5: Input the weighted vector of the chlorogenic acid hyperspectral cluster into the support vector machine regression model based on high-dimensional multi-scale, and output the predicted value of chlorogenic acid concentration; construct a chlorogenic acid concentration map based on the predicted value of chlorogenic acid concentration, and calculate the concentration anomaly value by the sliding window method. If the concentration anomaly value is greater than the preset threshold, an early warning will be issued, thereby realizing the detection of chlorogenic acid components.

2. The method for detecting chlorogenic acid using deep learning as described in claim 1, characterized in that: The process of capturing a hyperspectral image of chlorogenic acid from a sample using a hyperspectral imaging device includes the following specific steps: Place the chlorogenic acid sample stably on the loading plane of the electric translation stage, ensuring that its main surface area faces the lens. Turn on the adjustable halogen lamp light source, with its initial incident angle at a 45-degree angle to the normal of the chlorogenic acid sample. Start the hyperspectral camera and place the diffuse reflection reference white plate next to the sample on the loading stage as a reference for illumination calibration. The effect of specular reflection was eliminated by adjusting the incident angle of the halogen lamp source and the parameters of the hyperspectral camera using an electrically controlled rotary stage; finally, the calibrated hyperspectral image of chlorogenic acid was calculated. ; in, This represents a hyperspectral image of chlorogenic acid. This indicates that the chlorogenic acid sample is measured at pixels (u,v). The original spectral signal value corresponding to the wavelength, Indicates the diffuse reference whiteboard at The average value of the original signal of all spatial pixels at the specified wavelength. Indicates the diffuse reference whiteboard at Standard reflectance at wavelength; Finally, a hyperspectral image of chlorogenic acid was obtained. .

3. The method for detecting chlorogenic acid using deep learning as described in claim 2, characterized in that: The method of clustering chlorogenic acid hyperspectral images using a fuzzy C-means clustering algorithm based on several sub-class centers to obtain chlorogenic acid hyperspectral clusters includes the following steps: The chlorogenic acid hyperspectral image obtained in step S1 Transform into a two-dimensional matrix X∈ Where: N is the total number of pixels in the chlorogenic acid hyperspectral image, D is the number of spectral bands, and each row of the two-dimensional matrix X represents a different element. It is a D-dimensional vector. The d-th element It represents the center wavelength of the d-th band. The reflectance value at that location; Initialize the first fuzzy partitioning matrix Y, the second fuzzy partitioning matrix Z, and the class center matrix; by fixing Y, Z, and center, calculate the vector of each subclass center. This yields the subclass center matrix, centersun. With centers fixed at Y and Z, calculate the center of each cluster. : ; in, Let Q be the vector representing the c-th cluster center, Q be the number of sub-centers, and q be the sub-center index. Given fixed centersun, center, and Z, calculate the membership degree of each pixel to the subclass center: ; in, This represents the membership degree of the i-th pixel to the q-th subclass center. Indicates the first Vectors of subclass centers and ; With centersun, center, and Y fixed, calculate the membership degree of each subclass center to the cluster center using the following formula: ; in, This represents the membership degree of the q-th sub-cluster center to the c-th cluster center. Indicates the first The vector of cluster centers and ; Calculate the objective function value: ; Where J is the objective function value, Let be the weight of the i-th pixel. This is the adjustment coefficient; In each iteration, after sequentially updating the subclass center matrix centersun, the class center matrix center, the first fuzzy partition matrix Y, and the second fuzzy partition matrix Z, the objective function value of the DD-th iteration is calculated. Calculate the relative rate of change of the objective function value between two adjacent iterations. , For the objective function value of the DD-1th iteration, when If the number of iterations is less than the preset threshold, or the number of iterations exceeds the maximum number of iterations, When the threshold value is greater than or equal to the preset threshold, the fuzzy C-means clustering algorithm is determined to have converged and the iteration is terminated, finally yielding the chlorogenic acid hyperspectral cluster.

4. The method for detecting chlorogenic acid using deep learning as described in claim 3, characterized in that: Obtaining the subclass center matrix includes the following steps: Initialize the cluster center matrix center by using the K-means++ algorithm to cluster the standardized two-dimensional matrix X, and use the resulting cluster centers as the initial values ​​of the cluster center matrix Center to avoid the instability of random initialization of cluster centers. With Y, Z, and center fixed, calculate the vector of the center of each subclass. The subclass center matrix centersun is obtained: ; in, This vector represents the center of the q-th subclass, where i represents the pixel index. This represents the complete spectral curve of the i-th pixel. This represents the membership degree of the i-th pixel to the q-th subclass center. Here, r is the adjustment coefficient, and r is the fuzzy factor, which defaults to 2. This represents the membership degree of the q-th subclass center to the c-th cluster center. Let be the vector of the c-th cluster center, where C is the number of clusters; Calculate the vector of the center of each subclass one by one. All Arranged by rows, the final subclass center matrix is ​​obtained as centersun.

5. The method for detecting chlorogenic acid using deep learning as described in claim 4, characterized in that: The step of extracting features from the hyperspectral clusters of chlorogenic acid to obtain the chlorogenic acid spectral cluster feature vectors includes the following steps: For each cluster in the chlorogenic acid hyperspectral cluster, calculate its position within the chlorogenic acid characteristic band range. The pixel ratio, where DKEY is the number of keybands, and dkey is the index of the keyband, in each band If the reflectance values ​​are divided into B bins, then the c-th cluster in the band... The pixel ratio is: ; in, Indicates that cluster c is in the band The reflectance of the pixels in the b-th box is the proportion of the pixel reflectance. This represents the number of pixels in cluster c. Represents pixels ( In the band absolute reflectivity, This represents the maximum reflectivity. () is an indicator function, when When it equals b, The value is 1 if it is 1, otherwise it is 0. Here, b is the floor function, and b is the histogram bin index. Calculate the statistical measure of the pixel proportion in the characteristic band for each chlorogenic acid hyperspectral cluster: mean. ,variance skewness ,energy ; Finally, the color histogram features are obtained. =[ , , , ]; The gray-level co-occurrence matrix of the chlorogenic acid hyperspectral cluster is calculated, thus obtaining the probability matrix; then, based on the normalized probability matrix, the probability matrix for each direction is calculated. The four core directional moments include the second-order moment, the moment of inertia, the correlation, and the entropy. Finally, the four core directional moments of each of the four directions are concatenated in sequence to form a texture feature with a dimension of 16. ; For each chlorogenic acid hyperspectral cluster, the average spectral curve was subjected to continuum removal processing: for local maxima of the average spectral curve, cubic spline interpolation was used to fit a continuum envelope, and the average spectral curve was divided by this continuum envelope to obtain the normalized spectrum after continuum removal. Based on the normalized spectrum, four key absorption parameters were extracted, including absorption depth, absorption width, absorption area, and absorption symmetry. Finally, absorption depth, absorption width, absorption area, and absorption symmetry were used as the spectral absorption characteristics of each chlorogenic acid hyperspectral cluster. ; Finally, the color histogram features, texture features, and spectral absorption features are combined to obtain the chlorogenic acid hyperspectral cluster feature vector. =[ , , ].

6. The method for detecting chlorogenic acid using deep learning as described in claim 5, characterized in that: The step of calculating the mutual information of the chlorogenic acid hyperspectral cluster feature vectors to obtain the hyperspectral cluster feature weights includes the following steps: Collect chlorogenic acid samples and perform three key operations on each sample i in sequence: accurately measure the overall chlorogenic acid concentration using high-performance liquid chromatography. The chlorogenic acid sample was scanned using hyperspectral imaging technology to obtain hyperspectral images. Then, in the clustering and feature extraction stage, the chlorogenic acid hyperspectral images of each sample were spatially segmented using a fuzzy C-means clustering algorithm based on several sub-cluster centers to form several clusters. The chlorogenic acid hyperspectral cluster feature vector of each cluster was obtained, and all the chlorogenic acid hyperspectral cluster feature vectors were summed and averaged to finally obtain the hyperspectral feature vector of the sample to be tested. =[ , , ], This represents the color histogram features of the sample to be tested. This represents the texture features of the sample to be tested. Indicates the spectral absorption characteristics of the sample to be tested; For the hyperspectral feature vector of the sample to be tested Each feature and chlorogenic acid concentration Mutual information: ; in, The hyperspectral feature vector of the sample to be tested The h-th feature and chlorogenic acid concentration The mutual information is given by K, which represents the number of bins for the h-th feature in the hyperspectral feature vector of the sample to be tested, and L, which is the number of bins for the chlorogenic acid concentration. This represents the joint probability of the k-th bin in the hyperspectral feature vector of the sample under test with the l-th bin of chlorogenic acid concentration. Let h be the marginal probability of the h-th feature in the hyperspectral feature vector of the sample to be tested in the k-th bin. This represents the marginal probability of chlorogenic acid concentration in the l-th bin; Finally, the color histogram feature weights, texture feature weights, and spectral absorption feature weights are obtained: ; ; ; in, For the feature weights of the color histogram, For texture feature weights, Weights for spectral absorption characteristics. express Effective mutual information, express Effective mutual information, express Effective mutual information.

7. The method for detecting chlorogenic acid using deep learning as described in claim 6, characterized in that: The weighting vector for the chlorogenic acid hyperspectral cluster is: After Z-score normalization of the chlorogenic acid hyperspectral cluster feature vector, a weighted vector of the chlorogenic acid hyperspectral cluster is obtained by combining the chlorogenic acid hyperspectral cluster feature vector and the hyperspectral cluster feature weights. =[ , , ].

8. The method for detecting chlorogenic acid component by incorporating deep learning according to claim 7, characterized in that: The process of inputting the chlorogenic acid hyperspectral cluster weighted vector into a high-dimensional multi-scale support vector machine regression model to output the predicted chlorogenic acid concentration includes the following specific steps: Collect the training set of the support vector machine model, introduce the trapezoidal fuzzy membership function, and apply it to the i-th training sample in the training set. Assign membership degree ; Construct a multi-scale kernel function to map different types of features in the weighted vector of chlorogenic acid hyperspectral clusters: ; in, , Representing the first and the The comprehensive feature vector of each chlorogenic acid training sample. , The kernel function weights satisfy... + =1, used to balance the contribution of local and global features; For Gaussian kernel scaling parameters, It is the dot product of vectors; This is the polynomial kernel offset, which defaults to 1. d is the polynomial kernel degree, which controls the fitting order of the global features. The objective function of the support vector machine model is: ; in, Let b be the weight vector, and b be the offset. and As slack variables, For regularization parameters, For sample membership degree, The total number of samples in the training set is denoted by , and ittrain is the index of the training samples in the training set. The support vector machine model is trained using the training set to obtain a trained support vector machine model. For each weighted vector of the chlorogenic acid hyperspectral cluster to be tested, the weighted vector of the chlorogenic acid hyperspectral cluster to be tested is input into the trained high-dimensional multi-scale support vector machine regression model, and the predicted value of chlorogenic acid concentration is output. ; in, This represents the predicted chlorogenic acid concentration for the c-th hyperspectral cluster. Let m be the number of support vectors, and m be the index of the support vector. and For Lagrange multipliers, The weighted vector of the chlorogenic acid hyperspectral cluster of the sample to be tested. This is the weighted vector for the chlorogenic acid hyperspectral cluster corresponding to the m-th support vector. () represents the multi-scale kernel function. This represents the offset of the support vector machine model.

9. The method for detecting chlorogenic acid using deep learning as described in claim 8, characterized in that: The membership degree is: ; in, The Euclidean distance from the comprehensive feature vector of the i-th chlorogenic acid training sample to the centroid of the training set, where the centroid of the training set is the mean center of all comprehensive feature vectors in the training set, is used as an indicator to measure the importance of the sample. The 10th percentile of the distances from all composite feature vectors in the training set to the centroid of the training set. The 25th percentile of the distance from all composite feature vectors of the training set to the centroid of the training set. This is the 75th percentile of the distance from all composite feature vectors in the training set to the centroid of the training set. It is the 90th percentile of the distance from all the combined feature vectors of the training set to the centroid of the training set.

10. The method for detecting chlorogenic acid components by incorporating deep learning according to claim 9, characterized in that: The specific steps for calculating concentration outliers using the sliding window method are as follows: A chlorogenic acid concentration map is plotted based on the predicted chlorogenic acid concentration. The size and number of sliding windows are set. For each window w, the average concentration within the window is calculated based on the number of windows. and window concentration standard deviation Then, the overall average concentration of chlorogenic acid in the concentration map is calculated based on the number of windows. and overall concentration standard deviation Finally, the concentration outliers within the window are calculated: ; in, For the concentration outlier in the w-th window, Let w be the average concentration of the w-th window. Let w be the standard deviation of the concentration in the w-th window. This represents the overall average concentration of chlorogenic acid in the concentration graph. The overall concentration standard deviation of the chlorogenic acid concentration graph. The first adjustment coefficient, This is the second adjustment coefficient.

Citation Information

Patent Citations

  • Method for evaluating quality of cold-treating and cough-relieving granules by combining multi-index components with fingerprint spectrum

    CN112051350A

  • Tea variety classification method based on fuzzy linear machine learning

    CN112801174A

  • Hyperspectral image technology-based saxitoxin nondestructive rapid detection method and system

    CN114399674A

  • Method for detecting chlorogenic acid in honeysuckle based on infrared spectrum

    CN115901666A

  • Hyperspectral lithology intelligent identification method based on fuzzy clustering

    CN119915746A