Chromatographic fingerprint spectrum-based siraitia grosvenorii producing area traceability classification method

By constructing a method based on chromatographic fingerprinting, and utilizing an asymmetric least squares smoothing algorithm and a discriminant analysis model, the problem of data noise interference in the traceability of Luo Han Guo (monk fruit) origin was solved, achieving high confidence and robust origin classification.

CN122065142APending Publication Date: 2026-05-19GUILIN SANLENG BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUILIN SANLENG BIOTECH CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies for tracing the origin of monk fruit suffer from imbalanced feature weights at the data perception level, insufficient classification decision-making ability, inability to effectively identify trace features, and poor model generalization ability, making it difficult to achieve high-confidence origin tracing classification under complex and ever-changing data noise interference.

Method used

By constructing a method based on chromatographic fingerprinting, baseline noise is subtracted using an asymmetric least squares smoothing algorithm, retention time drift is corrected using a correlation optimization warping algorithm, and a discriminant analysis model is used to extract the origin-related prediction matrix. The importance of variable projection is calculated, origin markers are screened, and the confidence region boundary is determined by the Hotling statistic and residual distance. A set of sensing parameters is then generated for classification.

Benefits of technology

It effectively solves the problem of poor model generalization ability caused by complex background noise interference, enhances the sensitivity to minute differences, and ensures high confidence and robust origin tracing effect in complex noise environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065142A_ABST
    Figure CN122065142A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent sensing, in particular to a siraitia grosvenorii producing area traceability classification method based on a chromatographic fingerprint spectrum. Comprising the following steps: collecting a chromatographic signal of a siraitia grosvenorii training sample through an ultra-high performance liquid chromatograph, and constructing a chromatographic data matrix; deducting a baseline by using an asymmetric least square smoothing algorithm, correcting retention time drift through a related optimization warping algorithm, and outputting a fingerprint spectrum matrix; extracting a production area related prediction matrix through a discriminant analysis model, and screening out production area markers; extracting a corresponding peak area, calculating a Hotelling statistic critical value and a residual distance critical value, and determining a confidence domain boundary; and performing replacement test on the discriminant analysis model, and determining a classification boundary hyperplane. According to the method, chromatographic data with a high signal-to-noise ratio is provided for construction of the fingerprint spectrum through chromatographic pretreatment, the screening accuracy of the origin markers is ensured, confusion of strong background signals is avoided, and a traceability result has higher robustness and scientific credibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent sensing technology, specifically to a method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting. Background Technology

[0002] In the field of origin traceability and quality control of natural products (such as monk fruit), the separation and detection of chemical components in plant extracts using chromatographic analysis equipment is currently the main means of establishing material standards. However, traditional chromatographic detection terminals operate only as passive data acquisition sensors, lacking the ability to understand and process sensor data in real time. Existing technologies typically rely on offline workstations to use a comprehensive comparison method of full-component chromatographic elution profiles, that is, evaluating the homogeneity of the sample's chemical composition by comparing the overall overlap between the analyte and the standard reference in chromatographic retention time and response intensity.

[0003] However, existing technical solutions have significant shortcomings in terms of origin discrimination and classification decision-making capabilities. At the data perception level, the similarity algorithms used to evaluate quality suffer from feature weight imbalance and lack of classification decision boundaries. Existing algorithms are essentially dominated by high-response signals in vectors and cannot effectively identify trace features with small amplitudes but carrying key origin differences, making it difficult to effectively distinguish samples with similar overall outlines but key local differences. At the intelligent decision-making level, only unsupervised descriptive statistics can be performed, and automatic category assignment of samples from unknown sources cannot be achieved. Secondly, there is a problem with poor generalization ability of the traceability model. Since the differences in origin features of closely related species are easily masked by strong background signals with high similarity, it is difficult to achieve high-confidence origin traceability classification under complex and variable data noise interference.

[0004] Therefore, a method for tracing and classifying the origin of Luo Han Guo (monk fruit) based on chromatographic fingerprinting is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for tracing and classifying the origin of Luo Han Guo (Siraitia grosvenorii) based on chromatographic fingerprinting, comprising: Chromatographic signals of training samples of Luo Han Guo were collected by ultra-high performance liquid chromatography (UHPLC) to construct a chromatographic data matrix including retention time. The baseline of the chromatographic data matrix was subtracted using an asymmetric least squares smoothing algorithm, and the retention time drift was corrected by a related optimized warping algorithm to output a fingerprint matrix. The origin-related prediction matrix is ​​extracted from the fingerprint matrix using a discriminant analysis model. Based on the origin-related prediction matrix, the variable projection importance of chromatographic peaks in the fingerprint matrix is ​​calculated, and chromatographic peaks with a value greater than a preset threshold are selected as origin markers. Based on place of origin markers, the corresponding peak areas of the training samples of monk fruit were extracted, a characteristic peak training set was constructed, and the critical values ​​of the Hotelling statistic and the residual distance were calculated to determine the confidence region boundary. The discriminant analysis model was subjected to permutation test to determine the classification boundary hyperplane and generate a set of sensing parameters. Based on the sample of monk fruit to be tested and the set of sensor parameters, a standard fingerprint spectrum to be tested is generated and the characteristic peak vector to be tested is extracted. The Hotling statistic and the residual distance to be tested are calculated. When it is within the confidence region boundary, the characteristic peak vector to be tested is projected onto the classification boundary hyperplane, and the place of origin label and confidence level are output.

[0007] Preferably, the specific generation process of the chromatographic data matrix includes: an ultra-high performance liquid chromatograph (UHPLC) integrated into an intelligent sensor acquisition terminal; the UHPLC driving the Luo Han Guo training sample to be separated into component eluents through a chromatographic separation column and introduced into the detector flow cell; a full-band light source emitting a continuous spectral beam penetrating the detector flow cell to generate a monochromatic beam sequence; a photodiode array sensor receiving the monochromatic beam and converting it into an analog photoelectric voltage signal, outputting discrete light intensity through synchronous sampling and quantization processing; based on the discrete light intensity, generating absorbance values ​​according to the Lambert-Beer law; arranging the absorbance values ​​corresponding to different detection wavelengths in wavelength order to generate a spectral response vector; stacking the spectral response vectors according to the sampling time order to construct a chromatographic data matrix containing retention time, detection wavelength, and absorbance values; and uploading the chromatographic data matrix to a cloud server.

[0008] Preferably, the specific generation process of the fingerprint spectrum matrix includes: retrieving preset baseline smoothing parameters and a chromatographic data matrix; constructing an objective function containing a data fit term and a baseline smoothing term using an asymmetric least squares smoothing algorithm and performing iterative calculations; assigning different weights to data points in the baseline region and data points in the chromatographic peak region during the iteration process to fit a background drift curve; subtracting the background drift curve from the chromatographic data matrix to output a baseline correction matrix; retrieving a preset standard reference spectrum vector using a correlation optimization warp algorithm and dividing the baseline correction matrix into continuous segments along the time axis; performing scaling transformations on the continuous segments and calculating the correlation coefficient between the transformed segments and the corresponding intervals of the standard reference spectrum vector; performing cumulative determination based on the correlation coefficients and outputting an optimized path; reconstructing the baseline correction matrix based on the optimized path and outputting the fingerprint spectrum matrix.

[0009] Preferably, the specific construction process of the origin-related prediction matrix includes: using the fingerprint spectrum matrix as the independent variable dataset and retrieving a preset origin category label matrix as the dependent variable dataset; constructing a discriminant analysis model using an orthogonal partial least squares discriminant analysis algorithm; calculating the covariance matrix between the independent variable dataset and the dependent variable dataset using the discriminant analysis model; decomposing the fingerprint spectrum matrix into predictive structural components and orthogonal structural components based on the covariance matrix; removing the orthogonal structural components and performing regression fitting calculations on the predictive structural components to generate a regression coefficient array; applying the regression coefficient array to the fingerprint spectrum matrix for weighted reconstruction and outputting the origin-related prediction matrix.

[0010] Preferably, the specific generation process of the place of origin marker includes: retrieving the weight coefficient vector of each potential structural component in the place of origin related prediction matrix and the explanatory variance contribution to the place of origin category label matrix; performing importance quantification calculation on each feature variable contained in the fingerprint matrix: calculating the square value of each element in the weight coefficient vector, and multiplying the square value with the corresponding explanatory variance contribution to generate a weighted contribution component; performing a summation operation on all weighted contribution components and dividing by the total explanatory variance contribution of the place of origin category label matrix to generate a normalized importance intermediate value; performing a square root transformation on the normalized importance intermediate value to output the variable projection importance score corresponding to each feature variable; performing a numerical comparison between the variable projection importance score and a preset discrimination threshold, and when the preset discrimination threshold is exceeded, back-mapping the feature variable to the retention time axis, establishing the corresponding chromatographic peak as the place of origin marker, and recording the retention time positioning information of the place of origin marker.

[0011] Preferably, the specific process for generating the confidence region boundary includes: locking the corresponding chromatographic signal response interval from the fingerprint matrix based on the retention time positioning information; performing an integral accumulation operation on the absorbance values ​​within the chromatographic signal response interval to generate a peak area feature vector; collecting the peak area feature vectors of all Luo Han Guo training samples to construct a feature peak training dataset; performing dimensionality reduction decomposition on the feature peak training dataset to construct a latent variable space, extracting the principal component directions, and calculating the latent variable space covariance matrix; calculating the Mahalanobis distance from each training sample point to the center using the latent variable space covariance matrix, and deriving the upper limit threshold based on multivariate statistical distribution theory to output the critical value of the Hotelling statistic; calculating the vertical Euclidean distance of the training sample points from the latent variable space, setting the upper limit threshold based on the statistical error distribution, and outputting the critical value of the residual distance; combining the in-model variation range defined by the critical value of the Hotelling statistic with the out-of-model noise range defined by the critical value of the residual distance to output the confidence region boundary.

[0012] Preferably, the specific process of the classification boundary hyperplane includes: retrieving the fingerprint spectrum matrix and the origin category label matrix as permutation test input data; performing multiple randomized rearrangement operations on the origin category label matrix to generate a permutation label vector set; forming training pairs between the fingerprint spectrum matrix and each vector in the permutation label vector set to construct the corresponding random permutation discriminant model; calculating the explained variance and cross-validation prediction capability; constructing a random fitting parameter null distribution; performing a significance test on the discriminant analysis model in the random fitting parameter null distribution; if a preset confidence threshold is met, it is determined to be a discriminant model; extracting the regression coefficient array from the discriminant model; constructing a linear decision surface based on the regression coefficient array and outputting the classification boundary hyperplane; and summarizing the discriminant model, classification boundary hyperplane, confidence region boundary, baseline smoothing parameter, standard reference spectrum vector, principal component direction, retention time positioning information, and latent variable space covariance matrix to generate a sensing parameter set.

[0013] Preferably, the specific process of the origin label and confidence level includes: downloading the sensor parameter set from the cloud server to the intelligent sensor acquisition terminal; collecting the Luo Han Guo sample to be tested using an ultra-high performance liquid chromatogram (UHPLC) to generate the original chromatographic vector to be tested; calling the baseline smoothing parameters and the standard reference spectrum vector to perform signal preprocessing on the original chromatographic vector to be tested, and outputting the standard fingerprint spectrum to be tested; extracting the retention time point response values ​​from the standard fingerprint spectrum to be tested based on the retention time positioning information to construct the characteristic peak vector to be tested; calculating the Mahalanobis distance of the characteristic peak vector to be tested using the latent variable space covariance matrix, and outputting the Hotling statistic to be tested; calculating the vertical Euclidean distance of the characteristic peak vector to be tested using the principal component direction, and outputting the residual distance to be tested; comparing and verifying the Hotling statistic and the residual distance to be tested with the confidence region boundary; after successful verification, mapping the characteristic peak vector to be tested to the classification boundary hyperplane to perform linear discriminant operation, generating the discrimination result; outputting the origin label based on the sign attribute of the discrimination result and calculating and outputting the confidence level value based on the discrimination result.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By utilizing an asymmetric least squares smoothing algorithm to subtract the baseline from the chromatographic data matrix and then correcting retention time drift using a related optimized warping algorithm, the problem of poor model generalization ability caused by complex background noise interference is effectively solved. The asymmetric least squares smoothing algorithm adaptively fits the background, removing baseline noise without destroying the original chemical information; combined with the related optimized warping algorithm, the time axis is flexibly aligned, eliminating peak misalignment and ensuring high consistency in feature dimensions across different batches of data. Through a high-fidelity signal reconstruction mechanism, interference from highly similar background signals is removed from the data source, improving the signal-to-noise ratio of the fingerprint spectrum and laying a solid foundation for the model to identify the origin under complex noise conditions.

[0015] 2. A region-related prediction matrix is ​​extracted from the fingerprint spectral matrix using a discriminant analysis model. Based on this matrix, the variable projection importance of chromatographic peaks in the fingerprint spectral matrix is ​​calculated, and peaks exceeding a preset threshold are selected as region-related markers, effectively addressing the problem of easily masked regional differences within the same species. Orthogonal partial least squares discriminant analysis is used to orthogonally decompose the original data into region-related structures and orthogonal noise structures, actively filtering out interference from non-region-related factors. Quantitative screening based on variable projection importance allows for the identification of key trace components that contribute most to region classification from massive datasets, overcoming the weakness of traditional full-spectrum analysis which is susceptible to confusion by high-similarity background signals. This supervised feature extraction strategy enhances the model's sensitivity to subtle differences, constructing a region-related prediction matrix with stronger generalization capabilities.

[0016] 3. By calculating the critical values ​​of the Hotelling statistic and residual distance, the confidence region boundaries are determined. A permutation test is performed on the discriminant analysis model to determine the classification hyperplane, effectively solving the problems of low confidence and interference from outlier samples in classification accuracy. A statistically based dual-defense mechanism is constructed: using the Hotelling statistic and residual distance to define a strict applicable domain, outliers caused by excessive data noise or non-target samples can be identified and eliminated; simultaneously, permutation tests prevent model overfitting, ensuring the statistical significance of the classification hyperplane. When the target feature peak vector is projected onto the classification hyperplane, not only is the origin label output, but the confidence level is also validated based on the confidence region, ensuring higher robustness and scientific credibility of the traceability results in complex and ever-changing real-world testing environments.

[0017] 4. By constructing a chromatographic data matrix including retention times, a fingerprint matrix is ​​output; by extracting a region-related prediction matrix, chromatographic peaks exceeding a preset threshold are selected as region-of-origin markers; by determining the confidence region boundary, the region-of-origin label and confidence level are output. Through the coordinated processing of these three core steps, the problem of poor generalization ability caused by the inability of single-dimensional modeling to withstand complex data noise interference is solved. This overall architecture forms a progressively layered denoising closed loop: the front-end physical signal correction provides high-fidelity input for the mid-stage feature decoupling, ensuring the accuracy of region-of-origin marker selection and avoiding confusion from strong background signals; while feature extraction provides a clean variable space for the back-end confidence region boundary construction, preventing model overfitting. This tight connection from low-level data cleaning to high-level decision verification completely overcomes the disadvantage of weak differences in region-of-origin characteristics among closely related species, achieving a high confidence level and robust overall traceability effect even under variable interference environments. Attached Figure Description

[0018] Figure 1This is a flowchart of the method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting proposed in an embodiment of this invention application; Figure 2 This is a flowchart of the geometric deformation correction and feature decoupling proposed in an embodiment of this invention application; Figure 3 This is a flowchart of the manifold space confidence region and dual decision mechanism proposed in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figures 1-3 The method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting provided by this invention includes the following specific steps: Chromatographic signals of training samples of Luo Han Guo were collected by ultra-high performance liquid chromatography (UHPLC) to construct a chromatographic data matrix including retention time. The baseline of the chromatographic data matrix was subtracted using an asymmetric least squares smoothing algorithm, and the retention time drift was corrected by a related optimized warping algorithm to output a fingerprint matrix. The origin-related prediction matrix is ​​extracted from the fingerprint matrix using a discriminant analysis model. Based on the origin-related prediction matrix, the variable projection importance of chromatographic peaks in the fingerprint matrix is ​​calculated, and chromatographic peaks with a value greater than a preset threshold are selected as origin markers. Based on place of origin markers, the corresponding peak areas of the training samples of monk fruit were extracted, a characteristic peak training set was constructed, and the critical values ​​of the Hotelling statistic and the residual distance were calculated to determine the confidence region boundary. The discriminant analysis model was subjected to permutation test to determine the classification boundary hyperplane and generate a set of sensing parameters. Based on the sample of monk fruit to be tested and the set of sensor parameters, a standard fingerprint spectrum to be tested is generated and the characteristic peak vector to be tested is extracted. The Hotling statistic and the residual distance to be tested are calculated. When it is within the confidence region boundary, the characteristic peak vector to be tested is projected onto the classification boundary hyperplane, and the place of origin label and confidence level are output.

[0021] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.

[0022] Example 1 This application discloses a method for tracing and classifying the origin of Luo Han Guo (Siraitia grosvenorii) based on chromatographic fingerprinting. See also... Figure 1The specific steps proposed in this invention include: S1. Acquiring chromatographic signals of Luo Han Guo training samples using ultra-high performance liquid chromatography (UHPLC) to construct a chromatographic data matrix containing retention times; subtracting the baseline from the chromatographic data matrix using an asymmetric least squares smoothing algorithm, and correcting retention time drift using a related optimized warping algorithm to output a fingerprint matrix; S2. Extracting a region-related prediction matrix from the fingerprint matrix using a discriminant analysis model; calculating the variable projection importance of chromatographic peaks in the fingerprint matrix based on the region-related prediction matrix, and selecting chromatographic peaks greater than a preset threshold as region-related markers; S3. Based on the origin markers, extract the corresponding peak areas of the training samples of Luo Han Guo (monk fruit), construct a characteristic peak training set, and calculate the critical values ​​of the Hotling statistic and residual distance to determine the confidence region boundary; S4, perform a permutation test on the discriminant analysis model to determine the classification boundary hyperplane and generate a set of sensor parameters; S5, based on the Luo Han Guo samples to be tested and the sensor parameter set, generate the standard fingerprint spectrum to be tested and extract the characteristic peak vectors to be tested, calculate the Hotling statistic and residual distances to be tested; when within the confidence region boundary, project the characteristic peak vectors to be tested onto the classification boundary hyperplane and output the origin label and confidence level.

[0023] Furthermore, chromatographic signals of the Luo Han Guo training samples were acquired using ultra-high performance liquid chromatography (UHPLC) to construct a chromatographic data matrix including retention times. The baseline was subtracted from the chromatographic data matrix using an asymmetric least squares smoothing algorithm, and retention time drift was corrected using a related optimized warping algorithm to output a fingerprint matrix; this corresponds to step S1 above. The specific implementation process includes: An ultra-high performance liquid chromatograph (UHPLC) is integrated into an intelligent sensing and acquisition terminal. The UHPLC drives the Luo Han Guo training sample to be separated into component eluents through a chromatographic separation column and introduced into the detector flow cell. A full-band light source emits a continuous spectral beam that penetrates the detector flow cell, generating a monochromatic beam sequence. A photodiode array sensor receives the monochromatic beam and converts it into an analog photoelectric voltage signal. Through synchronous sampling and quantization, discrete light intensity is output. Based on the discrete light intensity, absorbance values ​​are generated according to the Lambert-Beer law. The absorbance values ​​corresponding to different detection wavelengths are arranged in wavelength order to generate a spectral response vector. The spectral response vectors are stacked according to the sampling time order to construct a chromatographic data matrix containing retention time, detection wavelength, and absorbance values. The chromatographic data matrix is ​​then uploaded to a cloud server.

[0024] Specifically, the formation process of the chromatographic data matrix is ​​as follows: In this embodiment, an ultra-high performance liquid chromatograph (UHPLC) serves as the core sensing terminal, with a built-in C18 reversed-phase column. The mobile phase system consists of ultrapure water containing 0.1% formic acid as phase A and acetonitrile containing 0.1% formic acid as phase B. The flow rate is controlled at a constant 0.4 mL / min, and the column temperature is maintained at 30°C. The monk fruit training sample is injected after being filtered through a microporous membrane, and a linear gradient elution program is executed. Initially (0.00 min), mobile phase A accounts for 95% and mobile phase B accounts for 5%. After maintaining linear gradient elution for 5.00 min, mobile phase A decreases to 90% and mobile phase B increases to 10%. The linear gradient continues until 15.00 min, when mobile phase A decreases to 65% and mobile phase B increases to 35%. Subsequently, the gradient is rapidly changed, and by 18.00 min, mobile phase A decreases to 10% and mobile phase B increases to 90% to elute impurities. Finally, at 22.00 min, the initial ratio (95% phase A, 5% phase B) is restored for equilibration. The key chemical components are spatially and temporally separated in the chromatographic column based on their polarity differences, forming sequentially eluted eluates that are continuously introduced into the detector flow cell. The 120 batches of Luo Han Guo training samples were selected from the core production area and surrounding non-core production areas, covering fruits from different harvest seasons (September to November) and different stages of maturity. Using the Kennard-Stone algorithm, these 120 batches of samples were divided into two parts at a 2:1 ratio: 80 samples were used as the training set for model construction and biomarker screening; the remaining 40 samples were used as the external test set for subsequent model validation.

[0025] In this embodiment, the detector uses a full-band deuterium lamp as the light source, emitting a continuous spectral beam covering a wavelength range of 190nm to 400nm. After passing through the flow cell, the beam illuminates the surface of the photodiode array sensor. The photodiode array sensor synchronously receives the transmitted beam at a sampling frequency of 20Hz, converts the optical signal into an analog voltage signal, and quantizes it via an analog-to-digital converter, outputting a series of discrete light intensity values. Downsampling integration is performed on the discrete light intensity values: a time window of 0.5 seconds (containing 10 original sampling points) is set, and the spectral data within the window are arithmetically averaged to generate a compressed spectral response vector. The calculation is performed according to the Lambert-Beer law, i.e., the absorbance value is equal to the common logarithm of the ratio of incident light intensity to transmitted light intensity.

[0026] In this embodiment, the characteristic absorption wavelength of mogrosides, 203 nm, was identified, and the absorbance value sequence at this wavelength was extracted. At each sampling time point, the absorbance values ​​of different detection wavelengths were arranged in ascending order of wavelength, generating a one-dimensional spectral response vector. Subsequently, according to the sampling time sequence, the 2640 consecutively acquired spectral response vectors were stacked in rows along the time dimension to construct a chromatographic data matrix containing the retention time dimension, the detection wavelength dimension, and the absorbance value (its dimension is defined as: number of samples × number of retention time points). This matrix, as a standardized digital fingerprint, was directly uploaded to the cloud server via an encryption protocol, serving as the sole input source for subsequent chemometric analysis.

[0027] The combination of ultra-high performance liquid chromatography (UHPLC) and photodiode array enables continuous spectral scanning of Luo Han Guo components across the entire spectrum, capturing weak characteristic signals and constructing a chromatographic data matrix that includes retention time, detection wavelength, and absorbance. This greatly enriches the dimensions of the characteristic data and provides high-fidelity, standardized underlying data support for subsequent processing.

[0028] The algorithm retrieves preset baseline smoothing parameters and chromatographic data matrices, constructs an objective function containing data fit and baseline smoothness terms using an asymmetric least squares smoothing algorithm, and performs iterative calculations. During the iteration process, different weights are assigned to data points in the baseline region and data points in the chromatographic peak region to fit the background drift curve. The background drift curve is subtracted from the chromatographic data matrix to output the baseline correction matrix. A correlation optimization warp algorithm retrieves a preset standard reference spectrum vector and divides the baseline correction matrix into continuous segments along the time axis. The algorithm performs scaling transformations on the continuous segments and calculates the correlation coefficient between the transformed segments and the corresponding intervals of the standard reference spectrum vector. Based on the correlation coefficients, an accumulation judgment is made, and an optimized path is output. Based on the optimized path, the baseline correction matrix is ​​reconstructed, and a fingerprint spectrum matrix is ​​output.

[0029] Specifically, the process of outputting the baseline correction matrix using the asymmetric least squares smoothing algorithm is as follows: In this embodiment, the single-column spectral response vector from the chromatographic data matrix is ​​first retrieved as the input variable. Maximum value normalization is then performed on the absorbance values ​​in the chromatographic data matrix, mapping them to a dimensionless interval of 0 to 1, and the baseline vector is initialized. The objective function constructed using the asymmetric least squares smoothing algorithm is defined as a weighted sum of two parts: the first part is the data fit term, i.e., the weighted sum of squares of the differences between the original signal and the estimated baseline, used to ensure the baseline closely follows the signal profile; the second part is the baseline smoothness term, i.e., the product of the sum of squares of the second-order differences of the baseline vector and the baseline smoothing parameter, used to penalize drastic fluctuations in the baseline. The baseline smoothing parameter determines the rigidity of the fitted baseline; the larger the baseline smoothing parameter, the smoother the baseline and the less susceptible it is to local noise; the smaller the baseline smoothing parameter, the more compliant the baseline. It is generally recommended that the baseline smoothing parameter range be between 10⁵ and 10⁷. In this embodiment, the baseline smoothing parameter is set to 10⁵ to ensure that the baseline has sufficient rigidity and is not affected by narrow peak interference. Minimizing the loss function is transformed into solving a system of linear equations, namely the sum of the products of the weight diagonal matrix, the baseline smoothing parameter, the second-order difference matrix, and its transpose. Multiplying this result by the baseline vector equals multiplying the weight diagonal matrix by the original signal vector. In the first iteration, all diagonal elements of the weight matrix are initialized to 1.

[0030] In this embodiment, an adaptive Laplace regularization constraint mechanism based on local signal complexity is introduced when constructing the objective function. By calculating the local second derivative variance of the chromatographic signal, the local weights of the smoothing parameters are dynamically adjusted to prevent overcorrection of the baseline fitting at narrow peaks. Specifically, before performing asymmetric least squares iteration, a sliding window is first defined to traverse the chromatographic data matrix. The width of the sliding window depends on the sampling frequency of the chromatographic system and the minimum half-peak width of the analyte, preferably set to cover a retention time of 0.5 to 1.0 seconds. In this embodiment, based on a sampling frequency of 20 Hz, the window width is set to 15 data points. Within each window, the second derivative variance of the original signal is calculated and normalized to a local complexity factor. Specifically, the local complexity factor is equal to 1 divided by the denominator, which is 1 plus the exponent of the natural constant, where the exponent is the negative value of the second derivative variance of the original signal within the sliding window centered on the current data point. Subsequently, the objective function is reconstructed, and a dynamic regularization term is introduced on top of the original baseline smoothness term. The corrected dynamic smoothing parameter is no longer a global constant (10⁵), but is defined as the sum of the base smoothing parameter multiplied by 1 plus the product of the enhancement coefficient and the local complexity factor. The base smoothing parameter is set to 10⁵, and the enhancement coefficient is adaptively mapped using the logarithm of the signal-to-noise ratio. Specifically, the enhancement coefficient equals a constant 8 minus the commonly used logarithm of the original signal-to-noise ratio. When the calculated result is less than 3, the enhancement coefficient is set to a lower limit of 3; when the calculated result is greater than 8, the enhancement coefficient is set to an upper limit of 8. In this embodiment, for the chromatographic data of the monk fruit extract with a signal-to-noise ratio of approximately 1000, the commonly used logarithm is 3. Based on the above logic, the enhancement coefficient is calculated to be 5.0. This adaptive mechanism ensures that the sensitivity of the baseline fitting to high-frequency narrow peaks remains in the optimal range under different instrument noise levels. When iteratively solving the linear equations, a diagonal weight matrix is ​​constructed, with its diagonal elements corresponding to the dynamic smoothing parameter at each time point. The improved objective function minimization problem is transformed into solving the following system of linear equations: the sum of the continuous products of the asymmetric weight diagonal matrix, the transpose of the second-order difference matrix, the dynamic smoothing parameter diagonal matrix, and the second-order difference matrix. The product of this sum matrix and the baseline vector to be determined is equal to the product of the asymmetric weight diagonal matrix and the original signal vector. Here, the baseline vector to be determined is an unknown quantity, the original signal vector is a known quantity, the asymmetric weight diagonal matrix is ​​iteratively updated based on the deviation between the signal and its baseline estimate, and the second-order difference matrix is ​​a constant matrix. When the signal is detected in a region of high-frequency, drastically changing chromatographic peaks (where the local complexity factor approaches 1), the local smoothing parameter automatically increases fivefold, forcing the baseline to remain rigid; while in flat baseline regions, the parameter maintains its baseline value, ensuring the compliance of the fit. This process solves the problem of narrow-peak baseline overcutting caused by fixed smoothing parameters and improves the integration accuracy of small chromatographic peaks.

[0031] In this embodiment, after calculating a new baseline estimate in each iteration, the relationship between the original data point and the baseline point is checked point by point. If the original data point is larger than the baseline point, it is determined that the point is in the chromatographic peak region, and its corresponding weight is updated to an asymmetric weight parameter. The asymmetric weight parameter determines the baseline's tolerance to negative deviation. The smaller the asymmetric weight parameter (usually less than 0.01), the more the baseline tends to fit the lower edge contour of the signal. In this embodiment, the parameter is set to 0.001 to reduce the influence of the chromatographic peak on the baseline fitting. If the original data point is smaller than or equal to the baseline point, it is determined that the point is in the baseline region, and the weight is updated to 1 minus the difference of the asymmetric weight parameter, i.e., 0.999, forcing the baseline to move closer to the data point. This process is repeated iteratively until the weight matrix of two adjacent iterations no longer changes or the preset maximum number of iterations (e.g., 10 times) is reached. The final output vector is the background drift curve. By performing matrix subtraction, the background drift curve is subtracted from the original chromatographic data matrix, and a clean baseline correction matrix is ​​output.

[0032] Specifically, the process of outputting the fingerprint matrix using the relevant optimized warp algorithm is as follows: In this embodiment, a set of positive samples confirmed as belonging to the core production area is selected from the baseline correction matrix. The average correlation coefficient of each sample in the set with all other samples is calculated, and the sample with the largest average correlation coefficient is selected as the standard reference spectral vector. Specifically, this standard reference spectral vector is a row vector containing 2640 data points. Its waveform characteristics are as follows: a maximum response value (normalized intensity of 1.0) at a retention time of 7.82 minutes, corresponding to the characteristic peak of mogroside V; a second strongest response (normalized intensity of approximately 0.56) at 12.52 minutes, corresponding to the characteristic peak of sympathomimetic acid I; and a characteristic peak with a normalized intensity of approximately 0.38 at 4.20 minutes. This embodiment stores the average spectral data. Preset parameters include fragment length and relaxation factor. In this embodiment, based on the average half-width of the mogroside chromatographic peak, the fragment length is set to 2 to 3 times the average half-width. In this embodiment, the measured average half-width at half-maximum (HWHM) is approximately 20 points. Therefore, the segment length is set to 50 data points, equivalent to a retention time window of 2.5 s. The relaxation factor is set to two data points, meaning that each segment can be stretched to 52 points or compressed to 48 points. The relevant optimization warp algorithm divides the spectral vector to be corrected and the reference vector into several continuous segments at equal intervals along the time axis, constructing an aligned grid.

[0033] In this embodiment, the process is based on the principle of dynamic programming. For each segment index, the Pearson correlation coefficient between the segment to be corrected and the reference segment is calculated within the allowed scaling range, i.e., when the length change does not exceed the relaxation factor. The Pearson correlation coefficient is obtained by dividing the covariance of the two sets of data by the product of the standard deviations of the two sets of data, and the value is between -1 and 1. A cumulative payoff matrix is ​​constructed to record the sum of the maximum cumulative correlation coefficients obtainable from the starting point to the current node. Specifically, for each node in the aligned grid, its cumulative payoff value is equal to the Pearson correlation coefficient of the current segment, plus the maximum cumulative payoff value among all reachable nodes in the previous column that satisfy the relaxation factor constraint. The cumulative payoff value of the first column of nodes is initialized to the correlation coefficient of the first segment corresponding to that node. By backtracking the path with the largest value in the cumulative payoff matrix, the optimal scaling amount for each segment is determined, forming an optimized path. If multiple paths have the same maximum cumulative payoff value, the path with the smallest scaling amount or the path closest to the diagonal is preferentially selected as the optimized path. Finally, using a linear interpolation algorithm, each segment of the vector to be corrected is mapped onto the standard time axis according to the optimized path. After performing this operation on all samples, the row-aligned, column-corresponding fingerprint matrix is ​​output.

[0034] The asymmetric least squares smoothing algorithm, through iterative computation and dynamic weight allocation, can intelligently distinguish between the true chromatographic peak signal and background drift trends. Without sacrificing effective peak area, it subtracts baseline noise caused by mobile phase gradient changes or instrument aging, thus improving the signal-to-noise ratio. Simultaneously, the related optimized warp algorithm utilizes a piecewise scaling transformation strategy to dynamically correct retention time drift caused by column temperature fluctuations or slight flow rate variations. This ensures the consistency of chromatographic peaks of the same chemical component across different batches of samples on the time axis, avoiding classification model failure due to peak mismatch. This step transforms the original chaotic chromatographic signal into a standardized fingerprint spectrum, laying a solid signal quality foundation for high-confidence origin traceability.

[0035] Furthermore, a place-of-origin prediction matrix is ​​extracted from the fingerprint matrix using a discriminant analysis model; corresponding to step S2 above; the specific implementation process includes: The fingerprint spectrum matrix is ​​used as the independent variable dataset, and a preset origin category label matrix is ​​retrieved as the dependent variable dataset. A discriminant analysis model is constructed using an orthogonal partial least squares discriminant analysis algorithm. The covariance matrix between the independent and dependent variable datasets is calculated using the discriminant analysis model. Based on the covariance matrix, the fingerprint spectrum matrix is ​​decomposed into predictive structural components and orthogonal structural components. The orthogonal structural components are removed, and regression fitting calculations are performed on the predictive structural components to generate a regression coefficient array. The regression coefficient array is applied to the fingerprint spectrum matrix for weighted reconstruction, and the origin-related prediction matrix is ​​output.

[0036] Specifically, the process of constructing a discriminant analysis model using the orthogonal partial least squares discriminant analysis algorithm and outputting the origin-related prediction matrix is ​​as follows: In this embodiment, the fingerprint matrix serves as the independent variable dataset, with its dimension defined as the number of samples multiplied by the number of retention time points. For example, the training set contains 80 monk fruit samples, with each sample's absorbance values ​​at 2640 time points forming a row. The dependent variable dataset is an 80-row, 1-column origin category label matrix, where core origins are labeled 1 and non-core origins are labeled 0. It includes 40 certified core origin samples and 40 samples from non-core origin samples, maintaining a 1:1 ratio of positive to negative samples to avoid class bias in the model. Before processing, the fingerprint matrix must be Pareto-scaling, i.e., subtracting the mean from each variable and dividing by the square root of the standard deviation. This aims to give low-content but significant trace components (such as symmancoside I) equal statistical weight to high-content components (such as mogroside V). Subsequently, the product of the transpose of the scaled fingerprint matrix and the origin category label matrix is ​​calculated to generate a covariance matrix, which quantifies the degree of linear correlation between the signal intensity change at the chromatographic retention time point and the origin category.

[0037] In this embodiment, the discriminant analysis model training employs the orthogonal partial least squares discriminant analysis algorithm, with its objective function set to maximize the sum of squared covariances between the predicted principal component score vector and the origin category vector. During the iteration process, a column of the origin category label matrix is ​​first selected as the initial score vector. The product of the transpose of the fingerprint spectrum matrix and the current score vector is calculated and normalized to obtain the weight vector. The product of the fingerprint spectrum matrix and the weight vector is then calculated to obtain a new score vector. This iteration continues until the magnitude of the difference between the score vectors generated in two adjacent iterations is less than a preset convergence tolerance. In this embodiment, to ensure the accuracy of the orthogonal decomposition, the convergence tolerance is set to 10⁻⁶. Subsequently, by calculating the residual between the score vector and the fingerprint spectrum matrix, the variant parts orthogonal to (i.e., uncorrelated) with the origin category label matrix vector are identified as orthogonal structural components. Orthogonal signal correction is then performed, and the variant information orthogonal to the origin category label is extracted iteratively until the preset number of principal components is reached. The number of principal components is determined through 7-fold cross-validation. Adding principal components stops when the decrease in the mean squared error of the prediction is less than 5%. In this embodiment, the optimal number of principal components is 2. Subtracting this orthogonal structural component from the original fingerprint matrix yields the pure predicted structural component. This process determines the optimal number of principal components through cross-validation, and the loss function minimizes the sum of squared predicted residuals, ensuring that the model effectively extracts origin features without overfitting.

[0038] In this embodiment, regression fitting calculations are performed based on the predicted structural components after removing orthogonal noise. Using the least squares method, a regression coefficient array that minimizes the prediction error is calculated. This array is a column vector of equal length to the number of retention time points, and its numerical value and sign directly indicate the direction and intensity of the contribution of the chromatographic peak at the corresponding retention time point to the origin classification. Finally, this regression coefficient array is used as a weighted filter and applied back to the original fingerprint matrix (i.e., performing matrix multiplication) to reconstruct the fingerprint matrix. The reconstructed matrix is ​​the origin-related prediction matrix, in which background interference unrelated to origin has been suppressed, while characteristic peak signals closely related to origin have been significantly enhanced.

[0039] By decomposing the predictive structural components and orthogonal structural components, variation information in the fingerprint spectrum that is orthogonal (i.e., unrelated) to the origin category label (such as systematic fluctuations caused by non-origin factors) can be removed from the model. By retaining only the predictive structural components that are highly correlated with origin and generating a regression coefficient array, this step enables the final origin-related prediction matrix to focus on essential differences. Even when samples are subjected to complex interferences due to storage conditions or batch differences, it can still maintain high sensitivity in capturing origin characteristics, thereby improving the accuracy and robustness of classification.

[0040] Furthermore, based on the origin-related prediction matrix, the variable projection importance of chromatographic peaks in the fingerprint matrix is ​​calculated, and chromatographic peaks exceeding a preset threshold are selected as origin markers; this corresponds to step S2 above; see [link to relevant documentation]. Figure 2 The specific implementation process includes: The system retrieves the weight coefficient vectors of each potential structural component in the origin-related prediction matrix and their explanatory variance contribution to the origin category label matrix. It then performs importance quantification calculations on each feature variable contained in the fingerprint matrix: calculating the square value of each element in the weight coefficient vector and multiplying the square value with the corresponding explanatory variance contribution to generate a weighted contribution component; summing all weighted contribution components and dividing by the total explanatory variance contribution of the origin category label matrix to generate a normalized importance median value; performing a square root transformation on the normalized importance median value to output the variable projection importance score corresponding to each feature variable; comparing the variable projection importance score with a preset discrimination threshold; when the score exceeds the preset discrimination threshold, the feature variable is back-mapped to the retention time axis, and the corresponding chromatographic peak is established as an origin marker, and the retention time location information of the origin marker is recorded.

[0041] Specifically, the process of quantifying importance and establishing geographical indications is as follows: In this embodiment, key statistics determining origin classification are extracted from the constructed and trained discriminant analysis model. First, the weight vectors corresponding to all potential structural components (including predicted principal components and orthogonal principal components) in the model are retrieved as weight coefficient vectors, and the sum of squared variances explained by each potential structural component to the dependent variable matrix (i.e., origin category labels) is used as the explained variance contribution. For the fingerprint matrix containing 2640 variables, the weight matrix is ​​obtained by loading the loading attributes in the model object, and the variance explained rate values ​​in the model statistics are extracted. These parameters quantify the contribution weight of signal fluctuations at each chromatographic retention time point to the final origin classification.

[0042] In this embodiment, for each feature variable (i.e., a specific retention time point) in the fingerprint matrix, the following arithmetic operations are performed: First, the square of the weight coefficient of the feature variable on each structural component is calculated, and the squared value is multiplied by the explanatory variance contribution of the corresponding component to generate a weighted contribution component; then, the weighted contribution components of the variable on all structural components are summed to obtain a summation result, and the summation result is multiplied by the total number of variables and divided by the total explanatory variance contribution of the model to the place of origin label to generate a normalized importance intermediate value. The purpose of this step is to normalize the average importance score of all feature variables to 1; finally, a square root transformation is performed on the normalized importance intermediate value to output the final variable projection importance score of the feature variable.

[0043] In this embodiment, a preset discrimination threshold of 1.0 is set, which is a common critical value for determining the significance of variables in chemometrics. The projected importance scores of all feature variables are iterated, and variables with scores greater than 1.0 are marked as significant features. Since chromatographic peaks consist of continuous scan points, the selected discrete significant variable indices are back-mapped to the original retention time axis. The retention time index corresponding to the feature variable is extracted, and temporally adjacent and continuous significant points are grouped into the same chromatographic peak interval. For example, if the scores of all data points within a continuous index range exceed the threshold, the chromatographic peak corresponding to that continuous interval is established as a geographical marker, and its start and end times are recorded as retention time location information. Considering the continuity of the chromatographic signal and the sampling frequency (20Hz in this embodiment), to avoid interference from isolated high-scoring points caused by noise, a sliding window clustering strategy is used to merge the retention time indices: a minimum peak width threshold of 40 sampling points (corresponding to a 2.0-second retention time) is set; all feature variable indices with scores greater than the threshold are iterated, and only when the number of consecutively adjacent indices exceeds the minimum peak width threshold is it defined as a valid chromatographic peak response interval. The first and last time points of the response interval are used as the retention time location information for the place of origin marker.

[0044] In this embodiment, a guided resampling-based stability verification is performed on the initially selected geographical indications. By reconstructing the training set with replacement using random perturbation, the probability that a marker maintains its variable projected importance value exceeding the limit in multiple resampling models is calculated, eliminating random features. Before establishing the chromatographic peak as a geographical indication, a guided verification loop containing 1000 iterations is constructed, with the seed parameter of the random number generator fixed at 42. In each iteration, based on the pseudo-random sequence generated by this fixed seed, random sampling with replacement is performed from the original 80 training samples to generate a new subset of independent variables and a corresponding subset of dependent variables of size 80. The orthogonal partial least squares discriminant analysis model is retrained based on each new subset of independent and dependent variables, and the variable projected importance score of each feature variable is recalculated. After the loop ends, each feature variable will have a distribution set of 1000 variable projected importance values. The 95% confidence interval of this distribution is calculated. The stability judgment logic is set as follows: a variable is considered statistically robust only when the lower limit of its confidence interval is greater than 1.0 (a preset discrimination threshold). If a chromatographic peak has a projected importance greater than 1.0 in the original model but a lower limit of its confidence interval is less than 0.9 in the resampling validation, its significance may originate from contributions from individual extreme samples, and it will be judged as a "spurious marker" and removed. Finally, only those chromatographic peaks that show significant discriminative power in 95% of the resampling models are retained as the final origin markers. This process eliminates feature misselection caused by individual outliers and enhances the model's generalization and anti-interference ability.

[0045] Specifically, this paper discusses in detail the actual selected geographical indications of monk fruit and their supporting data. In the analysis of 80 batches of training samples of monk fruit, three key time windows were identified based on the above process. Data shows that the chromatographic peak with a retention time between 7.75 and 7.90 minutes had a peak projection importance score of 3.82, far exceeding the threshold, and was identified by mass spectrometry as mogroside V, the primary marker distinguishing core and non-core production areas. The chromatographic peak with a retention time between 12.45 and 12.60 minutes scored 2.15 and was identified as cimanin I. Meanwhile, mogroside IV with a retention time between 4.15 and 4.25 minutes scored 1.45. These high-scoring chromatographic peak ranges were solidified as key targets in the sensor parameter set.

[0046] By comprehensively considering the contribution of weighted coefficient vectors and explained variance, the contribution to place of origin classification is quantified. In this process, common components that, despite their high content, lack discriminative power, as well as noise peaks, are removed. The selected feature variables are then mapped back onto the retained time axis to establish place of origin markers. This approach compresses data volume, improves computational efficiency, and ensures that subsequent classification operations are based on high-value net signals. This further strengthens the resistance to interference in complex background noise conditions and enhances classification confidence.

[0047] Furthermore, based on place-of-origin markers, the corresponding peak areas of the training samples of Luo Han Guo are extracted to construct a characteristic peak training set. The critical values ​​of the Hotelling statistic and the residual distance are calculated to determine the confidence region boundaries; this corresponds to step S3 above. The specific implementation process includes: Based on retention time location information, the corresponding chromatographic signal response interval is located from the fingerprint spectrum matrix. An integral accumulation operation is performed on the absorbance values ​​within the chromatographic signal response interval to generate a peak area feature vector. The peak area feature vectors of all Luo Han Guo training samples are collected to construct a feature peak training dataset. Dimensionality reduction decomposition is performed on the feature peak training dataset to construct a latent variable space, and the principal component directions are extracted. The latent variable space covariance matrix is ​​calculated. Within the latent variable space, the Mahalanobis distance from each training sample point to the center is calculated using the latent variable space covariance matrix. An upper limit threshold is derived based on multivariate statistical distribution theory, and the critical value of the Hotelling statistic is output. The vertical Euclidean distance of the training sample points from the latent variable space is calculated, and an upper limit threshold is set based on the statistical error distribution, outputting the critical value of the residual distance. The in-model variation range defined by the critical value of the Hotelling statistic is combined with the out-of-model noise range defined by the critical value of the residual distance to output the confidence region boundary.

[0048] Specifically, the process of constructing the feature peak training dataset is as follows: In this embodiment, based on retention time location information, such as the retention time window of mogroside V being 7.75 minutes to 7.90 minutes, the corresponding sub-matrix region, i.e., the chromatographic signal response interval, is cut out from the aligned fingerprint matrix. For each mogroside training sample, the trapezoidal integral method is used to perform an integral accumulation operation on all absorbance values ​​within the chromatographic signal response interval, converting the spectral response intensity within the interval into a single peak area value, which represents the relative content of a specific chemical component. All selected origin markers are traversed, and the peak areas of multiple markers for each sample are combined into a feature vector, i.e., a peak area feature vector, for example, a vector containing three dimensions: mogroside V, symmonoside I, and mogroside IV. The feature vectors of 80 training samples are collected to construct an 80-row, 3-column feature peak training dataset, and column centering is performed on this dataset, i.e., subtracting the arithmetic mean of each column to eliminate the influence of the intercept.

[0049] Specifically, the process for determining the confidence region boundary is as follows: In this embodiment, principal component analysis (PCA) is used to orthogonally decompose the centered dataset: the covariance matrix of the training dataset with eigenpeaks is calculated, and eigenvalue decomposition is performed on it. The two eigenvectors with the largest eigenvalues ​​are extracted as principal component directions to construct a two-dimensional latent variable space. The original high-dimensional features are projected onto this low-dimensional space to generate a score matrix and a loading matrix. The covariance matrix of the latent variable space is calculated, which describes the statistical distribution of the training samples in the principal component space. During this process, the cumulative variance contribution rate explained by the principal component model is recorded to ensure that it exceeds 95%, so as to ensure that the latent variable space fully retains the key origin characteristics of the original data.

[0050] In this embodiment, the Hotelling statistic is used to measure the Mahalanobis distance of a sample point from the model center within the latent variable space, reflecting the degree of variation of the sample within the model's coverage area. For any training sample, the calculation process for the critical value of the Hotelling statistic is described as follows: transpose the sample's score vector, multiply it by the inverse of the latent variable space covariance matrix, and then multiply by the score vector itself. To determine the in-model range of the confidence region boundary, the upper limit threshold is derived using the F-distribution based on multivariate statistical process control theory. The specific calculation formula is described as follows: multiply the number of principal components by the difference between the total number of samples and the number of principal components, divide by the difference between the total number of samples and the number of principal components, and then multiply the result by the quantile value of the F-distribution with degrees of freedom (first degree of freedom is the number of principal components, second degree of freedom is the number of samples minus the number of principal components) at a preset confidence level (e.g., 95%). In the experiment, based on 80 samples and 2 principal components, the calculated critical value of the Hotelling statistic is approximately 6.85. Regions below this value constitute the lower boundary of the confidence region.

[0051] In this embodiment, the residual distance measures the vertical Euclidean distance of a sample point from the latent variable space, reflecting anomalous structures or noise in the sample that are not explained by the model. First, the residual vector is obtained by subtracting the reconstructed vector in the principal component space from the original eigenvector. Then, the square root of the sum of squares of each element in the residual vector is calculated as the residual distance for that sample. To determine the upper limit threshold for noise, the Jackson-Mudholkar approximation method is used, which is based on the power sum estimation of the eigenvalues ​​of the residual matrix. Specifically: the residual matrix of the eigenvalue training set after reconstruction in the latent variable space is calculated, and the covariance matrix of this residual matrix is ​​calculated. Eigenvalue decomposition is performed on the covariance matrix, calculating the sum of the first, second, and third powers of all eigenvalues ​​respectively. The product of the first and third powers is multiplied by twice, and the result is divided by three times the squared value of the second power. The quotient is then subtracted from 1 to obtain the intermediate weight coefficients. The final residual distance critical value is calculated as follows: the sum of eigenvalues ​​is used as a benchmark, multiplied by the weight coefficient of the correction factor raised to the power of the sum of squares of the eigenvalues. The correction factor is composed of the weight coefficient, the ratio of the sum of squares of the eigenvalues ​​to the sum of the eigenvalues, and the quantile of the standard normal distribution at a pre-set confidence level. The quantile of the standard normal distribution at a 95% confidence level is 1.645 in this embodiment. In the experiment, this critical value was 2.15. Finally, the elliptical boundary defined by the Hotelling statistic critical value and the vertical height defined by the residual distance critical value together form a closed cylindrical geometric space, which is the confidence region boundary for tracing the origin of Luo Han Guo (monk fruit). This boundary is used to eliminate abnormal samples with severe deviations and other origins in subsequent steps.

[0052] The latent variable space constructed based on peak area feature vectors not only reflects the normal fluctuation range of the training set samples but also quantifies the deviation of the samples from the model center using Mahalanobis distance, which is used to identify extreme outliers within the model. Simultaneously, it quantifies the residuals of the samples outside the model space using vertical Euclidean distance, which is used to keenly capture exogenous interference noise that does not conform to the model structure. This confidence region boundary setting, which combines intra-model variation and extra-model noise, ensures that only samples that conform to the Luo Han Guo feature pattern and are of acceptable quality are included in the classification process. Outliers caused by severe noise interference or fabrication can be directly rejected, thus guaranteeing a high statistical confidence level in the output results.

[0053] Furthermore, a permutation test is performed on the discriminant analysis model to determine the classification boundary hyperplane and generate a set of sensing parameters; this corresponds to step S4 above; the specific implementation process includes: The fingerprint spectrum matrix and the origin category label matrix are retrieved as input data for the permutation test. Multiple randomization operations are performed on the origin category label matrix to generate a set of permutation label vectors. The fingerprint spectrum matrix is ​​paired with each vector in the permutation label vector set to construct a corresponding random permutation discriminant model. The explained variance and cross-validation prediction ability are calculated, and a random fitting parameter null distribution is constructed. The discriminant analysis model is subjected to a significance test on the random fitting parameter null distribution. If a pre-set confidence threshold is met, it is determined to be a discriminant model. A regression coefficient array is extracted from the discriminant model, and a linear decision surface is constructed based on the regression coefficient array, outputting a classification boundary hyperplane. The discriminant model, classification boundary hyperplane, confidence region boundary, baseline smoothing parameters, standard reference spectrum vector, principal component direction, retention time location information, and latent variable spatial covariance matrix are summarized to generate a sensor parameter set.

[0054] Specifically, the process of constructing the corresponding random permutation discriminant model and the random fitting parameter zero distribution is as follows: In this embodiment, the fingerprint matrix is ​​retrieved as the independent variable set, containing absorbance features of 80 Luo Han Guo training samples at 2640 retention time points. Simultaneously, the origin category label matrix is ​​retrieved as the dependent variable set, with core origins labeled as 1 and non-core origins labeled as 0. The core of the permutation test lies in disrupting the original correspondence between the independent and dependent variables. While keeping the fingerprint matrix unchanged, a random number generator is used to shuffle the row vectors of the origin category label matrix, generating a set of pseudo-label vectors. This process is repeated 200 times, generating a set of 200 permutation label vectors. Subsequently, the original fingerprint matrix is ​​paired with these 200 permutation label vectors to construct and train 200 random permutation discriminant models. For each random model, its explained variance for the dependent variable and its predictive power obtained through seven-fold cross-validation are calculated, thereby constructing a zero-distribution space for the random fitting parameters. This space describes the performance benchmark that the model can achieve under random noise conditions.

[0055] Specifically, the process of performing a significance test on the discriminant analysis model and outputting the classification boundary hyperplane is as follows: In this embodiment, a significance test is performed within the null distribution space. The number of times the predictive power value of one of the 200 random permutation models is greater than or equal to the predictive power value of the original model is counted. This number is divided by the total number of permutations (200) to obtain the p-value for the significance test. A two-dimensional coordinate system is constructed, with the horizontal axis representing the correlation coefficient between the permutation label vector and the original label vector, and the vertical axis representing the model's predictive power value. The coordinate points of the original model and the 200 random models are projected onto this coordinate system, and linear regression fitting is performed. The intersection of the regression line and the vertical axis (where the correlation coefficient is 0) is calculated and recorded as the validation intercept. A significance level threshold of 0.05 is set. If the p-value is less than this threshold and the validation intercept is negative (usually required to be less than 0.05), the discriminant analysis model is considered statistically significant and has not overfitted, thus being identified as a discriminant model. For example, in the experiment, an intercept of -0.273 was required for the model to pass the confidence test, confirming it as a discriminant model with genuine discriminative power, rather than a pseudo-pattern generated by data coupling. If the p-value of the significance test is greater than the preset significance level threshold (0.05), or the intercept term is positive, the model optimization closed-loop mechanism will be automatically triggered: It will backtrack to the step of generating the variable projection importance score, gradually increasing the preset discriminant threshold in increments of 0.1 (e.g., from 1.0 to 1.2), eliminating redundant variables with low contribution; simultaneously, it will re-examine the confidence region boundaries, eliminating abnormal training samples where the Hotelling statistic exceeds the limit. After completing the above adjustments, the discriminant analysis model will be reconstructed based on the updated feature subset, and the significance test will be performed again until the model statistical indicators meet the above confidence requirements.

[0056] In this embodiment, a regression coefficient array is extracted from the validated identification model. This array is a column vector with the same dimension as the number of retained time points, and its value quantifies the weighted contribution of each time point to the origin classification. The classification boundary hyperplane is a linear decision surface in a high-dimensional feature space used to separate different origin categories. Its mathematical equation is described as follows: calculate the geometric center vectors of the core origin sample set and the non-core origin sample set in the training set, respectively. In the binary classification scenario of this embodiment, the classification threshold is set to 0.5. The linear equation is solidified, and the classification boundary hyperplane parameters are output. The arithmetic mean vector of the two geometric center vectors (i.e., the midpoint) is calculated. The dot product (i.e., the inner product) of the midpoint vector and the regression coefficient array is calculated, and the negative of the dot product is taken and added to the preset classification threshold (0.5). The result is used as the intercept of the classification hyperplane. To address the problem of nonlinear drift of sample features caused by different years or maturity levels, a radial basis kernel function can be introduced to map the fingerprint spectrum matrix to a high-dimensional Hilbert space before constructing the linear decision surface. In this high-dimensional space, the samples exhibit linear separability. At this point, the regression coefficient array consists of linear weights within the kernel space. After asymmetric least squares smoothing and related optimization warping, the main nonlinear interferences have been removed, and the linear decision surface is sufficient to meet the pre-set confidence requirements. When the sample to be tested is projected onto one side of this hyperplane (i.e., the calculated score is greater than 0.5), it is classified as a core product; otherwise, it is classified as a non-core product. This hyperplane constitutes the core logical threshold for automated decision-making.

[0057] In this embodiment, a boundary margin adversarial test is performed on the generated classification boundary hyperplane. Synthetic perturbation samples are generated in the neighborhood of the hyperplane to evaluate the decision plane's sensitivity to small noise, and the hyperplane coefficients are fine-tuned by maximizing the minimum classification margin strategy. Before outputting the classification boundary hyperplane, critical sample points at the "support vector" level are automatically locked, with a Euclidean distance less than 0.05 (normalized units) from the hyperplane. Using a synthetic minority oversampling technique, 500 Gaussian white noise perturbation samples (noise variance of 0.01) are generated in extremely narrow bands on both sides of the hyperplane, using these critical samples as seeds. These perturbation samples are then substituted into the current linear decision equation for testing. The proportion of samples with label flipping is counted. If the flipping proportion is greater than 5%, it indicates that the hyperplane is too close to the edge of a certain class of sample clusters, posing a risk of overfitting. At this point, a hyperplane fine-tuning mechanism is triggered: a boundary margin penalty term is added to the loss function of the discriminant analysis model. The boundary margin penalty term is calculated as follows: the square of the Euclidean distance from each generated perturbed sample point to the current classification hyperplane is calculated, the reciprocal of the squared distance is taken, and the reciprocals of all perturbed sample distances are summed. Finally, this summation is multiplied by a preset penalty weight coefficient. In this embodiment, the penalty weight coefficient is set to 0.05. The penalty weight coefficient is selected through a cross-validation strategy. Specifically, a grid search is performed with a step size ranging from 0.01 to 0.5. For each candidate coefficient value, the classification accuracy of the model on the validation set and the stability index after adversarial perturbation are calculated. The coefficient value that minimizes the misclassification rate of boundary perturbed samples while maintaining high accuracy is selected. In this embodiment, when the coefficient is 0.05, the classification boundary exhibits optimal robustness. This process forces the hyperplane to move towards sparse sample regions to obtain a larger geometric classification margin. Repeat the above test until the flipping ratio is less than or equal to 5%, and finally lock and output the robustly enhanced classification boundary hyperplane coefficients. This process improves the tolerance of the classification decision surface in the sample boundary region, thereby reducing the misclassification rate of critically ambiguous samples.

[0058] Based on this, an independent blind sample test was performed on the constructed discrimination model using 40 reserved external test set samples. The specific steps of the test are as follows: The first step, parameter freezing and preprocessing, strictly maintains the baseline smoothing parameters, standard reference spectral vector, and related optimized warp path parameters generated during the training phase. The raw chromatographic data of 40 external test set samples are imported, and asymmetric least-squares smoothing and time axis alignment are performed using the frozen parameters to generate test set fingerprints.

[0059] The second step is feature extraction and statistical calculation: Based on the determined retention time location information of the origin markers, the characteristic peak area vectors of the test set samples are extracted. Subsequently, using the trained latent variable space covariance matrix and principal component directions, the Hotelling statistic and residual distance of each test sample are calculated.

[0060] The third step, confidence region interception and classification, involves comparing the statistics of the test sample with the determined confidence region boundaries (Hotling's critical value and residual distance critical value). If the test sample falls outside the confidence region boundary, it is recorded as an "outlier"; if it falls within the confidence region boundary, it is projected onto the classification boundary hyperplane, the classification score is calculated, and the predicted origin label is output.

[0061] The fourth step is accuracy evaluation: the predicted origin label is compared with the actual origin label of the test sample to calculate the classification accuracy. In this embodiment, the model is considered to have passed the final acceptance test only if the classification accuracy of the external test set is greater than a preset standard (e.g., 95%). If the accuracy does not meet the standard, the process returns to the feature selection step to adjust the selection threshold for origin markers.

[0062] The validation on the aforementioned independent test set confirms that the model not only performs excellently on the training set, but also has a stable ability to identify the place of origin on unseen external data.

[0063] Specifically, the process of generating the sensor parameter set is as follows: In this embodiment, to achieve rapid offline detection by the smart terminal, all trained and validated key parameters are packaged. This sensor parameter set is a standardized data structure file. As a specific implementation of the present invention, the core data structure of the sensor parameter set calculated based on the above 80 samples includes the following key parameters: baseline smoothing parameters (smoothing factor 105, asymmetric weight 0.001) and standard reference spectrum vector for signal preprocessing; location information of origin marker retention time (e.g., 7.75 minutes to 7.90 minutes) for feature extraction; principal component direction vector and latent variable space covariance matrix for projection calculation; critical value of Hotelling statistic (6.85) and critical value of residual distance (2.15) for anomaly detection; and classification boundary hyperplane regression coefficients for final decision, such as the intercept term of the classification boundary hyperplane. This parameter set is encrypted and sent to the smart sensor acquisition terminal integrated with the ultra-high performance liquid chromatograph, enabling it to directly trace the origin of the tested Luo Han Guo samples without retraining the model.

[0064] By repeatedly randomly rearranging labels and constructing a zero distribution for randomly fitted parameters, the statistical significance of the discrimination model is rigorously tested to ensure that the discovered patterns of origin are genuine chemical differences rather than data coincidences, effectively avoiding falsely high accuracy rates caused by noisy fitting. Furthermore, the rigorously validated discrimination model, classification boundary hyperplane, confidence region boundary, and various preprocessing parameters are aggregated into a standardized set of sensor parameters, achieving lightweight encapsulation of complex algorithms. This not only ensures consistency across different validation sets but also allows for direct adaptation to intelligent sensing terminals, ensuring stable and reliable origin traceability based on a fixed, high-standard parameter set in real-world application scenarios, even in the face of changing on-site environments.

[0065] Further, based on the sample of monk fruit to be tested and the set of sensor parameters, a standard fingerprint spectrum is generated and the characteristic peak vector is extracted. The Hotelling statistic and residual distance are calculated. When the peak vector is within the confidence region boundary, it is projected onto the classification boundary hyperplane, and the origin label and confidence level are output. This corresponds to step S5 above. See [link to relevant documentation]. Figure 3 The specific implementation process includes: The sensor parameter set is downloaded from the cloud server to the intelligent sensor acquisition terminal; the ultra-high performance liquid chromatograph collects the sample of monk fruit to be tested and generates the original chromatographic vector to be tested; the baseline smoothing parameter and the standard reference spectrum vector are called to perform signal preprocessing on the original chromatographic vector to be tested, and the standard fingerprint spectrum to be tested is output; the retention time point response value is extracted from the standard fingerprint spectrum to be tested based on the retention time positioning information to construct the characteristic peak vector to be tested; the Mahalanobis distance of the characteristic peak vector to be tested is calculated using the latent variable space covariance matrix, and the Hotling statistic to be tested is output; the vertical Euclidean distance of the characteristic peak vector to be tested is calculated using the principal component direction, and the residual distance to be tested is output; the Hotling statistic to be tested and the residual distance to be tested are compared and verified with the confidence region boundary; after the verification is passed, the characteristic peak vector to be tested is mapped to the classification boundary hyperplane to perform linear discriminant operation, generate the discrimination result, output the origin label according to the sign attribute of the discrimination result, and calculate and output the confidence value according to the discrimination result.

[0066] Specifically, the sampling and preprocessing process of the intelligent sensing acquisition terminal is as follows: In this embodiment, the intelligent sensing acquisition terminal first downloads a set of sensing parameters containing the latest model from the cloud server via an encrypted communication protocol. This parameter set encapsulates baseline smoothing parameters (smoothing factor set to 105, asymmetric weight set to 0.001), a standard reference spectral vector, and related optimized warping algorithm configurations (fragment length of fifty points, relaxation factor of two points). The ultra-high performance liquid chromatograph acquires the absorbance data of the monk fruit sample under standard conditions, generating the original chromatographic vector to be tested. The terminal processor then calls the baseline smoothing parameters, uses the asymmetric least squares algorithm to subtract background drift, and calls the standard reference spectral vector and warping parameters to perform a piecewise linear scaling transformation on the vector to be tested. This process forces the retention time axis of the sample to be tested to be completely aligned with the training model, outputting the standard fingerprint spectrum to be tested, ensuring spatial consistency for subsequent feature extraction.

[0067] Specifically, the process of generating the Hotelling statistic and the residual distance to be measured and performing a comparison and verification with the confidence region boundary is as follows: In this embodiment, a specific time window in the standard fingerprint spectrum to be tested is accurately located based on the retention time positioning information in the sensor parameter set. For example, absorbance response values ​​with retention times between 7.75 minutes and 7.90 minutes are automatically extracted, corresponding to the characteristic absorption peak of mogroside V; simultaneously, symmonoside I response values ​​between 12.45 minutes and 12.60 minutes are extracted. Trapezoidal integral operations are performed on the discrete light intensity values ​​within the above-mentioned locked intervals to calculate the absolute peak area of ​​each component. Subsequently, these peak area values ​​are arranged according to a predetermined characteristic order to construct a multidimensional characteristic peak vector to be tested. This vector is a digital abstraction of the physical sample, and its dimension is strictly consistent with the feature space dimension during model training.

[0068] In this embodiment, the Mahalanobis distance of the characteristic peak vector to be tested is calculated using the inverse matrix of the latent variable space covariance matrix, outputting the tested Hotelling statistic to quantify the variability of the sample within the model. Simultaneously, the projection length of the tested vector in the residual space is calculated using the principal component direction vectors, outputting the tested residual distance to quantify the noise components in the sample not explained by the model. These two statistics are compared and verified with preset critical values ​​for the Hotelling statistic (e.g., 6.85) and the residual distance (e.g., 2.15), respectively. Only when both the tested Hotelling statistic and the tested residual distance are less than the critical values ​​is the sample determined to fall within the confidence region boundary and be considered a valid Luo Han Guo sample; otherwise, it is directly intercepted and reported as an abnormal sample or a non-Luo Han Guo sample.

[0069] Specifically, the process of outputting the country of origin label and confidence level value is as follows: In this embodiment, after statistical verification, the feature peak vector to be tested is mapped to a high-dimensional classification space. The calculation process is described as follows: the feature peak vector to be tested is multiplied by the regression coefficient array of the classification boundary hyperplane (i.e., the corresponding elements are multiplied and then summed), and the intercept term is added to the result to obtain the discrimination result, which is the algebraic distance from the sample point to the classification boundary hyperplane. The origin label is output according to the relationship between the discrimination score and the classification threshold (0.5 in this embodiment): if the score is greater than 0.5, it is determined to be a core origin; otherwise, it is a non-core origin. Further, the discrimination score is converted into a probability value between 0 and 1, i.e., the confidence score, using an sigmoid function. The specific calculation logic is: the confidence score is equal to 1 divided by the denominator, which is 1 plus the exponent of the natural constant; the exponent part is calculated by multiplying the algebraic distance from the sample point to the classification boundary hyperplane by the slope parameter, adding the intercept parameter, and finally taking the negative value of the calculation result. In this embodiment, the decision values ​​(i.e., algebraic distances) output by the discriminant analysis model to the validation subset during the seven-fold cross-validation process are collected. Training pairs are constructed with the decision values ​​as independent variables and the corresponding true origin labels (0 or 1) as dependent variables. A negative log-likelihood objective function is constructed, and the minimum value of this objective function is solved using the Levenberg-Marquardt algorithm to fit the optimal Sigmoid function parameters, i.e., a slope parameter of 5.2 and an intercept parameter of 0. For example, the linear discriminant score is 0.88. Substituting into the above logical calculation: the exponential part is negative (5.2 multiplied by 0.88), approximately equal to -4.576, resulting in a final confidence level of 99.0%. Finally, the terminal interface displays a standardized result such as "Origin: Core Origin, Confidence Level: 99.0%", realizing the automated translation from chemical signals to decision information.

[0070] The intelligent terminal utilizes a distributed set of sensor parameters to perform standardized baseline correction and peak alignment on the collected raw data in real time, ensuring that the test data and training data are on the same reference plane and eliminating signal distortion caused by the on-site testing environment. Before outputting the origin label, compliance verification of the confidence region boundary is enforced. Projection discrimination is only performed when the Hotelling statistic and residual distance of the test sample are both within the threshold range. Through the logic of verification before classification, combined with the specific confidence value calculated based on the classification boundary distance, the test results are given interpretability. Not only can the origin be identified, but the reliability of the result can also be judged based on the confidence value, thus achieving high-reliability traceability.

[0071] This invention provides a method for tracing and classifying the origin of Luo Han Guo (Siraitia grosvenorii) based on chromatographic fingerprinting. Through a combination of asymmetric least squares smoothing and correlation-optimized warping algorithms in preprocessing, instrument drift and background noise are effectively eliminated at the source, retention time bias is corrected, and the alignment accuracy of multiple batches of chromatographic data is ensured, providing a high-quality data foundation for subsequent analysis. Origin markers are screened using variable projection importance, extracting the most discriminative characteristic peaks from massive amounts of redundant data and reducing noise interference introduced by irrelevant variables. A rigorous confidence region boundary is constructed by introducing a dual statistical monitoring mechanism of Hotelling's statistic and residual distance. This enables not only the differentiation of origins but also the identification of abnormal samples (such as inferior or adulterated products) outside the model's coverage, avoiding misclassification caused by the forced classification of traditional classifiers. Finally, the classification boundary hyperplane determined through permutation tests exhibits high robustness, providing quantitative confidence indices while outputting origin labels, achieving a dual improvement in the accuracy and reliability of traceability results under high-interference environments.

[0072] Example 2 This embodiment aims to demonstrate the specific application of the chromatographic fingerprint-based method for tracing the origin of monk fruit (Luo Han Guo) in an industrial raw material quality control scenario. Assume a large-scale traditional Chinese medicine (TCM) decoction piece manufacturer receives a batch of monk fruit raw materials, claiming to originate from a core production area, batch number YF2024-Batch09, totaling fifty bags. To verify the true origin of this batch of raw materials and eliminate inferior or adulterated products, the quality inspection department randomly selected five samples (numbered Sample-01 to Sample-05) and one control sample (numbered Sample-HN) whose origin is known to be outside the core production area.

[0073] First, during the data acquisition phase, the researchers sequentially placed five prepared sample solutions into the autosampler of the ultra-high performance liquid chromatograph (UHPLC). The generated sensor parameter set was invoked, and the flow rate was set to 0.4 mL / min, with a detection wavelength of 203 nm. After instrument startup, the full-spectrum light source penetrated the flow cell, and the photodiode array sensor recorded spectral information in real time at a frequency of 20 Hz. After a 22-minute elution process, a raw chromatographic data matrix containing retention time, detection wavelength, and absorbance values ​​was generated. Taking Sample-01 as an example, its raw data baseline showed a drift of up to 0.08 absorbance units at 18 minutes, and the retention time lagged by 0.04 minutes relative to the standard reference spectrum.

[0074] Upon entering the signal preprocessing stage, the intelligent terminal automatically calls the baseline smoothing parameter from the parameter set, i.e., the smoothing parameter equals 105 and the asymmetric weight parameter equals 0.001. The asymmetric least squares smoothing algorithm performs iterative operations on the matrix of Sample-01. After seven iterations, the weights converge, and the fitted background curve successfully captures the low-frequency drift caused by gradient elution. After subtracting this background, the baseline fluctuation amplitude is reduced to below 0.002 absorbance units. Next, a related optimization warp algorithm is introduced, aligning the chromatographic curve of Sample-01 with the standard reference spectrum vector segment by segment based on a preset fragment length of 50 and a relaxation factor of 2. The calculation results show that the full-spectrum correlation coefficient before alignment is only 0.85. After dynamic programming optimization and reconstruction, the correlation coefficient increases to 0.98, ensuring that the chromatographic peak of mogroside V accurately falls within the standard time window of 7.75 to 7.90 minutes.

[0075] In the feature extraction and statistical validation stage, fixed-point integration was performed on Sample-01 to Sample-05 and Sample-HN based on the retention time location information. The integration area of ​​Sample-01 in the mogroside V region was 12500 mAU, and the integration area in the sympathomimetic I region was 3200 mAU. After constructing the target feature peak vector, the Hotelling statistic and residual distance of each sample were calculated using the latent variable space covariance matrix. The results showed that the target Hotelling statistic of Sample-01 was 3.2, and the target residual distance was 1.1, both far below the determined critical values ​​(6.85 and 2.15, respectively), thus it was determined to be a qualified mogroside sample "within the confidence region". The target residual distance of Sample-03 was as high as 4.5, exceeding the upper limit threshold of 2.15. An alarm was immediately triggered, indicating "Abnormal sample: outside the model noise range". After manual review, it was found that the sample had mold, which caused the chemical composition to deviate significantly from the normal metabolic characteristics of monk fruit. The sample was directly rejected and not included in the subsequent origin determination.

[0076] Finally, the process proceeds to the origin determination and confidence level output stage. For Sample-01, Sample-02, Sample-04, Sample-05, and Sample-HN, which passed statistical validation, their feature peak vectors are mapped onto the classification boundary hyperplane. The regression coefficient array of the classification hyperplane is then multiplied by the feature vectors. Sample-01 has a linear discriminant score of 0.88, which is greater than the classification threshold of 0.5, and the output result is: "Origin: Core Origin, Confidence Level: 98.5%". Sample-02, Sample-04, and Sample-05 all have scores above 0.75 and are all identified as core origins. In stark contrast, although the Hotling statistic and residual distance of the non-core origin control sample Sample-HN are within the normal range (indicating that it is a normal monk fruit), its projection score on the classification hyperplane is 0.23, significantly lower than the threshold of 0.5. This is mainly because the ratio characteristics of its symmenioside I and mogroside V are fundamentally different from those of core origins. Based on this, the output was: "Origin: Non-core origin, confidence level: 92.1%". The final quality inspection report showed that there were a few moldy and rotten fruits in this batch of raw materials (intercepted by residual distance), while the remaining samples were all genuine monk fruit from the core origin (confirmed by the classification hyperplane).

[0077] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting, characterized in that, include: Chromatographic signals of training samples of Luo Han Guo were collected by ultra-high performance liquid chromatography (UHPLC) to construct a chromatographic data matrix including retention time. The baseline of the chromatographic data matrix was subtracted using an asymmetric least squares smoothing algorithm, and the retention time drift was corrected by a related optimized warping algorithm to output a fingerprint matrix. The origin-related prediction matrix is ​​extracted from the fingerprint matrix using a discriminant analysis model; Based on the origin-related prediction matrix, the variable projection importance of chromatographic peaks in the fingerprint matrix is ​​calculated, and chromatographic peaks with a value greater than a preset threshold are selected as origin markers. Based on place of origin markers, the corresponding peak areas of the training samples of monk fruit were extracted, a characteristic peak training set was constructed, and the critical values ​​of the Hotelling statistic and the residual distance were calculated to determine the confidence region boundary. Perform a permutation test on the discriminant analysis model to determine the classification boundary hyperplane and generate a set of sensing parameters. Based on the sample of monk fruit to be tested and the set of sensor parameters, a standard fingerprint spectrum to be tested is generated and the characteristic peak vector to be tested is extracted. The Hotling statistic and the residual distance to be tested are calculated. When the peak vector of the feature to be tested is projected onto the classification boundary hyperplane within the confidence region boundary, the origin label and confidence level are output.

2. The method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting according to claim 1, characterized in that, The specific generation process of the chromatographic data matrix includes: an ultra-high performance liquid chromatograph (UHPLC) integrated into an intelligent sensor acquisition terminal; the UHPLC driving the Luo Han Guo training sample to be separated into component eluents through a chromatographic separation column and introduced into the detector flow cell; a full-band light source emitting a continuous spectral beam penetrating the detector flow cell to generate a monochromatic beam sequence; a photodiode array sensor receiving the monochromatic beam and converting it into an analog photoelectric voltage signal, outputting discrete light intensity through synchronous sampling and quantization processing; based on the discrete light intensity, generating absorbance values ​​according to the Lambert-Beer law; arranging the absorbance values ​​corresponding to different detection wavelengths in wavelength order to generate a spectral response vector; stacking the spectral response vectors according to the sampling time order to construct a chromatographic data matrix containing retention time, detection wavelength, and absorbance values; and uploading the chromatographic data matrix to a cloud server.

3. The method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting according to claim 1, characterized in that, The specific generation process of the fingerprint matrix includes: retrieving preset baseline smoothing parameters and a chromatographic data matrix; constructing an objective function containing a data fit term and a baseline smoothing term using an asymmetric least squares smoothing algorithm and performing iterative calculations; assigning different weights to data points in the baseline region and data points in the chromatographic peak region during the iteration process to fit a background drift curve; subtracting the background drift curve from the chromatographic data matrix to output a baseline correction matrix; retrieving a preset standard reference spectrum vector using a correlation optimization warp algorithm and dividing the baseline correction matrix into continuous segments along the time axis; performing scaling transformations on the continuous segments and calculating the correlation coefficient between the transformed segments and the corresponding intervals of the standard reference spectrum vector; performing cumulative determination based on the correlation coefficients and outputting an optimized path; reconstructing the baseline correction matrix based on the optimized path and outputting the fingerprint matrix.

4. The method for tracing and classifying the origin of Luo Han Guo (Siraitia grosvenorii) based on chromatographic fingerprinting according to claim 1, characterized in that, The specific construction process of the origin-related prediction matrix includes: using the fingerprint spectrum matrix as the independent variable dataset and retrieving the preset origin category label matrix as the dependent variable dataset; constructing a discriminant analysis model using the orthogonal partial least squares discriminant analysis algorithm; calculating the covariance matrix between the independent variable dataset and the dependent variable dataset using the discriminant analysis model; decomposing the fingerprint spectrum matrix into predictive structural components and orthogonal structural components based on the covariance matrix; removing the orthogonal structural components and performing regression fitting calculations on the predictive structural components to generate a regression coefficient array; applying the regression coefficient array to the fingerprint spectrum matrix for weighted reconstruction and outputting the origin-related prediction matrix.

5. The method for tracing and classifying the origin of Luo Han Guo (Siraitia grosvenorii) based on chromatographic fingerprinting according to claim 1, characterized in that, The specific generation process of the origin marker includes: retrieving the weight coefficient vector of each potential structural component in the origin-related prediction matrix and its contribution to the explained variance of the origin category label matrix; performing importance quantification calculation on each feature variable contained in the fingerprint matrix: calculating the square value of each element in the weight coefficient vector, and multiplying the square value with the corresponding explained variance contribution to generate a weighted contribution component; performing a summation operation on all weighted contribution components and dividing by the total explained variance contribution of the origin category label matrix to generate a normalized importance intermediate value; performing a square root transformation on the normalized importance intermediate value to output the variable projection importance score corresponding to each feature variable; performing a numerical comparison between the variable projection importance score and a preset discrimination threshold, and when the value exceeds the preset discrimination threshold, back-mapping the feature variable to the retention time axis, establishing the corresponding chromatographic peak as the origin marker, and recording the retention time location information of the origin marker.

6. The method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting according to claim 1, characterized in that, The specific process for generating the confidence region boundary includes: locking the corresponding chromatographic signal response interval from the fingerprint matrix based on retention time positioning information; performing an integral accumulation operation on the absorbance values ​​within the chromatographic signal response interval to generate a peak area feature vector; collecting the peak area feature vectors of all Luo Han Guo training samples to construct a feature peak training dataset; performing dimensionality reduction decomposition on the feature peak training dataset to construct a latent variable space, extracting the principal component directions, and calculating the latent variable space covariance matrix; calculating the Mahalanobis distance from each training sample point to the center using the latent variable space covariance matrix, and deriving the upper limit threshold based on multivariate statistical distribution theory to output the critical value of the Hotelling statistic; calculating the vertical Euclidean distance of the training sample points from the latent variable space, setting the upper limit threshold based on the statistical error distribution, and outputting the critical value of the residual distance; combining the in-model variation range defined by the critical value of the Hotelling statistic with the out-of-model noise range defined by the critical value of the residual distance to output the confidence region boundary.

7. The method for tracing and classifying the origin of monk fruit based on chromatographic fingerprinting according to claim 1, characterized in that, The specific process of the classification boundary hyperplane includes: retrieving the fingerprint spectrum matrix and the origin category label matrix as input data for the permutation test; performing multiple randomized rearrangement operations on the origin category label matrix to generate a set of permutation label vectors; forming training pairs between the fingerprint spectrum matrix and each vector in the set of permutation label vectors to construct the corresponding random permutation discriminant model; calculating the explained variance and cross-validation prediction capability; constructing a randomized fitting parameter null distribution; performing a significance test on the discriminant analysis model in the randomized fitting parameter null distribution; if the pre-set confidence threshold is met, it is determined to be a discriminant model; extracting the regression coefficient array from the discriminant model; constructing a linear decision surface based on the regression coefficient array and outputting the classification boundary hyperplane; and summarizing the discriminant model, classification boundary hyperplane, confidence region boundary, baseline smoothing parameter, standard reference spectrum vector, principal component direction, retention time positioning information, and latent variable space covariance matrix to generate a sensing parameter set.

8. The method for tracing and classifying the origin of Luo Han Guo (Siraitia grosvenorii) based on chromatographic fingerprinting according to claim 1, characterized in that, The specific process of the origin label and confidence level includes: downloading the sensor parameter set from the cloud server to the intelligent sensor acquisition terminal; collecting the sample of monk fruit to be tested using an ultra-high performance liquid chromatogram (UHPLC) to generate the original chromatographic vector to be tested; calling the baseline smoothing parameters and the standard reference spectrum vector to perform signal preprocessing on the original chromatographic vector to be tested, and outputting the standard fingerprint spectrum to be tested; extracting the retention time point response value from the standard fingerprint spectrum to be tested based on the retention time positioning information to construct the characteristic peak vector to be tested; calculating the Mahalanobis distance of the characteristic peak vector to be tested using the latent variable space covariance matrix, and outputting the Hotelling statistic to be tested; calculating the vertical Euclidean distance of the characteristic peak vector to be tested using the principal component direction, and outputting the residual distance to be tested; comparing and verifying the Hotelling statistic and the residual distance to be tested with the confidence region boundary; after successful verification, mapping the characteristic peak vector to be tested to the classification boundary hyperplane to perform linear discriminant operation, generating the discrimination result; outputting the origin label based on the sign attribute of the discrimination result and calculating and outputting the confidence level value based on the discrimination result.