Purchase wine authenticity identification method fusing multivariate statistics and metabolic feature extraction
By integrating multivariate statistical methods with metabolic feature extraction, the structural heterogeneity and complexity of spirits are quantified, and the threshold is dynamically adjusted. This solves the problem of misjudgment in traditional spirits identification methods, which struggle to distinguish between samples from different processes using the same raw materials. This enables high-precision identification and traceability of the authenticity of spirits.
Patent Information
- Application Number
- CN202511048502.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for identifying spirits are insufficient to accurately distinguish between samples from the same source of raw materials but with different processing techniques. Traditional methods rely on peak area or mass-to-charge ratio, which are insufficient to capture subtle structural differences, resulting in a high false positive rate.
By integrating multivariate statistical and metabolic feature extraction methods, including full-scan extraction of signal peaks, construction of a set of non-dominant structural parameters, calculation of the comprehensive heterogeneity index FCY, and embedding the OPLS-DA model, the structural heterogeneity and complexity of samples are quantified, the cluster boundary threshold is dynamically adjusted, and atypical samples are identified.
It significantly improves the accuracy of distinguishing between different batches of the same type of spirits, reduces misjudgments, provides visual feedback on structural differences, enhances the fine-grainedness and accuracy of identifying complex spirits, and adapts to samples with market volatility.
Smart Images

Figure CN120948679A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of authenticity identification technology for spirits, specifically a method for authenticity identification of spirits that integrates multivariate statistics and metabolic feature extraction. Background Technology
[0002] In recent years, with the widespread application of GC-MS in complex food systems, researchers have gradually introduced it into the identification of origin, detection of adulteration, and evaluation of processing consistency of spirits products. In this type of analysis, conventional target components are easily affected by the variety of raw materials, making it difficult to accurately reflect deep structural differences. Therefore, how to extract non-obvious features with discriminative value from complex metabolic profiles has become one of the key challenges in the refined identification of the authenticity of spirits.
[0003] Existing models for identifying and classifying spirits typically rely on multivariate statistical methods such as principal component analysis (PCA), discriminant analysis (DA), or partial least squares regression (PLSR), distinguishing based solely on the position and intensity of metabolic peaks. While these methods are effective in differentiating between different types or brands of spirits, their accuracy drops significantly for samples from the same source of raw materials that differ only in processing techniques.
[0004] The root cause of the aforementioned identification difficulties lies in the fact that the metabolic profiles of spirits contain a large number of structurally similar but functionally isomers of metabolic fragments. These structural variants usually do not significantly change the mass-to-charge ratio or peak intensity, but they exhibit subtle but crucial differences in the internal molecular structure, such as "ring structure rearrangement, aromatic configuration conversion, and side chain breakage." Traditional models based on peak area differences struggle to capture these deep variations, especially when sample batches are similar and origin labels are ambiguous. This can easily lead to situations where "highly similar counterfeit spirits are misjudged as genuine" or "genuine samples are misidentified as suspicious," seriously affecting the credibility of quality control and regulatory oversight. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction, thus solving the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction, comprising the following steps:
[0007] S1. By performing a full scan extraction on the original spectrum obtained from the sample analysis by HS-SPME-GC-MS, all signal peaks are identified, and background peaks and co-elution noise are removed to obtain the initial screening set of metabolic peaks, Praw.
[0008] S2. Match the mass-to-charge ratio MZi of each effective peak i in the initial screening set of metabolic peaks Praw to the mass spectrometry database to obtain the corresponding inferred molecular structure. Construct a molecular graph structure using cheminformatics tools and extract the set of non-dominant structural parameters.
[0009] S3. Construct a structural parameter vector xP based on the set of non-obvious structural parameters, and obtain the structural heterogeneity variance Svar by calculating the degree of variation of the structural parameter vector xP within the same sample.
[0010] S4. Integrate the graph complexity factor Dp and structural heterogeneity variance Svar in the set of non-explicit structural parameters to construct the comprehensive heterogeneity index FCY.
[0011] S5. Embed the comprehensive heterogeneity index FCY into the orthogonal partial least squares discriminant analysis space of the samples, observe the offset of samples with high comprehensive heterogeneity index FCY values in the space, and obtain the sample authenticity deviation distance DFL.
[0012] S6. Compare the obtained sample authenticity deviation distance DFL with the set cluster boundary threshold Tdf to determine whether it is an atypical sample.
[0013] Preferably, S1 includes S11 and S12;
[0014] S11. Identify and export all peak information from the original GC-MS spectrum in the acquisition software ChemStation, and use the peak detection algorithm to extract all signal peaks, and obtain the peak mass-to-charge ratio MZ, peak area A and peak retention time rT.
[0015] The mass-to-charge ratio MZ of the peak is obtained from the ionic fragments after the sample molecules are broken up in the ionization source by the signal peak detected by GC-MS.
[0016] Acquisition method: After GC separation, the molecules enter the MS section; ionization methods such as electron bombardment and chemical ionization break the molecules apart; the fragments are accelerated into a quadrupole or flight tube and separated according to M / Z; at each time point, the system records all detected ion signals and their peak mass-to-charge ratios (MZ);
[0017] Peak area A represents the integral value of the chromatographic peak, which indicates the relative content or abundance of the substance in the sample.
[0018] Acquisition method: Each compound was separated by GC and sequentially introduced into a mass spectrometer according to its retention time; an ion current chromatogram was generated showing the change of ion current over time in the chromatogram.
[0019] Peak retention time rT represents the time it takes for a substance to travel from injection to reaching the detector;
[0020] Acquisition method: The chromatographic column separates each compound in the sample according to its physicochemical properties; each component elutes in a certain order, and the elution time of each peak is recorded as the peak retention time rT;
[0021] S12. The obtained peak mass-to-charge ratio MZ, peak area A, and peak retention time rT are processed to construct the peak intensity variation index VMZ, and invalid and duplicate peaks caused by ionization interference, solvent baseline shift, and co-eluting substances are removed from the GC-MS data.
[0022] The peak intensity variation index (VMZ) is obtained using the following formula:
[0023] ;
[0024] In the formula, VMZi represents the peak intensity variation index of the i-th effective peak, σt(Ai) represents the local standard deviation of the area of the i-th effective peak, pA represents the average peak area, log represents the logarithmic function, rTi represents the peak retention time of the i-th effective peak, and prT represents the average retention time of all peaks.
[0025] The noise was filtered out by the obtained peak intensity variation index VMZ, and the mass-to-charge ratio MZ, peak area A and peak retention time rT of the filtered peaks were fitted to obtain the initial set of metabolic peaks Praw={MZi|VMZ<TV, and A>TA}.
[0026] The exclusion criteria are:
[0027] The peak intensity variation index VMZ < TV, and the peak area A > TA;
[0028] In the formula, TV represents the variation index threshold, and TA represents the minimum peak area threshold.
[0029] Preferably, S2 includes S21 and S22;
[0030] S21. Perform structural information matching on the mass-to-charge ratio MZi of each effective peak i to obtain the corresponding speculative molecular structure.
[0031] The mass-to-charge ratio MZi of each effective peak i is used as the search keyword and input into the mainstream mass spectrometry MassBank database platform. During the search process, the peak area A and peak retention time rT are simultaneously introduced to form a three-dimensional feature description tuple of the sample peak. This is used to filter the interference of background and co-elution peaks and to select the most likely structure from the candidate results in the database.
[0032] Among the multiple candidate compounds returned by the database, the candidates were sorted and selected in the following two ways: retention time window comparison and chromatographic behavior similarity score;
[0033] The final output is a description of the chemical structure corresponding to the mass-to-charge ratio MZi of each valid peak i, presented in a standardized structural representation: including linear symbols representing the molecular topology SMILES and the international chemical structural bond InChIKey;
[0034] Using the RDKit cheminformatics tool, the SMILES format molecules with linear symbolic representations of molecular topology are converted into molecular graph structures Gi with effective peak i.
[0035] Preferably, S22 extracts non-dominant structural parameter indices from the molecular graph structure Gi of each effective peak i, including the number of ring structures Rc, aromaticity factor Af, unsaturation factor Us, and spectral complexity factor Dp, and fits them into a set of non-dominant structural parameters.
[0036] The number of ring structures, Rc, is obtained by identifying the molecular diagram structure Gi of the effective peak i using a ring detection algorithm.
[0037] The aromaticity factor Af is obtained through the following formula:
[0038] ;
[0039] In the formula, Nari represents the number of aromatic rings in the i-th effective peak, a1 represents the aromatic ring adjustment factor with a value of 2.5, Ncoi represents the number of conjugated double bonds in the i-th effective peak, and a2 represents the conjugation adjustment factor with a value of 1.2.
[0040] The unsaturation factor Us is obtained using the following formula:
[0041] ;
[0042] In the formula, Ci represents the number of carbon atoms in the i-th effective peak, Hi represents the number of hydrogen atoms in the i-th effective peak, Xi represents the total number of halogen atoms in the i-th effective peak, and Ni represents the number of nitrogen atoms in the i-th effective peak.
[0043] The graph complexity factor Dp is obtained using the following formula:
[0044] ;
[0045] In the formula, Dpi represents the spectral complexity factor of the i-th effective peak, (vj, vk) represents the j-th and k-th atomic nodes in the molecular graph structure G, and path-length represents the shortest path length in the molecular structure graph; that is, the number of hops through covalent bonds.
[0046] Preferably, S3 includes S31 and S32;
[0047] S31. Construct a structural parameter vector xP=[Rc, Af, Us] based on the obtained set of non-explicit structural parameters;
[0048] Extract n fragments from each sample and construct the corresponding structure parameter matrix ZxP:
[0049] The structural parameter matrix ZxP is obtained using the following formula:
[0050] ;
[0051] S32. Based on the structural parameter matrix ZxP, calculate the overall variability of the structural parameter vector within the same sample, obtain the structural parameter mean vector pZxP, and analyze it to obtain the structural heterogeneity variance Svar.
[0052] The structural parameter mean vector pZxP is obtained using the following formula:
[0053] ;
[0054] In the formula, ZxPi represents the structure parameter vector of the i-th effective peak;
[0055] The structural heterogeneity variance Svar is obtained using the following formula:
[0056] .
[0057] Preferably, S4 includes S41 and S42;
[0058] S41. Integrate and process the obtained graph complexity factor Dp to obtain the total sample complexity index Dpsum;
[0059] The total sample complexity index Dpsum is obtained by summing the spectral complexity factors Dp of all valid peaks i in the sample;
[0060] By constructing a nonlinear penalty function, the total sample complexity index Dpsum is analyzed to obtain the complexity adjustment coefficient Zdp;
[0061] The complexity adjustment factor Zdp is obtained using the following formula:
[0062] ;
[0063] In the formula, log represents the logarithmic function.
[0064] Preferably, the obtained structural heterogeneity variance Svar is integrated with the complexity adjustment coefficient Zdp to construct a comprehensive heterogeneity index FCY;
[0065] The composite heterogeneity index FCY is obtained using the following formula:
[0066] ;
[0067] In the formula, c1 represents the exponential adjustment coefficient;
[0068] The heterogeneity of the samples is determined by obtaining the comprehensive heterogeneity index FCY.
[0069] When the comprehensive heterogeneity index FCY ≤ 1.2, it indicates a low heterogeneity range; the metabolic structural fragments in the sample are highly consistent in the three dimensions of "number of ring structures, aromaticity, and degree of unsaturation";
[0070] The overall complexity of the graph is low, indicating that its structure is highly clustered in space;
[0071] This indicates that their processing sources, raw material processes, fermentation metabolic pathways, etc., are highly consistent;
[0072] When 1.2 < FCY (Comprehensive Heterogeneity Index) ≤ 2.5, it indicates a moderate heterogeneity range; there is a certain degree of non-dominant structural variation in the fragments within the sample; there are slight fluctuations in process consistency, or some different brewing source components have been added; the spectral complexity is in the moderate range, and some structural fragments show slightly non-standard patterns.
[0073] When 2.5 < FCY (Comprehensive Heterogeneity Index), it indicates a high heterogeneity range; the non-dominant structural differences of the structural fragments within the sample are significant and highly discrete; the same sample bottle may contain metabolic fragments from multiple sources, with a chaotic structural distribution; the spectral complexity is high, but the structures are irregular and show a non-convergent trend.
[0074] Preferably, S5 includes S51 and S52;
[0075] S51. The comprehensive heterogeneity index FCY is used as an additional quantitative variable, which is combined with the original metabolic peak principal component and embedded into the OPLS-DA multidimensional discriminant space to obtain a new input dataset Zinput=[XPCA,FCY].
[0076] Where XPCA represents the principal component score matrix, a two-dimensional matrix with dimensions n×k, and XPCA∈R n×k This represents the top k principal component components extracted from each sample;
[0077] The newly acquired input dataset Zinput is fed into the OPLS-DA model to construct a discriminant vector space containing class labels. Orthogonal interference removal factors are introduced into the principal factor subspace to achieve noise degradation of non-discriminative dimensions.
[0078] The FCY exponent variable is used as the "structure-driven feature quantity" of the projected principal component of the model and is given explicit weights in the principal space of the model.
[0079] By introducing this variable, the OPLS-DA model can be more interpretable for "structurally heterogeneous dominant deviation samples".
[0080] Compared to traditional OPLS-DA discriminant models based on peak area or abundance features, this structure is more suitable for identifying hidden mutation paths in samples that deviate from the true labels but have structural rationality.
[0081] Preferably, in step S52, the principal component spatial distance Dpc and the heterogeneity index perturbation factor Dfcy are constructed using the principal component score matrix XPCA and the comprehensive heterogeneity index FCY.
[0082] The principal factor spatial distance Dpc is obtained using the following formula:
[0083] ;
[0084] In the formula, XPCAsa represents the projection coordinates of the principal component score matrix in the OPLS principal space, and XPCAre represents the center point of the standard genuine sample space;
[0085] The heterogeneity exponential perturbation factor Dfcy is obtained using the following formula:
[0086] ;
[0087] In the formula, FCYsa represents the projection value of the comprehensive heterogeneity index in the OPLS principal space, and pFCYre represents the average value of the comprehensive heterogeneity index in the standard sample;
[0088] The obtained principal factor spatial distance Dpc and heterogeneous index perturbation factor Dfcy are combined to obtain the sample authenticity deviation distance DFL;
[0089] The Deviation from True Value (DFL) is obtained using the following formula:
[0090] ;
[0091] In the formula, D1 represents the heterogeneous perturbation adjustment factor.
[0092] Preferably, S6, all true deviation distances in the sample set are integrated to obtain the deviation distance set QDFL={DFL1, DFL2, ..., DFLj}, and a probability density function f(DFL) is established.
[0093] The probability density function f(DFL) is obtained by the following formula:
[0094] ;
[0095] In the formula, m represents the total number of samples, h represents the bandwidth parameter, K() represents the kernel function, and DFLj represents the deviation distance of the j-th sample;
[0096] The obtained probability density function f (DFL) is combined with the interquartile range to obtain the cluster boundary threshold Tdf;
[0097] The cluster boundary threshold Tdf is obtained using the following formula:
[0098] TDF = Q3 + β * IQR;
[0099] In the formula, Q3 represents the third quartile, IQR represents the interquartile range, IQR = Q3 - Q1, Q1 represents the first quartile, and β represents the adjustment factor;
[0100] The obtained sample authenticity deviation distance (DFL) is compared with the cluster boundary threshold (Tdf) to determine whether the sample is an atypical sample.
[0101] The determination method is obtained through the following means:
[0102] When the sample authenticity deviation distance DFL is less than or equal to the cluster boundary threshold Tdf, it means that the sample has not been identified as an atypical structural offset sample.
[0103] When the sample authenticity deviation distance DFL is greater than the cluster boundary threshold Tdf, it indicates that the sample is identified as an atypical structural offset sample.
[0104] For atypical structural migration samples, all structural fragments are traversed and the SMILES expressions of the structural fragments are extracted. They are classified according to the peak mass-to-charge ratio MZ, the number of ring structures Rc, and the aromaticity factor Af, and then mapped to the two-dimensional structural fingerprint space. The corresponding processing source labels, such as place of origin and distillation temperature range, are superimposed.
[0105] This invention provides a method for identifying the authenticity of spirits by integrating multivariate statistical analysis and metabolic feature extraction, which has the following beneficial effects:
[0106] (1) By incorporating non-dominant structural differences in metabolic profiles into the analytical model, the limitations of traditional methods that rely solely on peak area or mass-to-charge ratio for differentiation are overcome. Traditional methods struggle to distinguish metabolic differences caused by variations in different processes or environments. However, this method, through multi-dimensional non-targeted metabolic feature analysis, can accurately capture the impact of processing technology on the metabolic structure of wines, significantly improving the differentiation accuracy between different batches of the same variety and wines made from the same raw materials but processed using different methods.
[0107] This invention constructs a comprehensive heterogeneity index (FCY) by combining structural heterogeneity variance and spectral complexity factor, effectively solving the difficulty in judging alcohol samples due to their complex structure and latent variations. Traditional methods often fail to accurately distinguish between real and fake samples due to small sample variations, while the introduction of the comprehensive heterogeneity index can quantitatively describe the differences in sample structure, providing a more quantifiable basis for distinguishing real and fake samples.
[0108] (2) This invention constructs a comprehensive heterogeneity index (FCY) that combines structural complexity with heterogeneity, enabling a quantitative description of the complexity and differences of metabolic fragments within a sample. Compared to traditional methods, this method not only focuses on the intensity and position of metabolic peaks but also deeply explores the structural heterogeneity of metabolites, significantly improving the ability to comprehensively identify the authenticity of samples.
[0109] By comparing the sample authenticity deviation distance (DFL) with the cluster boundary threshold (Tdf), this embodiment avoids the errors of manually setting fixed thresholds in traditional methods. It can adaptively adjust the boundary of anomaly detection based on sample characteristics, thereby improving the flexibility and accuracy of the model. This dynamic mechanism improves the ability to accurately distinguish between "typical samples" and "atypical samples" in different batches of samples, and is especially suitable for samples with a certain degree of market volatility.
[0110] (3) By calculating the graph complexity factor and integrating it into the total sample complexity index, and combining it with the nonlinear penalty function to adjust the sample complexity adjustment coefficient, a quantitative evaluation of the impact of structural complexity on sample heterogeneity is provided. This new approach can effectively solve the problem in traditional methods of not being able to distinguish between "complex but consistent" wine samples and "simple but highly different" wine samples, enhance the fine-grained identification ability of complex wines, and avoid misjudgment caused by simple indicators.
[0111] The introduction of the Comprehensive Heterogeneity Index (FCY) combines the influencing factors of structural heterogeneity and spectral complexity to form a comprehensive and quantitative standard for judging the degree of heterogeneity of samples. By setting three ranges—low, medium, and high—it can effectively distinguish different types of spirits samples, provide more accurate heterogeneity determination, and help regulatory or quality control personnel better identify suspicious samples in actual operation.
[0112] (4) By introducing an orthogonal interference removal factor, noise degradation is performed in the principal factor space of the OPLS-DA model, ensuring that the model focuses on the most discriminative features rather than irrelevant noise variables. This approach reduces data redundancy and interference from irrelevant factors on the discrimination results, improving the stability and accuracy of the model. In particular, it can effectively improve the reliability of the model when facing complex datasets or diverse samples.
[0113] This invention quantifies the distance between a sample and the center of a standard sample by calculating the principal component spatial distance (DPC) and the heterogeneity index perturbation factor (DFcy), thereby calculating the sample authenticity deviation distance. This comprehensive indicator can quantitatively describe the degree of sample heterogeneity and, through numerical deviation distance, more intuitively reflect the degree of authenticity deviation of the sample, facilitating the rapid identification of potentially problematic samples. Attached Figure Description
[0114] Figure 1 This is a schematic diagram illustrating the steps of a method for identifying the authenticity of spirits that integrates multivariate statistics and metabolic feature extraction, according to the present invention. Detailed Implementation
[0115] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0116] Example 1
[0117] This invention provides a method for identifying the authenticity of spirits by integrating multivariate statistical methods and metabolic feature extraction. Please refer to [link / reference]. Figure 1 This includes the following steps:
[0118] S1. By performing a full scan extraction on the original spectrum obtained from the sample analysis by HS-SPME-GC-MS, all signal peaks are identified, and background peaks and co-elution noise are removed to obtain the initial screening set of metabolic peaks, Praw.
[0119] S2. Match the mass-to-charge ratio MZi of each effective peak i in the initial screening set of metabolic peaks Praw to the mass spectrometry database to obtain the corresponding inferred molecular structure. Construct a molecular graph structure using cheminformatics tools and extract the set of non-dominant structural parameters.
[0120] S3. Construct a structural parameter vector xP based on the set of non-obvious structural parameters, and obtain the structural heterogeneity variance Svar by calculating the degree of variation of the structural parameter vector xP within the same sample.
[0121] S4. Integrate the graph complexity factor Dp and structural heterogeneity variance Svar in the set of non-explicit structural parameters to construct the comprehensive heterogeneity index FCY.
[0122] S5. Embed the comprehensive heterogeneity index FCY into the orthogonal partial least squares discriminant analysis space of the samples, observe the offset of samples with high comprehensive heterogeneity index FCY values in the space, and obtain the sample authenticity deviation distance DFL.
[0123] S6. Compare the obtained sample authenticity deviation distance DFL with the set cluster boundary threshold Tdf to determine whether it is an atypical sample.
[0124] In this embodiment, by incorporating non-dominant structural differences in metabolic profiles into the analytical model, the limitations of traditional methods that rely solely on peak area or mass-to-charge ratio for differentiation are overcome. Traditional methods struggle to distinguish metabolic differences arising from variations in different processes or environments, while this method, through multi-dimensional non-targeted metabolic feature analysis, can accurately capture the impact of processing techniques on the metabolic structure of wines, significantly improving the accuracy of distinguishing between different batches of the same variety and wines made from the same raw materials but processed using different methods.
[0125] This invention constructs a comprehensive heterogeneity index (FCY) by combining structural heterogeneity variance and spectral complexity factor, effectively solving the difficulty in judging alcohol samples due to their complex structure and latent variations. Traditional methods often fail to accurately distinguish between real and fake samples due to small sample variations, while the introduction of the comprehensive heterogeneity index can quantitatively describe the differences in sample structure, providing a more quantifiable basis for distinguishing real and fake samples.
[0126] Traditional deviation assessments typically rely on fixed thresholds, which are susceptible to fluctuations in sample size and changes in the experimental environment. In contrast, this method dynamically constructs a cluster boundary threshold (TDF) and adaptively assesses deviations based on the true deviation distribution of samples, significantly enhancing the ability to identify atypical samples and avoiding the misjudgments or omissions that may occur under static threshold settings.
[0127] In this method, highly heterogeneous structural fragments deviating from the sample are visualized, directly listing the SMILES form, mass-to-charge ratio, and possible processing origins of the structural fragments. This function not only improves the accuracy of sample tracing but also provides data support for optimizing production processes. Traditional methods struggle to trace minute structural differences in complex wine profiles, while visual feedback of structural differences can effectively identify the sources of counterfeit wines and non-standard processing.
[0128] Example 2
[0129] This embodiment is an explanation based on Embodiment 1. Please refer to it. Figure 1 Specifically: S1 includes S11 and S12;
[0130] S11. Identify and export all peak information from the original GC-MS spectrum in the acquisition software ChemStation, and use the peak detection algorithm to extract all signal peaks, and obtain the peak mass-to-charge ratio MZ, peak area A and peak retention time rT.
[0131] The mass-to-charge ratio MZ of the peak is obtained from the ionic fragments after the sample molecules are broken up in the ionization source by the signal peak detected by GC-MS.
[0132] Peak area A represents the integral value of the chromatographic peak, which represents the relative content or abundance of the substance in the sample;
[0133] Peak retention time rT represents the time it takes for a substance to travel from injection to reaching the detector;
[0134] S12. The obtained peak mass-to-charge ratio MZ, peak area A, and peak retention time rT are processed to construct the peak intensity variation index VMZ, and invalid and duplicate peaks caused by ionization interference, solvent baseline shift, and co-eluting substances are removed from the GC-MS data.
[0135] The peak intensity variation index (VMZ) is obtained using the following formula:
[0136] ;
[0137] In the formula, VMZi represents the peak intensity variation index of the i-th effective peak, σt(Ai) represents the local standard deviation of the area of the i-th effective peak, pA represents the average peak area, log represents the logarithmic function, rTi represents the peak retention time of the i-th effective peak, and prT represents the average retention time of all peaks.
[0138] The noise was filtered out by the obtained peak intensity variation index VMZ, and the mass-to-charge ratio MZ, peak area A and peak retention time rT of the filtered peaks were fitted to obtain the initial set of metabolic peaks Praw={MZi|VMZ<TV, and A>TA}.
[0139] The exclusion criteria are:
[0140] The peak intensity variation index VMZ < TV, and the peak area A > TA;
[0141] In the formula, TV represents the variation index threshold, and TA represents the minimum peak area threshold.
[0142] S2 includes S21 and S22;
[0143] S21. Perform structural information matching on the mass-to-charge ratio MZi of each effective peak i to obtain the corresponding speculative molecular structure.
[0144] The mass-to-charge ratio MZi of each effective peak i is used as the search keyword and input into the mainstream mass spectrometry MassBank database platform. During the search process, the peak area A and peak retention time rT are introduced simultaneously to form a three-dimensional feature description tuple of the sample peak.
[0145] Among the multiple candidate compounds returned by the database, the candidates were sorted and selected in the following two ways: retention time window comparison and chromatographic behavior similarity score;
[0146] The final output is a description of the chemical structure corresponding to the mass-to-charge ratio MZi of each valid peak i, presented in a standardized structural representation: including linear symbols representing the molecular topology SMILES and the international chemical structural bond InChIKey;
[0147] Using the RDKit cheminformatics tool, the SMILES format molecular topology, represented by linear symbols, is converted into the molecular graph structure Gi of each effective peak i. S22 extracts non-dominant structural parameters from the molecular graph structure Gi of each effective peak i, including the number of ring structures Rc, aromaticity factor Af, unsaturation factor Us, and spectral complexity factor Dp, and fits them into a set of non-dominant structural parameters.
[0148] The number of ring structures, Rc, is obtained by identifying the molecular diagram structure Gi of the effective peak i using a ring detection algorithm.
[0149] The aromaticity factor Af is obtained through the following formula:
[0150] ;
[0151] In the formula, Nari represents the number of aromatic rings in the i-th effective peak, a1 represents the aromatic ring adjustment factor, Ncoi represents the conjugated double bond count in the i-th effective peak, and a2 represents the conjugation adjustment factor.
[0152] The unsaturation factor Us is obtained using the following formula:
[0153] ;
[0154] In the formula, Ci represents the number of carbon atoms in the i-th effective peak, Hi represents the number of hydrogen atoms in the i-th effective peak, Xi represents the total number of halogen atoms in the i-th effective peak, and Ni represents the number of nitrogen atoms in the i-th effective peak.
[0155] The graph complexity factor Dp is obtained using the following formula:
[0156] ;
[0157] In the formula, Dpi represents the spectral complexity factor of the i-th effective peak, (vj, vk) represents the j-th and k-th atomic nodes in the molecular graph structure G, and path-length represents the shortest path length in the molecular structure graph.
[0158] This embodiment introduces non-dominant metabolic features, including the number of ring structures (Rc), the aroma factor (Af), and the degree of unsaturation factor (Us), and combines them with cheminformatics tools to effectively identify wine samples from the same raw material source but with different processing techniques. This improvement solves the problem faced by traditional methods in distinguishing samples of the same variety from different production lines, especially when the raw materials are similar and there are no obvious differences in surface characteristics, enabling accurate differentiation through deeper metabolic feature identification.
[0159] In step S1, the peak intensity variation index (VMZ) is introduced to effectively screen out invalid and duplicate peaks caused by ionization interference, co-eluting substances, etc. This pretreatment step ensures that only true and valid metabolic peaks are used in subsequent analyses, reducing the interference of noise on the analytical results and significantly improving the accuracy and stability of the analysis.
[0160] By matching the mass-to-charge ratio of each effective peak to a database, the corresponding molecular structure can be accurately inferred, and a set of non-dominant structural parameters, such as the number of ring structures, aromaticity factor, and unsaturation factor, can be extracted. These structural parameters can accurately reflect subtle differences in metabolites and provide more refined structural feature analysis than traditional methods, thereby improving the accuracy of identifying the authenticity of alcoholic beverages.
[0161] This invention constructs a comprehensive heterogeneity index (FCY), combining structural complexity with heterogeneity, to quantitatively describe the complexity and variability of metabolic fragments within a sample. Compared to traditional methods, this approach not only focuses on the intensity and location of metabolic peaks but also deeply explores the structural heterogeneity of metabolites, significantly improving the comprehensive ability to identify the authenticity of samples.
[0162] By comparing the sample authenticity deviation distance (DFL) with the cluster boundary threshold (Tdf), this embodiment avoids the errors of manually setting fixed thresholds in traditional methods. It can adaptively adjust the boundary of anomaly detection based on sample characteristics, thereby improving the flexibility and accuracy of the model. This dynamic mechanism improves the ability to accurately distinguish between "typical samples" and "atypical samples" in different batches of samples, and is especially suitable for samples with a certain degree of market volatility.
[0163] By combining structural fingerprint cloud maps with origin labels for visualization, this method can not only determine whether a sample is atypical, but also provide specific SMILES forms, mass-to-charge ratios, and possible processing sources for each deviation sample. This visualization function provides direct technical evidence for quality supervision, helping inspection personnel quickly locate problems and conduct further traceability investigations, thus improving technical transparency and operability.
[0164] One of the core innovations of this embodiment is a non-targeted identification method based on metabolic characteristics. This method is not only applicable to traditional spirits samples but can also be extended to the authenticity testing of emerging or non-standard products. This approach can provide more precise technical means for the quality supervision of alcoholic beverages in the global market, offering strong support for combating counterfeit and substandard products and protecting consumer rights.
[0165] Example 3
[0166] This embodiment is an explanation based on Embodiment 2. Please refer to it. Figure 1 Specifically: S3 includes S31 and S32;
[0167] S31. Construct a structural parameter vector xP=[Rc, Af, Us] based on the obtained set of non-explicit structural parameters;
[0168] Extract n fragments from each sample and construct the corresponding structure parameter matrix ZxP:
[0169] The structural parameter matrix ZxP is obtained using the following formula:
[0170] ;
[0171] S32. Based on the structural parameter matrix ZxP, calculate the overall variability of the structural parameter vector within the same sample, obtain the structural parameter mean vector pZxP, and analyze it to obtain the structural heterogeneity variance Svar.
[0172] The structural parameter mean vector pZxP is obtained using the following formula:
[0173] ;
[0174] In the formula, ZxPi represents the structure parameter vector of the i-th effective peak;
[0175] The structural heterogeneity variance Svar is obtained using the following formula:
[0176] .
[0177] S4 includes S41 and S42;
[0178] S41. Integrate and process the obtained graph complexity factor Dp to obtain the total sample complexity index Dpsum;
[0179] The total sample complexity index Dpsum is obtained by summing the spectral complexity factors Dp of all valid peaks i in the sample;
[0180] By constructing a nonlinear penalty function, the total sample complexity index Dpsum is analyzed to obtain the complexity adjustment coefficient Zdp;
[0181] The complexity adjustment factor Zdp is obtained using the following formula:
[0182] ;
[0183] In the formula, log represents the logarithmic function.
[0184] S42. Integrate the obtained structural heterogeneity variance Svar with the complexity adjustment coefficient Zdp to construct the comprehensive heterogeneity index FCY;
[0185] The composite heterogeneity index FCY is obtained using the following formula:
[0186] ;
[0187] In the formula, c1 represents the exponential adjustment coefficient;
[0188] The heterogeneity of the samples is determined by obtaining the comprehensive heterogeneity index FCY.
[0189] When the comprehensive heterogeneity index FCY ≤ 1.2, it indicates a low heterogeneity range;
[0190] When 1.2 < Composite Heterogeneity Index (FCY) ≤ 2.5, it indicates the intermediate heterogeneity range;
[0191] When 2.5 < the composite heterogeneity index FCY, it indicates a high heterogeneity range.
[0192] This embodiment introduces a set of non-dominant structural parameters and constructs a structural parameter vector based on these features, enabling a more detailed capture of structural differences in alcoholic beverage samples. This is particularly true for spirits of the same variety and origin, where subtle structural differences caused by processing techniques and production batches can be identified. Compared to traditional methods that rely solely on surface features such as aroma and color, this method deeply mines the metabolic structure of alcoholic beverage samples, accurately identifying non-dominant changes in complex alcoholic beverages.
[0193] In step S3, by calculating and filtering the peak intensity variation index (VMZ) of metabolic peaks for each sample, invalid and duplicate peaks caused by factors such as ionization interference and solvent baseline shift can be effectively eliminated. This processing method significantly improves the accuracy of subsequent data analysis, ensuring that only high-quality and reliable metabolic peak data are used, thereby reducing errors caused by noise contamination and improving the stability and accuracy of the final model.
[0194] This embodiment calculates the graph complexity factor and integrates it into a total sample complexity index. By combining this with a nonlinear penalty function to adjust the sample complexity adjustment coefficient, it provides a quantitative evaluation of the impact of structural complexity on sample heterogeneity. This novel approach effectively solves the problem in traditional methods of failing to distinguish between "complex but consistent" wine samples and "simple but highly varied" wine samples, enhancing the fine-grained identification capability of complex wines and avoiding misjudgments caused by simple indicators.
[0195] The introduction of the Comprehensive Heterogeneity Index (FCY) combines the influencing factors of structural heterogeneity and spectral complexity to form a comprehensive and quantitative standard for judging the degree of heterogeneity of samples. By setting three ranges—low, medium, and high—it can effectively distinguish different types of spirits samples, provide more accurate heterogeneity determination, and help regulatory or quality control personnel better identify suspicious samples in actual operation.
[0196] The comprehensive heterogeneity index in this embodiment not only provides a fixed calculation result, but also dynamically adjusts the threshold to better adapt the model to different batches and sources of spirits. By introducing an index adjustment coefficient, the sensitivity of the model can be adaptively adjusted to effectively cope with the diversity of structural differences in different samples, avoiding the errors or inadequacies of traditional static threshold settings.
[0197] The comprehensive heterogeneity index in this embodiment not only helps identify the heterogeneity of samples but also provides a basis for the subsequent "structural difference visualization" module. Visualization tools can display information such as the SMILES structure and mass-to-charge ratio of highly variable structural fragments. Combined with possible processing origins, this enhances the interpretability of the analysis results and the operability of practical operations, providing more transparent technical support for quality control by regulatory authorities and consumer choice.
[0198] Example 4
[0199] This embodiment is an explanation based on Embodiment 3. Please refer to it. Figure 1 Specifically: S5 includes S51 and S52;
[0200] S51. The comprehensive heterogeneity index FCY is used as an additional quantitative variable, which is combined with the original metabolic peak principal component and embedded into the OPLS-DA multidimensional discriminant space to obtain a new input dataset Zinput=[XPCA,FCY].
[0201] Where XPCA represents the principal component score matrix, a two-dimensional matrix with dimensions n×k, and XPCA∈R n×k This represents the top k principal component components extracted from each sample;
[0202] The newly acquired input dataset Zinput is fed into the OPLS-DA model to construct a discriminant vector space containing class labels. Orthogonal interference removal factors are introduced into the principal factor space to achieve noise degradation of non-discriminative dimensions.
[0203] S52. Construct the principal factor spatial distance Dpc and the heterogeneity index perturbation factor Dfcy using the principal component score matrix XPCA and the comprehensive heterogeneity index FCY.
[0204] The principal factor spatial distance Dpc is obtained using the following formula:
[0205] ;
[0206] In the formula, XPCAsa represents the projection coordinates of the principal component score matrix in the OPLS principal space, and XPCAre represents the center point of the standard genuine sample space;
[0207] The heterogeneity exponential perturbation factor Dfcy is obtained using the following formula:
[0208] ;
[0209] In the formula, FCYsa represents the projection value of the comprehensive heterogeneity index in the OPLS principal space, and pFCYre represents the average value of the comprehensive heterogeneity index in the standard sample;
[0210] The obtained principal factor spatial distance Dpc and heterogeneous index perturbation factor Dfcy are combined to obtain the sample authenticity deviation distance DFL;
[0211] The Deviation from True Value (DFL) is obtained using the following formula:
[0212] ;
[0213] In the formula, D1 represents the heterogeneous perturbation adjustment factor.
[0214] This embodiment incorporates the comprehensive heterogeneity index (FCY) as an additional quantitative variable, synergistically with the original metabolic peak principal component (XPCA), and embeds it into the multidimensional discriminant space of the OPLS-DA model. Compared to traditional methods that rely solely on the simple statistical characteristics of metabolic peaks, combining the FCY index enables more comprehensive and accurate sample classification and discrimination capabilities through multidimensional information fusion. This method not only captures the metabolic patterns of samples but also incorporates heterogeneity features into the analysis, significantly improving the accuracy of distinguishing complex samples.
[0215] This embodiment introduces an orthogonal interference removal factor to degrade noise in the principal factor space of the OPLS-DA model, ensuring that the model focuses on the most discriminative features rather than irrelevant noise variables. This approach reduces data redundancy and interference from irrelevant factors on the discrimination results, improving the model's stability and accuracy. It is particularly effective in enhancing the model's reliability when dealing with complex datasets or diverse samples.
[0216] This invention quantifies the distance between a sample and the center of a standard sample by calculating the principal component spatial distance (DPC) and the heterogeneity index perturbation factor (DFcy), thereby calculating the sample authenticity deviation distance. This comprehensive indicator can quantitatively describe the degree of sample heterogeneity and, through numerical deviation distance, more intuitively reflect the degree of authenticity deviation of the sample, facilitating the rapid identification of potentially problematic samples.
[0217] In this embodiment, the calculated deviation distance allows for dynamic determination of whether a sample is atypical, without relying on a fixed discrimination threshold. This adaptive discrimination mechanism avoids the errors caused by using a fixed threshold in traditional methods, and can more flexibly adapt to the natural fluctuations between different samples, improving its adaptability and accuracy to diverse samples.
[0218] Example 5
[0219] This embodiment is an explanation based on Embodiment 4. Please refer to it. Figure 1 Specifically: S6. Integrate all true deviation distances in the sample set to obtain the deviation distance set QDFL={DFL1, DFL2, ..., DFLj}, and establish the probability density function f(DFL);
[0220] The probability density function f(DFL) is obtained by the following formula:
[0221] ;
[0222] In the formula, m represents the total number of samples, h represents the bandwidth parameter, K() represents the kernel function, and DFLj represents the deviation distance of the j-th sample;
[0223] The obtained probability density function f (DFL) is combined with the interquartile range to obtain the cluster boundary threshold Tdf;
[0224] The cluster boundary threshold Tdf is obtained using the following formula:
[0225] TDF = Q3 + β * IQR;
[0226] In the formula, Q3 represents the third quartile, IQR represents the interquartile range, IQR = Q3 - Q1, Q1 represents the first quartile, and β represents the adjustment factor;
[0227] The obtained sample authenticity deviation distance (DFL) is compared with the cluster boundary threshold (Tdf) to determine whether the sample is an atypical sample.
[0228] The determination method is obtained through the following means:
[0229] When the sample authenticity deviation distance DFL is less than or equal to the cluster boundary threshold Tdf, it means that the sample has not been identified as an atypical structural offset sample.
[0230] When the sample authenticity deviation distance DFL is greater than the cluster boundary threshold Tdf, it indicates that the sample is identified as an atypical structural offset sample.
[0231] For atypical structural migration samples, all structural fragments are traversed and the SMILES expressions of the structural fragments are extracted. The fragments are classified according to the peak mass-to-charge ratio MZ, the number of ring structures Rc, and the aromaticity factor Af. The fragments are then mapped to the two-dimensional structural fingerprint space and the corresponding processing source labels are superimposed.
[0232] This embodiment dynamically calculates the cluster boundary threshold by combining the probability density function of the deviation distance with the interquartile range, avoiding the errors that may be caused by statically setting a fixed threshold in traditional methods. Through this method, the model can adaptively adjust the threshold according to the distribution characteristics of the samples, effectively avoiding misjudgments caused by sample fluctuations or changes in external factors, and enhancing the ability to identify atypical samples.
[0233] In step S6, by calculating the deviation distance of the sample, the degree of deviation in metabolic characteristics and processing technology can be quantitatively described. Combining the probability density function and quartile discrimination logic, it is possible to accurately assess whether a sample belongs to the typical range. This processing method provides a data-driven method for determining authenticity, overcoming the instability caused by traditional manually set judgment criteria, and improving the scientific rigor and accuracy of sample discrimination.
[0234] For wine samples identified as atypical structural shifts, the example process involves traversing all structural fragments and extracting their SMILES expressions. These fragments are then categorized based on mass-to-charge ratio, number of ring structures, and aromaticity factor, and further mapped to a two-dimensional structural fingerprint space for analysis. This function not only identifies counterfeit wine samples but also provides specific structural information, such as relevant chemical characteristics and possible processing origins, offering reliable data support for subsequent traceability, quality control, and production process optimization.
[0235] This embodiment enhances the interpretability of sample traceability by overlaying structural fingerprints and processing origin tags, enabling inspectors to clearly identify the specific source and processing characteristics of deviation samples. This function provides transparent and accurate technical support for quality supervision, product verification, and market risk assessment, effectively preventing counterfeit and substandard products from entering the market and protecting consumer rights. By combining kernel density estimation and quartile statistical methods, this embodiment can more precisely capture subtle differences between samples, especially when dealing with diverse sample sets. It can accurately classify different categories, particularly significantly improving sensitivity to minor processing differences or variations in alcoholic beverages, thereby reducing misclassification rates and improving the accuracy of the classification model.
[0236] This method does not rely on fixed rules or manually set thresholds. Instead, it uses data-driven dynamic analysis to flexibly adapt to different spirits samples and has good universality. Whether dealing with complex spirits or samples from different origins and batches, it can effectively adapt to the differences in their metabolic characteristics, ensuring the wide applicability of the method.
[0237] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction, characterized in that: Includes the following steps: S1. By performing a full scan extraction on the original spectrum obtained from the sample analysis by HS-SPME-GC-MS, all signal peaks are identified, and background peaks and co-elution noise are removed to obtain the initial screening set of metabolic peaks, Praw. S2. Match the mass-to-charge ratio MZi of each effective peak i in the initial screening set of metabolic peaks Praw to the mass spectrometry database to obtain the corresponding inferred molecular structure. Construct a molecular graph structure using cheminformatics tools and extract the set of non-dominant structural parameters. S3. Construct a structural parameter vector xP based on the set of non-obvious structural parameters, and obtain the structural heterogeneity variance Svar by calculating the degree of variation of the structural parameter vector xP within the same sample. S4. Integrate the graph complexity factor Dp and structural heterogeneity variance Svar in the set of non-explicit structural parameters to construct the comprehensive heterogeneity index FCY. S5. Embed the comprehensive heterogeneity index FCY into the orthogonal partial least squares discriminant analysis space of the samples, observe the offset of samples with high comprehensive heterogeneity index FCY values in the space, and obtain the sample authenticity deviation distance DFL. S6. Compare the obtained sample authenticity deviation distance DFL with the set cluster boundary threshold Tdf to determine whether it is an atypical sample.
2. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 1, characterized in that: S1 includes S11 and S12; S11. Identify and export all peak information from the original GC-MS spectrum in the acquisition software ChemStation, and use the peak detection algorithm to extract all signal peaks, and obtain the peak mass-to-charge ratio MZ, peak area A and peak retention time rT. The mass-to-charge ratio MZ of the peak is obtained from the ionic fragments after the sample molecules are broken up in the ionization source by the signal peak detected by GC-MS. Peak area A represents the integral value of the chromatographic peak, which represents the relative content or abundance of the substance in the sample; Peak retention time rT represents the time it takes for a substance to travel from injection to reaching the detector; S12. The obtained peak mass-to-charge ratio MZ, peak area A, and peak retention time rT are processed to construct the peak intensity variation index VMZ, and invalid and duplicate peaks caused by ionization interference, solvent baseline shift, and co-eluting substances are removed from the GC-MS data. The peak intensity variation index (VMZ) is obtained using the following formula: ; In the formula, VMZi represents the peak intensity variation index of the i-th effective peak, σt(Ai) represents the local standard deviation of the area of the i-th effective peak, pA represents the average peak area, log represents the logarithmic function, rTi represents the peak retention time of the i-th effective peak, and prT represents the average retention time of all peaks. The noise was filtered out by the obtained peak intensity variation index VMZ, and the mass-to-charge ratio MZ, peak area A and peak retention time rT of the filtered peaks were fitted to obtain the initial set of metabolic peaks Praw={MZi|VMZ<TV, and A>TA}. The exclusion criteria are: The peak intensity variation index VMZ < TV, and the peak area A > TA; In the formula, TV represents the variation index threshold, and TA represents the minimum peak area threshold.
3. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 2, characterized in that: S2 includes S21 and S22; S21. Perform structural information matching on the mass-to-charge ratio MZi of each effective peak i to obtain the corresponding inferred molecular structure; The mass-to-charge ratio MZi of each effective peak i is used as the search keyword and input into the mainstream mass spectrometry MassBank database platform. During the search process, the peak area A and peak retention time rT are introduced simultaneously to form a three-dimensional feature description tuple of the sample peak. Among the multiple candidate compounds returned by the database, the candidates were sorted and selected in the following two ways: retention time window comparison and chromatographic behavior similarity score; The final output is a description of the chemical structure corresponding to the mass-to-charge ratio MZi of each valid peak i, presented in a standardized structural representation: including linear symbols representing the molecular topology SMILES and the international chemical structural bond InChIKey; Using the RDKit cheminformatics tool, the SMILES format molecules with linear symbolic representations of molecular topology are converted into molecular graph structures Gi with effective peak i.
4. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 3, characterized in that: S22 extracts non-dominant structural parameter indices from the molecular graph structure Gi of each effective peak i, including the number of ring structures Rc, aromaticity factor Af, unsaturation factor Us, and spectral complexity factor Dp, and fits them into a set of non-dominant structural parameters. The number of ring structures, Rc, is obtained by identifying the molecular diagram structure Gi of the effective peak i using a ring detection algorithm. The aromaticity factor Af is obtained through the following formula: ; In the formula, Nari represents the number of aromatic rings in the i-th effective peak, a1 represents the aromatic ring adjustment factor, Ncoi represents the number of conjugated double bonds in the i-th effective peak, and a2 represents the conjugation adjustment factor. The unsaturation factor Us is obtained using the following formula: ; In the formula, Ci represents the number of carbon atoms in the i-th effective peak, Hi represents the number of hydrogen atoms in the i-th effective peak, Xi represents the total number of halogen atoms in the i-th effective peak, and Ni represents the number of nitrogen atoms in the i-th effective peak. The graph complexity factor Dp is obtained using the following formula: ; In the formula, Dpi represents the spectral complexity factor of the i-th effective peak, (vj, vk) represents the j-th and k-th atomic nodes in the molecular graph structure G, and path-length represents the shortest path length in the molecular structure graph.
5. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 4, characterized in that: S3 includes S31 and S32; S31. Construct a structural parameter vector xP=[Rc, Af, Us] based on the obtained set of non-explicit structural parameters; Extract n fragments from each sample and construct the corresponding structure parameter matrix ZxP: The structural parameter matrix ZxP is obtained using the following formula: ; S32. Based on the structural parameter matrix ZxP, calculate the overall variability of the structural parameter vector within the same sample, obtain the structural parameter mean vector pZxP, and analyze it to obtain the structural heterogeneity variance Svar. The structural parameter mean vector pZxP is obtained using the following formula: ; In the formula, ZxPi represents the structure parameter vector of the i-th effective peak; The structural heterogeneity variance Svar is obtained using the following formula: 。 6. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 5, characterized in that: S4 includes S41 and S42; S41. Integrate and process the obtained graph complexity factor Dp to obtain the total sample complexity index Dpsum; The total sample complexity index Dpsum is obtained by summing the spectral complexity factors Dp of all valid peaks i in the sample; By constructing a nonlinear penalty function, the total sample complexity index Dpsum is analyzed to obtain the complexity adjustment coefficient Zdp; The complexity adjustment factor Zdp is obtained using the following formula: ; In the formula, log represents the logarithmic function.
7. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 6, characterized in that: S42. Integrate the obtained structural heterogeneity variance Svar with the complexity adjustment coefficient Zdp to construct the comprehensive heterogeneity index FCY; The composite heterogeneity index FCY is obtained using the following formula: ; In the formula, c1 represents the exponential adjustment coefficient; The heterogeneity of the samples is determined by obtaining the comprehensive heterogeneity index FCY. When the comprehensive heterogeneity index FCY ≤ 1.2, it indicates a low heterogeneity range; When 1.2 < Composite Heterogeneity Index (FCY) ≤ 2.5, it indicates the intermediate heterogeneity range; When 2.5 < the composite heterogeneity index FCY, it indicates a high heterogeneity range.
8. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 7, characterized in that: S5 includes S51 and S52; S51. The comprehensive heterogeneity index FCY is used as an additional quantitative variable, which is combined with the original metabolic peak principal component and embedded into the OPLS-DA multidimensional discriminant space to obtain a new input dataset Zinput=[XPCA,FCY]. Where XPCA represents the principal component score matrix, a two-dimensional matrix with dimensions n×k, and XPCA∈R n×k This represents the top k principal component components extracted from each sample; The newly acquired input dataset Zinput is fed into the OPLS-DA model to construct a discriminant vector space containing class labels. Orthogonal interference removal factors are introduced into the principal factor space to achieve noise degradation of non-discriminative dimensions.
9. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 8, characterized in that: S52. Construct the principal factor spatial distance Dpc and the heterogeneity index perturbation factor Dfcy using the principal component score matrix XPCA and the comprehensive heterogeneity index FCY. The principal factor spatial distance Dpc is obtained using the following formula: ; In the formula, XPCAsa represents the projection coordinates of the principal component score matrix in the OPLS principal space, and XPCAre represents the center point of the standard genuine sample space; The heterogeneity exponential perturbation factor Dfcy is obtained using the following formula: ; In the formula, FCYsa represents the projection value of the comprehensive heterogeneity index in the OPLS principal space, and pFCYre represents the average value of the comprehensive heterogeneity index in the standard sample; The obtained principal factor spatial distance Dpc and heterogeneous index perturbation factor Dfcy are combined to obtain the sample authenticity deviation distance DFL; The Deviation from True Value (DFL) is obtained using the following formula: ; In the formula, D1 represents the heterogeneous perturbation adjustment factor.
10. The method for identifying the authenticity of spirits by integrating multivariate statistics and metabolic feature extraction according to claim 9, characterized in that: S6. Integrate all true deviations in the sample set to obtain the deviation set QDFL={DFL1, DFL2, ..., DFLj}, and establish the probability density function f(DFL). The probability density function f(DFL) is obtained by the following formula: ; In the formula, m represents the total number of samples, h represents the bandwidth parameter, K() represents the kernel function, and DFLj represents the deviation distance of the j-th sample; The obtained probability density function f (DFL) is combined with the interquartile range to obtain the cluster boundary threshold Tdf; The cluster boundary threshold Tdf is obtained using the following formula: TDF = Q3 + β * IQR; In the formula, Q3 represents the third quartile, IQR represents the interquartile range, IQR = Q3 - Q1, Q1 represents the first quartile, and β represents the adjustment factor; The obtained sample authenticity deviation distance (DFL) is compared with the cluster boundary threshold (Tdf) to determine whether the sample is an atypical sample. The determination method is obtained through the following means: When the sample authenticity deviation distance DFL is less than or equal to the cluster boundary threshold Tdf, it means that the sample has not been identified as an atypical structural offset sample. When the sample authenticity deviation distance DFL is greater than the cluster boundary threshold Tdf, it indicates that the sample is identified as an atypical structural offset sample. For atypical structural migration samples, all structural fragments are traversed and the SMILES expressions of the structural fragments are extracted. The fragments are classified according to the peak mass-to-charge ratio MZ, the number of ring structures Rc, and the aromaticity factor Af. The fragments are then mapped to the two-dimensional structural fingerprint space and the corresponding processing source labels are superimposed.