A method for detecting and analyzing the content of liquid flavor substances in beverages

By employing multi-source signal acquisition and correlation mapping, data preprocessing fusion, and model building, this method addresses the issues of incomplete signal features, insufficient data purity, and inadequate model generalization ability in existing beverage flavor substance detection technologies. It enables accurate detection and systematic analysis of beverage flavor substance content, thereby improving the comprehensiveness and reliability of the detection.

CN121364222BActive Publication Date: 2026-04-03SICHUAN SHIYU ZHIHUI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies for detecting flavor compounds in beverages suffer from problems such as incomplete signal characteristics, insufficient data purity, inadequate model generalization ability, and imperfect result analysis. These issues result in insufficient comprehensiveness, accuracy, and practicality of the detection results, making it difficult to meet the needs of precise detection.

Method used

A method involving multi-source signal capture and correlation mapping, data preprocessing and fusion, model building and deviation analysis is adopted. Characteristic parameters are obtained through dielectric response and surface acoustic wave signals. Combined with ensemble empirical mode decomposition, Bayesian network and gradient boosting tree algorithm, data purification and feature fusion are performed to build a flavor substance content calculation model. Temporal consistency verification and deviation analysis are also performed.

Benefits of technology

It enables comprehensive and accurate detection and systematic analysis of flavor substance content in beverages, improving the comprehensiveness and reliability of detection, enhancing the generalization ability and applicability of the model, and providing precise quality control guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121364222B_ABST
    Figure CN121364222B_ABST
Patent Text Reader

Abstract

This invention discloses a method for detecting and analyzing the content of liquid flavor substances in beverages, belonging to the field of beverage quality testing technology. The method acquires the dielectric response signal and surface acoustic wave propagation signal of flavor substances from beverage samples, integrates them through correlation mapping to form a structured raw data set; preprocesses the data using ensemble empirical mode decomposition, integrates features through a Bayesian network multi-source data fusion algorithm, and constructs an optimized content calculation model based on a gradient boosting tree algorithm; after verifying the temporal consistency and dimensionality matching of the dataset, it is input into the model to generate prediction results; relevant content values ​​are extracted and deviations from preset standard ranges are analyzed, the contribution weights of signal features are calculated, and a test report including compliance judgment and deviation tracing is output. This method can comprehensively characterize the physicochemical features of flavor substances, achieve accurate prediction and deviation tracing, improve the comprehensiveness, reliability, and efficiency of detection, and provide strong support for beverage quality control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of beverage quality testing technology, and in particular to a method for detecting and analyzing the content of liquid flavor substances in beverages. Background Technology

[0002] In the field of beverage production and quality control, flavor compounds, as core factors determining the consistency of taste, aroma, and quality, have always been a key focus of the industry for content detection and analysis. With the large-scale development of the food industry and the increasing demands of consumers for beverage quality, flavor compound content detection technology is constantly evolving towards greater precision, efficiency, and systematization. Currently, the industry has developed various technical pathways for flavor compound detection, covering key aspects such as signal acquisition, data processing, model prediction, and result analysis. In signal acquisition, single or multiple signal detection technologies, such as dielectric response, surface acoustic waves, and spectroscopy, are commonly used to obtain characteristic parameters related to the physicochemical properties of flavor compounds. Data processing incorporates methods such as empirical mode decomposition and filtering algorithms to optimize data quality. Model construction often relies on machine learning algorithms such as gradient boosting trees and support vector machines to predict content. Result analysis often combines preset standard ranges for compliance judgment, providing a reference for beverage quality control. The application of these technologies has initially realized the detection of flavor compound content in beverages, supporting the basic quality control needs in the beverage production process.

[0003] However, existing technologies for detecting flavor compounds in beverages still face numerous unresolved issues, directly impacting the comprehensiveness, accuracy, and practicality of the results. First, current technologies largely rely on single-signal detection or simple multi-signal superposition, failing to fully explore the unique response characteristics of different signals to flavor compounds. This results in incomplete characterization of the physicochemical properties of flavor compounds, making it difficult to fully reflect the true state of various flavor compounds within a complex matrix. Second, data preprocessing offers limited effectiveness in removing noise and matrix interference, and the lack of a scientific weighting mechanism for multi-source data fusion leads to insufficient purity of feature data, affecting the reliability of subsequent model construction. Third, predictive model construction often lacks a systematic optimization process, failing to fully consider the attribute stratification and data distribution characteristics of flavor compounds, resulting in insufficient model generalization ability and predictive accuracy that cannot meet the demands of precise detection. Furthermore, the data verification process is inadequate, lacking systematic verification of temporal consistency and dimensional matching, potentially leading to invalid data input into the model and further reducing predictive reliability. Finally, result analysis often focuses only on the compliance of individual substance contents, lacking deviation analysis of flavor synergistic grouping and calculation of signal feature contribution weights, making it difficult to trace deviations and provide precise guidance for quality problem rectification. These problems make it difficult for existing testing technologies to achieve comprehensive, accurate detection and systematic analysis of the flavor substance content in beverages, thus hindering the improvement of beverage quality control. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for detecting and analyzing the content of liquid flavor substances in beverages.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] A method for detecting and analyzing the content of liquid flavor substances in beverages is provided, the method comprising the following steps:

[0007] S1. Obtain the dielectric response signal of flavor substances and the surface acoustic wave propagation signal of the beverage sample to be tested, record the dielectric constant, loss tangent, wave velocity offset, and attenuation coefficient, establish a correlation mapping based on the specific response characteristics of the two types of signals to flavor substances, and integrate them to form a structured raw data set.

[0008] S2. Preprocess the original dataset by using ensemble empirical mode decomposition to eliminate noise and mode mixing, integrate the two types of feature data through a Bayesian network multi-source data fusion algorithm, retain the sensitive signal dimension, and construct and optimize the flavor substance content calculation model based on the gradient boosting tree algorithm.

[0009] S3. Perform temporal consistency and dimensionality matching checks on the preprocessed dataset. After passing the checks, input the dataset into the model to generate prediction results for the content of each individual flavor substance.

[0010] S4. Extract individual, flavor synergy group, and total content values, perform deviation analysis with preset standard ranges, calculate signal feature contribution weights, determine whether the beverage meets quality requirements, and output a test analysis report including compliance determination and deviation traceability.

[0011] Furthermore, step S1 specifically includes the following sub-steps:

[0012] S1.1 The beverage samples to be tested are pretreated by a combination of centrifugation and solid-phase microextraction to remove suspended particles and macromolecular impurities, adsorb volatile and semi-volatile flavor substances, and avoid matrix interference.

[0013] S1.2 Simultaneously acquire preliminary signals from different regions with a set number of samples, calculate the dielectric constant variation coefficient and wave velocity stability index. If both meet the preset threshold, proceed to the next step; otherwise, repeat the preprocessing.

[0014] S1.3 Acquire two types of signals under a set frequency range and excitation conditions, construct an association matrix based on the signal-specific response characteristics, and integrate them to form a structured original data set.

[0015] Furthermore, step S2 specifically includes the following sub-steps:

[0016] S2.1 uses ensemble empirical mode decomposition to decompose the two types of data, eliminating the interference from the beverage matrix and obtaining purified data;

[0017] S2.2 Construct a data probabilistic network structure, determine the dependency relationship of feature nodes and calculate the confidence level, allocate fusion weights based on the confidence level, and superimpose them to form a unified feature dataset;

[0018] S2.3 The training and validation sets are divided into hierarchical groups according to the chemical properties of flavor substances. The gradient boosting tree algorithm is used to select key dimensions. The parameters are adjusted by the validation set error and the coefficient of determination to obtain the optimized model.

[0019] Furthermore, in step S2, noise removal employs a local outlier algorithm to eliminate global outlier data points, and data calibration uses a partial least squares regression calibration model constructed with a series of concentration gradient standard flavor substance mixed solutions to eliminate system errors caused by equipment drift and matrix interference.

[0020] Furthermore, in step S2, when constructing the model, leave-one-out cross-validation is used to divide the training set and validation set, and the model is stratified according to the functional attribute categories of flavor substances. L2 regularization term and early stopping mechanism are introduced to avoid overfitting and improve generalization ability.

[0021] Furthermore, step S3 specifically includes the following sub-steps:

[0022] S3.1 performs format standardization, feature integrity and time sequence consistency checks on the dataset, adapts it to the range of flavor substance content in beverages, ensures that no key features are missing and that the similarity with the training data meets the standard, and if it fails, it is reprocessed.

[0023] S3.2 Input the validated dataset into the model to derive the content, and simultaneously calculate the confidence interval based on bootstrap resampling to characterize the reliable range of the prediction results.

[0024] Furthermore, step S4 specifically includes the following sub-steps:

[0025] S4.1 Divide flavor synergy groups according to the synergistic relationship between sweetness and aroma, and sourness and mouthfeel, and retrieve the corresponding preset standard range;

[0026] S4.2 Calculate the contribution ratio of the relative deviation of individual content and the group deviation. If all indicators meet the standards, the beverage is deemed qualified. Otherwise, mark the unqualified items and trace the cause, and output a complete test analysis report.

[0027] Furthermore, in step S4, the analysis results are grouped according to both chemical structure similarity and flavor synergy. The coefficient of variation of content within each group and the coefficient of synergy between groups are calculated to form a five-level analysis result. The report simultaneously presents the data and the correlation analysis with the flavor quality of the beverage.

[0028] Furthermore, the fusion weight is determined by calculating the signal-to-noise ratio of dielectric characteristic signals, the integrity index of surface acoustic wave characteristic signals, and the mutual information values ​​of the two types of data. It is dynamically adjusted to retain the sensitive signal dimension and leverage the complementary value of the data.

[0029] Furthermore, the model update sets two triggering conditions: the accumulation of prediction error and the drift of feature importance. Standard sample data of beverages within a set time period are selected, and after preprocessing and fusion, incremental training is performed, updating only the key parameters of tree depth and learning rate to complete the iteration.

[0030] The beneficial effects of this invention are:

[0031] (1) By capturing and mapping multi-source signals, preprocessing and fusing data, building models and analyzing deviations, the physicochemical characteristics of flavor substances are fully characterized, so as to achieve accurate prediction of content and traceability of deviations, and ensure the comprehensiveness and reliability of detection.

[0032] (2) By using signal purification, abnormal data removal and calibration techniques, interference from matrix, noise and other systematic errors can be effectively eliminated, improving data purity and accuracy, and providing high-quality feature input for model construction;

[0033] (3) By using model hierarchical training, regularization optimization and dynamic update mechanism, the generalization ability and long-term applicability of the model are enhanced, the stable prediction accuracy is maintained, and the dynamic needs of different beverage detection scenarios are adapted. Attached Figure Description

[0034] Figure 1 A flowchart illustrating the steps of a method for detecting and analyzing the content of liquid flavor substances in a beverage.

[0035] Figure 2 The following is a flowchart illustrating the implementation steps of a method for detecting and analyzing the content of liquid flavor substances in a beverage, provided as an example. Detailed Implementation

[0036] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Example 1

[0038] See Figure 1 This paper provides a method for detecting and analyzing the content of liquid flavor substances in beverages, which includes the following steps:

[0039] S1. Obtain the dielectric response signal of flavor substances and the surface acoustic wave propagation signal of the beverage sample to be tested, record the dielectric constant, loss tangent, wave velocity offset, and attenuation coefficient, establish a correlation mapping based on the specific response characteristics of the two types of signals to flavor substances, and integrate them to form a structured raw data set.

[0040] S2. Preprocess the original dataset by using ensemble empirical mode decomposition to eliminate noise and mode mixing, integrate the two types of feature data through a Bayesian network multi-source data fusion algorithm, retain the sensitive signal dimension, and construct and optimize the flavor substance content calculation model based on the gradient boosting tree algorithm.

[0041] S3. Perform temporal consistency and dimensionality matching checks on the preprocessed dataset. After passing the checks, input the dataset into the model to generate prediction results for the content of each individual flavor substance.

[0042] S4. Extract individual, flavor synergy group, and total content values, perform deviation analysis with preset standard ranges, calculate signal feature contribution weights, determine whether the beverage meets quality requirements, and output a test analysis report including compliance determination and deviation traceability.

[0043] In some embodiments, the dielectric response signal and near-infrared spectral signal of flavor substances in the beverage sample to be tested are acquired, and the dielectric constant, loss tangent, characteristic peak intensity, and absorbance are recorded. A mapping rule is established based on the specific response relationship of the two types of signals to flavor substances, and a standardized raw data set is formed. Wavelet packet decomposition is used to eliminate noise and mode mixing. The two types of feature data are integrated through a weighted fusion algorithm, prioritizing the retention of signal dimensions sensitive to flavor substance content. A flavor substance content calculation model is constructed and optimized based on the random forest algorithm. The preprocessed dataset is checked for data integrity and dimensionality adaptability. After the check passes, the data is input into the model to generate prediction results for the content of each individual flavor substance. The individual, flavor function group, and total content values ​​are extracted, and the degree of deviation from the preset standard range is evaluated. The influence weight of each signal feature on the content deviation is calculated, and it is determined whether the beverage quality standard is met. A test report containing the conclusion of compliance and the analysis of the reasons for the deviation is output.

[0044] Step S1 specifically includes the following sub-steps:

[0045] S1.1 The beverage samples to be tested are pretreated by a combination of centrifugation and solid-phase microextraction to remove suspended particles and macromolecular impurities, adsorb volatile and semi-volatile flavor substances, and avoid matrix interference.

[0046] S1.2 Simultaneously acquire preliminary signals from different regions with a set number of samples, calculate the dielectric constant variation coefficient and wave velocity stability index. If both meet the preset threshold, proceed to the next step; otherwise, repeat the preprocessing.

[0047] S1.3 Acquire two types of signals under a set frequency range and excitation conditions, construct an association matrix based on the signal-specific response characteristics, and integrate them to form a structured original data set.

[0048] In some embodiments, the beverage samples to be tested undergo a combined pretreatment of ultrasonic-assisted extraction and solid-phase extraction. Ultrasonic extraction enhances the release of flavor substances, while solid-phase extraction columns adsorb the target flavor substances, removing large molecular interferences such as proteins and polysaccharides, as well as fine suspended impurities. Preliminary signals from multiple uniformly distributed regions of the sample are simultaneously acquired, and the fluctuation coefficient of the dielectric signal and the consistency index of the surface acoustic wave signal are calculated. The sample's performance is then comprehensively assessed based on the signal intensity compliance, and if it does not meet the requirements, the pretreatment is repeated. Two types of signals are acquired under optimized frequency ranges and excitation parameters, respectively. A feature mapping table is constructed based on the signal response to flavor substances, and the data is integrated into a structured raw data set according to a preset data structure to ensure the orderliness and traceability of the data.

[0049] Step S2 specifically includes the following sub-steps:

[0050] S2.1 uses ensemble empirical mode decomposition to decompose the two types of data, eliminating the interference from the beverage matrix and obtaining purified data;

[0051] S2.2 Construct a data probabilistic network structure, determine the dependency relationship of feature nodes and calculate the confidence level, allocate fusion weights based on the confidence level, and superimpose them to form a unified feature dataset;

[0052] S2.3 The training and validation sets are divided into hierarchical groups according to the chemical properties of flavor substances. The gradient boosting tree algorithm is used to select key dimensions. The parameters are adjusted by the validation set error and the coefficient of determination to obtain the optimized model.

[0053] In some embodiments, a combination of empirical mode decomposition (EMD) and wavelet denoising is used to process the two types of data. First, EMD separates the matrix interference components from the effective signal, and then wavelet denoising is used to further purify the data to obtain high-purity feature data. A fusion model based on DS evidence theory is constructed to determine the correlation strength between feature nodes and calculate the credibility. Fusion weights are assigned according to the credibility, and the two types of feature data are nonlinearly fused to form a unified feature dataset. The training set and validation set are hierarchically divided according to the functional attribute categories of flavor substances. The dataset is input into the support vector machine algorithm, and key feature dimensions are selected by recursive feature elimination. The model parameters are adjusted based on the prediction accuracy and stability index of the validation set to obtain a performance-optimized flavor substance content calculation model.

[0054] In step S2, noise removal uses a local outlier algorithm to remove global outlier data points, and data calibration uses a partial least squares regression calibration model constructed with a series of concentration gradient standard flavor substance mixed solutions to eliminate systematic errors caused by equipment drift and matrix interference.

[0055] In some embodiments, noise removal employs an isolated forest algorithm to identify and remove outlier data points from the original dataset. Multiple isolated trees are constructed to score the data points in isolation, and outlier data is determined and removed based on the scoring results. Data calibration uses standard flavor substance solutions with different concentration gradients, and their corresponding characteristic signals are collected to construct a multiple linear regression calibration model. The actual content of the standard substances is used as a reference, and the original data is corrected through calibration equations. This effectively compensates for systematic errors caused by factors such as equipment parameter drift and interference from beverage matrix components, thereby improving the accuracy and reliability of the data.

[0056] In step S2, when building the model, leave-one-out cross-validation is used to divide the training set and the validation set, and the model is stratified according to the functional attribute categories of flavor substances. L2 regularization and early stopping mechanism are introduced to avoid overfitting and improve generalization ability.

[0057] In some embodiments, when constructing the flavor compound content calculation model, the k-fold cross-validation method is used to divide the training set and the validation set, and the samples are stratified according to the chemical structure type of the flavor compounds to ensure that the samples at each level are evenly distributed in the training set and the validation set. An L1 regularization term is introduced to constrain the model parameters, and the influence of redundant features is reduced by sparsifying the parameters. At the same time, the dropout mechanism is combined to randomly deactivate some model nodes, reducing the model's over-reliance on training data. By monitoring the error change trend of the training set and the validation set, the regularization strength and dropout probability are dynamically adjusted to further improve the model's generalization ability and adaptability to different beverage samples.

[0058] Step S3 specifically includes the following sub-steps: S3.1 Standardize the format of the dataset, verify the integrity of features and the consistency of time series, adapt it to the range of flavor substance content in beverages, ensure that there are no missing key features and that the similarity with the training data meets the standard, and if it fails, reprocess it; S3.2 Input the dataset that has passed the verification into the model to derive the content, and simultaneously calculate the confidence interval based on the bootstrap method resampling to characterize the reliable range of the prediction results.

[0059] In some embodiments, the preprocessed dataset undergoes data format standardization and key feature integrity verification. This standardizes the representation and magnitude range of different types of features, checks the completeness of features that play a crucial role in predicting flavor substance content, and returns the dataset to the preprocessing stage for reprocessing if any features are missing or have abnormal formats. The verified dataset is then input into the optimized model for content derivation. Simultaneously, the Jackknife resampling method is used to calculate the confidence interval for each individual flavor substance content. By sequentially removing each sample and performing predictions, the confidence interval is determined based on the distribution characteristics of multiple prediction results, clearly characterizing the reliable range of the prediction results and providing a reference for subsequent result analysis.

[0060] Step S4 specifically includes the following sub-steps: S4.1 Divide the flavor synergy groups according to the synergistic relationship between sweetness and aroma, and sourness and taste, and retrieve the corresponding preset standard range; S4.2 Calculate the contribution ratio of the relative deviation of the individual content and the group deviation. If all indicators meet the standards, the beverage is judged to be qualified. Otherwise, mark the unqualified items and trace the cause, and output a complete test analysis report.

[0061] In some embodiments, flavor compounds are grouped according to their volatility and flavor type, such as a volatile aroma-sweetness synergistic group and a non-volatile acidity-richness synergistic group. Preset content standard ranges and individual substance standard intervals are retrieved for each group. The deviation of each individual flavor compound content from the standard center value is calculated, as well as the contribution ratio of each substance within the group to the total deviation of the group. Combined with the correlation analysis of signal characteristics and content deviation, if all individual, group, and total contents are within the standard range and the deviation indicators meet the requirements, the beverage is deemed qualified. If there are unqualified items, the abnormal items are clearly marked and the reasons for the deviation are analyzed, such as fluctuations in raw material composition or deviations in production process parameters. A test report containing detailed analysis conclusions and improvement suggestions is output.

[0062] In step S4, the analysis results are grouped according to both chemical structure similarity and flavor synergy. The coefficient of variation of content within each group and the coefficient of synergy between groups are calculated to form a five-level analysis result. The report simultaneously presents the data and the correlation analysis with the flavor quality of the beverage.

[0063] In some embodiments, the test results are analyzed by dual grouping based on the polarity and functional synergy of flavor substances, such as a strong polar flavor substance synergy group and a weak polar aroma substance synergy group. The cumulative content value, the uniformity index of the content of substances within the group, and the functional complementarity coefficient between groups are calculated for each group. The uniformity index within the group reflects the balance of the content of each substance within the same group, and the functional complementarity coefficient between groups reflects the synergistic contribution effect of different groups to the overall flavor of the beverage. A four-level analysis result is formed: "individual substance - uniformity within the group - cumulative content of groups - functional complementarity between groups - total content". The test report simultaneously presents the data at each level, the correlation logic between the data and the flavor quality of the beverage, and the analysis of the direction of flavor optimization, providing a more detailed reference for the quality control of beverages.

[0064] The fusion weight is determined by calculating the signal-to-noise ratio of dielectric characteristic signals, the integrity index of surface acoustic wave characteristic signals, and the mutual information values ​​of the two types of data. It is dynamically adjusted to retain the sensitive signal dimension and give full play to the complementary value of the data.

[0065] In some embodiments, the fusion weight is determined by calculating the feature correlation coefficient of dielectric feature data, the signal stability score of surface acoustic wave feature data, and the redundancy index of the two types of data. The feature correlation coefficient characterizes the correlation strength between dielectric features and flavor substance content, the signal stability score assesses the reliability of surface acoustic wave signals, and the redundancy index reflects the degree of information overlap between the two types of data. The fusion weight is dynamically allocated based on the comprehensive evaluation results of these three indicators, and higher weights are given to features that are strongly correlated with flavor substance content, have high stability, and low redundancy, so as to ensure that the fused data can give full play to the complementary advantages of the two types of signals and improve the comprehensiveness and accuracy of feature representation.

[0066] The model update is set with two triggering conditions: cumulative prediction error and feature importance drift. Standard sample data of beverages within a set time period are selected, and incremental training is performed after preprocessing and fusion. Only the key parameters of tree depth and learning rate are updated to complete the iteration.

[0067] In some embodiments, the model update is set with a single trigger condition for the cumulative prediction error of continuously detected samples. When the cumulative error reaches a preset threshold, the model update process is initiated. Standard beverage sample data within a set time period is screened, and a data timeliness assessment step is added to remove sample data that has exceeded the validity period or has significantly different detection conditions. The screened valid sample data is merged with historical core training data, and after preprocessing and feature fusion, the model is fine-tuned using transfer learning. The focus is on updating the output layer parameters and feature weights of the model. Model performance optimization can be achieved without full retraining, ensuring that the model can adapt to the dynamic changes in beverage sample characteristics and maintain stable prediction accuracy.

[0068] Example 2

[0069] See Figure 2This embodiment provides a specific implementation process for a method of detecting and analyzing the content of liquid flavor substances in beverages. The main process includes multi-source signal acquisition, data preprocessing and fusion, model building and prediction, and result analysis and report output. The specific implementation steps are as follows:

[0070] S1. Multi-source signal acquisition and raw data set construction:

[0071] S1.1 Sample preprocessing operations:

[0072] Before conducting beverage sample testing, a basic condition check must be performed on the samples to confirm that there are no abnormalities such as spoilage, layering, or contamination, and to rule out changes in flavor substances that may occur due to improper storage conditions or transportation, so as to provide a basic guarantee for the reliability of subsequent test data.

[0073] Pretreatment was carried out using a combination of centrifugation and solid-phase microextraction. First, centrifugation was initiated, and appropriate centrifugation parameters were set according to the characteristics of the beverage matrix. Centrifugal force was used to separate and remove suspended particles and macromolecular impurities from the beverage. Suspended particles included common insoluble solids found in beverages, while macromolecular impurities included high-molecular-weight substances that could interfere with detection signals. Centrifugation effectively reduced the negative impact of these impurities on subsequent signal detection.

[0074] After centrifugation, solid-phase microextraction (SPE) was used to enrich the target flavor compounds. Based on the volatility, polarity, and other physicochemical properties of the flavor compounds in the beverage, a suitable type of extraction fiber coating was selected to ensure efficient adsorption of the target compounds. The selected extraction fiber was inserted into the beverage sample, and the temperature and time of the extraction process were strictly controlled. The extraction temperature needed to closely match the beverage's normal storage temperature to avoid changes in the properties of flavor compounds or excessive volatilization due to high temperatures. The extraction time was rationally set based on the concentration and diffusion rate of the flavor compounds in the sample to ensure that the extraction fiber could fully adsorb volatile and semi-volatile flavor compounds while avoiding excessive adsorption of matrix components, thus ensuring the purity of the subsequent detection signal. After the extraction operation, the extraction fiber was desorbed to prepare for subsequent signal acquisition.

[0075] S1.2 Verification of sample homogeneity and stability:

[0076] After sample pretreatment, its homogeneity and stability need to be verified to ensure that the subsequently acquired signals can accurately reflect the overall flavor characteristics of the sample. The pretreated sample is gently shaken, and detection points are selected in different areas of the sample container. The distribution of detection points follows the principle of uniform coverage to ensure a comprehensive representation of the overall sample condition and avoid result bias caused by detection in localized areas.

[0077] Using signal acquisition equipment, preliminary dielectric response signals and surface acoustic wave propagation signals at each detection point are acquired synchronously. During the acquisition process, all equipment parameters are kept consistent to avoid introducing additional errors due to parameter fluctuations. The dielectric constant variation coefficient is calculated from the acquired dielectric signals. The formula for this coefficient is: ,in, The standard deviation of the dielectric constant for each region is given. The average dielectric constant of each region is the coefficient that directly reflects the uniformity of flavor substances in the sample. Meanwhile, the wave velocity stability index is calculated for the surface acoustic wave signal. This index is obtained by dividing the sum of the absolute values ​​of the deviations between the wave velocity values ​​of each region and the overall average wave velocity by the number of detection points, and is used to evaluate the consistency of the surface acoustic wave propagation speed.

[0078] The calculated dielectric constant variation coefficient and wave velocity stability index are compared with preset thresholds. These preset thresholds are determined statistically based on detection data from a large number of standard homogeneous samples and can effectively distinguish the homogeneity of the samples. If both indicators meet the preset threshold requirements, it indicates that the homogeneity and stability of the sample meet the detection standards, and the process can proceed to the subsequent signal acquisition steps. If either indicator fails to meet the preset requirements, it indicates that the sample may have problems such as local impurity aggregation, uneven distribution of flavor substances, or fluctuations in matrix composition. The sample needs to return to S1.1 for reprocessing, and if necessary, a stirring step can be added followed by centrifugation until the homogeneity and stability of the sample meet the detection requirements.

[0079] S1.3 Multi-source signal acquisition and raw data integration:

[0080] After the samples pass the homogeneity and stability verification, the formal multi-source signal acquisition process is initiated to ensure that the signals can comprehensively and accurately reflect the physicochemical characteristics of flavor substances.

[0081] Within a defined frequency range, dielectric response signals were acquired from beverage samples, recording the dielectric constant and loss tangent at different frequencies. The dielectric constant reflects the degree of polarization of a substance under an electric field and is closely related to the molecular structure and number of polar groups of flavor substances. The loss tangent characterizes the energy loss properties of a substance under an electric field and can indirectly reflect the concentration and interaction state of flavor substances. These two parameters together constitute the core features of the dielectric response signal.

[0082] For surface acoustic wave (SAW) propagation signals, SAW is excited under set excitation conditions and propagated in a beverage sample, with wave velocity shift and attenuation coefficient recorded simultaneously. Wave velocity shift is the difference between the speed of SAW propagation in the sample and its speed under no-load conditions, and is related to the density, viscosity, and acoustic impedance of the flavor substances. The attenuation coefficient characterizes the proportion of energy loss of the SAW during propagation and reflects the distribution of flavor substances. These two parameters reflect the characteristics of flavor substances from different dimensions.

[0083] After acquiring the dielectric response signal and the surface acoustic wave (SAW) propagation signal, a correlation mapping is established based on the specific response characteristics of these two types of signals to flavor substances. First, key features are extracted from both types of signals: peak dielectric constant and valley loss tangent are extracted from the dielectric response signal; maximum wave velocity shift and average attenuation coefficient are extracted from the SAW signal. Then, the correspondence between the two types of features is analyzed to identify feature combinations that respond to the same type of flavor substance. Based on the identification results, a correlation matrix is ​​constructed, where elements represent the correlation strength between different signal features. Finally, using a pre-defined data structure, the raw data of the two types of signals, the extracted key features, and the correlation matrix are integrated to form a structured raw dataset, providing basic data support for subsequent preprocessing and fusion operations.

[0084] S2. Data preprocessing, fusion, and prediction model construction:

[0085] S2.1 Data Cleaning Processing:

[0086] Although the original dataset contains effective signals of flavor substances, it inevitably contains irrelevant signals such as beverage matrix, equipment electronic noise, and environmental interference. These interference signals will distort the characteristics of the effective signals and affect the accuracy of subsequent data fusion and model construction. Therefore, targeted purification processing is required.

[0087] This embodiment employs an ensemble empirical mode decomposition algorithm to purify dielectric characteristic data and surface acoustic wave characteristic data. First, based on the signal complexity and sampling frequency, a reasonable number of decomposition layers is set to ensure sufficient separation of signals with different frequency components. After the original signal is input into the algorithm, the algorithm automatically decomposes the signal into several intrinsic mode functions and one residual component. The intrinsic mode functions are oscillation components at different frequencies, and the residual component is the trend term of the signal.

[0088] During the decomposition process, the focus is on separating signal interference from matrices such as sugars and proteins in the beverage. These matrices have relatively stable physicochemical properties, and their corresponding signal components typically operate within specific frequency ranges, exhibiting fluctuation patterns significantly different from those of flavor compounds. By observing the frequency distribution and fluctuation characteristics of each intrinsic mode function, mode functions containing matrix interference can be accurately identified. For intrinsic mode functions exhibiting mode aliasing (superposition of different frequency components), a secondary decomposition is required, adjusting the decomposition parameters to ensure each mode function becomes a single frequency component, further purifying the effective signal.

[0089] Finally, intrinsic mode functions containing interfering components are removed, while effective mode functions and residual components related to flavor substances are retained. The effective mode functions and residual components are then reconstructed to obtain purified data with noise and mode aliasing eliminated, providing high-quality input for subsequent data fusion.

[0090] S2.2 Multi-source data fusion processing:

[0091] After data purification, dielectric characteristic data and surface acoustic wave characteristic data need to be integrated using a multi-source data fusion algorithm to fully utilize the complementary information of the two types of signals and improve the comprehensiveness and accuracy of feature representation. This embodiment uses a Bayesian network multi-source data fusion algorithm, and the specific operation process is as follows:

[0092] First, feature node screening is performed. From the purified dielectric and surface acoustic wave feature data, key features sensitive to flavor substance content are selected as network nodes. The screening criterion is the correlation between the feature and the flavor substance content. By calculating the correlation coefficient between the feature and the known content in the standard sample, features with an absolute value of the correlation coefficient higher than a preset threshold are selected as nodes to ensure that the nodes can effectively reflect changes in flavor substance content.

[0093] Subsequently, the conditional dependencies between feature nodes were determined, and the independence between nodes was analyzed using the chi-square test. If the chi-square value of two nodes is greater than a critical value, it indicates that the two nodes have a conditional dependency; otherwise, they are independent nodes. Based on the conditional dependencies, a Bayesian network topology was constructed, forming a network structure with flavor substance content as the target node and dielectric feature nodes and surface acoustic wave feature nodes as observation nodes, ensuring that the network can accurately characterize the association logic between features and targets.

[0094] The confidence level of each feature node is calculated by estimating the maximum a posteriori probability. The feature data of the standard sample and the corresponding flavor substance content are used as training data to construct a likelihood function. The posterior probability is calculated by combining the prior probability. The posterior probability is the confidence level of the node, which represents the reliability of the node feature in detecting the flavor substance content.

[0095] The fusion weights are assigned based on the confidence level of each node, following the principle of "higher confidence level, greater weight". The weight calculation formula is as follows: ,in, For the first The fusion weight of each node, For the first The confidence level of each node. The sum of the confidence scores of all nodes is used. At the same time, signal dimensions that are sensitive to flavor substances are retained first, and nodes with confidence scores below a preset threshold are downweighted or removed to avoid redundant information affecting the fusion effect.

[0096] Dielectric feature data and surface acoustic wave feature data are weighted and superimposed using assigned fusion weights. For numerical features, the sum of the products of each feature value and its corresponding weight is directly calculated to obtain the fused feature value. For curvilinear features, weighted fusion is performed one by one along the corresponding dimension to obtain the fused feature curve, ultimately forming a unified feature dataset.

[0097] The determination of fusion weights requires dynamic adjustment of multiple auxiliary indicators to ensure optimal fusion results: Calculating the signal-to-noise ratio (SNR) of dielectric characteristic data (the ratio of peak dielectric constant to mean noise) to characterize the proportion of effective signal to noise; calculating the signal integrity index (signal-to-noise ratio of wave velocity offset) of surface acoustic wave characteristic data to assess data integrity and reliability; and calculating the mutual information value between the two types of data to characterize the degree of overlap and complementarity of the two signal information types. The product of the SNR, signal integrity index, and mutual information value is used as the basis for correcting the node confidence, dynamically adjusting the fusion weights of each node to ensure that the integrated data can fully leverage the complementary value of the two signal types.

[0098] S2.3 Construction and optimization of the calculation model for flavor compound content:

[0099] After the unified feature dataset is constructed, a flavor substance content calculation model is built and optimized based on the gradient boosting tree algorithm. The algorithm explores the nonlinear relationship between features and flavor substance content to achieve accurate prediction of flavor substance content.

[0100] The flavor compound content calculation model adopts a gradient boosting tree ensemble structure, which is divided into four layers: input layer, feature processing layer, decision tree ensemble layer, and output layer.

[0101] Input layer: Receives a unified feature dataset containing all key feature dimensions after the fusion of dielectric features and surface acoustic wave features, providing basic input data for the model;

[0102] Feature processing layer: Standardizes and normalizes the input feature data to eliminate differences in units, performs preliminary screening of feature importance, removes redundant features, and outputs simplified feature data.

[0103] Decision tree ensemble layer: It consists of multiple CART regression trees. Each decision tree is trained based on the prediction residuals of the previous decision tree. The outputs of each decision tree are aggregated through weighted fusion to form the output of the ensemble layer.

[0104] Output layer: Receives the fusion results from the decision tree ensemble layer and outputs the predicted content values ​​of each individual flavor substance, completing the mapping from feature data to content results.

[0105] The training steps for the model are as follows:

[0106] 1. Dataset Partitioning: A stratified sampling approach is used to divide the unified feature dataset into a training set and a validation set. The stratification is based on the chemical attribute categories of flavor substances, ensuring that the proportion of feature data from each category is consistent between the training and validation sets. This avoids bias in model training towards a particular flavor substance due to uneven data distribution. During the partitioning process, sample equalization is incorporated. If a certain category has too many samples, undersampling is used to randomly remove some redundant samples; if a certain category has too few samples, oversampling is used to generate new samples based on the feature distribution of existing samples, thus balancing the sample distribution across all levels.

[0107] 2. Feature Importance Selection: The processed training set is input into the feature processing layer of the model. The feature importance evaluation function built into the gradient boosting tree algorithm is used to calculate the number of splits and split gain of each feature in all decision trees. The importance of each feature dimension is ranked, and the features with the highest importance ranking are selected as the input dimensions of the model. This further simplifies the model structure and improves training efficiency and prediction accuracy.

[0108] 3. Model Parameter Initialization: Based on the feature dimension and the number of samples, set the initial parameters of the model: the tree depth is reasonably set according to the feature dimension and the number of samples to ensure that the model can fully explore the feature correlation and avoid overfitting; the initial learning rate is set to a small value to ensure stable convergence of the model training process; the number of iterations is set according to the size of the training set, with the standard of fully exploring the correlation between features and the target; at the same time, the regularization coefficient is set, with the initial value set based on empirical values, to control the complexity of the decision tree.

[0109] 4. Iterative Training and Parameter Tuning: The model is iteratively trained based on the training set. In each iteration, a new decision tree is built, trained using the prediction residuals from the previous iteration. The residuals are the difference between the actual sample size and the predicted value from the previous iteration. The model's predictive performance is evaluated in real-time using a validation set, with mean absolute error (MAE) as the core evaluation metric. The calculation formula is as follows: ,in To verify the actual content of the sample set, These are the model's predicted values. The validation set sample size represents the average deviation between predicted and actual values. Simultaneously, the model's explanatory power for data variations is considered (the closer this power is to 1, the better the predictive performance). Based on the feedback from the validation set, key model parameters are dynamically adjusted: if the mean absolute error is too large and the explanatory power is insufficient, it indicates underfitting; in this case, the tree depth, number of iterations, or learning rate can be appropriately increased. If the validation set error is much larger than the training set error, it indicates overfitting; in this case, the tree depth, learning rate, or regularization coefficient can be decreased.

[0110] 5. Cross-validation optimization: During model training, leave-one-out cross-validation is used to further optimize the partitioning of the training and validation sets. By stratifying the model according to the functional attribute categories of flavor substances, one sample is selected from each stratum as the validation set, and the remaining samples are used as the training set. This process is iterated until all samples participate in the validation. This approach fully utilizes limited sample data, avoids evaluation bias caused by a single partition, ensures consistent sample distribution across strata, and effectively prevents overfitting. Simultaneously, L2 regularization and an early stopping mechanism are introduced to improve the model's generalization ability. The L2 regularization term controls model complexity by constraining the sum of squares of model parameters. The early stopping mechanism monitors changes in the validation set error and sets a threshold for the number of consecutive iterations. When the validation set error no longer decreases after a set number of iterations, training automatically stops, preventing the model from experiencing a decline in generalization ability due to overfitting to training set noise in the later stages of training.

[0111] The model's input is a unified feature dataset, which includes key features fused from dielectric response signals and surface acoustic wave propagation signals. These features comprehensively reflect the physicochemical properties and content correlation of flavor substances. The model's output is the predicted content of each individual flavor substance. The output results are highly consistent with the business logic of beverage testing and can directly provide data support for subsequent result analysis and compliance determination.

[0112] S2.4 Noise Removal and Data Calibration:

[0113] In the preprocessing process, in addition to data purification through ensemble empirical mode decomposition, noise removal is specifically performed on outlier data points, and systematic errors are eliminated through data calibration to ensure that the data can truly reflect the actual content characteristics of flavor substances.

[0114] Noise removal employs a local outlier factor algorithm, which identifies outliers by calculating the local density deviation of each data point relative to its neighborhood data. First, the neighborhood size is determined based on the number and distribution density of data points. For each data point, the distances to its nearest neighbors are calculated to obtain the local reachability density. Then, the local outlier factor for that data point is calculated, which is the average ratio of the local reachability density of each neighboring data point to the local reachability density of that data point. Data points with local outlier factors exceeding a preset threshold are identified as global outliers and removed. These outliers are often caused by sudden fluctuations in the beverage matrix, momentary equipment malfunctions, or electromagnetic interference during signal acquisition; failure to remove them would severely impact the accuracy of model training.

[0115] Data calibration employs a series of standard flavor compound mixtures at varying concentration gradients. These mixtures contain common flavor compounds found in beverages, and the concentration gradients cover the range of content likely encountered in actual testing, ensuring the applicability of the calibration model. Dielectric response and surface acoustic wave propagation signals of the standard solutions at each concentration gradient are acquired. Using the actual content of the standard compounds as the dependent variable and the key characteristics of the corresponding signals as independent variables, a partial least squares regression calibration model is constructed. Regression analysis is used to determine the correlation between the independent and dependent variables, obtaining regression parameters. The original dataset is then input into the calibration model for mapping and transformation, effectively eliminating systematic errors caused by equipment parameter drift and beverage matrix interference, resulting in corrected data that more accurately reflects the actual content characteristics of the flavor compounds.

[0116] S3. Data Validation and Content Prediction:

[0117] S3.1 Dataset Validation:

[0118] Before inputting the preprocessed dataset into the model, a comprehensive validity check must be performed to ensure that the data format, completeness, and temporal features all meet the model input requirements, so as to avoid distortion of prediction results due to data problems.

[0119] Data format standardization is used to unify the units and magnitudes of dielectric characteristic data and surface acoustic wave characteristic data. First, the original units of the two types of data are reviewed, and the units of parameters such as dielectric constant (unitless), loss tangent (unitless), wave velocity offset (m / s), and attenuation coefficient (dB / m) are uniformly recorded and labeled. Then, a normalization method is used to unify the magnitudes of the data, ensuring that the normalized data value range matches the content range of flavor substances in beverages. This avoids imbalances in feature weight allocation during model training due to differences in units or magnitudes, ensuring that each feature can participate fairly in model training.

[0120] Feature integrity verification requires examining each feature dimension in the unified feature dataset one by one, comparing it with the key feature list determined in the feature selection stage, and determining whether any features crucial for predicting flavor substance content are missing. During the verification process, if the data missing rate of a certain key feature dimension is found to be lower than a preset threshold, the missing data is supplemented using the mean or median imputation method; if the missing rate is higher than the preset threshold, it indicates a problem in the data preprocessing or fusion process, and it is necessary to return to S2 for reprocessing and fusion, supplementing the key feature data to ensure the integrity of the dataset. At the same time, it is necessary to check whether there are outliers in the data of each feature dimension, and to reconfirm the outliers. If they are data collection errors, they need to be corrected; if they are genuine outliers, they need to be removed using noise removal methods.

[0121] Temporal consistency verification is implemented using a dynamic time warping algorithm to ensure that the temporal features of the preprocessed dataset are consistent with those of the model training data (beverage batch detection data), thus adapting to the model's temporal prediction logic. First, the temporal features of the preprocessed dataset are extracted, including the trend of feature values ​​over time, fluctuation frequency, and peak occurrence time. Then, the temporal features of the model training data are extracted as a reference template. The similarity between the two types of temporal features is calculated using the dynamic time warping algorithm. This algorithm finds the optimal time alignment path to minimize the cumulative distance between the two types of sequences; the smaller the cumulative distance, the higher the similarity. If the similarity is higher than a preset threshold, it indicates that the temporal features of the preprocessed dataset are consistent with the training data, and the subsequent prediction steps can proceed. If the similarity does not meet the threshold, it indicates abnormal temporal features, possibly due to temporal interference or sample batch fluctuations. The dataset needs to be returned to S2 for data purification and fusion, and if necessary, the signal should be re-acquired until the temporal consistency requirements are met.

[0122] S3.2 Content Prediction and Confidence Interval Calculation:

[0123] After the validity verification is passed, the dataset is input into the optimized flavor compound content calculation model, triggering the model's feature association calculation process. The model first calls the optimal parameters determined during training to extract and process the features of the input data layer by layer. Through the filtering and optimization of the feature processing layer, the core features are input into the decision tree ensemble layer. Each decision tree in the decision tree ensemble layer performs feature splitting and association analysis to explore the mapping relationship between features and flavor compound content. The output results of each decision tree are fused through weights. Finally, the output layer integrates and corrects the fused results to generate the content prediction results of each individual flavor compound. The prediction results include specific content values ​​and corresponding units.

[0124] To intuitively characterize the reliability range of the prediction results and aid in result analysis and judgment, confidence intervals for the content of each individual flavor substance were simultaneously calculated based on bootstrapping resampling. First, the number of resampling attempts was determined based on the sample size and data distribution characteristics to ensure a sufficient reflection of the data distribution features. Multiple samplings with replacement were performed on the validated dataset, each yielding a resampled sample set with the same sample size as the original dataset. Each resampled sample set was input into the optimized model to obtain the corresponding content prediction results. The prediction results of all resampled sample sets were statistically analyzed, and the mean and standard deviation were calculated. Combining the confidence level requirements for beverage testing, the corresponding distribution quantiles were determined. The upper and lower limits of the confidence interval were calculated using the mean, standard deviation, quantiles, and number of resampling attempts, clearly reflecting the reliability range of the prediction results and providing a reference for subsequent result analysis.

[0125] S4. Results Analysis and Report Output:

[0126] S4.1 Flavor Synergy Grouping and Standard Range Retrieval:

[0127] From the model's output prediction results, the content values ​​of individual flavor compounds, the cumulative content values ​​of flavor synergy groups, and the total flavor compound content value are extracted. Flavor synergy groups are divided based on the synergistic relationship of flavor compounds. Substances that have a synergistic effect in the presentation of beverage flavor and jointly influence the overall flavor quality are grouped together, such as the sweetness-aroma synergy group, the sourness-mouth synergy group, etc. Flavor compounds within the same group complement each other in taste and olfaction perception, jointly determining the flavor harmony of the beverage.

[0128] During the grouping process, synergistic groups were initially formed based on the physicochemical properties and flavor mechanisms of flavor substances. Then, the grouping structure was adjusted in conjunction with the correlation logic of flavor perception to ensure that the grouping accurately reflects the synergistic relationship of flavor substances. After grouping, the cumulative content value (the sum of the contents of all individual flavor substances within the group) and the total content value (the sum of the contents of all individual flavor substances within the group) of each group were calculated to provide basic data for subsequent deviation analysis.

[0129] Based on the quality control requirements of beverages, pre-defined standard ranges for the content of individual flavor substances, synergistic group content, and total content are retrieved. These standard ranges are established based on relevant industry standards and specifications, beverage flavor and quality requirements, and production process parameters, and are expressed in interval form, including upper and lower limits. For some key flavor substances, standard center values ​​are also set to calculate relative deviations. During the retrieval process, the appropriate standard range version must be selected according to the type of beverage and the purpose of testing to ensure the applicability of the standard range.

[0130] S4.2 Deviation Analysis and Compliance Judgment:

[0131] A comprehensive deviation analysis was conducted on the extracted content values ​​to assess the degree of difference between the actual content and the standard requirements. First, the relative deviation of each individual flavor compound content value from the corresponding preset standard range center value was calculated using the formula... Calculations are performed, in which, This represents the predicted content value of a single flavor compound. The relative deviation represents the standard center value corresponding to the substance, and it characterizes the degree of deviation of the individual content from the standard center value, reflecting the accuracy of the substance content.

[0132] Simultaneously, the contribution percentage of each individual substance within the co-group to the grouping deviation is calculated: First, the deviation between the actual cumulative content value of the group and the preset grouping standard center value is calculated. This deviation is the difference between the sum of the content values ​​of each individual substance within the group and the grouping standard center value. Then, the absolute value of the deviation between the content value of each individual substance within the group and the corresponding standard center value is calculated. By the ratio of the absolute value of the deviation of an individual substance to the sum of the absolute values ​​of the deviations of all individual substances within the group, the contribution percentage of that individual substance to the grouping deviation is obtained. This indicator can identify the main contributors to the grouping deviation and provide key evidence for deviation tracing.

[0133] The content values ​​of each type are compared with the corresponding preset standard ranges, and a comprehensive judgment is made by combining the relative deviation and contribution ratio indicators: if the content values ​​of all individual flavor substances are within the corresponding individual standard range, the cumulative content value of the synergistic group is within the group standard range, the total content value is within the total standard range, and the relative deviation of all individual substances and the contribution ratio of group deviation meet the preset requirements, then the beverage sample is judged to be qualified; if any indicator fails to meet the standard, the corresponding unqualified item is marked, and the key individual substance causing the group abnormality is located by the contribution ratio. Combined with the beverage matrix interference, signal characteristic changes, preprocessing parameters, and other factors, the cause of the deviation is traced, such as raw material quality fluctuations, changes in production process parameters, matrix interference, equipment calibration deviations, etc.

[0134] When using dual grouping based on chemical structure similarity and flavor synergy, it is necessary to further calculate the coefficient of variation (COP) of substance content within each group and the coefficient of synergy between groups. The COP is calculated as the ratio of the standard deviation to the mean of each individual substance content within the group, reflecting the uniformity of flavor substance distribution within the group. The coefficient of synergy between groups is calculated as the ratio of the covariance to the product of the standard deviations of the substance contents of the two groups, reflecting the degree of synergistic contribution of different groups to the overall flavor of the beverage. These calculations form a five-level analytical result: "individual substance - intra-group variation - group cumulative - inter-group synergy - total content". The test analysis report simultaneously presents the five levels of data and the correlation analysis conclusions between each level and the flavor quality of the beverage, providing a comprehensive and accurate reference for quality control.

[0135] S4.3 Model Update Process:

[0136] To ensure the long-term effectiveness of the method and to adapt to the impact of factors such as beverage production process optimization, raw material batch changes, and fluctuations in the testing environment, a mechanism for regular model updates should be established.

[0137] The model update is set with two trigger conditions, and the update process is initiated when either condition is met: First, the cumulative prediction error threshold for continuously detected samples. By statistically analyzing the sum of prediction errors of a set number of continuously detected samples, when the cumulative error reaches the preset threshold, it indicates that the model's prediction accuracy can no longer meet the detection requirements and an update is necessary. Second, the feature importance drift threshold. By comparing the feature importance ranking of the current detected samples with the feature importance ranking of the initial training data, the ranking difference is calculated. When the difference reaches the preset threshold, it indicates that the sample feature distribution has changed significantly, and the model needs to be updated to adapt to the new feature distribution.

[0138] After initiating the model update process, the first step is to screen standard beverage sample data within a set timeframe. These samples undergo rigorous quality control to ensure standardized testing procedures, accurate and reliable data, and the ability to reflect the actual characteristics of the beverages. The validity of the sample data is assessed through a data quality score, based on three indicators: signal integrity, calibration bias, and detection repeatability. A comprehensive score is calculated using preset weights: signal integrity is assessed by the percentage of effective signal; a higher percentage results in a higher score. Calibration bias is assessed by the difference between the calibrated actual value and the predicted value; a smaller bias results in a higher score. Detection repeatability is assessed by the coefficient of variation of multiple tests on the same sample; a smaller coefficient of variation results in a higher score. Invalid samples with a comprehensive score below a set threshold are removed, and valid sample data is retained.

[0139] The effective sample data is merged with the historical core training data (which reflects the flavor characteristics of different batches and types of beverages) filtered by feature transfer. The merged dataset is then processed through the S2 preprocessing and fusion process: data is purified by ensemble empirical mode decomposition, multi-source features are fused by Bayesian network, outliers are removed by local outlier algorithm, and data is calibrated by partial least squares regression model to obtain a new unified feature dataset, ensuring the quality and consistency of the updated data.

[0140] Incremental training is used to update the model, eliminating the need for full retraining. Only key parameters such as tree depth and learning rate are updated, improving update efficiency and ensuring model stability. The new unified feature dataset is divided into incremental training and incremental validation sets using stratified sampling to maintain a balanced data distribution. The parameters of the original optimized model are used as initial parameters, and the incremental training set is input into the model for training. During training, some parameters of the lower-level feature extraction layers are frozen, and only the parameters of the decision tree ensemble layer and output layer are updated to avoid the model forgetting historical knowledge due to a full update. The performance of the updated model is evaluated using the incremental validation set, employing a comprehensive evaluation based on mean absolute error and prediction stability metrics. If the updated model outperforms the original model, the updated model parameters are saved, and model iteration is completed. If the performance does not meet the requirements, the incremental training parameters are adjusted or additional valid sample data is added, and incremental training is repeated until the model performance meets the detection requirements.

[0141] The method for detecting and analyzing the content of liquid flavor substances in beverages provided in this embodiment, through complete process design and detailed implementation, achieves comprehensive detection and accurate analysis of the content of liquid flavor substances in beverages. This method utilizes the combined capture of dielectric response signals and surface acoustic wave propagation signals, fully leveraging the complementary characteristics of these two types of signals to comprehensively characterize the physicochemical features of flavor substances, avoiding the information bias caused by single signal detection. The preprocessing method, combining empirical mode decomposition and local outlier factor algorithms, effectively eliminates the influence of interference factors such as beverage matrix and equipment noise, improving data purity and reliability. The multi-source data fusion algorithm based on Bayesian networks achieves efficient integration of the two types of feature data, retaining key signal dimensions sensitive to flavor substances and providing high-quality input data for model construction. The flavor substance content calculation model, through a reasonable structure... The well-designed and scientific training process and parameter optimization accurately uncovered the correlation between features and flavor substance content. Combined with cross-validation and regularization mechanisms, the model's generalization ability and prediction accuracy were significantly improved. Multi-dimensional data verification, including temporal consistency and feature integrity, ensured the validity of the input model data and guaranteed the reliability of the prediction results. Deviation analysis and group collaborative analysis were used to comprehensively evaluate the compliance of beverage flavor substance content with standards, accurately trace the causes of deviations, and provide strong support for quality control. The model update mechanism ensured the continued effectiveness of the method in long-term use and could adapt to the impact of various external factors.

[0142] Among them, the scheme of multi-source signal fusion and gradient boosting tree model construction can accurately predict the content of different types of flavor substances in beverages. It effectively overcomes the problems of incomplete information, large interference, and insufficient prediction accuracy in traditional detection methods, and provides scientific and reliable technical support for beverage quality control. It helps to improve the quality control level in the beverage production process and ensure the stability and consistency of beverage flavor quality.

[0143] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for detecting and analyzing the content of liquid flavor substances in a beverage, characterized in that, Includes the following steps: S1. Obtain the dielectric response signal of flavor substances and the surface acoustic wave propagation signal of the beverage sample to be tested, record the dielectric constant, loss tangent, wave velocity offset, and attenuation coefficient, establish a correlation mapping based on the specific response characteristics of the two types of signals to flavor substances, and integrate them to form a structured raw data set. S2. Preprocess the original dataset by using ensemble empirical mode decomposition to eliminate noise and mode mixing, integrate the two types of feature data through a Bayesian network multi-source data fusion algorithm, retain the sensitive signal dimension, and construct and optimize the flavor substance content calculation model based on the gradient boosting tree algorithm. S3. Perform temporal consistency and dimensionality matching checks on the preprocessed dataset. After passing the checks, input the dataset into the model to generate prediction results for the content of each individual flavor substance. S4. Extract individual, flavor synergy group and total content values, perform deviation analysis with preset standard range, calculate signal feature contribution weight, determine whether it meets beverage quality requirements and output a test analysis report including compliance judgment and deviation traceability; Step S1 specifically includes the following sub-steps: S1.1 The beverage samples to be tested are pretreated by a combination of centrifugation and solid-phase microextraction to remove suspended particles and macromolecular impurities, adsorb volatile and semi-volatile flavor substances, and avoid matrix interference. S1.2 Simultaneously acquire preliminary signals from different regions with a set number of samples, calculate the dielectric constant variation coefficient and wave velocity stability index. If both meet the preset threshold, proceed to the next step; otherwise, repeat the preprocessing. S1.3 Acquire two types of signals under a set frequency range and excitation conditions, construct an association matrix based on the signal-specific response characteristics, and integrate them to form a structured raw data set; Step S2 specifically includes the following sub-steps: S2.1 uses ensemble empirical mode decomposition to decompose the two types of data, eliminating the interference from the beverage matrix and obtaining purified data; S2.2 Construct a data probabilistic network structure, determine the dependency relationship of feature nodes and calculate the confidence level, allocate fusion weights based on the confidence level, and superimpose them to form a unified feature dataset; S2.3 The training set and validation set are divided into hierarchical groups according to the chemical attribute categories of flavor substances. The gradient boosting tree algorithm is used to select key dimensions. The parameters are adjusted by the validation set error and the coefficient of determination to obtain the optimized model. In step S2, noise removal uses a local outlier algorithm to remove global outlier data points, and data calibration uses a partial least squares regression calibration model constructed with a series of concentration gradient standard flavor substance mixed solutions to eliminate system errors caused by equipment drift and matrix interference.

2. The method according to claim 1, characterized in that, In step S2, when building the model, leave-one-out cross-validation is used to divide the training set and the validation set, and the model is stratified according to the functional attribute categories of flavor substances. L2 regularization and early stopping mechanism are introduced to avoid overfitting and improve generalization ability.

3. The method according to claim 1, characterized in that, Step S3 specifically includes the following sub-steps: S3.1 performs format standardization, feature integrity and time sequence consistency checks on the dataset, adapts it to the range of flavor substance content in beverages, ensures that no key features are missing and that the similarity with the training data meets the standard, and if it fails, it is reprocessed. S3.2 Input the validated dataset into the model to derive the content, and simultaneously calculate the confidence interval based on bootstrap resampling to characterize the reliable range of the prediction results.

4. The method according to claim 1, characterized in that, Step S4 specifically includes the following sub-steps: S4.1 Divide flavor synergy groups according to the synergistic relationship between sweetness and aroma, and sourness and mouthfeel, and retrieve the corresponding preset standard range; S4.2 Calculate the contribution ratio of the relative deviation of individual content and the group deviation. If all indicators meet the standards, the beverage is deemed qualified. Otherwise, mark the unqualified items and trace the cause, and output a complete test analysis report.

5. The method according to claim 1, characterized in that, In step S4, the analysis results are grouped according to both chemical structure similarity and flavor synergy. The coefficient of variation of content within each group and the coefficient of synergy between groups are calculated to form a five-level analysis result. The report simultaneously presents the data and the correlation analysis with the flavor quality of the beverage.

6. The method according to claim 1, characterized in that, The fusion weight is determined by calculating the signal-to-noise ratio of dielectric characteristic signals, the integrity index of surface acoustic wave characteristic signals, and the mutual information values ​​of the two types of data. It is dynamically adjusted to retain the sensitive signal dimension and give full play to the complementary value of the data.

7. The method according to claim 1, characterized in that, The model update is set with two triggering conditions: cumulative prediction error and feature importance drift. Standard sample data of beverages within a set time period are selected, and incremental training is performed after preprocessing and fusion. Only the key parameters of tree depth and learning rate are updated to complete the iteration.

Citation Information

Patent Citations

  • Method for predicting mass spectrum ionization efficiency of compound based on quantitative structure-activity relationship model of COSMO-RS and ANN algorithms

    CN115862762A

  • Method for predicting dynamic flavor of traditional fermented trachinotus ovatus based on multi-modal fusion deep learning

    CN120072096A