A method for constructing a true and false wine discrimination model

By extracting multimodal data and performing feature fusion and hierarchical extraction of deep learning models, combining transfer learning and adversarial training, the generalization ability of the authentic and fake wine discrimination model is optimized, and the problem of decreased accuracy of cross-brand and cross-batch detection is solved, and an efficient, intelligent and low-cost authentic and fake wine identification solution is achieved.

CN119808010BActive Publication Date: 2025-06-10SHANDONG BAIMAIQUAN LIQUEUR IND CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510295564.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-10
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing genuine and fake wine discrimination model lacks generalization ability when cross-brand and batches, resulting in a decrease in the accuracy of the model's detection of wine in new brands or new batches.

Method used

By obtaining wine samples from multiple brands and batches, multi-modal data is extracted using multiple non-destructive testing methods and standardized processing is carried out. Feature fusion and hierarchical extraction are performed based on deep learning models to extract distribution similarity features and domain offset robust features. The detection accuracy of the model on unseen brands and batches was evaluated through cross-validation using transfer learning and adversarial training optimization models.

Benefits of technology

It improves the adaptability and generalization ability of the authentic and fake wine discrimination model, enhances its applicability and detection reliability in complex market environments, ensures the accurate judgment results of the wine to be tested, and provides confidence scores to ensure the credibility of the test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808010B_ABST
    Figure CN119808010B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a true and false wine discrimination model, specifically relating to the technical field of liquor detection; by non-destructively detecting and extracting the composition data, spectral curves, odor fingerprints and texture information of the liquor, standardizing and feature-fusing the multi-modal data, extracting the distribution similarity features and domain shift robustness features, constructing a comprehensive feature vector, and using a machine learning model to evaluate the generalization ability value of the true and false wine discrimination model; for the low generalization ability model, a transfer learning method is adopted for multi-stage training and optimization, and cross-validation is combined to evaluate its detection accuracy on unseen brands and batches; finally, the multi-modal features of the liquor to be tested are input into the optimized model, and the true and false wine discrimination result and confidence score are output; the present invention can significantly improve the stability and adaptability of the true and false wine discrimination model, ensuring its high accuracy, high generalization ability and wide market applicability in the detection of liquors of different brands and batches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of liquor detection, and particularly relates to a method for constructing a true and false liquor discrimination model. Background Art

[0002] With the development of social economy, the types of liquor products are becoming increasingly rich. The appearance of fake liquor in the market has brought serious threats to consumers' health, safety and brand reputation. Traditional methods for identifying fake liquor mainly rely on means such as manual sensory evaluation, chemical analysis or spectral detection. However, these methods usually have problems such as long detection cycle, high cost, and strong dependence on professional equipment and personnel. In recent years, with the development of big data and artificial intelligence technologies, using data-driven methods to construct true and false liquor discrimination models has become a research hotspot. By analyzing the physical and chemical properties, spectral characteristics or fingerprint information of liquor, and combining machine learning algorithms, more efficient and accurate fake liquor detection can be achieved.

[0003] The existing technologies have the following deficiencies:

[0004] When constructing a true and false liquor discrimination model, the generalization ability of the model is insufficient due to the differences in the components of liquor across different brands and batches. Since different brands of liquor use different raw materials, brewing processes and storage methods, even for genuine liquor, there may be significant differences in their component characteristics. In addition, for the same brand of liquor, the detection characteristics may change due to fluctuations in trace components between different batches, resulting in a decrease in the detection accuracy of the trained model for new batches or new brands of liquor. This problem causes the existing discrimination models to be prone to overfitting to the sample data of specific brands or batches, and it is difficult to be popularized and applied in a wide market environment, restricting their practicality and commercial value. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for constructing a true and false liquor discrimination model to solve the deficiencies in the background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a true and false liquor discrimination model, comprising:

[0007] Obtain liquor samples of multiple brands and batches, and extract multi-modal data based on a variety of non-destructive detection means, including component data, spectral curves, odor fingerprints and texture information;

[0008] Perform standardization processing on the collected multi-modal data, and perform feature fusion and hierarchical extraction on the feature data of different modalities based on a pre-constructed deep true and false liquor discrimination model, and respectively extract the distribution similarity features and domain shift robustness features in the multi-modal data;

[0009] Convert the extracted distribution similarity features and domain shift robustness features into a comprehensive feature vector, and use the comprehensive feature vector as the input of a machine learning model. Output the generalization ability value of the true and false wine discrimination model through the machine learning model;

[0010] For the true and false wine discrimination model with low generalization ability, use the transfer learning method to perform multi-stage training and optimization based on datasets of different brands and batches, and evaluate the detection accuracy of the true and false wine discrimination model on unseen brands and batches through cross-validation;

[0011] Input the multi-modal features of the wine to be tested into the optimized true and false wine discrimination model, output the true and false wine discrimination result, and give the corresponding confidence score.

[0012] Preferably, multi-modal data is extracted based on a variety of non-destructive testing methods, and the variety of non-destructive testing methods includes near-infrared spectroscopy, Raman spectroscopy, ultraviolet-visible spectroscopy, fluorescence spectroscopy, Fourier transform infrared spectroscopy, gas chromatography-mass spectrometry, electronic nose, and high-resolution microscopy, which are used to extract the composition data, spectral curves, odor fingerprints, and texture information of the wine respectively.

[0013] Preferably, after analyzing the distribution similarity features of wine samples of different brands and batches in the feature space, a distribution similarity deviation index is generated. The acquisition method of the distribution similarity deviation index is as follows:

[0014] Suppose there are two wine samples of brands or batches, and their feature vectors are respectively: , representing the features of the wine sample of the first brand / batch, and the number of samples is m; , representing the features of the wine sample of the second brand / batch, and the number of samples is n; where each sample and are both high-dimensional feature vectors;

[0015] Define a kernel function k(x,y) to measure the similarity between two sample points: ; where σ is the bandwidth parameter, and calculate the mean distribution deviation degree value MD of the two samples in the high-dimensional feature space. The expression is: , compare the calculated MD value with the maximum MD value calculated in all brand or batch combinations, and calculate the distribution similarity deviation index. The expression is: ; In the formula, is the maximum MD value calculated in all brand or batch combinations, is the distribution similarity deviation index.

[0016] Preferably, after analyzing the obtained domain shift robustness features, a domain shift robustness anomaly index is generated. The acquisition method of the domain shift robustness anomaly index is as follows:

[0017] Obtain the source brand / batch training set and the target brand / batch test set ; Construct a denoising autoencoder or variational autoencoder for training so that it can accurately reconstruct the input features from the source brand / batch data ;

[0018] The autoencoder structure includes: Encoder: Map the high-dimensional input features to the low-dimensional latent space z;

[0019] Decoder: Remap the low-dimensional latent variable z back to the original feature space;

[0020] Loss function: Calculate the error between the input data x and the reconstructed data using the mean squared error;

[0021] After training is completed, the autoencoder should be able to accurately reconstruct the sample features of the source brand / batch;

[0022] Input the target brand / batch data into the trained autoencoder model to generate reconstructed features and calculate the average reconstruction error of the target brand / batch , and the expression is: ; Calculate the domain shift robustness anomaly index DSRAI, and the expression is: ; In the formula, is the average reconstruction error of the source brand / batch, is the maximum reconstruction error among all brand / batch combinations.

[0023] Preferably, convert the distribution similarity deviation index and the domain shift robustness anomaly index into a comprehensive feature vector, use the comprehensive feature vector as the input of the machine learning model, use the generalization ability value label of the true and false wine discrimination model predicted by each group of comprehensive feature vectors as the prediction target, use minimizing the sum of the prediction errors of the generalization ability value labels of all true and false wine discrimination models as the training target, train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training, and determine the generalization ability value of the true and false wine discrimination model according to the model output result, where the machine learning model is a polynomial regression model.

[0024] Preferably, compare the generalization ability value of the obtained true and false wine discrimination model with the reference threshold of the pre-set generalization ability value. If the generalization ability value of the true and false wine discrimination model is greater than or equal to the reference threshold of the pre-set generalization ability value, it indicates that the generalization ability of the false wine discrimination model is high, and it is marked as a true and false wine discrimination model with high generalization ability; if the generalization ability value of the true and false wine discrimination model is less than the reference threshold of the pre-set generalization ability value, it indicates that the generalization ability of the false wine discrimination model is low, and it is marked as a true and false wine discrimination model with low generalization ability.

[0025] Preferably, the cross-validation adopts the K-fold cross-validation method, divides the target brand or batch dataset into K parts, trains and tests the model respectively, calculates the average accuracy of all K rounds of experiments. If the average accuracy is greater than the preset threshold, it indicates that the transfer learning optimization is effective.

[0026] Preferably, when the optimized true and false wine discrimination model discriminates the liquid to be tested, it performs inference based on the deep learning model of multi-modal feature fusion and outputs a confidence score. The confidence score is based on the feature similarity of the model to the sample to be tested. If the confidence level ≥ 90%, the determination result is reliable. If the confidence level is lower than 70%, additional detection means are added to assist in decision-making.

[0027] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0028] 1. The present invention uses non-destructive detection means such as near-infrared spectroscopy, Raman spectroscopy, gas chromatography-mass spectrometry (GC-MS), electronic nose, and high-resolution microscopic imaging to extract the composition data, spectral curves, odor fingerprints, and texture information of the wine liquid, and standardizes the data to eliminate the measurement deviation between brands. The true and false wine discrimination model constructed based on deep learning performs feature fusion through the self-attention mechanism and cross-modal feature alignment, extracts the distribution similarity feature and domain shift robustness feature, calculates the distribution similarity deviation index and domain shift robustness anomaly index, forms a comprehensive feature vector, and uses the polynomial regression model to predict the generalization ability value of the true and false wine discrimination model. For the model with low generalization ability, transfer learning and adversarial training are used for optimization, and its adaptability to unseen brands and batches is evaluated through K-fold cross-validation to ensure the stability of its detection accuracy on new brands / batches.

[0029] 2. The present invention not only improves the adaptability and generalization ability of the true and false wine discrimination model, but also enhances its applicability and detection reliability in complex market environments. After receiving the multi-modal features of the wine to be tested, the optimized model can accurately output the true and false wine discrimination results and give a confidence score to ensure the credibility and interpretability of the detection results. When the confidence level ≥ 90%, the determination result is highly reliable; when the confidence level is below 70%, the system can recommend adding additional detection means or combining expert evaluation to assist in decision-making. Compared with traditional false wine detection methods, the present invention realizes an efficient, intelligent, and low-cost true and false wine identification solution, which is applicable to the quality control and market supervision of liquor products of different brands and batches, and has broad commercial value and industry application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0031] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0033] Embodiment, please refer to Figure 1 As shown, a method for constructing a true and false wine discrimination model in this embodiment includes:

[0034] Obtain wine liquid samples of multiple brands and batches, and extract multi-modal data based on a variety of non-destructive detection means, including component data, spectral curves, odor fingerprints, and texture information;

[0035] Perform standardization processing on the collected multi-modal data, and perform feature fusion and hierarchical extraction on the feature data of different modalities based on a pre-constructed deep true and false wine discrimination model, and respectively extract the distribution similarity features and domain shift robustness features in the multi-modal data;

[0036] Convert the extracted distribution similarity features and domain shift robustness features into a comprehensive feature vector, and use the comprehensive feature vector as the input of a machine learning model to output the generalization ability value of the true and false wine discrimination model through the machine learning model;

[0037] For the true and false wine discrimination model with low generalization ability, use the transfer learning method to perform multi-stage training and optimization based on datasets of different brands and batches, and evaluate the detection accuracy of the true and false wine discrimination model on unseen brands and batches through cross-validation;

[0038] Input the multi-modal features of the wine sample to be tested into the optimized true and false wine discrimination model, output the true and false wine discrimination result, and give the corresponding confidence score.

[0039] The multi-modal data extraction steps include: Select wine samples of multiple brands and different batches to ensure that the samples cover different alcohol contents, brewing processes, and raw material sources to increase the applicability of the model. Ensure that the sample storage environment is consistent to avoid data deviation caused by external factors (such as temperature and humidity). Use near-infrared spectroscopy (NIR) or Raman spectroscopy to perform non-destructive detection on the main chemical components of the wine (such as ethanol content, organic acids, esters, sugars). Combine gas chromatography-mass spectrometry (GC-MS) method to analyze trace chemical components and extract chemical fingerprint data. Record the component change characteristics between different brands and batches and perform standardization processing.

[0040] Use ultraviolet-visible spectroscopy (UV-Vis) and fluorescence spectroscopy to obtain the absorption and emission characteristics of the wine at different wavelengths to form a spectral curve. Analyze the molecular vibration information of the wine through Fourier transform infrared spectroscopy (FTIR) to extract molecular structure features. Normalize the spectral data to eliminate the systematic error of the measurement device and improve the feature stability. Use an electronic nose (E-nose) system, and use metal oxide semiconductor sensors (MOS) or surface acoustic wave sensors (SAW) to sense the odor characteristics of the volatile components of the wine. Extract odor fingerprint data to form a volatile organic compound (VOC) feature matrix. Combine principal component analysis (PCA) or linear discriminant analysis (LDA) to reduce the data dimension and improve the model training efficiency.

[0041] Use high-resolution microscopy or optical coherence tomography (OCT) to analyze the surface texture or precipitation characteristics of the wine sample. Extract the texture pattern of the wine through image processing techniques (such as edge detection and morphological analysis) to improve the ability to identify adulterated components or impurities. Use deep learning methods (such as CNN) to automatically extract key visual features.

[0042] After standardizing the collected multi-modal data and performing feature fusion and hierarchical extraction through a pre-constructed true and false wine discrimination model, the distribution similarity features and domain shift robustness features are obtained to enhance the generalization ability of the model.

[0043] To ensure that data from different sources and different modalities are learned on the same scale, the following standardization methods are adopted:

[0044] Z-score standardization is adopted, which is applicable to the chemical components of liquor (such as ethanol, esters, organic acids, etc.), to eliminate the measurement scale differences between different brands and batches.

[0045] Max-min normalization is adopted, which is applicable to ultraviolet-visible spectroscopy (UV-Vis), fluorescence spectroscopy, Fourier transform infrared spectroscopy (FTIR), etc., to eliminate equipment measurement errors and improve feature stability.

[0046] Odor fingerprint data processing (output of electronic nose sensors): Principal component analysis (PCA) is used for dimensionality reduction to remove redundant information and retain the main odor features. The odor sensor response values are normalized to make the outputs of different sensors consistent.

[0047] Image data standardization (texture information): Gaussian normalization is used to adjust the image contrast to make the visual features of different batches consistent. Deep learning algorithms (CNN) are used for feature extraction, and the high-dimensional feature vectors are normalized.

[0048] After the multi-modal data is standardized, a pre-constructed deep true and false wine discrimination model is used for feature fusion and hierarchical extraction.

[0049] The Transformer network or self-attention mechanism (Self-Attention) is used to perform global feature fusion on the multi-modal data to improve the model's understanding ability of different modal features.

[0050] Cross-modal feature alignment is adopted. Through self-supervised learning or contrast learning methods, the component, spectral, odor, and texture features are made consistent in the high-dimensional space.

[0051] Through adaptive weighted fusion, the weights of different modal features are dynamically adjusted to adapt to different brands and batches of liquor.

[0052] Hierarchical feature extraction includes:

[0053] Low-level features: Local patterns (such as spectral peaks, odor response curves, etc.) are extracted from the original multi-modal data.

[0054] Mid-level features: Extract time series features (such as the change of odor signals over time) through a convolutional neural network (CNN) or a long short-term memory network (LSTM).

[0055] High-level features: Use a deep neural network (DNN) or a Transformer to extract a global feature vector and generate the final feature representation for true and false wine discrimination.

[0056] Distribution similarity feature extraction is used to measure the similarity of wine samples of different brands and batches in the feature space and improve the generalization ability of the model. Calculate the Fréchet distance between different brands / batches. A lower similarity value indicates that the feature distributions between different brands / batches are closer, and the model has better generalization ability.

[0057] Domain shift robustness feature extraction is to measure the adaptability of the model to unseen brands or batches and prevent the performance of the model from degrading on new samples.

[0058] When training the model, two models (the baseline model and the multi-domain adaptation model) are constructed using single-brand data and multi-brand data respectively. Test on new brands or batches, calculate the performance degradation ratio, and take the reciprocal of the performance degradation ratio as the domain shift robustness feature. The higher the value, the stronger the adaptability of the model to new brands or batches and the better the generalization ability.

[0059] After analyzing the distribution similarity features of wine samples of different brands and batches in the feature space, a distribution similarity deviation index is generated. The method for obtaining the distribution similarity deviation index is as follows:

[0060] Suppose there are two brands or batches of wine samples, and their feature vectors are respectively: , representing the features of the wine samples of the first brand / batch, and the number of samples is m; , representing the features of the wine samples of the second brand / batch, and the number of samples is n; where each sample and are high-dimensional feature vectors, usually composed of deep features after multi-modal data fusion.

[0061] When calculating, a kernel function k(x,y) needs to be defined to measure the similarity between two sample points. Common kernel functions include: Gaussian kernel (RBF kernel): ; where σ is the bandwidth parameter, controlling the scale of the kernel function.

[0062] Calculate the mean distribution deviation degree value MD of the two samples in the high-dimensional feature space. The expression is: , compare the calculated MD value with the maximum MD value calculated in all brand or batch combinations, and calculate the distribution similarity deviation index. The expression is: ; where is the maximum MD value calculated from all brand or batch combinations, and

[0063] After analyzing the obtained domain shift robustness features, a domain shift robustness anomaly index is generated. The method for obtaining the domain shift robustness anomaly index is as follows:

[0064] Obtain the source brand / batch training set , and the target brand / batch test set ; where each sample x is a high-dimensional feature vector, usually obtained by fusing multimodal features (such as spectral curves, odor fingerprints, composition data).

[0065] Construct a denoising autoencoder or variational autoencoder for training so that it can accurately reconstruct the input features on the source brand / batch data .

[0066] The autoencoder structure includes: Encoder: Maps the high-dimensional input features to a low-dimensional latent space z;

[0067] Decoder: Remaps the low-dimensional latent variable z back to the original feature space;

[0068] Loss function: Calculates the error between the input data x and the reconstructed data using the mean squared error (MSE);

[0069] After training, the autoencoder should be able to accurately reconstruct the sample features of the source brand / batch.

[0070] Input the target brand / batch data into the trained autoencoder model to generate reconstructed features , and calculate the average reconstruction error of the target brand / batch. The expression is: ; If is higher than the reconstruction error of the source brand / batch, it indicates that the model has poor adaptability to the target brand / batch and there is a serious domain shift. Calculate the domain shift robustness anomaly index DSRAI, and the expression is: ; where is the average reconstruction error of the source brand / batch, is the maximum reconstruction error among all brand / batch combinations (for normalization).

[0071] DSRAI close to 0: The reconstruction error of the target brand / batch is similar to that of the source brand / batch, indicating that the model adapts well to the new brand / batch and the domain shift is small.

[0072] The DSRAI is close to 1: the reconstruction error of the target brand / batch is much higher than that of the source brand / batch, indicating that the adaptability of the model on the new brand / batch is poor and there is a serious domain shift problem.

[0073] Convert the extracted distribution similarity features and domain shift robustness features into a comprehensive feature vector, use the comprehensive feature vector as the input of the machine learning model, and output the generalization ability value of the true and false wine discrimination model through the machine learning model. Specifically:

[0074] Convert the distribution similarity deviation index and the domain shift robustness anomaly index into a comprehensive feature vector, use the comprehensive feature vector as the input of the machine learning model, and use the machine learning model to predict the generalization ability value label of the true and false wine discrimination model for each group of comprehensive feature vectors as the prediction target, and minimize the sum of the prediction errors of the generalization ability value labels of all true and false wine discrimination models as the training target. Train the machine learning model until the sum of the prediction errors converges and then stop the model training. Determine the generalization ability value of the true and false wine discrimination model according to the model output result. Among them, the machine learning model is a polynomial regression model.

[0075] The method for obtaining the generalization ability value of the true and false wine discrimination model is: obtain the corresponding function expression from the comprehensive feature vector training data of the trained machine learning model: ; where F is the output function of the model, is the distribution similarity deviation index, DSRAI is the domain shift robustness anomaly index, is the generalization ability value of the true and false wine discrimination model.

[0076] Compare the obtained generalization ability value of the true and false wine discrimination model with the pre-set reference threshold of the generalization ability value. If the generalization ability value of the true and false wine discrimination model is greater than or equal to the pre-set reference threshold of the generalization ability value, it indicates that the generalization ability of the false wine discrimination model is high, and mark it as a true and false wine discrimination model with high generalization ability; if the generalization ability value of the true and false wine discrimination model is less than the pre-set reference threshold of the generalization ability value, it indicates that the generalization ability of the false wine discrimination model is low, and mark it as a true and false wine discrimination model with low generalization ability.

[0077] For the true and false wine discrimination model with low generalization ability, use the transfer learning method to perform multi-stage training and optimization based on datasets of different brands and batches, and evaluate the detection accuracy of the true and false wine discrimination model on unseen brands and batches through cross-validation;

[0078] Let the source brand / batch dataset be DS, its feature distribution be PS(X), and the label distribution be ; Let the target brand / batch dataset be DT, its feature distribution be PT(X), and the label distribution be .

[0079] The goal is to transfer the model trained using DS to DT and optimize its performance. The expression is: ; where: θ are the model parameters, E is the loss function (such as cross-entropy loss, mean squared error), represents the expected loss under the target brand / batch data distribution. When PS(X) ≠ PT(X) or the model may have insufficient generalization ability, so transfer learning methods need to be adopted for optimization.

[0080] The initial model trained from the source brand / batch dataset DS is used for transfer learning, that is: ; where: are the initial model parameters on the source brand / batch, are the optimized model parameters on the target brand / batch, and Δθ is the amount of parameter adjustment during the transfer process.

[0081] To make the model adapt to the target brand / batch data, self-supervised domain adaptation is adopted for optimization to minimize the distribution deviation between the source domain DS and the target domain DT:

[0082] During the transfer learning process, adversarial training is introduced, such as the Gradient Reversal Layer (GRL), to minimize the brand / batch deviation of the true and false wine discrimination features:

[0083] To evaluate the optimized model for detection accuracy on unseen brands / batches, K-Fold Cross-Validation is used for evaluation.

[0084] The target brand / batch dataset DT is divided into K parts for K rounds of training and testing: in the k-th round, use the part in DT as the training set, and the remaining 1 part as the test set;

[0085] Calculate the model accuracy Acck for each round;

[0086] Calculate the average accuracy as the generalization ability index of the model on the target brand / batch:

[0087] If the average accuracy is greater than the pre-set accuracy threshold, it indicates that the transfer learning optimization is effective and the model has better adaptability on new brands / batches.

[0088] After completing the transfer learning and generalization ability optimization of the model, the multi-modal features of the wine sample to be tested can be input into the optimized real / fake wine discrimination model to output the real / fake wine discrimination result and give the corresponding confidence score, ensuring the accuracy and reliability of the detection.

[0089] The feature data of the wine sample to be tested is obtained through non-destructive testing, including but not limited to: composition data (such as ethanol content, organic acids, esters, etc.); spectral curves (ultraviolet-visible light spectra, infrared spectra, etc.); odor fingerprints (characteristics of volatile organic compounds detected by electronic nose sensors); texture information (high-resolution image analysis). These features are standardized and feature-extracted, and converted into the input format required by the optimized real / fake wine discrimination model.

[0090] The optimized real / fake wine discrimination model receives the input features and performs inference to determine whether the wine sample to be tested is real or fake.

[0091] The input features are analyzed through a deep learning model or a machine learning classifier (such as a convolutional neural network CNN, random forest RF) and compared with the existing data in the model.

[0092] Combined with the generalization ability optimization across brands and batches, the adaptability of the model to different wine samples is ensured.

[0093] Based on the fusion result of the multi-modal features, the final real / fake wine discrimination result is output.

[0094] To increase the credibility of the detection result, the confidence score is calculated simultaneously to measure the reliability of the discrimination result.

[0095] The confidence score is based on the feature similarity of the model to the sample to be tested and measures its matching degree with the known real / fake wine samples in the database.

[0096] The score usually ranges from 0% to 100%. The higher the value, the more certain the model is about the judgment of the sample. Confidence ≥ 90%: High-confidence discrimination, the model is highly certain about the real / fake wine result.

[0097] Confidence 70% - 90%: Medium confidence, it is recommended to combine other detection methods for auxiliary judgment.

[0098] Confidence ≤ 70%: Low confidence, it may be necessary to increase the training data or further optimize the model.

[0099] Finally, the system outputs the real / fake wine discrimination result and provides the confidence score. For example: Result: Fake wine; Confidence: 95%; indicating that the spectral features and odor fingerprints of the sample are highly similar to the known fake wine samples, and the model has a high confidence in judging it as fake wine.

[0100] If the confidence level is low, the system can provide suggestions for further detection, such as adding additional detection means (such as chemical analysis) or combining expert evaluation to improve the reliability of the final decision.

[0101] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0102] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0103] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0104] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A method for constructing a model for distinguishing genuine and fake wines, characterized in that: include: Obtain wine samples from multiple brands and batches, and extract multimodal data based on a variety of non-destructive testing methods, including composition data, spectral curves, odor fingerprints, and texture information; The collected multimodal data is standardized, and feature fusion and hierarchical extraction of feature data of different modalities are performed based on the pre-built deep wine discrimination model, and the distribution similarity features and domain shift robustness features in the multimodal data are extracted respectively; This includes testing on new brands or batches, calculating the performance degradation ratio, and taking the inverse of the performance degradation ratio as a domain shift robustness feature; The extracted distribution similarity features and domain shift robustness features are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as inputs of the machine learning model, and the generalization ability value of the real and fake wine discrimination model is output through the machine learning model; For the genuine and fake wine discrimination model with low generalization ability, we use the transfer learning method to conduct multi-stage training and optimization based on data sets of different brands and batches, and evaluate the detection accuracy of the genuine and fake wine discrimination model on unseen brands and batches through cross-validation; The multimodal features of the wine to be tested are input into the optimized model for distinguishing genuine and fake wines, and the results of distinguishing genuine and fake wines are output, along with the corresponding confidence scores.

2. The method for constructing a genuine wine discrimination model according to claim 1, characterized in that: Multimodal data is extracted based on a variety of non-destructive testing methods, including near-infrared spectroscopy, Raman spectroscopy, UV-visible spectroscopy, fluorescence spectroscopy, Fourier transform infrared spectroscopy, gas chromatography-mass spectrometry, electronic nose and high-resolution microscopy, which are used to extract the composition data, spectral curves, odor fingerprints and texture information of the wine respectively.

3. The method for constructing a genuine wine discrimination model according to claim 2, characterized in that: After analyzing the distribution similarity characteristics of different brands and batches of wine samples in the feature space, the distribution similarity deviation index is generated. The method for obtaining the distribution similarity deviation index is as follows: ; Define a kernel function k(x,y) to measure the similarity between two sample points: 。 4. The method for constructing a genuine wine discrimination model according to claim 3, characterized in that: After analyzing the acquired domain shift robustness features, a domain shift robustness anomaly index is generated. The method for obtaining the domain shift robustness anomaly index is as follows: ; The autoencoder structure includes: Encoder: maps high-dimensional input features to a low-dimensional latent space z; Decoder: remaps the low-dimensional latent variable z back to the original feature space; Loss function: Use mean square error to calculate the error between input data x and reconstructed data; After training, the autoencoder should be able to accurately reconstruct the sample features of the source brand / batch; 。 5. The method for constructing a genuine wine discrimination model according to claim 4, characterized in that: The distribution similarity deviation index and the domain shift robustness anomaly index are converted into comprehensive feature vectors, which are used as inputs of the machine learning model. The machine learning model predicts the generalization ability value label of the genuine and fake wine discrimination model with each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of prediction errors of the generalization ability value labels of all genuine and fake wine discrimination models as the training target. The machine learning model is trained until the sum of prediction errors converges, and the model training is stopped. The generalization ability value of the genuine and fake wine discrimination model is determined according to the model output results, wherein the machine learning model is a polynomial regression model.

6. The method for constructing a genuine wine discrimination model according to claim 5, characterized in that: The obtained generalization ability value of the genuine and fake wine discrimination model is compared with a pre-set reference threshold of the generalization ability value. If the generalization ability value of the genuine and fake wine discrimination model is greater than or equal to the pre-set reference threshold of the generalization ability value, it indicates that the fake wine discrimination model has high generalization ability, and it is marked as a genuine and fake wine discrimination model with high generalization ability. If the generalization ability value of the genuine and fake wine discrimination model is less than a pre-set reference threshold of the generalization ability value, it means that the generalization ability of the fake wine discrimination model is low, and it is marked as a genuine and fake wine discrimination model with low generalization ability.

7. The method for constructing a genuine wine discrimination model according to claim 6, characterized in that: The cross-validation adopts the K-fold cross-validation method, which divides the target brand or batch data set into K parts, trains and tests the model separately, and calculates the average accuracy of all K rounds of experiments. If the average accuracy is greater than a preset threshold, it means that the transfer learning optimization is effective.

8. The method for constructing a genuine wine discrimination model according to claim 7, characterized in that: When distinguishing the wine to be tested, the optimized model for distinguishing genuine and fake wine uses a deep learning model based on multimodal feature fusion for inference and outputs a confidence score. The confidence score is based on the feature similarity of the model to the sample to be tested. If the confidence is ≥90%, the judgment result is reliable. If the confidence is lower than 70%, additional detection methods are added to assist decision-making.

Citation Information

Patent Citations

  • Method and system for measuring total acid content and total ester content of wine

    CN116735811A

  • Self-adaptive feature fusion physiological signal identification method and related equipment

    CN117668604A