Spectrum-based liquor year identification method, system and device and storage medium
Through a spectrum-based method, three-dimensional fluorescence spectral data and prediction model are used to solve the problem of identifying the year of the sauce-flavored liquor, and the accurate identification of the year of the sauce-flavored liquor is achieved, ensuring the authenticity of the year of the liquor.
Patent Information
- Application Number
- CN202510313501.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology lacks effective annual testing technology for sauce-flavored liquor, which has led to the uneven quality of sauce-flavored liquors on the market, and false vintage liquors are emerging one after another, making it difficult to accurately identify the year of liquor.
Using a spectral-based method, we use parallel factor analysis and neural networks or random forest regression models to establish a prediction model and identify the year of the liquor by obtaining three-dimensional fluorescence spectral data of sauce-flavored liquor, and using parallel factor analysis and neural networks or random forest regression models.
The accurate identification of the year of sauce-flavored liquor has been achieved, the accuracy and reliability of the identification have been improved, and the authenticity of the year of liquor has been ensured.
Smart Images

Figure CN120277550A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of detection technology, and in particular to a method, system, device and storage medium for identifying the vintage of Baijiu based on spectroscopy. Background Art
[0002] Jiangxiang-flavor Baijiu is the Baijiu with the most complex brewing process among the twelve major Baijiu flavors in China. Aged Daqu Jiangxiang-flavor Baijiu is scarce and expensive. The price of Daqu Jiangxiang-flavor Baijiu varies greatly depending on the vintage, and fake vintage Baijiu is emerging in an endless stream. Therefore, accurately identifying the vintage of Baijiu is of great significance. At the same time, due to the lack of a mature detection technology and system for the vintage of Jiangxiang-flavor Baijiu, the quality of Jiangxiang-flavor vintage Baijiu on the market is uneven, and even counterfeit and shoddy products are prevalent. Therefore, there is an urgent need in the market for a new technology that can detect the vintage of Jiangxiang-flavor Baijiu. Summary of the Invention
[0003] In view of this, the purpose of the embodiments of the present invention is to provide a method, system, device and storage medium for identifying the vintage of Baijiu based on spectroscopy, which can accurately identify the vintage of Jiangxiang-flavor Baijiu.
[0004] On the one hand, the embodiments of the present invention provide a method for identifying the vintage of Baijiu based on spectroscopy, including:
[0005] Obtaining the three-dimensional fluorescence spectrum data of the to-be-detected Jiangxiang-flavor Baijiu, and determining several characteristic component score values of the to-be-detected Jiangxiang-flavor Baijiu according to the three-dimensional fluorescence spectrum data of the to-be-detected Jiangxiang-flavor Baijiu;
[0006] Inputting the several characteristic component score values of the to-be-detected Jiangxiang-flavor Baijiu into a preset prediction model to identify the vintage of the to-be-detected Jiangxiang-flavor Baijiu; the prediction model is obtained by training according to sample data, and the sample data includes three-dimensional fluorescence spectrum data samples and vintage samples of several Jiangxiang-flavor Baijius.
[0007] Optionally, the prediction model is obtained by training through the following method:
[0008] Decomposing the three-dimensional fluorescence spectrum data sample of each Jiangxiang-flavor Baijiu in the sample data to obtain several characteristic component score sample values of each Jiangxiang-flavor Baijiu; the data decomposition includes parallel factor analysis;
[0009] Taking the several characteristic component score value samples of each Jiangxiang-flavor Baijiu as the input and the vintage sample as the output, training and testing a preset model until the requirements are met, determining the parameters of the preset model, and obtaining the prediction model; the preset model includes a forward neural network regression model or a random forest regression model.
[0010] Optionally, determining a plurality of characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor includes:
[0011] Determining a plurality of characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data and a plurality of characteristic component score sample values of a plurality of Maotai-flavor liquors.
[0012] Optionally, determining a plurality of characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data and a plurality of characteristic component score sample values of a plurality of Maotai-flavor liquors includes:
[0013] Establishing a relationship between the spectral data parameters and the characteristic component score parameters according to the characteristic component score sample parameters;
[0014] Substituting the three-dimensional fluorescence spectrum data and the plurality of characteristic component score sample values of a plurality of Maotai-flavor liquors into the relationship respectively, and determining a plurality of characteristic component score values of the to-be-detected Maotai-flavor liquor by using the least squares method for fitting.
[0015] Optionally, the method further includes:
[0016] Selecting a plurality of sample data by stratified sampling for least squares method fitting to obtain the least squares method fitting residuals;
[0017] Comparing the fitting residuals of parallel factor analysis with the least squares method fitting residuals, and verifying the effectiveness of a plurality of characteristic component score values according to the result of the similarity comparison.
[0018] Optionally, the method further includes:
[0019] Performing outlier analysis on a plurality of sample data, and removing the outliers if there are outliers.
[0020] Optionally, the method further includes:
[0021] Removing Raman scattering and Rayleigh scattering from the three-dimensional fluorescence spectrum data sample and the three-dimensional fluorescence spectrum data.
[0022] On the other hand, an embodiment of the present invention provides a liquor vintage identification system based on spectrum, including:
[0023] A first module, configured to obtain three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor, and determine a plurality of characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor;
[0024] A second module is configured to input the score values of several characteristic components of the to-be-detected Jiangxiang-flavor Baijiu into a preset prediction model to identify the vintage of the to-be-detected Jiangxiang-flavor Baijiu; the prediction model is trained based on sample data, and the sample data includes three-dimensional fluorescence spectrum data samples and vintage samples of several Jiangxiang-flavor Baijiu.
[0025] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above method when executed by the processor.
[0026] On the other hand, an embodiment of the present invention provides a Baijiu vintage identification system based on spectroscopy, including a fluorescence spectrometer and a computer device connected to the fluorescence spectrometer; wherein,
[0027] The fluorescence spectrometer is configured to collect the fluorescence spectrum of Jiangxiang-flavor Baijiu;
[0028] The computer device includes:
[0029] At least one processor;
[0030] At least one memory for storing at least one program;
[0031] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0032] Implementing the embodiments of the present invention includes the following beneficial effects: In this embodiment, a prediction model is first trained based on three-dimensional fluorescence spectrum data samples and vintage samples of several Jiangxiang-flavor Baijiu, and then the score values of several characteristic components of the to-be-detected Jiangxiang-flavor Baijiu are determined according to the three-dimensional fluorescence spectrum data of the to-be-detected Jiangxiang-flavor Baijiu, and the score values of several characteristic components of the to-be-detected Jiangxiang-flavor Baijiu are input into the prediction model to identify the vintage of the Baijiu. By analyzing the characteristic components of Jiangxiang-flavor Baijiu, the vintage of Jiangxiang-flavor Baijiu can be accurately identified. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic flow chart of the steps of a Baijiu vintage identification method based on spectroscopy provided by an embodiment of the present invention;
[0034] Figure 2 is a PARAFAC analysis process of the three-dimensional fluorescence spectrum of a Jiangxiang-flavor Baijiu provided by an embodiment of the present invention;
[0035] Figure 3 is a three-dimensional fluorescence spectrum diagram of a Jiangxiang-flavor Baijiu provided by an embodiment of the present invention;
[0036] Figure 4It is the PARAFAC analysis component matrix diagram of a Maotai-flavor aged liquor provided by an embodiment of the present invention;
[0037] Figure 5 It is the comparison diagram of the characteristic component score values of a modeling sample provided by an embodiment of the present invention;
[0038] Figure 6 It is the prediction result diagram of a BPNN regression model provided by an embodiment of the present invention;
[0039] Figure 7 It is the prediction result diagram of an RF regression model provided by an embodiment of the present invention;
[0040] Figure 8 It is the structural block diagram of a liquor age identification system based on spectroscopy provided by an embodiment of the present invention;
[0041] Figure 9 It is another structural block diagram of a liquor age identification system based on spectroscopy provided by an embodiment of the present invention. Detailed implementation manners
[0042] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0043] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the description, claims and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0045] Some technical terms in this embodiment will be explained below.
[0046] Three-dimensional fluorescence spectroscopy is a non-destructive detection method. There are more than 1,400 flavor substances in Maotai-flavor liquor. Although the main flavor components are not yet clear, a series of chemical reactions such as oxidation, hydrolysis, and esterification occur during the aging process of Maotai-flavor liquor. The flavor substances are in dynamic change and gradually reach equilibrium. Therefore, by measuring the three-dimensional fluorescence spectroscopy of Maotai-flavor liquor and combining mathematical analysis methods, the identification of the age of Maotai-flavor liquor can be achieved.
[0047] Maotai-flavor liquor is a complex system. The fluorescence signal generated by the liquor comes from the combination of individual signals of different intrinsic fluorescent molecules and is also affected by physicochemical environments such as temperature, color, and pH. Therefore, in order to handle this complexity, a multivariate and multi-way data processing method can be adopted. Parallel factor analysis (PARAFAC) is a powerful tool for measuring fluorescent components, which can extract the excitation and emission spectra of fluorescent molecules and their relative concentrations in each sample. The information obtained by PARAFAC can be used to establish robust and reliable classification and discrimination models, such as support vector machine and partial least squares discriminant analysis models.
[0048] PARAFAC is a multi-dimensional data decomposition method that can be used for the analysis of high-dimensional data. Different from the traditional two-dimensional PCA (principal component analysis), PARAFAC has no rotation problem, can recover pure spectra from multivariate spectral data, and the model construction is simple and intuitive, with wide applications in fields such as chemometrics, signal processing, psychology, and bioinformatics. The excitation-emission spectral data of liquor samples contain a large amount of information about the compounds in the liquor, which is complex and difficult to interpret. Through PARAFAC, the complex fluorescence signal is decomposed into potential individual fluorescences, obtaining a score set and loadings of potential components to describe the original data.
[0049] The random forest regression model improves the accuracy and robustness of prediction by constructing multiple decision trees (usually called base learners or weak learners). Each decision tree is trained on different data subsets and feature subsets, and finally, a more accurate and robust prediction is generated by voting or averaging the prediction results. Increasing the number of trees can improve the accuracy and robustness of the model because the voting of multiple trees can reduce the risk of overfitting. The parameters of the random forest regression model include the number of decision trees, the maximum depth of each tree, leaf nodes, etc. The number of trees in the random forest regression model can be determined by various methods, such as cross-validation method or error analysis method.
[0050] Refer to Figure 1 , the embodiment of the present invention provides a method for identifying the age of liquor based on spectroscopy, including:
[0051] S100. Obtain the three-dimensional fluorescence spectral data of the to-be-detected Jiangxiang-type Baijiu, and determine the score values of several characteristic components of the to-be-detected Jiangxiang-type Baijiu according to the three-dimensional fluorescence spectral data of the to-be-detected Jiangxiang-type Baijiu;
[0052] S200. Input the score values of several characteristic components of the to-be-detected Jiangxiang-type Baijiu into a preset prediction model to identify the production year of the to-be-detected Jiangxiang-type Baijiu; the prediction model is trained according to sample data, and the sample data includes the three-dimensional fluorescence spectral data samples and production year samples of several Jiangxiang-type Baijius.
[0053] Specifically, first, decompose the three-dimensional fluorescence spectral data samples of several Jiangxiang-type Baijius with known production years to obtain the score samples of the characteristic components of several Jiangxiang-type Baijius; then, use the score samples of the characteristic components of several Jiangxiang-type Baijius as the input of the model, and use the production year samples of several Jiangxiang-type Baijius as the output of the model. Train the model according to the sample data. After the training is completed, obtain the preset prediction model; then, analyze the three-dimensional fluorescence spectral data of the to-be-detected Jiangxiang-type Baijiu to obtain the score values of several characteristic components, and input the score values of several characteristic components into the preset prediction model to identify the production year of the to-be-detected Jiangxiang-type Baijiu.
[0054] It should be noted that the number of characteristic components is determined according to actual applications, and no specific limitation is made in this embodiment, such as 3 characteristic components. The prediction model is determined according to actual applications, and no specific limitation is made in this embodiment.
[0055] Due to the unique production technology and blending technology of Baijiu, it is difficult to collect real aged Baijiu on the market. Therefore, according to the market definition of the production year of Baijiu, the production year of the Baijiu samples used in this embodiment refers to the storage year of Jiangxiang-type Baijiu, that is, starting from the second year after the production of Xiasha, to ensure the authenticity of the production year of the Baijiu samples.
[0056] In this embodiment, a total of 154 known-year Daqu Jiangxiang-type Baijiu samples from 38 Baijiu production enterprises of different scales in the core production area of Jiangxiang-type Baijiu are collected as sample data. Since the production cycle of Jiangxiang-type Baijiu is 1 year, the actual production year of the Baijiu is calculated as "2024 - the production start year of the Baijiu - 1". The production year distribution of the Baijiu samples is shown in Table 1. Select 2 / 3 of the samples to train the production year prediction regression model, and 1 / 3 of the samples are used to test the performance of the regression model.
[0057] Table 1
[0058] Year Number of wine samples 1 - 5 years 18 6 - 10 years 55 11 - 15 years 52 16 - 20 years 21 21 - 25 years 6 26 - 30 years 2
[0059] Optionally, the method further includes:
[0060] S001. Remove Raman scattering and Rayleigh scattering from the three-dimensional fluorescence spectral data samples and the three-dimensional fluorescence spectral data.
[0061] Raman scattering and Rayleigh scattering affect the linear structure data of the three-dimensional fluorescence spectrum data sample and the three-dimensional fluorescence spectrum data, prepare for subsequent calculations, and remove Raman scattering and Rayleigh scattering.
[0062] Optionally, the method further includes:
[0063] S002. Perform outlier analysis on a number of sample data. If there are outliers, remove the outliers.
[0064] Outliers will affect the accuracy of training and testing the regression model of sample data. Therefore, removing the outliers in a number of sample data can improve the accuracy of the prediction model.
[0065] Optionally, the prediction model is trained by the following method:
[0066] S010. Decompose the three-dimensional fluorescence spectrum data sample of each Maotai-flavor liquor in the sample data to obtain a number of characteristic component score sample values for each Maotai-flavor liquor; the data decomposition includes parallel factor analysis;
[0067] S020. Take the number of characteristic component score values of each Maotai-flavor liquor as the input and the year sample as the output, train and test the preset model until the requirements are met, determine the parameters of the preset model, and obtain the prediction model; the preset model includes a forward neural network regression model or a random forest regression model.
[0068] Parallel factor analysis (PARAFAC) is a classical multi-way data decomposition method, which is widely used in the mathematical decomposition of fluorescence data matrices. A unique decomposition result with clear chemical significance can be obtained by decomposing the data matrix. In this study, the DOMFluor toolbox was used to perform PARAFAC analysis on the three-dimensional fluorescence data.
[0069] The three-dimensional fluorescence spectrum is obtained from a series of fluorescence (emission) spectra measured at different excitation wavelengths. The fluorescence data of a series of liquor samples are stacked into a three-dimensional excitation-emission data array (EEM, Excitation Emission Matrix) with a size of I×J×K (number of samples × number of emission wavelengths × number of excitation wavelengths). After removing Raman scattering and Rayleigh scattering from the fluorescence data matrix, PARAFAC analysis is performed on it. PARAFAC uses the following equation to perform modeling analysis on the three-dimensional data.
[0070]
[0071] where, x ijk is the fluorescence intensity of sample i at emission wavelength j and excitation wavelength k, F is the number of fluorescence components, εijk To minimize the sum of squared residuals. Through decomposition, three matrices with clear chemical meanings are obtained, including the characteristic component score matrix a, the emission spectrum matrix b, and the excitation spectrum matrix c. The characteristic component score matrix a can be combined with other pattern recognition algorithms to build a year prediction model.
[0072] In a specific embodiment, refer to Figure 2 , the three-dimensional spectral data obtained by a fluorescence spectrometer is decomposed into a three-dimensional excitation-emission data array (EEMs), and through PARAFAC analysis, the characteristic component score values of C1 / C2 / C3 are obtained.
[0073] The training process of the backpropagation neural network (BPNN) regression model is as follows: The three characteristic component score values obtained by PARAFAC analysis of the three-dimensional spectral data of the modeling samples are used to build a BP neural network regression prediction model. With the help of a neural network toolbox or code, using the characteristic component score values of Baijiu from different years as inputs and the Baijiu year as the output, a two-layer BP neural network with sigmoid hidden neurons and linear output neurons is constructed, and parameter optimization is performed according to the actual situation during the training process.
[0074] The training process of the random forest (RF) regression model is as follows: Determine parameters such as the number of trees and the number of leaf nodes of the random forest regression model. For example, the number of trees in the random forest regression model is 100, and the minimum number of leaf nodes is 2; using the characteristic component score values of Baijiu from different years as inputs and the Baijiu year as the output. Parameter optimization is performed according to the actual situation during the training process.
[0075] Optionally, several characteristic component score values are determined according to the three-dimensional fluorescence spectral data, including:
[0076] S110. Determine the characteristic component score values of the to-be-detected Maotai-flavor Baijiu according to the three-dimensional fluorescence spectral data of the to-be-detected Maotai-flavor Baijiu and the characteristic component score sample values of several Maotai-flavor Baijiu.
[0077] The PARAFAC analysis results are obtained based on batch samples, and the addition of new samples will change the PARAFAC analysis results. Therefore, in order to use the built neural network model to predict the year of unknown Baijiu samples, the component scores of the samples need to be obtained first. In this embodiment, according to the characteristic component score samples of several Maotai-flavor Baijiu, several characteristic component score values that best match the three-dimensional fluorescence spectral data of the to-be-detected Maotai-flavor Baijiu are determined.
[0078] Optionally, determining the score values of several characteristic components of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data and the score sample values of several characteristic components of several Maotai-flavor liquors includes:
[0079] S111. Establish a relationship between the spectral data parameters and the characteristic component score parameters according to the characteristic component score sample parameters;
[0080] S112. Substitute the three-dimensional fluorescence spectrum data and the score sample values of several characteristic components of several Maotai-flavor liquors into the relationship respectively, and use the least squares method to fit and determine the score values of several characteristic components of the to-be-detected Maotai-flavor liquor.
[0081] The relationship between the spectral data parameters and the characteristic component score parameters is as follows:
[0082] y = a1*Z1 + a2*Z2 +... + an*Zn + D
[0083] Wherein, y represents the three-dimensional fluorescence data matrix of the to-be-detected Maotai-flavor liquor, and Z1, Z2,... Zn are the score values of several characteristic components obtained from the PARAFAC results respectively. By using the least squares method to fit and solve, a1, a2,... an are obtained to minimize D. a1, a2,... an respectively represent multiple proportionality coefficients of the to-be-detected Maotai-flavor liquor. According to a1, a2,... an and Z1, Z2,... Zn, the final score values of several characteristic components of the to-be-detected Maotai-flavor liquor are determined, and the year of the to-be-detected Maotai-flavor liquor is predicted by using the score values of several characteristic components.
[0084] Optionally, the method further includes:
[0085] S121. Select several sample data by stratified sampling for least squares fitting to obtain the least squares fitting residual;
[0086] S122. Compare the fitting residual of parallel factor analysis with the least squares fitting residual, and verify the effectiveness of the score values of several characteristic components according to the result of the similarity comparison.
[0087] To ensure the effectiveness of the score values of the characteristic components of the to-be-detected Maotai-flavor liquor obtained by least squares fitting, in this embodiment, several modeling samples are selected by stratified sampling for least squares fitting, and the least squares fitting residual is compared with the PARAFAC fitting residual for similarity. To retain the internal structural information of the matrix, the cosine similarity of the two matrices is calculated by using the inner product and norm of the matrix. The calculation formula is as follows:
[0088]
[0089] ∑ i,j Ai,j B i,j is the sum of the products of the corresponding elements of matrix A (the least squares fitting residual matrix) and B (the PARAFAC fitting residual matrix) (also known as the Frobenius inner product); and are the Frobenius norms of matrices A and B, respectively.
[0090] It should be noted that all code implementations are carried out in the MATLAB environment, using tools such as the plotting software Origin 2019b and Adobe Illustrator CC 2018.
[0091] The following uses a specific embodiment to illustrate the process of liquor vintage identification based on spectroscopy.
[0092] Step 1: EEM fluorescence spectra and data preprocessing of Maotai-flavor vintage liquor.
[0093] The original EEM fluorescence spectra of vintage liquor were obtained in the excitation range of 250–550 nm and the emission range of 300–650 nm, as Figure 3 shown, Figure 3 where (A) in Figure 3 represents a three-dimensional stereogram, and (B) in
[0094] represents a contour plot. Since Raman scattering and Rayleigh scattering affect the linear data structure of EEM, this provides a basis for the identification of the vintage of Maotai-flavor liquor.
[0094] Step 2: PARAFAC analysis.
[0095] The size of the three-dimensional data matrix of the Maotai-flavor liquor vintage prediction model is 100×36×31. After preprocessing, in the non-negative constrained data analysis mode, 1 abnormal liquor sample was deducted through outlier analysis. After model verification, it was determined that the three-component PARAFAC model is applicable to the analysis of Maotai-flavor vintage liquor. Figure 4 shows the score values, emission spectral matrix, and excitation spectral matrix of the three characteristic components obtained by PARAFAC analysis of Maotai-flavor vintage liquor. Figure 4 where (A) in Figure 4 represents the characteristic component score matrix, (B) in Figure 4 represents the score of the first characteristic component, (C) in Figure 4 represents the score of the second characteristic component, and (D) in Figure 5As shown Figure 5 In (A), it represents the scores of each sample component Figure 5 In (B), it represents the average score of the samples within the year interval. From Figure 5 In (B), it can be seen that as the age of the Baijiu increases, the score values of the characteristic components generally show an increasing trend, which provides a basis for the prediction of the age of Baijiu
[0096] Step 3: Establish a regression prediction model for the age of Daqu Maotai-flavor Baijiu based on different strategies
[0097] Forward neural network (BPNN) regression model: Through PARAFAC analysis, three characteristic components of 99 Daqu Maotai-flavor aged Baijiu samples and the score values of each liquor sample were obtained. With the help of the neural network toolbox or code, using the component score values as the input and the age of Baijiu as the output, a two-layer BP neural network with sigmoid hidden neurons and linear output neurons was constructed. The number of neurons in the hidden layer of the neural network is 9, the number of nodes in the input layer is 3, the number of nodes in the output layer is 1, the training function is trainlm, and the transfer function, learning rate, target error, and number of training iterations all adopt the default parameters of the toolbox. The relevant parameters of the trained neural network model are as Figure 6 shown in Table 2 Figure 6 In (A), it represents the comparison chart of the actual value and the predicted value results Figure 6 In (B), it represents the error histogram. The correlation coefficient R of the training set is 0.9, the determination coefficient R2 is 0.82, the validation set R is 0.89, R2 is 0.76, the test set R is 0.9, and R2 is 0.79. Among the absolute values of the error (actual year - network predicted year) results output by the network, the proportion within 3 years is 80%, the proportion within 5 years is 95%, and the average predicted year difference is 2 years
[0098] Random forest (RF) regression model: Using the score values of the three characteristic components of 99 Daqu Maotai-flavor aged Baijiu samples as the input and the age of Baijiu as the output, a random forest regression model was constructed. The results show that referring to Figure 7 , the regression value R of the training set is 0.94, R2 is 0.82, the validation set R is 0.9, R2 is 0.79, the test set R is 0.84, and R2 is 0.68. Among the absolute values of the error (actual year - network predicted year) results, the proportion within 3 years is 80%, the proportion within 5 years is 96%, and the average predicted year difference is 1.9 years
[0099] Table 2
[0100]
[0101] Step 4: Predict the age of Daqu Maotai-flavor Baijiu based on the regression model
[0102] To ensure the effectiveness of the individual samples a1, a2, and a3 obtained by fitting using the least squares method, in this embodiment, 20 modeling samples are selected by stratified sampling for fitting. The similarity between the least squares fitting residuals and the PARAFAC fitting residuals is compared, and the comparison results are shown in Table 3. From the similarity comparison results, it can be seen that the residuals obtained by fitting using the least squares method have good similarity, and the average difference in predicted years between the two is 0.3 years. Therefore, the scores of samples with unknown years can be obtained by fitting using the least squares method, and the established neural network model can be used to predict their years.
[0103] Fifty-four year-old wine samples were fitted using the least squares method, and their scores were used to predict the years to test the performance of the established neural network year prediction model. From the test results, it can be seen that the difference between the predicted year of the BPNN neural network regression model and the actual year of the sample accounts for 72% within 3 years and 85% within 5 years, the MSE is 14.26, and the average predicted year difference is 2.5 years. The difference between the predicted year of the RF regression model and the actual year of the sample accounts for 70% within 3 years and 87% within 5 years, the MSE is 12.25, and the average predicted year difference is 2.4 years.
[0104] Table 3
[0105]
[0106]
[0107] This embodiment proposes a complete method for detecting the age of Maotai-flavor Baijiu (including EEM measurement, data preprocessing, PARAFAC analysis, regression model establishment, and prediction of the age of Baijiu samples). From the results of PARAFAC analysis, it can be seen that the scores of the three components obtained by decomposing Baijiu samples of different years are different. In response to this result, this study proposes a strategy for constructing year prediction regression models of BPNN and RF. At the same time, to solve the problem that the addition of new samples will change the PARAFAC analysis results of the modeling samples, this study proposes a method for solving the scores of Baijiu components with unknown years by fitting using the least squares method. The obtained scores can be used by the BPNN and RF regression models to predict the age of Baijiu. From the results, it can be seen that the established BPNN and RF regression models have good performance. The average predicted year differences for the modeling samples are 2 years and 1.9 years respectively, and the average predicted year differences for the 54 test samples are 2.5 years and 2.4 years respectively. The above results strongly support that the excitation-emission matrix fluorescence spectroscopy combined with chemometric methods can become a powerful tool for identifying the age of Daqu Maotai-flavor Baijiu.
[0108] Implementing the embodiments of the present invention includes the following beneficial effects: First, a prediction model is trained based on the three-dimensional fluorescence spectrum data samples and year samples of several kinds of Maotai-flavor Baijiu. Then, several characteristic component score values of the Maotai-flavor Baijiu to be measured are determined according to the three-dimensional fluorescence spectrum data of the Maotai-flavor Baijiu to be measured, and the several characteristic component score values are input into the prediction model to identify the year of the Baijiu. By analyzing the characteristic components of the Maotai-flavor Baijiu, the year of the Maotai-flavor Baijiu can be accurately identified.
[0109] Refer to Figure 8 , the embodiments of the present invention provide a Baijiu year identification system based on spectroscopy, including:
[0110] A first module, configured to obtain the three-dimensional fluorescence spectrum data of the Maotai-flavor Baijiu to be measured, and determine several characteristic component score values of the Maotai-flavor Baijiu to be measured according to the three-dimensional fluorescence spectrum data of the Maotai-flavor Baijiu to be measured;
[0111] A second module, configured to input several characteristic component score values of the Maotai-flavor Baijiu to be measured into a preset prediction model to identify the year of the Maotai-flavor Baijiu to be measured; the prediction model is trained according to sample data, and the sample data includes three-dimensional fluorescence spectrum data samples and year samples of several kinds of Maotai-flavor Baijiu.
[0112] It can be seen that the content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0113] The embodiments of the present invention also provide a computer-readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to implement the above method when executed by the processor.
[0114] It can be understood that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0115] Referring to Figure 9 , an embodiment of the present invention provides a liquor vintage identification system based on spectroscopy, including a fluorescence spectrometer and a computer device connected to the fluorescence spectrometer; wherein,
[0116] The fluorescence spectrometer is used to collect the fluorescence spectrum of Maotai-flavor liquor;
[0117] The computer device includes:
[0118] At least one processor;
[0119] At least one memory for storing at least one program;
[0120] When at least one program is executed by at least one processor, the at least one processor implements the above method.
[0121] Specifically, a fluorescence spectrometer (also known as a fluorescence spectrophotometer) is an instrument that realizes qualitative and quantitative analysis by detecting the fluorescence signal emitted by a substance after excitation. The fluorescence spectrometer includes a light source, a monochromator, a sample cell, a detector, a data processing system, etc. In this embodiment, a steady-state fluorescence spectrometer equipped with a 150W continuous xenon light source is used to obtain the excitation-emission matrix (EEM, Excitation Emission Matrix) of white liquor. The instrument setting parameters are 250-550 nm, with a step size of 10 nm; the emission wavelength range is 300-650 nm, with a step size of 10 nm, a spectral bandwidth of 5 nm, a scanning speed of 100 nm / s, and an integration time of 100 s. For the computer device, it can be different types of electronic devices, including but not limited to terminals such as desktop computers and laptops.
[0122] It can be seen that the content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented in the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0123] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (piece) of the following" or its similar expression refers to any combination of these items, including any combination of single item (piece) or plural items (pieces). For example, at least one (piece) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0124] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0125] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for identifying the vintage of Chinese liquor based on spectroscopy, characterized in that, Including: Obtain the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor, and determine several characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor; Input several characteristic component score values of the to-be-detected Maotai-flavor liquor into a preset prediction model to identify the vintage of the to-be-detected Maotai-flavor liquor; the prediction model is trained according to sample data, and the sample data includes three-dimensional fluorescence spectrum data samples and vintage samples of several Maotai-flavor liquors.
2. The method according to claim 1, wherein The prediction model is trained by the following method: Decompose the three-dimensional fluorescence spectrum data sample of each Maotai-flavor liquor in the sample data to obtain several characteristic component score sample values of each Maotai-flavor liquor; the data decomposition includes parallel factor analysis; Use several characteristic component score sample values of each Maotai-flavor liquor as input and the vintage sample as output to train and test a preset model until the requirements are met, determine the parameters of the preset model, and obtain a prediction model; the preset model includes a forward neural network regression model or a random forest regression model.
3. The method according to claim 2, wherein The step of determining several characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor includes: Determine several characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data and several characteristic component score sample values of several Maotai-flavor liquors.
4. The method according to claim 3, wherein The step of determining several characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data and several characteristic component score sample values of several Maotai-flavor liquors includes: Establish a relationship between the spectral data parameters and the characteristic component score parameters according to the characteristic component score sample parameters; Substitute the three-dimensional fluorescence spectrum data and several characteristic component score sample values of several Maotai-flavor liquors into the relationship respectively, and use the least squares method to fit to determine several characteristic component score values of the to-be-detected Maotai-flavor liquor.
5. The method according to claim 4, characterized in that, The method further includes: Select several sample data by stratified sampling for least squares fitting to obtain the least squares fitting residuals; Compare the fitting residuals of parallel factor analysis with the least squares fitting residuals, and verify the effectiveness of several characteristic component score values according to the result of the similarity comparison.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Perform outlier analysis on several sample data, and if there are outliers, remove the outliers.
7. The method according to any one of claims 1-5, characterized in that, The method further includes: Remove Raman scattering and Rayleigh scattering from the three-dimensional fluorescence spectrum data sample and the three-dimensional fluorescence spectrum data.
8. A liquor vintage identification system based on spectroscopy, characterized in that, Including: A first module for obtaining the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor and determining several characteristic component score values of the to-be-detected Maotai-flavor liquor according to the three-dimensional fluorescence spectrum data of the to-be-detected Maotai-flavor liquor; A second module, configured to input the score values of the several characteristic components of the to-be-detected Maotai-flavor liquor into a preset prediction model to identify the vintage of the to-be-detected Maotai-flavor liquor; the prediction model is obtained by training based on sample data, and the sample data includes three-dimensional fluorescence spectrum data samples and vintage samples of several Maotai-flavor liquors.
9. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute the method according to any one of claims 1-7.
10. A liquor age identification system based on spectroscopy, characterized in that, It includes a fluorescence spectrometer and a computer device connected to the fluorescence spectrometer; wherein, The fluorescence spectrometer is used to collect the fluorescence spectrum of Maotai-flavor liquor. The computer device includes: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-7.