White spirit intelligent distillation process parameter optimization method
By introducing industrial IoT networks and machine learning models into the baijiu distillation process, the problems of distillation control precision and data management have been solved, enabling precise control and process optimization of the distillation process, thereby improving production efficiency and the quality of the raw liquor.
Patent Information
- Application Number
- CN202511473669.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-02-13
AI Technical Summary
The existing baijiu distillation process has shortcomings in terms of control precision, data management and optimization methods, resulting in low distillation efficiency, unstable quality of raw liquor, and a lack of scientific operating standards.
Industrial Internet of Things (IIoT) platforms are used to collect distillation process data, build machine learning models, and establish process optimization parameter manuals through data cleaning, correlation analysis, and dimensionality reduction techniques to achieve real-time monitoring and dynamic adjustment.
It improved the precision of distillation control, optimized process parameters, reduced operational differences, improved production efficiency and consistency of raw wine quality, and reduced carbon emissions.
Smart Images

Figure CN121528334A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of liquor brewing, and particularly relates to a method for optimizing process parameters of liquor intelligent distillation. BACKGROUND
[0002] Distillation is a core link in liquor brewing production. At present, the industry generally adopts a "mixed distillation and mixed burning" process to extract ethanol and flavor substances in fermented grains, while achieving the gelatinization of new grains and the control of acidity. The typical characteristics of this process are: a large number of influencing factors, strong coupling between variables, complex control process, and difficulty in quantifying process parameters. In batch distillation production, there is a phenomenon of steam coupling between workshops and between stills, which leads to frequent fluctuations in parameters such as steam pressure and flow. For example, in the prior art, the liquor receiving temperature and valve opening degree are manually controlled, which has poor accuracy. The steam pressure is measured by a mechanical pressure gauge, which has low measurement accuracy, and often appears the phenomenon of "same pressure, different flow", which seriously affects the distillation efficiency.
[0003] The traditional distillation process highly depends on the experience of operators, and the process parameters meeting the production requirements are obtained through repeated "trial and error". This mode has the following significant defects: 1) Insufficient control accuracy: When manually adjusting key parameters such as steam flow and liquor receiving temperature, precise control cannot be achieved. For example, the control deviation of the liquor receiving temperature is usually more than ±5℃, which leads to unstable distillation of flavor substances in the raw liquor.
[0004] 2) Difficulty in data quantification: The measurement accuracy of mechanical instruments is low, and there is a lack of real-time collection means for physicochemical indicators such as acidity, moisture, and starch of grains, making it difficult to establish a quantitative correlation between process parameters and raw liquor quality.
[0005] 3) Low efficiency of process optimization: The traditional optimization method based on simple statistical analysis cannot effectively handle the nonlinear coupling relationship between multiple parameters in the distillation process. For example, when controlling a single variable for trial and error experiments, a large amount of manpower and resources are required, and the prerequisite for parameter change is difficult to guarantee, resulting in weak reference of statistical data and the possibility of local optimal results rather than global optimal results.
[0006] 4) Low standardization: The operation mode relying on human experience cannot form a scientific operation procedure, and the operation differences between different workshops and different operators are significant, leading to uncontrollable grains entering the cellar, large carbon emissions in the production process, and prominent problems such as non-standard operation.
[0007] In addition, in the prior art, the thermal parameters (such as steam temperature, pressure) and the data of the state of the fermented grains (such as acidity, residual sugar) of the distillation process are collected in a lack of systematicness, and the breadth and granularity of the data records are insufficient, which cannot provide sufficient data support for process optimization. The existing control method does not realize real-time monitoring and dynamic adjustment of the distillation process, and it is difficult to cope with parameter fluctuations and disturbances in production. For example, when the steam flow suddenly changes, the traditional system cannot quickly adjust the water inflow of the circulating water, causing the wine receiving temperature to deviate greatly from the target value, affecting the consistency of the base liquor quality.
[0008] In summary, the existing white spirit distillation process has obvious deficiencies in control precision, data management, optimization method, etc., and it is urgent to introduce digital and intelligent technology to realize precise control and process optimization of the distillation process. SUMMARY
[0009] The purpose of the present application is to overcome the deficiencies of the prior art, and to provide an optimization method for white spirit intelligent distillation process parameters. The present application digs into the mechanisms of heat transfer and mass transfer in the white spirit distillation process, and integrates advanced digital technology. By collecting key distillation process parameters, the state of the fermented grains and the quality data of the base liquor, a brewing process database is constructed. Based on the analysis of the mechanism of heat flow and the change of related rules, the characteristic data causality and correlation direction are determined, and a machine learning model is constructed using high-dimensional and differentiated data. The model supports interpretable analysis, and the correlation between process parameters and product yield and quality is analyzed, which is used to establish a process manual to guide actual production.
[0010] To achieve the above purpose, the technical scheme adopted by the present application is as follows: An optimization method for white spirit intelligent distillation process parameters, comprising the following steps: S1, an industrial Internet of Things network platform covering the whole process of white spirit distillation is established, and variable characteristics related to process thermal parameters and physicochemical properties of fermented grains are obtained , and index characteristics related to the base liquor are recorded ; all variable characteristics are uniformly included in the variable input set , and all index characteristics are uniformly included in the index output set , and ; S2, based on the sampling time , the variable characteristics and the index characteristics are tabulated to obtain a white spirit distillation data sample list ; each row of the white spirit distillation data sample list represents a sample, and the th sample contains All variable characteristic values obtained in real time and index characteristic values , , indicates the total number of samples in the liquor distillation data sample list Each column of the liquor distillation data sample list corresponds to a characteristic field respectively.
[0011] S3, the liquor distillation data sample list Implement data cleaning processing, including identifying and deleting outliers, and using data filling method to process missing values, finally obtain the cleaned liquor distillation data sample list .
[0012] S4, based on the liquor distillation data sample list Carry out data analysis, and mine the correlation between each index characteristic and all variable characteristics , to filter out the associated variable characteristics which have significant influence on the corresponding index characteristics from the variable input set .
[0013] S5, for the target index to be predicted, based on the data analysis results, extract the associated variable characteristics from the variable input set to construct a candidate variable characteristic set , combined with data dimension reduction technology to eliminate irrelevant and redundant variable characteristics, obtain a prediction modeling variable characteristic set ; wherein, , , , .
[0014] S6, for each target index to be predicted, based on the prediction modeling variable characteristic set and the liquor distillation data sample list , construct a prediction model training sample; according to the preset proportion, the prediction model training sample is divided into training set, validation set and test set, and the support vector machine regression model is trained to machine learning prediction model.
[0015] S7, based on the trained machine learning prediction model, carry out process optimization, and construct a liquor distillation process optimization parameter manual for guiding production practice.
[0016] Preferably, in step S1, the variable characteristics include 6 process variables related to process operation and 14 non-process variables which cannot be controlled and intervened by human beings, that is Among them: process variables Includes steam flow Flowing time Duration of the closing toast Gelatinization time Circulating water inflow and the temperature of the wine Non-process variables Including steam temperature Steam pressure Boiling time Steaming time upper part temperature of the cooler Temperature at the bottom of the cooler Circulating water inlet temperature Circulating water outlet temperature Acidity of fermented mash , water content of fermented mash , fermented mash and residual sugar , fermented mash starch Tail wine addition amount Add concentration to tail sake .
[0017] Preferably, in step S1, the number of indicator features It covers the physicochemical indicators of raw spirits, the taste evaluation indicators of raw spirits, the physicochemical indicators of mash, and yield; among which, the data related to the physicochemical indicators of raw spirits includes alcohol content. Alcohol content Total lipids Total acid Ethyl hexanoate Ethyl lactate Ethyl acetate Ethyl butyrate Ethyl valerate n-Propanol Acetaldehyde acetal furfural sec-butanol Isobutanol n-Butanol and isoamyl alcohol Data related to the evaluation indicators of the original spirit's taste includes taste ratings. These are ordered category values graded according to three levels; data related to the physicochemical indicators of the fermented mash include the degree of gelatinization. ,acidity , moisture Residual sugar ,starch Data related to production volume includes raw wine production. .
[0018] Preferably, in step S3, the feature field is set to... This is a sample list of data for baijiu distillation. The Middle The list parameter of the column, let the feature value This is a sample list of data for baijiu distillation. The Middle Liede The list parameter values for rows, Identifying and removing outliers includes: for list parameters that are normally distributed. pass The principle is to perform anomaly checks on variable values, specifically for the list of baijiu distillation data samples. The first in List parameters of columns Calculate the mean and standard deviation ; Sample list of baijiu distillation data The Middle Liede List of row parameter values ,like Then determine the list parameter value These are outliers. For list parameters that are not normally distributed. Anomaly detection is performed using box plots, specifically by obtaining a list of samples of baijiu distillation data. The first in Column list parameters lower quartile and upper quartiles Take its interquartile range ; Sample list of baijiu distillation data The Middle Liede List of row parameter values ,like or ,but Decision list parameter values These are outliers. After deleting outliers, they are marked as missing values and used in the missing value imputation process.
[0019] Preferably, in step S3, the data imputation method for handling missing values includes the following steps: processing the list of baijiu distillation data samples... The missing value rate of each feature field in the list is calculated, and the missing value rate of the feature values is deleted. Feature fields, retaining the feature value missing rate in the list. By analyzing the feature fields, we can obtain a sample list of baijiu distillation data. From the list of data samples of baijiu distillation Extract complete samples corresponding to all feature fields to construct a complete parameter field sample set. And using a complete parameter field sample set A random forest prediction model is built for each feature field. (This is based on a sample list of baijiu distillation data.) For each sample with missing feature values, all its non-missing feature values are input into a random forest prediction model to predict and fill in the missing feature values in the corresponding sample, ultimately obtaining a complete list of baijiu distillation data samples. .
[0020] Preferably, in step S4, the method used is... The algorithm applies each indicator feature With all variable characteristics The correlation between them is measured, that is, for the list of baijiu distillation data samples. Each individual indicator feature exists in Iterate through the list of baijiu distillation data samples in sequence. All variable characteristics present in Each iteration includes the following steps: S41, List of Baijiu Distillation Data Samples At the same time Obtained variable eigenvalues and indicator characteristic values Establish variable indicator parameter pairs , obtain Sampling pool of group variable index parameter pairs .
[0021] S42, based on sampling pool Implementation Round-by-round self-service random sampling, generating Subsample set , Each subset Includes Individual variable indicator parameter pairs, .
[0022] S43, based on each individual subset Calculate variable characteristics With indicator characteristics correlation coefficient , obtain The set of coefficients of each correlation coefficient ;Right now , .
[0023] S45, based on coefficient set Establish correlation coefficient distribution curve Statistical coefficient set The probability that the correlation coefficient is less than 0 The expression is ; For a single correlation coefficient The probability density function.
[0024] S46, determine whether it satisfies the condition. , The threshold for a one-sided significance test is set; if so, the current variable characteristic is determined. In order to match the current indicator characteristics Identify the characteristics of significantly correlated variables and calculate the strength of their correlation. If not, then determine the characteristics of the current variable. In order to match the current indicator characteristics There are no significantly associated non-related variables.
[0025] Preferably, in step S43, the correlation coefficient for Correlation coefficient, mutual information coefficient, or accuracy metrics for predictions based on machine learning models .
[0026] Preferably, in step S46, when the current variable characteristics are determined... In order to match the current indicator characteristics When there are significantly correlated features in the associated variables, the current variable features Compared with current indicator characteristics The formula for calculating the correlation strength is as follows: ; Indicates the characteristics of the current variable Compared with current indicator characteristics The mathematical expectation of the correlation coefficient between them.
[0027] Preferably, in step S5, a feature set of predictive modeling variables is obtained. Includes the following steps: S51, based on the required prediction Target Indicators Construct a set of prediction targets S52, for the set of prediction targets Each individual target metric in Based on the data analysis results, input variables into the set. Extract features from the corresponding related variables to construct a feature set of related variables. ; Feature set of candidate variables For the prediction target set middle Target Indicators Corresponding associated variable feature set The union of the features of candidate variables. S53, based on data dimensionality reduction techniques that maximize mutual information, from the feature set of candidate variables. Selecting variable features to construct a predictive modeling variable feature set .
[0028] Preferably, in step S53, the feature set of prediction modeling variables is constructed. Includes the following steps: S531, Initialize the number of iterations Select the variable feature set and the feature set of remaining candidate variables Set the maximum number of iterations. .
[0029] S532, from the feature set of remaining candidate variables Filter out the set of prediction targets The variable features with the strongest overall correlation The expression is ; For the feature set of the remaining candidate variables Characteristics of candidate variables in; Mutual information is represented by probability distribution for discrete variables and kernel density estimation for continuous variables.
[0030] S533, the currently selected variable features Add to selected feature set and from the remaining candidate feature set delete.
[0031] S534, Calculate the characteristics of the currently selected variables. Information increment for predicting target indicators The expression is ; Indicates The selected variable feature set after the next iteration; Select the set of feature variables Any selected variable feature in the context.
[0032] S535, determine whether any of the following termination conditions are met: ① number of iterations ② The information increment of two consecutive iterations satisfies and If not, then let Then, return to step S532; if so, output the selected feature variable set. As a set of features for predictive modeling variables .
[0033] Preferably, in step S6, the method for constructing training samples for the prediction model is as follows: based on the prediction modeling feature set... In line with target indicators All associated variable characteristics From the list of data samples of baijiu distillation Extract all feature values from the corresponding values. Standardization or normalization methods are used to eliminate dimensional differences based on sampling time. target indicators eigenvalues Features of all associated variables eigenvalues Associated, obtain Training samples for the prediction model of a set of data samples.
[0034] Preferably, in step S6, a method based on... Random sampling or 5-fold cross-validation can be used to divide the dataset into training, validation, and test sets.
[0035] Preferably, in step S6, when training the machine learning prediction model, the construction of the support vector machine regression model is based on structural risk minimization. This is achieved by finding the optimal hyperplane in the high-dimensional feature space to minimize the deviation between the data points and the hyperplane, with the objective loss function being... ; Indicates the model weight parameters; Represents the weight parameters Norm square; Indicates paranoia; Represents the regularization parameter; Indicates tolerance level; This represents the total number of data samples in the training set. .
[0036] Preferably, in step S6, when training the machine learning prediction model, for data inputs and outputs with non-linear relationships, a kernel function is used to map the input to a high-dimensional space; the kernel function includes linear kernel functions, polynomial kernel functions, radial basis function kernel functions, and... At least one of the kernel functions.
[0037] Preferably, in step S6, when training the machine learning prediction model, the evaluation metrics of the machine learning prediction model include at least one of mean absolute error, mean square error, root mean square error, log-root mean square error, mean absolute percentage error, or coefficient of determination.
[0038] Preferably, step S7, constructing the manual for optimizing the baijiu distillation process parameters, includes the following steps: S71, based on The additive model interpretation method interprets and analyzes the results of the trained machine learning prediction model, calculates the contribution of each variable feature to the target indicator, and obtains the global importance of the feature and the individual conditional expectation curve.
[0039] S72, based on the individual condition expectation curve, generates a partial dependency curve to analyze the correlation trend between variable feature values and target indicators.
[0040] S73, based on the partial dependency curve and the baseline level of the target indicator, divide the values of the variable characteristics into suitable and unsuitable ranges.
[0041] S74 compiles the appropriate value ranges of each variable characteristic into a process optimization parameter manual to guide production practice.
[0042] The beneficial effects of this invention are: 1) This technical solution solves the problem of "insufficient control precision" in traditional baijiu distillation. Compared with the traditional method of manually adjusting parameters such as steam flow rate and distillation temperature, which often results in deviations exceeding the expected range, this technical solution accurately acquires 20 variable features through a full-process industrial IoT network, combines this with a support vector machine regression model to predict the influence of parameters, and corrects key indicators in real time, resulting in control deviations far lower than those of traditional processes.
[0043] 2) This technical solution overcomes the pain point of "difficulty in data quantification" in traditional processes. Compared to traditional methods that rely on low-precision instruments and lack multi-dimensional data, this technical solution records 24 raw liquor-related indicators (including physicochemical properties, taste, mash characteristics, and yield), using... The algorithm quantifies the correlation between variables and indicators, and processes data through standardized procedures to ensure integrity, providing a reliable foundation for analysis.
[0044] 3) This technical solution improves process optimization efficiency. Compared to the traditional "single-variable trial and error" method, which is labor-intensive and prone to local optima, this technical solution uses "mutual information maximization" to reduce dimensionality and select core features, supports vector machines to handle multi-parameter nonlinear coupling, and then... Analysis quickly locates the appropriate range for variables, significantly shortening the optimization cycle.
[0045] 4) This technical solution establishes a unified process standard. Compared with the traditional reliance on technician experience, which results in large differences in operation, this technical solution compiles the appropriate value range of variables into a process manual, replacing experience-based judgment, adapting to different stills and workshops, reducing errors, and also reducing carbon emissions and avoiding non-standard operations through parameter optimization.
[0046] 5) This technical solution achieves closed-loop control throughout the entire process. Compared with the fragmented and unresponsive nature of traditional data acquisition, this technical solution collects data in real time throughout the entire distillation process, captures the impact of parameter fluctuations based on the model, and dynamically corrects the associated process variables to ensure consistent receiving temperature and raw spirit quality, thus solving the problem of "different flow rates under the same pressure". Attached Figure Description
[0047] Figure 1 This is a basic implementation flowchart of the technical solution; Figure 2 A sample list based on baijiu distillation data A schematic diagram illustrating the principles of data analysis. Figure 3 A schematic diagram for training a machine learning prediction model; Figure 4 A schematic diagram illustrating the principle of using 5-fold cross-validation to divide the training samples for the prediction model; Figure 5 This is a schematic diagram illustrating the process optimization based on a trained machine learning prediction model. Detailed Implementation
[0048] To make the purpose, technical solution and advantages of the invention clearer, the technical solution of the invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the invention, but not all embodiments.
[0049] Therefore, the following detailed description of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0050] Example 1 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu (Chinese liquor). As a preferred embodiment of the present invention, such as... Figure 1 As shown, it includes the following steps: S1. Establish an industrial Internet of Things (IoT) network platform covering the entire process of baijiu distillation to acquire data related to process thermal parameters and the physicochemical properties of the mash. Each variable has characteristics, and corresponding records are kept. Individual indicators and characteristics related to the raw liquor; all variable characteristics are uniformly incorporated into the variable input set. Using variable characteristics express, ; Incorporate all indicator characteristics into the indicator output set Using indicator features express, .
[0051] S2, based on sampling time The variable characteristics and indicator characteristics are summarized in a table to obtain a sample list of baijiu distillation data. List of sample data for baijiu distillation Each row represents a sample, i.e., the first row. The sample contains All variable eigenvalues obtained at time 1 and indicator characteristic values , , List of sample data for baijiu distillation Total number of samples; Sample list of baijiu distillation data Each column corresponds to a feature field.
[0052] S3, a sample list of baijiu distillation data. The data cleaning process involves identifying and removing outliers, then using data imputation methods to handle missing values, ultimately resulting in a cleaned list of baijiu distillation data samples. .
[0053] S4, Sample list based on baijiu distillation data Conduct data analysis to uncover the characteristics of each indicator. With all variable characteristics The relationships between them, from the input set of variables Screening for corresponding indicator characteristics Characteristics of associated variables with significant impact.
[0054] S5, for those who need to predict Target Indicators Based on the data analysis results, input variables into the set Extract features of related variables to construct a set of candidate variable features By combining data dimensionality reduction techniques to remove irrelevant and redundant variable features, a set of predictive modeling variable features is obtained. ;in, , , .
[0055] S6, for each target indicator that needs to be predicted Based on the feature set of predictive modeling variables List of data samples for baijiu distillation Construct training samples for the prediction model; divide the training samples into training set, validation set and test set according to a preset ratio, and train the machine learning prediction model based on the support vector machine regression model.
[0056] S7, based on the trained machine learning prediction model, carries out process optimization and constructs a manual of optimization parameters for the baijiu distillation process to guide production practice.
[0057] Example 2 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu. As a preferred embodiment of the present invention, based on embodiment 1, in step S1, a total of 20 data fields related to process thermal parameters and physicochemical properties of mash are collected from the industrial Internet of Things network platform. Specifically, the variable features include 6 process variables related to process operation. And 14 non-process variables that cannot be controlled or intervened in. ,Right now Among them: process variables Includes steam flow Flowing time Duration of the closing toast Gelatinization time Circulating water inflow and the temperature of the wine Non-process variables Including steam temperature Steam pressure Boiling time Steaming time upper part temperature of the cooler Temperature at the bottom of the cooler Circulating water inlet temperature Circulating water outlet temperature Acidity of fermented mash , water content of fermented mash , fermented mash and residual sugar , fermented mash starch Tail wine addition amount Add concentration to tail sake These non-process variables come from upstream inputs or sensor detection information and cannot be controlled or intervened by humans.
[0058] As shown in Table 1 below, the physicochemical indicators of the raw spirit during the distillation process were obtained from the industrial IoT network platform. Evaluation indicators of raw wine taste Physicochemical indicators of fermented mash and output Related There are several data fields, including, except for the original wine taste evaluation index. Apart from the ordered categorical parameters for "Special Grade," "Superior Grade," and "Regular Grade," all other indicator parameters are continuous. Data related to the physicochemical properties of the raw spirit includes alcohol content. Alcohol content Total lipids Total acid Ethyl hexanoate Ethyl lactate Ethyl acetate Ethyl butyrate Ethyl valerate n-Propanol Acetaldehyde acetal furfural sec-butanol Isobutanol n-Butanol and isoamyl alcohol Data related to the evaluation indicators of the original spirit's taste includes taste ratings. These are ordered category values graded according to three levels; data related to the physicochemical indicators of the fermented mash include the degree of gelatinization. ,acidity , moisture Residual sugar ,starch Data related to production volume includes raw wine production. .
[0059]
[0060] Based on this, in step S2, the list of baijiu distillation data samples... As shown in Table 2 below:
[0061] Example 3 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu. As a preferred implementation of the present invention, based on embodiment 1 or 2, the data cleaning in step S3 aims to improve data quality by processing anomalies (errors, invalidities) and missing data in the original data, ensuring the accuracy, integrity and consistency of the data, thereby laying the foundation for subsequent accurate and reliable data analysis and modeling.
[0062] First, identify the original list of baijiu distillation data samples. This involves identifying potential outliers in the table. On one hand, obvious errors can be corrected or deleted based on established rules and experience. On the other hand, potential outliers can be identified and deleted by combining the descriptive statistical characteristics of each feature field in the table. Specifically, the feature fields... This is a sample list of data for baijiu distillation. The Middle The list parameter of the column, let the feature value This is a sample list of data for baijiu distillation. The Middle Liede The list parameter values for rows, The identification and removal of outliers includes: For list parameters that are normally distributed pass The principle is to perform anomaly checks on variable values, specifically for the list of baijiu distillation data samples. The first in List parameters of columns Calculate the mean and standard deviation ; Sample list of baijiu distillation data The Middle Liede List of row parameter values ,like Then determine the list parameter value This is an outlier.
[0063] For list parameters that are not normally distributed Anomaly detection is performed using box plots, specifically by obtaining a list of samples of baijiu distillation data. The first in Column list parameters lower quartile and upper quartiles Take its interquartile range ; Sample list of baijiu distillation data The Middle Liede List of row parameter values ,like or ,but Decision list parameter values This is an outlier.
[0064] After the above checks, a list of baijiu distillation data samples will be compiled. All data identified as outliers are deleted and marked as missing values, then used in the missing value imputation process.
[0065] Example 4 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu. As a preferred embodiment of the present invention, based on any one of embodiments 1 to 3, in step S3, the missing values are processed using a data imputation method on the baijiu distillation data sample list. The missing value rate of each feature field in the list is calculated, and the missing value rate of the feature values is deleted. Feature fields, retaining the feature value missing rate in the list. By analyzing the feature fields, we can obtain a sample list of baijiu distillation data. List of data samples for baijiu distillation Missing values are filled in.
[0066] Traditional missing value imputation methods are mostly based on the sample information of the field to be imputed, including: replacement with the nearest neighbor sample values; replacement with the mean, median, and mode; linear interpolation; and nearest neighbor sample interpolation.
[0067] In this technical solution, considering the potential coupling between different feature fields, simply considering information from a single field for imputation may lead to the loss of correlation information between feature fields in the imputed data, resulting in reduced accuracy in analysis and modeling. Therefore, this technical solution considers using a random forest-based approach... The method for imputing missing values in multiple missing data is as follows: From the list of data samples of baijiu distillation Extract complete samples corresponding to all feature fields to construct a complete parameter field sample set. And using a complete parameter field sample set A random forest prediction model is built for each of the feature fields.
[0068] Therefore, if a sample has missing values in a certain field, random forests combined with data from other fields can be used to predict and impute the missing values. This applies to a sample list of baijiu distillation data. For each sample with missing feature values, all its non-missing feature values are input into a random forest prediction model to predict and fill in the missing feature values in the corresponding sample, ultimately obtaining a complete list of baijiu distillation data samples. .
[0069] Example 5 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu. As a preferred embodiment of the present invention, based on any one of embodiments 1 to 4, step S4 involves using... ( The correlation coefficient algorithm based on self-service random sampling is used for each indicator feature. With all variable characteristics The correlation between them is measured in the sample list of baijiu distillation data. Each individual indicator feature exists in Iterate through the list of baijiu distillation data samples in sequence. All variable characteristics present in ,like Figure 2 As shown, each traversal includes the following steps: S41, List of Baijiu Distillation Data Samples At the same time Obtained variable eigenvalues and indicator characteristic values Establish variable indicator parameter pairs , obtain Sampling pool of group variable index parameter pairs .
[0070] S42, based on sampling pool Implementation Round-by-round self-service random sampling, generating Subsample set , Each subset Includes Individual variable indicator parameter pairs, .
[0071] S43, based on each individual subset Calculate variable characteristics With indicator characteristics correlation coefficient , obtain The set of coefficients of each correlation coefficient ;Right now , .
[0072] S45, based on coefficient set Establish correlation coefficient distribution curve Statistical coefficient set The probability that the correlation coefficient is less than 0 The expression is ; For a single correlation coefficient The probability density function.
[0073] S46, determine whether it satisfies the condition. , The threshold for a one-sided significance test is set; if so, the current variable characteristic is determined. In order to match the current indicator characteristics Identify the characteristics of significantly correlated variables and calculate the strength of their correlation. If not, then determine the characteristics of the current variable. In order to match the current indicator characteristics There are no significantly associated non-related variables. The quality of field data was low in the early stages of the project; therefore, the threshold for one-sided significance testing can be lowered. Set to a relatively large value; as data quality continues to improve, the threshold for one-sided significance testing will be gradually reduced. To obtain more reliable results.
[0074] Example 6 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu. As a preferred embodiment of the present invention, based on embodiment 5, in step S43, the correlation coefficient... for Correlation coefficient, mutual information coefficient, or accuracy metrics for predictions based on machine learning models .
[0075] Example 7 This embodiment discloses a method for optimizing the process parameters of intelligent distillation of baijiu. As a preferred embodiment of the present invention, based on embodiment 5 or 6, in step S46, when the current variable characteristics are determined... In order to match the current indicator characteristics When there are significantly correlated features in the associated variables, the current variable features Compared with current indicator characteristics The formula for calculating the correlation strength is as follows: ; Indicates the characteristics of the current variable Compared with current indicator characteristics The mathematical expectation of the correlation coefficient between them.
[0076] Example 8 This embodiment discloses a method for optimizing the intelligent distillation process parameters of baijiu. As a preferred embodiment of the present invention, based on any one of embodiments 5 to 7, according to the data analysis results of step S4, it can be seen that for each index characteristic in each distillation stage... This allows us to find a set of key variable features composed of multiple associated variable features, if the subsequent target indicator... quantity This involves The overall data sample formed by the union of the feature sets of key variables, if directly used for model construction, will greatly increase the dimensionality of the variables substituted into the prediction, making the model more prone to overfitting and reduced generalization ability with a limited sample size, i.e., the Hughes phenomenon. ).
[0077] Therefore, based on a limited data sample, data dimensionality reduction techniques can be used to further filter out variables closely related to the prediction target, eliminating interference from other irrelevant and redundant variables without losing the model's predictive information and accuracy. Common data dimensionality reduction methods can be divided into two categories according to whether feature data is transformed: feature selection and feature extraction. Feature selection only filters fields without changing values; while independent principal component analysis (ICP-A)... Feature extraction methods such as [missing information] will transform the data, which may lead to irreversible loss of data prediction information.
[0078] This technical solution adopts a conditional mutual information maximization approach based on causal features. The data dimensionality reduction technique ensures sufficient data for accurate prediction while preserving the interpretability between the predictor variables and the target. Specifically, in step S5, the feature set of the predictor modeling variables is obtained. Includes the following steps: S51, based on the need for prediction Target Indicators Construct a set of prediction targets .
[0079] S52, for the set of prediction targets Each individual target metric in Based on the data analysis results, input variables into the set. Extract features from the corresponding related variables to construct a feature set of related variables. ; Feature set of candidate variables For the prediction target set middle Target Indicators Corresponding associated variable feature set The union of .
[0080] S53, a data dimensionality reduction technique based on maximizing mutual information, from the feature set of candidate variables... Selecting variable features to construct a predictive modeling variable feature set Specifically, it includes the following steps: S531, Initialize the number of iterations Select the variable feature set and the feature set of remaining candidate variables Set the maximum number of iterations. .
[0081] S532, from the feature set of remaining candidate variables Filter out the set of prediction targets The variable features with the strongest overall correlation The expression is ; For the feature set of the remaining candidate variables Characteristics of candidate variables in; Mutual information is represented by probability distribution for discrete variables and kernel density estimation for continuous variables.
[0082] S533, the currently selected variable features Add to selected feature set and from the remaining candidate feature set delete.
[0083] S534, Calculate the characteristics of the currently selected variables. Information increment for predicting target indicators The expression is ; Indicates The selected variable feature set after the next iteration; Select the set of feature variables Any selected variable feature in the context.
[0084] S535, determine whether any of the following termination conditions are met: ① number of iterations ② The information increment of two consecutive iterations satisfies and If not, then let Then, return to step S532; if so, output the selected feature variable set. As a set of features for predictive modeling variables .
[0085] Example 9 This embodiment discloses a method for optimizing the parameters of a smart distillation process for baijiu (Chinese liquor). As a preferred embodiment of the present invention, based on embodiment 8, the method for constructing training samples for the prediction model in step S6 is as follows: based on the prediction modeling feature set... In line with target indicators All associated variable characteristics From the list of data samples of baijiu distillation Extract all feature values from the corresponding values. Standardization or normalization methods are used to eliminate dimensional differences based on sampling time. target indicators eigenvalues Features of all associated variables eigenvalues Associated, obtain Training samples for the prediction model of a set of data samples.
[0086] Example 10 This embodiment discloses a method for optimizing the parameters of a smart distillation process for baijiu (Chinese liquor). As a preferred embodiment of the present invention, based on any one of embodiments 1-9, in step S6, the relationship between the model parameters of the feature data during the training of the machine learning prediction model is as follows: Figure 3 As shown: Model hyperparameters The model's structure and other information are determined in advance and do not require learning; the model's internal parameters... (such as weight) Bias and threshold (etc.) then predicts the output by continuously comparing it. and actual output The model is obtained through reverse optimization, and this process is called model training.
[0087] When the amount of data is sufficient, this technical solution will predict the data in the model training samples. The data samples serve as a validation set, used to adjust hyperparameters related to model structure, etc. Using a 5-fold cross-validation method The training and test sets are randomly divided according to a certain ratio, such as... Figure 4 As shown. Therefore, the ratio of the validation set, training set, and test set in the prediction model training samples is: In actual project implementation, if the data volume is relatively small, then a method based on... The 5-fold cross-validation method is replaced by random sampling. An equal number of samples with replacement are randomly drawn from the data samples other than the validation set as the training set, and the other samples are used as the test set.
[0088] By repeatedly performing random sampling, training, and testing, the mean absolute error of the model's predictions is obtained. ), mean square error ( ), root mean square error ( ), log-root mean square error ( Mean absolute percentage error ( Coefficient of determination The accuracy and error index distributions are used to comprehensively and meticulously evaluate the predictive performance of the model for different target values.
[0089] Example 11 This embodiment discloses a method for optimizing the parameters of a smart distillation process for baijiu (Chinese liquor). As a preferred embodiment of the present invention, based on any one of embodiments 1-10, in step S6, when training the machine learning prediction model, the construction of the support vector machine regression model is based on minimizing structural risk. This is achieved by finding the optimal hyperplane in the high-dimensional feature space to minimize the deviation between the data points and the hyperplane. When using... When performing regression learning, the first step is to define a threshold. To determine the qualified (in training data) For sample points, the prediction error for these sample points is not included in the loss function, while the difference between the predicted value and the true value exceeds a threshold. When the time is right, it is included in the loss function.
[0090] SVM attempts to find a decision boundary such that as many sample points as possible fall on either side of the boundary threshold. Within the determined range, i.e., the target loss function: ; in, Indicates the model weight parameters; Represents the weight parameters Norm squared is used to control model complexity; Indicates paranoia; This represents the regularization parameter, used to control whether the predicted value of a sample exceeds a threshold. The scope and severity of punishment; This represents the tolerance level, which is a penalty term for sample prediction errors exceeding a threshold. This represents the total number of data samples in the training set. .
[0091] In the process of training a machine learning prediction model, when the input and output have a simple linear relationship, a linear approach can be used. Regression directly finds the optimal hyperplane in the corresponding dimension; however, when the data is complex and exhibits a certain degree of nonlinearity, a kernel function can be used. ) will input Mapping to a high-dimensional space indirectly helps find the corresponding optimal linear hyperplane, which corresponds to nonlinearity. Common methods for calculating the kernel function between two samples include: linear kernel function (…). ), polynomial kernel function ( ), radial basis kernel function ( ), Kernel function ( )wait.
[0092] Example 13 This embodiment discloses a method for optimizing the parameters of a smart distillation process for baijiu (Chinese liquor). As a preferred embodiment of the present invention, following any of the embodiments 1-12, during the use of the model, it is necessary to study in depth the influence of different input values on the output. In step S7 of this technical solution, constructing a baijiu distillation process optimization parameter manual includes the following steps: S71, adopting a method based on... Additive model interpretation method ( S71) Interpret and analyze the results of the trained machine learning prediction model, calculate the contribution of each variable feature to the target indicator, and obtain the global importance and individual conditional expectation curves of the features. S72) Generate partial dependency curves based on the individual conditional expectation curves to analyze the correlation trend between variable feature values and the target indicator. S73) Based on the partial dependency curves and the baseline level of the target indicator, divide the variable feature values into suitable and unsuitable ranges. S74) Compile the suitable value ranges of each variable feature into a process optimization parameter manual to guide production practice.
[0093] This embodiment uses linearity. Taking the internally used linear model as an example The calculation process will be explained as follows: First, assuming it passes In the sample The following linear regression prediction model is obtained from these features: ; Indicates the weighting coefficient. Let represent the eigenvalue. Then the ... The global impact of each feature on the prediction result ( That is, the baseline level of the score explained by this feature. for: ; Indicates weight With eigenvalues The mathematical expectation, Indicates the first The first sample Features of each feature ; Furthermore, in the The first sample The local contribution score of each feature to the prediction ( ) It equals the predicted result minus the baseline level, i.e.: Therefore, on this sample, the contribution of all features is... ,Right now: .
[0094] The method will incorporate the above process Extending this calculation to models other than linear regression involves traversing all feature combinations and calculating their corresponding probabilities, resulting in high computational complexity when the sample size is large. Therefore, alternative methods not based on a specific model can be used. and based on the underlying characteristics of the model (For neural networks) (For decision tree models) Implements approximate calculation of feature contribution analysis.
[0095] Finally, by synthesizing the model interpretation results across all samples, we obtain the global importance and individual conditional expectation of each feature. ) and partial dependency curves ( This allows us to obtain the impact of different values of each input feature on the change of the output target. In the nth sample Features The curve is calculated by preserving other features of the sample. The value of remains unchanged, and the value of is randomly changed. The values of each feature are obtained. Model conditional predictions With eigenvalue The change, i.e., the curve Based on this, and taking into account all bar sample The curve can then be used to obtain the first Each feature predicts the target value. Impact Curve, that is: .
[0096] Curves are used to help develop more reliable and optimized process manufacturing manuals, such as in... Figure 5 In the example shown, based on the prediction objective regarding the first... Partial dependency curve of each feature and the overall baseline level of the prediction results The first one related to the operation process The values of each feature are divided into suitable and unsuitable ranges according to corresponding thresholds, and these ranges are compiled into the process optimization parameter manual. Thus, in actual production, to achieve higher target values, the process optimization parameter manual will recommend adjusting the values of the first feature... Each feature value is limited to a corresponding threshold range. This process optimization method is easy for frontline personnel to understand and use.
Claims
1. A method for optimizing the parameters of a smart distillation process for baijiu (Chinese liquor), characterized in that, Includes the following steps: S1. Establish an industrial Internet of Things (IoT) network platform covering the entire process of baijiu distillation to acquire data related to process thermal parameters and the physicochemical properties of the mash. Each variable has characteristics, and corresponding records are kept. Individual indicators and characteristics related to the raw liquor; all variable characteristics are uniformly incorporated into the variable input set. Using variable characteristics express, ; Incorporate all indicator characteristics into the indicator output set Using indicator features express, ; S2, based on sampling time By summarizing the variable and indicator characteristics in a table, a sample list of baijiu distillation data is obtained. List of sample data for baijiu distillation Each row represents a sample, i.e., the first row. The sample contains All variable eigenvalues obtained at time 1 and indicator characteristic values , , List of sample data for baijiu distillation Total number of samples; Sample list of baijiu distillation data Each column corresponds to a feature field; S3, a sample list of baijiu distillation data. The data cleaning process involves identifying and removing outliers, then using data imputation methods to handle missing values, ultimately obtaining a cleaned list of baijiu distillation data samples. ; S4, Sample list based on baijiu distillation data Conduct data analysis to uncover the characteristics of each indicator. With all variable characteristics The relationships between them, from the input set of variables Screening for corresponding indicator characteristics Characteristics of associated variables with significant impact; S5, for those who need to predict Target Indicators Based on the data analysis results, input variables into the set Extract features of related variables to construct a set of candidate variable features By combining data dimensionality reduction techniques to remove irrelevant and redundant variable features, a set of predictive modeling variable features is obtained. ;in, , , ; S6, for each target indicator that needs to be predicted Based on the feature set of predictive modeling variables List of data samples for baijiu distillation Construct training samples for the prediction model; The training samples of the prediction model are divided into training set, validation set and test set according to a preset ratio, and the machine learning prediction model is trained based on the support vector machine regression model. S7, based on the trained machine learning prediction model, carries out process optimization and constructs a manual of optimization parameters for the baijiu distillation process to guide production practice.
2. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 2, characterized in that, In step S1, the variable features include six process variables related to the process operation. And 14 non-process variables that cannot be controlled or intervened in. ,Right now Among them: process variables Includes steam flow Flowing time Duration of the closing toast Gelatinization time Circulating water inflow and the temperature of the wine Non-process variables Including steam temperature Steam pressure Boiling time Steaming time upper part temperature of the cooler Temperature at the bottom of the cooler Circulating water inlet temperature Circulating water outlet temperature Acidity of fermented mash , water content of fermented mash , fermented mash and residual sugar , fermented mash starch Tail wine addition amount Add concentration to tail sake .
3. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S1, the number of indicator features It covers the physicochemical indicators of raw spirits, the taste evaluation indicators of raw spirits, the physicochemical indicators of mash, and yield; among which, the data related to the physicochemical indicators of raw spirits includes alcohol content. Alcohol content Total lipids Total acid Ethyl hexanoate Ethyl lactate Ethyl acetate Ethyl butyrate Ethyl valerate n-Propanol Acetaldehyde acetal furfural sec-butanol Isobutanol n-Butanol and isoamyl alcohol Data related to the evaluation indicators of the original spirit's taste includes taste ratings. These are ordered category values graded according to three levels; data related to the physicochemical indicators of the fermented mash include the degree of gelatinization. ,acidity , moisture Residual sugar ,starch Data related to production volume includes raw wine production. .
4. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S3, let the feature field This is a sample list of data for baijiu distillation. The Middle The list parameter of the column, let the feature value This is a sample list of data for baijiu distillation. The Middle Liede The list parameter values for rows, ; The identification and removal of outliers includes: For list parameters that are normally distributed pass The principle is to perform anomaly checks on variable values, specifically for the list of baijiu distillation data samples. The first in List parameters of columns Calculate the mean and standard deviation ; Sample list of baijiu distillation data The Middle Liede List of row parameter values ,like Then determine the list parameter value This is an outlier; For list parameters that are not normally distributed Anomaly detection is performed using box plots, specifically by obtaining a list of samples of baijiu distillation data. The first in Column list parameters lower quartiles and upper quartiles Take its interquartile range ; Sample list of baijiu distillation data The Middle Liede List of row parameter values ,like or ,but Decision list parameter values This is an outlier; After deleting outliers, mark them as missing values and substitute them into the missing value filling process.
5. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S3, the data imputation method for handling missing values includes the following steps: List of sample data for baijiu distillation The missing value rate of each feature field in the list is calculated, and the missing value rate of the feature values is deleted. Feature fields, retain the feature value missing rate in the list. By analyzing the feature fields, we can obtain a list of samples of baijiu distillation data. ; From the list of data samples of baijiu distillation Extract complete samples corresponding to all feature fields to construct a complete parameter field sample set. And using a complete parameter field sample set For each feature field, a random forest prediction model is built; List of data samples for baijiu distillation For each sample with missing feature values, all its non-missing feature values are input into a random forest prediction model to predict and fill in the missing feature values in the corresponding sample, ultimately obtaining a complete list of baijiu distillation data samples. .
6. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S4, the following method is used: The algorithm applies each indicator feature With all variable characteristics The correlation between them is measured, that is, for the list of baijiu distillation data samples. Each individual indicator feature exists in Iterate through the list of baijiu distillation data samples in sequence. All variable characteristics present in Each iteration includes the following steps: S41, List of Baijiu Distillation Data Samples At the same time Obtained variable eigenvalues and indicator characteristic values Establish variable indicator parameter pairs , obtain Sampling pool of group variable index parameter pairs ; S42, based on sampling pool Implementation Round-by-round self-service random sampling, generating Subsample set , Each subset Includes Individual variable indicator parameter pairs, ; S43, based on each individual subset Calculate variable characteristics With indicator characteristics correlation coefficient , obtain The set of coefficients of each correlation coefficient ;Right now , ; S45, based on coefficient set Establish correlation coefficient distribution curve Statistical coefficient set The probability that the correlation coefficient is less than 0 The expression is ; For a single correlation coefficient The probability density function; S46, determine whether it satisfies the condition. , The threshold for a one-sided significance test is set; if so, the current variable characteristic is determined. In order to match the current indicator characteristics Identify the characteristics of significantly correlated variables and calculate the strength of their correlation. If not, then determine the characteristics of the current variable. In order to match the current indicator characteristics There are no significantly associated non-related variables.
7. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 6, characterized in that, In step S43, the correlation coefficient for Correlation coefficient, mutual information coefficient, or accuracy metrics for predictions based on machine learning models .
8. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 6, characterized in that, In step S46, when the current variable characteristics are determined... In order to match the current indicator characteristics When there are significantly correlated features of related variables, the current variable features Compared with current indicator characteristics The formula for calculating the correlation strength is as follows: ; Indicates the characteristics of the current variable Compared with current indicator characteristics The mathematical expectation of the correlation coefficient between them.
9. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 6, characterized in that, In step S5, the feature set of prediction modeling variables is obtained. Includes the following steps: S51, based on the need for prediction Target Indicators Construct a set of prediction targets ; S52, for the set of prediction targets Each individual target metric Based on the data analysis results, input variables into the set. Extract features from the corresponding related variables to construct a feature set of related variables. ; Feature set of candidate variables For the prediction target set middle Target Indicators Corresponding associated variable feature set The union of; S53, a data dimensionality reduction technique based on maximizing mutual information, from the feature set of candidate variables... Selecting variable features to construct a predictive modeling variable feature set .
10. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 9, characterized in that, In step S53, the feature set of prediction modeling variables is constructed. Includes the following steps: S531, Initialize the number of iterations Select the variable feature set and the feature set of remaining candidate variables Set the maximum number of iterations. ; S532, from the feature set of remaining candidate variables Filter out the set of prediction targets The variable features with the strongest overall correlation The expression is ; For the feature set of the remaining candidate variables Characteristics of candidate variables in; Mutual information is represented by probability distribution for discrete variables and kernel density estimation for continuous variables. S533, the currently selected variable features Add to selected feature set and from the remaining candidate feature set delete; S534, Calculate the characteristics of the currently selected variables. Information increment for predicting target indicators The expression is ; Indicates The selected variable feature set after the next iteration; Select the set of feature variables Any selected variable feature in the; S535, determine whether any of the following termination conditions are met: ① number of iterations ② The information increment of two consecutive iterations satisfies and If not, then let Then, return to step S532; if so, output the selected feature variable set. As a set of features for predictive modeling variables .
11. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 9, characterized in that, In step S6, the method for constructing training samples for the prediction model is as follows: based on the prediction modeling feature set... In line with target indicators All associated variable characteristics From the list of data samples of baijiu distillation Extract all feature values from the corresponding values. Standardization or normalization methods are used to eliminate dimensional differences based on sampling time. target indicators eigenvalues Features of all associated variables eigenvalues Associated, obtain with Training samples for the prediction model of a set of data samples.
12. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S6, a method based on... Random sampling or 5-fold cross-validation can be used to divide the dataset into training, validation, and test sets.
13. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S6, when training the machine learning prediction model, the support vector machine regression model is constructed based on structural risk minimization. This is achieved by finding the optimal hyperplane in the high-dimensional feature space to minimize the deviation between the data points and the hyperplane, with the objective loss function being... ; Indicates the model weight parameters; Represents the weight parameters Norm square; Indicates paranoia; Represents the regularization parameter; Indicates tolerance level; This represents the total number of data samples in the training set. .
14. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S6, when training the machine learning prediction model, for data inputs and outputs with non-linear relationships, a kernel function is used to map the input to a high-dimensional space; the kernel function includes linear kernel functions, polynomial kernel functions, radial basis function kernel functions, and... At least one of the kernel functions.
15. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 11, characterized in that, In step S6, when training the machine learning prediction model, the evaluation metrics of the machine learning prediction model include at least one of mean absolute error, mean square error, root mean square error, log-root mean square error, mean absolute percentage error, or coefficient of determination.
16. The method for optimizing the intelligent distillation process parameters of baijiu as described in claim 1, characterized in that, In step S7, constructing the manual for optimizing the baijiu distillation process includes the following steps: S71, based on The additive model interpretation method interprets and analyzes the results of the trained machine learning prediction model, calculates the contribution of each variable feature to the target indicator, and obtains the global importance of the feature and the individual conditional expectation curve. S72, Generate a partial dependency curve based on the individual condition expectation curve, and analyze the correlation trend between variable feature values and target indicators; S73, based on the partial dependency curve and the baseline level of the target indicator, divide the values of the variable characteristics into suitable and unsuitable ranges; S74 compiles the appropriate value ranges of each variable characteristic into a process optimization parameter manual to guide production practice.