White spirit stack wine yield prediction method, model construction method and system

By using an improved partial least squares regression algorithm, combined with the characteristic data of Maotai-flavor liquor piles and information on the year and workshop, a liquor production prediction model was constructed. This model solved the problem of accurately predicting the liquor production of Maotai-flavor liquor in the early stage of pile fermentation, enabling early intervention and fine-grained management, and improving the timeliness and accuracy of the prediction.

CN122024880APending Publication Date: 2026-05-12KWEICHOW MOUTAI COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KWEICHOW MOUTAI COMPANY
Filing Date
2026-01-07
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies lack refined management of the production volume of Maotai-flavor liquor, especially in the early stage of fermentation, where accurate prediction is difficult to achieve. Furthermore, they fail to effectively consider the differences between years and workshops, resulting in coarse prediction granularity, insufficient timeliness, and unstable models.

Method used

An improved partial least squares regression algorithm is adopted. By collecting and preprocessing the characteristic data of the pile, including process and physicochemical indicators, and combining the encoding and processing of year and workshop information, a partial least squares regression model is constructed to predict the wine production and achieve precise management of the pile-level unit.

Benefits of technology

It enables accurate prediction of alcohol production in the early stage of stack fermentation, supports early intervention, provides finer prediction granularity and higher model capabilities, is suitable for online updates and simultaneous prediction of multiple stacks, and improves the timeliness and accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024880A_ABST
    Figure CN122024880A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of Baijiu production prediction, and relates to a Baijiu stack wine production prediction method, a model construction method and a system. The invention discloses a model construction method for predicting the liquor yield of liquor stacks. The model construction method comprises the following steps: collecting characteristic data of a plurality of stacks and corresponding liquor yields; preprocessing the characteristic data of the stacks of the same production round to form an input matrix X and a wine yield vector Y; performing coding processing or continuous characterization processing on the year and workshop information to which the heap belongs to obtain a characteristic variable capable of representing the systematic deviation of the year and the workshop relative to the overall average level, and adding the characteristic variable into the input matrix X; and establishing a regression prediction model from the input matrix X to the wine yield vector Y based on a partial least square regression model according to the input matrix X and the wine yield vector Y. According to the method, stack-level yield prediction is realized in the earlier stage of stack fermentation, and a coding variable capable of representing systematic deviation of the input matrix X relative to the overall average level is added into the input matrix X to realize early warning and accurate prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of liquor production prediction technology, and relates to a method, model construction method and system for predicting the liquor production of a liquor pile. Background Technology

[0002] Baijiu is a traditional Chinese distilled spirit, made from grains such as sorghum and wheat through a complex process involving solid-state fermentation, distillation, aging, and blending. It boasts a long history and profound cultural heritage. Based on aroma type, baijiu is mainly divided into sauce-aroma, strong-aroma, and light-aroma types, among which sauce-aroma baijiu is highly regarded for its unique brewing process and flavor characteristics.

[0003] Maotai, a type of soy sauce-flavored baijiu, is represented by Guizhou Maotai. It employs the "12987" process (i.e., a one-year production cycle, two rounds of feeding, nine rounds of steaming, eight rounds of fermentation, and seven rounds of distillation). It relies on natural microbial communities for open solid-state fermentation and is highly sensitive to climate, water quality, fermentation pit environment, and operational experience. Its yield is not only affected by the raw material ratio but also closely related to dynamic factors such as temperature, humidity, and microbial succession during the fermentation process.

[0004] Currently, most liquor companies rely on manual experience or simple statistical models to estimate liquor production, and existing technologies lack the practical need for refined management of the complex brewing process. Summary of the Invention

[0005] The purpose of this invention is to provide a method, model construction method and system for predicting the production volume of baijiu (Chinese liquor) piles.

[0006] The first aspect of this application provides a model construction method for predicting the production volume of baijiu (Chinese liquor) from a distillery, including: Collect feature data from multiple piles and their corresponding wine production; The characteristic data of the piles in the same production cycle are preprocessed to form an input matrix X and a wine production vector Y; the year and workshop information of the pile are encoded or continuously characterized to obtain characteristic variables that can characterize the systematic deviation of the year and workshop relative to the overall average level, and the characteristic variables are added to the input matrix X; Based on the input matrix X and the wine production vector Y, a regression prediction model from the input matrix X to the wine production vector Y is established using a partial least squares regression model.

[0007] In some embodiments of this application, the feature data includes process index data and physicochemical index data; In some embodiments of this application, the physicochemical indicators include temperature, moisture, acidity, sugar content, ethanol, starch, and acetic acid data measured at the pile surface, core, and bottom of the pile; In some embodiments of this application, the process index data includes stacking time, spreading and drying time, mixing temperature, temperature at which the starter culture is placed on the stack, and room temperature data of the drying hall when the starter culture is placed on the stack.

[0008] In some embodiments of this application, the encoding process is dummy variable processing, including: Select a base year and a base workshop; For each year, a 0-1 variable is set for each year other than the base year. The value is 1 when the year to which the heap belongs is that year, and 0 otherwise. For each workshop, a 0-1 variable is set for each workshop other than the reference workshop. The value is 1 when the pile belongs to the workshop, and 0 otherwise. The base year and base workshop are implicitly represented by corresponding dummy variables all being 0.

[0009] In some embodiments of this application, the continuous characterization process includes: constructing a continuous environmental index that reflects the overall fermentation environment of a year or workshop based on the statistical characteristics or principal component analysis scores of historical processes and physicochemical indicators, and using the continuous environmental index as a characteristic variable characterizing the systematic deviation of that year or workshop.

[0010] In some embodiments of this application, the encoding process includes effect encoding: setting an offset variable for a year or workshop with a number equal to the number of its categories minus one, and learning the offset strength of each year or workshop relative to the overall mean during model training.

[0011] In some embodiments of this application, the preprocessing includes cleaning; the cleaning includes: identifying abnormal data in each group where each indicator deviates from three times the standard deviation of the mean or is missing; if a single pile has two or more abnormal indicators, the pile is removed; if only a single indicator is abnormal, the indicator value corresponding to the pile with the smallest overall distance from the abnormal pile among all piles with normal indicators in the same group is used for replacement.

[0012] In some embodiments of this application, the preprocessing includes standardization, which includes: standardizing the process indicators and physicochemical indicators by subtracting the mean and then dividing by the standard deviation; centering the dummy variable indicators generated by converting year and workshop information by subtracting the mean; and standardizing the wine production vector by subtracting the mean and then dividing by the standard deviation.

[0013] The second aspect of this application provides a method for predicting the output of a liquor production pile; The steps include: obtaining the characteristic data of the pile to be predicted, including process indicators, physicochemical indicators, and information on the year and workshop to which it belongs; The year and workshop information of the pile are encoded or continuously characterized to obtain feature variables that can characterize the systematic deviation of the year and workshop relative to the overall average level, and an input vector is formed based on the feature data and the feature variables. The input vector is input to a pre-built wine production prediction model for the same production round, and the wine production prediction model outputs the predicted wine production value for that batch. The wine production prediction model is obtained by the model construction method as described in any one of the first aspects.

[0014] The third aspect of this application provides a system for constructing a predictive model for the production volume of a liquor pile. include: The data acquisition module is used to collect characteristic data of multiple piles in the same production cycle and their corresponding wine production. The data preprocessing module is used to preprocess the feature data, converting the year and workshop information of the pile into feature variables that can characterize its systematic deviation relative to the overall average level, and forming an input matrix X and a wine production vector Y. The model building module is used to construct a regression prediction model from X to Y based on the input matrix X and the wine production vector Y, using the partial least squares regression method, as the wine production prediction model for this production cycle.

[0015] The fourth aspect of this application provides a system for predicting the production volume of a liquor pile. The feature acquisition module is used to acquire feature data of the pile to be predicted. The feature data includes process indicators, physicochemical indicators, and information on the year and workshop to which it belongs. The encoding conversion module is used to convert the year and workshop information into feature variables that can characterize their systematic deviation relative to the overall average level, and generate an input vector based on the feature data and the feature variables; The prediction execution module is used to input the input vector into a pre-built wine production prediction model for the same production round, and output the predicted wine production value of the pile. The wine production prediction model is a partial least squares regression model obtained by the model construction method described in any one of the first aspects.

[0016] The method, model construction method, and system for predicting the production volume of a liquor production pile disclosed in this application have the following characteristics: 1. Earlier prediction timing: Supports prediction in the early stages of accumulation and fermentation. This application does not rely on complete fermentation cycle data, especially not on data from different stages within the fermentation pit, thus avoiding the limitation of only being able to retrospectively analyze yield data after the fact. This invention can predict yield in the early stages of pile fermentation, enabling timely early warning before the fermentation process. Based on the predicted yield, it allows for effective production intervention within a sufficient timeframe, real-time mitigation of any abnormal situations that may occur during fermentation, and true forward-looking judgment of the production process.

[0017] 2. Finer prediction granularity: Supports heap-level cell prediction This application avoids using the batch level as a unit, thus avoiding the difficulty in capturing differences between workshops, work groups, and piles. This invention uses the "pile" as the modeling unit to achieve differentiated modeling and precise management between piles. In the early stages of fermentation, knowing the output of each pile in all workshops and work groups across the entire company in its respective batch allows for rapid identification of abnormal fermentation piles, enabling manual intervention to change the fermentation status; it also allows for rapid identification of piles with abnormal fermentation, recording production operations, and summarizing beneficial practices. Knowing the output of each pile in all workshops and work groups across the entire company in its respective batch in the early stages of fermentation allows for clear understanding of the output of each work group, each workshop, and the entire company in that batch, providing a comprehensive understanding of the production situation of Maotai-flavor liquor; it allows for comparative analysis of different work groups and workshops in the same year within that batch, as well as comparative analysis of work group, workshop, and company-level outputs in different years within that batch; it allows for learning from the superior operations of high-yield work groups and workshops, while being wary of the inferior operations of low-yield work groups and workshops, thus better guiding production; and it provides progressive, comprehensive, and detailed management at each level, facilitating assessment and better managing the entire production process.

[0018] 3. Enhanced modeling capabilities: Incorporating dummy variables representing year and workshop differences. This application improves the PLS (Partial Least Squares Regression) algorithm by adding dummy variable indicators representing year-to-year and workshop-to-workshop differences, and extracts principal components through dimensionality reduction. This not only reflects year-to-year and workshop-to-workshop differences but also solves the problems of variable redundancy and collinearity, thereby improving the model's interpretability and stability while achieving a higher fitting ability.

[0019] 4. Higher deployment efficiency: Lightweight structure, efficient training, and rapid inference. The PLS+ (Improved Partial Least Squares Regression) model used in this application has a simple structure, making it easy to embed and deploy in factory digital platforms; it has a fast inference speed, making it suitable for online updates and simultaneous prediction of multiple heaps.

[0020] Therefore, the method of this application, based on an improved partial least squares regression algorithm, effectively solves the problems of coarse prediction granularity, insufficient timeliness, failure to consider the year, significant differences between workshops, and unstable models in the existing technology, and provides reliable technical support for the digital brewing and capacity optimization of Maotai-flavor liquor. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.

[0022] Figure 1 It is the overall structural block diagram of the method of the present application; Figure 2 It is the PLS+ prediction flowchart of the method of the present application; Figure 3 It is the schematic diagram comparing the predicted liquor yield and the actual liquor yield in an embodiment of the method of the present application. Detailed implementation manners

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope protected by the present application.

[0024] It should be noted that in the specific implementation manners of the present application, the fermentation of sauce-flavored Baijiu is used as an example to explain the present application, but the fermentation of the present application is not limited to the fermentation of sauce-flavored Baijiu.

[0025] The process of sauce-flavored Baijiu includes: steaming the fermented grains (sorghum) in a container called a steamer. After steaming, it is spread out on a drying floor for airing. When the temperature of the fermented grains drops to a certain temperature, the aired fermented grains are mixed with koji and then gathered together to form a pile similar to a semi-ellipsoid. A pile includes the fermented grains (sorghum) of multiple steamers. The pile is piled and fermented on the drying floor. After fermentation to a certain extent, it is put into the cellar. Finally, a complete pile contains the fermented grains of multiple steamers, so such a process needs to be repeated many times. In this process, starting and piling up include starting the pile and piling up. Starting the pile can be understood as the first steamer, a state where the pile starts from nothing. Piling up can be understood as the subsequent second steamer, third steamer, etc., continuously covering the fermented grains on it, and the state where the pile grows from small to large.

[0026] Term explanation: Starting and piling-up temperature: The starting pile temperature refers to the temperature of the fermented grains after the first steamer of fermented grains is aired and mixed with koji, when the aired fermented grains are gathered and piled up into a pile similar to a semi-ellipsoid. The piling-up temperature refers to the temperature when the fermented grains of the subsequent steamers are gathered and piled up.

[0027] The ambient temperature in the drying room when the fermented mash is piled up into a semi-elliptical shape: the temperature inside the production room when the fermented mash is piled up into a semi-elliptical shape.

[0028] Output of a pile of stills = Total output of alcohol from the pile / Total number of stills in the pile.

[0029] Some studies on predicting the yield of Maotai-flavor baijiu have focused on the yield of mash and base liquor from different fermentation batches. They analyzed the moisture, acidity, sugar, and starch content during the fermentation and extraction stages, and linearly fitted the relationship between these indicators and the base liquor yield. While this study fitted the relationship between the physicochemical indicators of Maotai-flavor baijiu and the base liquor yield, providing a reference for subsequent research, it is based on mash entering or leaving the fermentation pit and only linearly fitted the yield for each fermentation batch. Therefore, it cannot solve the problem of predicting the yield of the entire fermentation pile in the early stages of the pile fermentation process.

[0030] Some studies have used backpropagation (BP) neural networks to predict the alcohol content in fermentation mash over time by measuring alcohol data at various stages of fermentation in fermentation pits. This study clarified the change in alcohol content within the fermentation pit over time, but it was still based on the pit's internal stages, not the early stages of pile fermentation, and it predicted only alcohol content, not a specific yield.

[0031] Some researchers have used temperature data, moisture content, and acidity within the fermentation pits to construct a solid-state baijiu production prediction model based on the XGBoost model, comparing it with support vector machines, backpropagation neural networks, least squares support vector machines, and least squares support vector machines optimized by particle swarm optimization. This research demonstrates the application of various machine learning models in baijiu production prediction; however, it remains based on the fermentation pit stage and is specific to strong-aroma baijiu, failing to address the problem of predicting the production of sauce-aroma baijiu in the early stages of fermentation, which is the focus of this invention.

[0032] Some studies, based on data from the entire baijiu production process, first derive strong rules between various production intervals and baijiu quality using association rules. Then, machine learning is applied to baijiu quality classification, yield regression prediction, and production process parameter optimization. Finally, addressing the issue of large yield regression prediction errors in machine learning models, a deep learning algorithm based on fully connected feature models, shortcut connection feature models, and convolutional feature models is proposed for baijiu quality classification and yield prediction. This research combines association rules, machine learning, and deep learning to classify baijiu quality and predict yield, but it relies on data from the entire production process and cannot predict the early stages of fermentation.

[0033] Some methods provide a way to predict the production of baijiu in each cellar by constructing a GBR (gradient boosting regression) model based on the stacking time of the mash and the fermentation temperature of the mash at different cellar depths. However, this method is based on the fermentation stage in the cellar, at which point it is impossible to intervene too much in the mash in the cellar to control the production, and it is also impossible to make predictions in the early stage of stacking and fermentation.

[0034] Some methods offer a fusion model that combines random forest, gradient boosting, and XGBoost models using stacking techniques. This method uses physicochemical data of the fermented mash and starter culture at the team level, including moisture, acidity, starch, and reducing sugar, to predict the yield of Maotai-flavor liquor. While this method integrates multiple black-box machine learning models and has a certain level of accuracy, it suffers from poor interpretability and weak generalization. Furthermore, since the input data consists of physicochemical data of the fermented mash at the team level, not the yield prediction data at the pile level, it still cannot solve the problem of predicting the yield of each pile in the early stages of fermentation.

[0035] Some methods offer an XGBoost regression model to predict the alcohol yield of a pile during the fermentation stage, using data such as temperature, acidity, and moisture content at the time of fermentation. Although this method predicts at the pile level, it still only predicts at the fermentation stage and cannot solve the problem of predicting the alcohol yield of each pile in the early stages of fermentation, nor can it address the need to adjust the alcohol yield by manipulating the pile during the fermentation stage.

[0036] Current technologies have the following characteristics: 1. Coarse prediction granularity: Current research is based on "round-level" data for training and prediction, which cannot predict individual workshops, work groups, or piles, making it difficult to meet the actual needs of refined management. 2. Insufficient prediction timeliness: Most modeling methods rely on complete data for the entire fermentation cycle, especially information after fermentation and even after distillation. Therefore, they cannot predict alcohol production in the early stages of fermentation, limiting the foresight of the pile adjustment process. 3. Failure to consider differences between years and workshops: Due to adjustments in overall production strategies, there are significant differences in alcohol production between different years; differences in workshop management personnel also lead to differences between workshops. Most modeling methods do not use multi-year, multi-workshop data or do not consider differences between years and workshops, resulting in poor adaptability or low accuracy of the prediction model. 4. Significant variable redundancy and multicollinearity issues: The high correlation between various physicochemical indicators (temperature, moisture, sugar content, etc.) during fermentation easily leads to multicollinearity problems in modeling, affecting model accuracy and robustness.

[0037] Therefore, how to construct a single-round liquor production modeling method at the pile level that can both predict the early stage of pile fermentation and take into account interpretability and prediction accuracy has become a key technical challenge in the current digital brewing of Maotai-flavor liquor.

[0038] include: Collect feature data from multiple piles and their corresponding wine production; The characteristic data of the piles in the same production cycle are preprocessed to form an input matrix X and a wine production vector Y; the year and workshop information of the pile are encoded or continuously characterized to obtain characteristic variables that can characterize the systematic deviation of the year and workshop relative to the overall average level, and the characteristic variables are added to the input matrix X; Based on the input matrix X and the wine production vector Y, a regression prediction model from the input matrix X to the wine production vector Y is established using a partial least squares regression model.

[0039] In some embodiments of this application, the feature data includes process index data and physicochemical index data; In some embodiments of this application, the physicochemical indicators include temperature, moisture, acidity, sugar content, ethanol, starch, and acetic acid data measured at the pile surface, core, and bottom of the pile; In some embodiments of this application, the process index data includes stacking time, spreading and drying time, mixing temperature, temperature at which the starter culture is placed on the stack, and room temperature data of the drying hall when the starter culture is placed on the stack.

[0040] In some embodiments of this application, the encoding process is dummy variable processing, including: Select a base year and a base workshop; For each year, a 0-1 variable is set for each year other than the base year. The value is 1 when the year to which the heap belongs is that year, and 0 otherwise. For each workshop, a 0-1 variable is set for each workshop other than the reference workshop. The value is 1 when the pile belongs to the workshop, and 0 otherwise. The base year and base workshop are implicitly represented by corresponding dummy variables all being 0.

[0041] In some embodiments of this application, the continuous characterization process includes: constructing a continuous environmental index that reflects the overall fermentation environment of a year or workshop based on the statistical characteristics or principal component analysis scores of historical processes and physicochemical indicators, and using the continuous environmental index as a characteristic variable characterizing the systematic deviation of that year or workshop.

[0042] In some embodiments of this application, the encoding process includes effect encoding: setting an offset variable for a year or workshop with a number equal to the number of its categories minus one, and learning the offset strength of each year or workshop relative to the overall mean during model training.

[0043] In some embodiments of this application, the preprocessing includes cleaning; the cleaning includes: identifying abnormal data in each group where each indicator deviates from three times the standard deviation of the mean or is missing; if a single pile has two or more abnormal indicators, the pile is removed; if only a single indicator is abnormal, the indicator value corresponding to the pile with the smallest overall distance from the abnormal pile among all piles with normal indicators in the same group is used for replacement.

[0044] In some embodiments of this application, the preprocessing includes standardization, which includes: standardizing the process indicators and physicochemical indicators by subtracting the mean and then dividing by the standard deviation; centering the dummy variable indicators generated by converting year and workshop information by subtracting the mean; and standardizing the wine production vector by subtracting the mean and then dividing by the standard deviation.

[0045] The second aspect of this application provides a method for predicting the output of a liquor production pile; The steps include: obtaining the characteristic data of the pile to be predicted, including process indicators, physicochemical indicators, and information on the year and workshop to which it belongs; The year and workshop information of the pile are encoded or continuously characterized to obtain feature variables that can characterize the systematic deviation of the year and workshop relative to the overall average level, and an input vector is formed based on the feature data and the feature variables. The input vector is input to a pre-built wine production prediction model for the same production round, and the wine production prediction model outputs the predicted wine production value for that batch. The wine production prediction model is obtained by the model construction method as described in any one of the first aspects.

[0046] The third aspect of this application provides a system for constructing a predictive model for the production volume of a liquor pile. include: The data acquisition module is used to collect characteristic data of multiple piles in the same production cycle and their corresponding wine production. The data preprocessing module is used to preprocess the feature data, converting the year and workshop information of the pile into feature variables that can characterize its systematic deviation relative to the overall average level, and forming an input matrix X and a wine production vector Y. The model building module is used to construct a regression prediction model from X to Y based on the input matrix X and the wine production vector Y, using the partial least squares regression method, as the wine production prediction model for this production cycle.

[0047] The fourth aspect of this application provides a system for predicting the production volume of a liquor pile. The feature acquisition module is used to acquire feature data of the pile to be predicted. The feature data includes process indicators, physicochemical indicators, and information on the year and workshop to which it belongs. The encoding conversion module is used to convert the year and workshop information into feature variables that can characterize their systematic deviation relative to the overall average level, and generate an input vector based on the feature data and the feature variables; The prediction execution module is used to input the input vector into a pre-built wine production prediction model for the same production round, and output the predicted wine production value of the pile. The wine production prediction model is a partial least squares regression model obtained by the model construction method described in any one of the first aspects.

[0048] In some specific embodiments, the improved partial least squares regression model is constructed using an improved partial least squares regression method based on the characteristic data of historical piles and their corresponding wine production. The improved partial least squares regression method includes: forming an input matrix X and a wine production vector Y based on the characteristic data of historical piles and their corresponding wine production; iteratively extracting comprehensive variables based on the input matrix X and the wine production vector Y; performing unit projection correction on the weight vector determined in each iteration during the extraction of comprehensive variables; and establishing a regression prediction model from the input matrix X to the wine production vector Y based on the comprehensive variables obtained in the final iteration.

[0049] Example The embodiments and comparative examples provided below illustrate the implementation of this application in more detail. Various tests and evaluations were conducted according to the methods described below. Furthermore, unless otherwise specified, "parts" and "%" are quality standards.

[0050] Example 1: A model construction method for predicting the production volume of baijiu (Chinese liquor) from a distillery. See Figure 1 and Figure 2 , Figure 1 and Figure 2 A flowchart illustrating the method of this application is provided.

[0051] 1. Data collection and processing (1) Collect and organize the basic information of the pile, including the rotation number, year, workshop number, shift number, and pile number; (2) Collect and organize the production process data of the pile, including the piling time, spreading and drying time, mixing temperature, lifting and piling temperature and the room temperature data of the drying hall when lifting and piling; (3) Collect and organize the physicochemical index data of the pile in the early stage of pile fermentation, including temperature, moisture, acidity, sugar, ethanol, starch and acetic acid data of the top, center and bottom of the pile; (4) Collect and organize data on the production of alcohol from the piles, including the total production of alcohol from the piles, the total number of stills, and the production of alcohol from each still.

[0052] 2. Data Cleaning Data cleaning preprocessing is performed on the same batch of data, including: (1) Remove invalid data, including pile data with extremely low or extremely high distillation yield; (2) Identify and clean abnormal data in the process and physicochemical indicators: Considering the differences between years and workshops, the cleaning data should be identified using "year + workshop" as the unit. First, determine whether each indicator for each pile under "Year + Workshop" is abnormal: Represents a pile of Indicators, if Not here Internal or If missing, determine the next item in the "year + workshop" list. of The indicator data is abnormal, among which For the "year + workshop" The mean of the indicator, For the "year + workshop" Standard deviation of the indicator.

[0053] Then, the abnormal data under "Year + Workshop" is processed: if the pile If two or more indicators are abnormal, then the pile will be... Data is deleted; otherwise, the heap is deleted. The abnormal indicators were corrected.

[0054] Finally, the abnormal data under "Year + Workshop" was corrected: The abnormal data was found within the clusters where all indicators for "Year + Workshop" were normal. Most similar piles That is, distance ( The smallest heap, using this most similar normal heap. of Replace the abnormal heap with the value of the indicator of The value of the indicator. Anomaly heap. With normal pile The distances between them are as follows, where , which represents the total number of process and physicochemical indicators.

[0055] ; In the above formula, For distance; ; Represents a pile of index.

[0056] 3. Construct a PLS+ (Improved Partial Least Squares Regression) prediction model A PLS+ (improved partial least squares regression) early prediction model based on the pile level was constructed. The process and physicochemical indicators that affect the yield of alcohol in a certain round of pile, such as moisture, acidity, and stacking time, as well as dummy variable indicators that reflect the differences in year and workshop, were input. The predicted value of the still yield of the pile in that round was output, as well as the degree of influence of each process and physicochemical indicator on the predicted still yield of the pile (VIP score, Variable Importance in Projection), as follows: (1) Define dummy variable indicators For each pile, dummy variables representing the year and workshop are added based on the pile's year and workshop. These dummy variables are encoded as 0 or 1; 1 represents a pile belonging to that year, and 0 represents a pile not belonging to that year. Similarly, 1 represents a pile belonging to that workshop, and 0 represents a pile not belonging to that workshop. To avoid multicollinearity between dummy variables representing years and workshops, one year and one workshop are selected as the baseline. That is, the number of dummy variables representing years is one less than the current number of years, and the number of dummy variables representing workshops is one less than the current number of workshops.

[0057] Let's take an example. Suppose we have data from 2021 to 2023, for workshops 1-3, with 2023 and workshop 2 as the baseline. In this case, there are two dummy variables representing the years: one for 2021 and one for 2022. If the pile belongs to 2023, then the dummy variables for 2021 and 2022 are both 0; if the pile belongs to 2022, then the dummy variable for 2022 is 1 and the dummy variable for 2021 is 0; if the pile belongs to 2021, then the dummy variable for 2021 is 1 and the dummy variable for 2022 is 0. In this case, there are also two dummy variables representing the workshops: one for workshop 1 and one for workshop 3. If the pile belongs to workshop 2, it means that the dummy variables of workshops 1 and 3 are all 0; if the pile belongs to workshop 1, it means that the dummy variable of workshop 1 is 1 and the dummy variable of workshop 3 is 0; if the pile belongs to workshop 3, it means that the dummy variable of workshop 3 is 1 and the dummy variable of workshop 1 is 0.

[0058] In addition, the analysis needs to be based on the baseline year and the baseline workshop, that is, how other years affect the forecast of the output of the stacking still relative to the baseline year; and how other workshops affect the forecast of the output of the stacking still relative to the baseline workshop.

[0059] (2) Organize the input dataset The process indicators of a certain round of fermentation in the early stage of pile fermentation, including pile duration, spreading and cooling duration, mixing temperature, temperature when lifting the pile, and room temperature of the cooling hall when lifting the pile, as well as the physicochemical indicators in the early stage of fermentation, including moisture, acidity, sugar content, ethanol, starch, acetic acid, and dummy variable indicators representing the year and workshop, are converted into a multivariate indicator input data matrix X, or simply indicator X. Each row represents a pile sample, and each column represents an indicator.

[0060] The output of each pile in index X is transformed into an input matrix Y, or simply pile output Y, where each row represents a pile sample and a single column represents the output of the pile.

[0061] The process, physicochemical, and dummy variable indicators participating in the model have different units and data magnitudes. To avoid the impact of this situation on the model, the indicator data needs to be processed before being input into the model. The specific processing method is shown in the following formula: process and physicochemical indicators ( Standardize the data to represent dummy variable indicators for year and workshop. ) Perform mean removal processing, where This represents the total number of indicators involved in the forecast; see the formula below for details.

[0062] ; in express The value of the indicator This indicates the total number of indicators involved in the forecast. Indicates process and physicochemical properties. , representing dummy variable indicators for year and workshop, This indicates the output of the steamer. , This indicates the output of the processed data and the output of the steamer. , This represents the average output of each indicator and the number of steamers in the modeling dataset. , This represents the standard deviation of each indicator and the output of the stacked steamer in the modeling dataset.

[0063] (3) Establish a PLS+ (improved partial least squares regression) production prediction model ① As described above, input the prepared multivariate index input data matrix X and the pile steamer output input matrix Y into the model.

[0064] ② Simultaneously considering the relationships between indicators and the relationship between indicators and the output of the steamer, a set of latent variables is extracted using dimensionality reduction. and Replacing the original index X and the output of the steamer Y, making , .

[0065] ③ By maximizing and The degree of correlation, i.e. Solve to determine the weight matrix of index X. and latent variable matrix The weight matrix of the output Y of the steamer pile and latent variable matrix .

[0066] ④ Based on the extracted latent variables With indicators Production of the stacked steamer linear relationship between , + Solving using the least squares method, neglecting the residual term, determines... and .

[0067] ⑤ The regression relationship in the latent variable space cannot be directly used for prediction; it must be mapped back to the original variable space X. According to... Ignoring the residual term, multiply both sides to the right by . ,Right now And because It is invertible, thus allowing for equivalent coordinate transformations within the latent variable subspace, yielding... .Will Substitution ,Right now Thus, the regression coefficients are determined. Finally, the predicted output of the stack steamer was obtained. ( ).

[0068] ⑥ Determine the number of latent variables to be extracted Since PLS (Partial Least Squares Regression) is a method of reducing dimensions to extract latent variables to represent the index X and the output of the steamer Y, it is necessary to determine the optimal number of latent variables to be extracted in this round based on the prediction results.

[0069] This invention sets the following condition: when the contribution of extracting more latent variables to the prediction of the still's output is less than 0.5%, the current number of latent variables is determined to be the optimal number of latent variables to be extracted. The contribution rate of extracting K latent variables to the prediction of the hot pot output = The meanings of the symbols in the formula are explained above.

[0070] ⑦ Obtain the PLS+ (Improved Partial Least Squares Regression) yield prediction model Once the final number of latent variables extracted is determined, the final PLS+ (improved partial least squares regression) yield prediction model for that round can be obtained.

[0071] ⑧ Calculate the importance score of the measurement indicator to output forecast (VIP score). In the PLS+ (Improved Partial Least Squares Regression) model, the VIP (Variable Importance in Projection) score integrates the weights of indicators across multiple latent variables with the explanatory power of these latent variables on the yield of the steamer. It measures the importance of each indicator in the PLS+ model for that round, and the degree of influence of each indicator on the prediction of steamer yield. Indicators The VIP score calculation formula is as follows: First calculate the indicators. VIP points At this time It integrates process, physicochemical, and dummy variable indicators; to facilitate the assessment of the importance of process and physicochemical indicators to the output of the stack steamer, it will... The values ​​are limited to 14 process and physicochemical indicators, converted into .

[0072] ; ; in the formula Indicates the indicator, here Only the VIP score for physicochemical and process indicators is calculated; This indicates the total number of indicators involved in the forecast, including 14 process and physicochemical indicators, as well as dummy variable indicators representing workshops and years; This indicates the final number of latent variables extracted, which is selected based on the specific round.

[0073] After the VIP score is calculated, the importance of the indicator to the prediction of the still's output can be determined based on the VIP value. If... : Indicates indicators As an important variable, it has a significant impact on the prediction of the output of the steam pot; if : Indicates indicators Less important, may consider removing.

[0074] ⑨ Output prediction results Once the PLS+ (Improved Partial Least Squares Regression) model is established, data from the current pile in the early stage of fermentation can be input in real time during practical applications to complete advance predictions and provide decision support.

[0075] ⑩ Results Feedback Based on the output of the still's yield prediction results and the degree of influence of various processes and physicochemical indicators on the still's yield prediction, feedback is provided for subsequent still regulation.

[0076] Example 2: A method for predicting the yield of liquor from a liquor pile The model obtained in Example 1 was used for actual prediction. The results show that this application can operate stably on data from multiple years and multiple workshops. For example, the prediction of the pile yield in the early stage of the third round of distillation during the stack fermentation process highly matched the actual yield, with a relative absolute error of only 4.8% and an average absolute error of 3.6 kg. See details below. Figure 3 .

[0077] Example 3: A system for constructing a predictive model for the yield of baijiu (Chinese liquor) from a distillery. include: The preprocessing data module is used to collect feature data from multiple piles and their corresponding wine production; and to clean and standardize the feature data from the same round to form an input matrix X and a wine production vector Y. The model building module is used to build a prediction model based on the input matrix X and the wine production vector Y using an improved partial least squares regression method. The method of constructing a prediction model using the improved partial least squares regression method includes: iteratively extracting comprehensive variables based on the input matrix X and the wine production vector Y; performing unit projection correction on the weight vector determined in each iteration during the extraction of comprehensive variables; and establishing a regression prediction model from the input matrix X to the wine production vector Y based on the comprehensive variables obtained in the final iteration.

[0078] The data flow between the above modules is automatically scheduled by the data interface and control logic, realizing full-process automation from "data input" to "production forecast output".

[0079] Example 4: A system for constructing a predictive model for the yield of baijiu (Chinese liquor) from a distillery. include: The data acquisition module is used to acquire the characteristic data of the heap. The prediction module is used to input the feature data into the improved partial least squares regression model and output the predicted wine production value of each pile in this round. The improved partial least squares regression model is constructed using an improved partial least squares regression method based on the characteristic data of historical piles and their corresponding wine production. The improved partial least squares regression method includes: forming an input matrix X and a wine production vector Y based on the characteristic data of historical piles and their corresponding wine production; iteratively extracting comprehensive variables based on the input matrix X and the wine production vector Y; performing unit projection correction on the weight vector determined in each iteration during the extraction of comprehensive variables; and establishing a regression prediction model from the input matrix X to the wine production vector Y based on the comprehensive variables obtained in the final iteration.

[0080] The data flow between the above modules is automatically scheduled by the data interface and control logic, realizing full-process automation from "data input" to "production forecast output".

[0081] The method, model construction method, and system for predicting the production volume of a liquor production pile disclosed in this application have the following characteristics: 1. Earlier prediction timing: Supports prediction in the early stages of accumulation and fermentation. This application does not rely on complete fermentation cycle data, especially not on data from different stages within the fermentation pit, thus avoiding the limitation of only retrospectively analyzing yield data. In this invention, the process indicators involved in the prediction are piling time, spreading time, mixing temperature, temperature at which the starter culture is lifted from the pile, and the room temperature in the drying hall at the time of lifting. The physicochemical indicators involved in the prediction are moisture, acidity, sugar content, ethanol, starch, and acetic acid. These indicators can be obtained in the early stages of piling fermentation. Therefore, this invention can complete yield prediction in the early stages of piling fermentation, achieving timely early warning before the fermentation process. Based on the predicted yield, production intervention can be effectively carried out within a sufficient timeframe, real-time compensation for any abnormal situations that may occur during fermentation, and achieving forward-looking judgment of the actual production process.

[0082] 2. Finer prediction granularity: Supports heap-level cell prediction This application avoids using the batch level as a unit, thus avoiding the difficulty in capturing differences between workshops, work groups, and piles. This invention uses the "pile" as the modeling unit to achieve differentiated modeling and precise management between piles. In the early stages of fermentation, knowing the output of each pile in all workshops and work groups across the entire company in its respective batch allows for rapid identification of abnormal fermentation piles, enabling manual intervention to change the fermentation status; it also allows for rapid identification of piles with abnormal fermentation, recording production operations, and summarizing beneficial practices. Knowing the output of each pile in all workshops and work groups across the entire company in its respective batch in the early stages of fermentation allows for clear understanding of the output of each work group, each workshop, and the entire company in that batch, providing a comprehensive understanding of the production situation of Maotai-flavor liquor; it allows for comparative analysis of different work groups and workshops in the same year within that batch, as well as comparative analysis of work group, workshop, and company-level outputs in different years within that batch; it allows for learning from the superior operations of high-yield work groups and workshops, while being wary of the inferior operations of low-yield work groups and workshops, thus better guiding production; and it provides progressive, comprehensive, and detailed management at each level, facilitating assessment and better managing the entire production process.

[0083] 3. Enhanced modeling capabilities: Incorporating dummy variables representing year and workshop differences. This application improves the PLS (Partial Least Squares Regression) algorithm by adding dummy variable indicators representing year-to-year and workshop-to-workshop differences, and extracts principal components through dimensionality reduction. This not only reflects year-to-year and workshop-to-workshop differences but also solves the problems of variable redundancy and collinearity, thereby improving the model's interpretability and stability while achieving a higher fitting ability.

[0084] 4. Higher deployment efficiency: Lightweight structure, efficient training, and rapid inference. The PLS+ (Improved Partial Least Squares Regression) model used in this application has a simple structure, making it easy to embed and deploy in factory digital platforms; it has a fast inference speed, making it suitable for online updates and simultaneous prediction of multiple heaps.

[0085] In exploring the proposed solution, the applicant employed other algorithms for modeling and found that neural network modeling was prone to overfitting. Using PCA (Principal Component Analysis) + neural network, the extraction direction might be unrelated to yield because the dependent variable, yield, was not considered. Using a time-series model, such as LSTM (Long Short-Term Memory) / GRU (Gated Recurrent Unit), was problematic because it required continuous time-series data from the heap, which was difficult to collect, and existing heap fermentation data was often unstructured and discontinuous, making it unsuitable for application. While fuzzy logic reasoning systems could be used for modeling, they were difficult to quantify and exhibited weak generalization ability and limited scalability.

[0086] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present invention are protected by the present invention.

Claims

1. A method for constructing a model for predicting the production volume of baijiu (Chinese liquor) from a distillery, characterized in that, include: Collect feature data from multiple piles and their corresponding wine production; The characteristic data of the piles in the same production cycle are preprocessed to form an input matrix X and a wine production vector Y; the year and workshop information of the pile are encoded or continuously characterized to obtain characteristic variables that can characterize the systematic deviation of the year and workshop relative to the overall average level, and the characteristic variables are added to the input matrix X; Based on the input matrix X and the wine production vector Y, a regression prediction model from the input matrix X to the wine production vector Y is established using a partial least squares regression model.

2. The model construction method according to claim 1, characterized in that, The feature data includes process index data and physicochemical index data; Preferably, the physicochemical indicators include temperature, moisture, acidity, sugar content, ethanol, starch, and acetic acid data measured at the pile surface, core, and bottom of the pile; Preferably, the process index data includes the stacking time, spreading and drying time, mixing temperature, temperature at which the starter culture is placed on the stack, and the room temperature of the drying hall when the starter culture is placed on the stack.

3. The model construction method according to claim 1, characterized in that, The encoding process is a dummy variable processing, including: Select a base year and a base workshop; For each year, a 0-1 variable is set for each year other than the base year. The value is 1 when the year to which the heap belongs is that year, and 0 otherwise. For each workshop, a 0-1 variable is set for each workshop other than the reference workshop. The value is 1 when the pile belongs to the workshop, and 0 otherwise. The base year and base workshop are implicitly represented by corresponding dummy variables all being 0.

4. The model construction method according to claim 1, characterized in that, The continuous characterization process includes: constructing a continuous environmental index that reflects the overall fermentation environment of a year or workshop based on the statistical characteristics or principal component analysis scores of historical processes and physicochemical indicators, and using the continuous environmental index as a characteristic variable to characterize the systematic deviation of that year or workshop.

5. The model construction method according to claim 1, characterized in that, The encoding process includes effect encoding: setting an offset variable for each year or workshop with a number equal to the number of its categories minus one, and learning the offset strength of each year or workshop relative to the overall mean during model training.

6. The model construction method according to any one of claims 1-3, characterized in that, The preprocessing includes cleaning; the cleaning includes: identifying abnormal data in each group where each indicator deviates from the mean by three times the standard deviation or is missing; if a single pile has two or more abnormal indicators, the pile is removed; if only a single indicator is abnormal, the corresponding indicator value of the pile with the smallest overall distance from the abnormal pile among all piles with normal indicators in the same group is used to replace it.

7. The model construction method according to any one of claims 1-3, characterized in that, The preprocessing includes standardization, which includes: standardizing the process indicators and physicochemical indicators by subtracting the mean and then dividing by the standard deviation; centering the dummy variable indicators generated by converting year and workshop information by subtracting the mean; and standardizing the wine production vector by subtracting the mean and then dividing by the standard deviation.

8. A method for predicting the yield of liquor from a liquor production pile, characterized in that, Includes the following steps: Acquire the feature data of the pile to be predicted, including process indicators, physicochemical indicators, and information on the year and workshop to which it belongs; The year and workshop information of the pile are encoded or continuously characterized to obtain feature variables that can characterize the systematic deviation of the year and workshop relative to the overall average level, and an input vector is formed based on the feature data and the feature variables. The input vector is input into a pre-built wine production prediction model for the same production round, and the wine production prediction model outputs the predicted wine production value of the pile. The wine production prediction model is obtained by the model construction method as described in any one of claims 1 to 7.

9. A system for constructing a predictive model for the yield of liquor from a liquor production pile, characterized in that, The system includes: include: The data acquisition module is used to collect characteristic data of multiple piles in the same production cycle and their corresponding wine production. The data preprocessing module is used to preprocess the feature data, converting the year and workshop information of the pile into feature variables that can characterize its systematic deviation relative to the overall average level, and forming an input matrix X and a wine production vector Y. The model building module is used to construct a regression prediction model from X to Y based on the input matrix X and the wine production vector Y, using the partial least squares regression method, as the wine production prediction model for this production cycle.

10. A system for predicting the yield of liquor from a liquor production pile, characterized in that, The system includes: include: The feature acquisition module is used to acquire feature data of the pile to be predicted. The feature data includes process indicators, physicochemical indicators, and information on the year and workshop to which it belongs. The encoding conversion module is used to convert the year and workshop information into feature variables that can characterize their systematic deviation relative to the overall average level, and generate an input vector based on the feature data and the feature variables; The prediction execution module is used to input the input vector into a pre-built wine production prediction model for the same production round, and output the predicted wine production value of the pile. The wine production prediction model is a partial least squares regression model obtained by the model construction method described in any one of claims 1 to 7.