Tobacco leaf quality feature library-based threshing and redrying digital formula design method and system
By constructing a multi-dimensional tobacco leaf quality characteristic library and combining it with a hybrid optimization strategy, the problems of the one-sidedness of the tobacco leaf quality characteristic library and the multiple constraints of formulation design were solved, thus achieving the comprehensiveness of the tobacco leaf quality characteristic library and the efficient multi-objective response of formulation design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies have limitations in constructing tobacco leaf quality characteristic databases, including limitations in comprehensiveness and application. They cannot fully cover the multi-dimensional indicators of tobacco leaves, and the formulation design methods fail to take into account multiple production constraints, resulting in low efficiency and limited practicality.
A multi-dimensional tobacco leaf quality characteristic library was constructed, including inherent, appearance, physical, chemical and sensory evaluation indicators. A hybrid optimization strategy combining linear programming solver and particle swarm optimization algorithm was used to generate multi-objective formulation schemes, which were then predicted and calibrated using XGBoost and neural network models.
It improves the comprehensiveness of the tobacco leaf quality characteristic database and the responsiveness of formulation design, enabling it to meet users' multi-objective needs and improve the efficiency and adaptability of formulation design.
Smart Images

Figure CN121659708A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tobacco production technology, specifically to a digital formulation design method and system based on a tobacco leaf quality feature library for threshing and re-drying. More specifically, it is a formulation design method and system based on a feature library constructed from multiple indicators such as tobacco leaf appearance, physical indicators, chemical indicators, and sensory evaluation, combined with a multi-objective hybrid strategy intelligent algorithm. Background Technology
[0002] In the technical system of tobacco leaf re-drying production, formula design is the core link that determines the quality of the final finished tobacco leaf. The rationality of the formula ratio directly affects the chemical composition, sensory quality, and production cost of the product. Traditional formula design relies heavily on the personal experience of the formula designer, determining the grade and proportion of tobacco leaves by manually reviewing historical data and repeatedly trying different combinations. This method is inefficient, inconsistent, and difficult to quantify and pass on.
[0003] In recent years, with the development of digital technology, the industry has begun to explore the use of tobacco leaf quality characteristic libraries and mathematical models to assist in formulation design. However, existing technical solutions still have many limitations, mainly in the following two aspects: The construction of tobacco leaf quality characteristic databases suffers from limitations and application bottlenecks: Existing technologies for digital characterization of tobacco leaf quality often focus on a single or a few indicators, resulting in an insufficiently comprehensive and in-depth abstraction and expression of the overall quality of tobacco leaves. For example, patent CN104568823 primarily relies on near-infrared spectroscopy to reflect the chemical characteristics of tobacco leaves. While this method is fast, it lacks the inclusion of physical characteristics of tobacco leaf appearance and sensory evaluation indicators that determine the final product's taste. This leads to an incomplete characteristic database, unable to support the prediction and optimization of the overall quality of finished tobacco leaves, especially sensory quality. Patent CN118245836, although attempting to integrate appearance, physicochemical, and sensory data to form a multi-dimensional characteristic system, heavily relies on costly and inefficient manual evaluation for its sensory data. This results in long sample acquisition cycles, high costs, and difficulty in large-scale expansion and updates to the characteristic database, creating a serious application bottleneck in practice. Therefore, at the feature library level, the core problem of existing technologies is that they cannot build a tobacco quality feature library that can comprehensively cover multiple dimensions of tobacco leaf characteristics, such as inherent, appearance, physical, chemical and sensory aspects, while also overcoming the heavy reliance on manual evaluation and possessing good scalability.
[0004] The optimization objectives of the formula design method are too singular and fail to take into account the multiple constraints in actual production: In the formula calculation stage, the optimization objectives pursued by the existing optimization algorithms are too singular and cannot meet the urgent need for the coordinated optimization of multiple factors such as cost and inventory in modern refined production.
[0005] Patent CN111543668 uses near-infrared spectroscopy based on tobacco leaves to find similar and alternative tobacco leaves. Its optimization goal is mainly to minimize the error between the finished product's chemical indicators and the target values. This method completely ignores other equally crucial factors in the formulation process, such as: Inventory constraints: Failure to consider the actual inventory of each raw material tobacco leaf may result in the calculated optimal formula being unable to be put into production due to insufficient raw materials.
[0006] Cost control: Failure to take the cost differences of each raw tobacco leaf as an optimization target may result in a formula that meets the chemical indicators but has poor economic efficiency and high production costs.
[0007] Historical experience utilization: There is a lack of effective learning and utilization of historical successful formula data, and the algorithm's "intelligence" is insufficient, making it more like a mathematical tool than a carrier of experience inheritance and intelligent decision-making.
[0008] Therefore, at the formulation algorithm level, the core problem of existing technologies is that they are "single-objective" optimization methods that are detached from production reality. They cannot respond to the complex needs of multiple objectives such as cost control and inventory consumption while ensuring quality. Their decision-making dimensions are narrow and their practicality is limited. Summary of the Invention
[0009] The purpose of this invention is to provide a digital formulation design method and system for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library, addressing the aforementioned problems.
[0010] The technical solution of the present invention is as follows: A digital formulation design method based on a tobacco leaf quality characteristic library, comprising the following steps: Based on multi-dimensional indicator data of tobacco raw materials and finished tobacco products, a tobacco quality characteristic database is constructed; the multi-dimensional indicator data includes at least inherent indicators, appearance indicators, physical indicators, chemical indicators and sensory evaluation indicators. Based on the raw tobacco leaf quality characteristic data and finished tobacco leaf quality characteristic data in the tobacco leaf quality characteristic database, as well as the proportion of raw tobacco leaves in historical formulas, at least two different types of prediction models are trained and integrated to construct a meta-model for tobacco leaf quality prediction. The historical formula data takes the proportion and quality characteristics of raw tobacco leaves as input and the quality characteristics of finished tobacco leaves as output. It receives multi-objective optimization requirements set by users, and calculates the formula based on the meta-model for predicting tobacco quality using a hybrid optimization strategy that combines linear programming solver and particle swarm optimization algorithm, generating and outputting an ordered set of solutions for tobacco raw material ratio schemes. Based on the meta-model for predicting tobacco quality, the finished product indicators are predicted for the blending schemes in the ordered solution set, and the prediction results are output.
[0011] Using the above methods, a multi-dimensional comprehensive tobacco quality characteristic library was constructed based on the appearance, physical, chemical, sensory evaluation, and inherent indicators of tobacco leaves, providing data support for more fully utilizing historical formulation experience in the formulation design process.
[0012] Furthermore, the establishment of a tobacco quality characteristic database based on various indicator data of raw tobacco leaves and finished tobacco leaves specifically includes: Determine the indicator data: Research the indicator data for both raw tobacco leaves and finished tobacco products. The indicator data includes: inherent indicators, appearance indicators, physical indicators, chemical indicators, and sensory evaluation indicators; specifically as follows: Inherent indicators: tobacco leaf origin, tobacco leaf year, tobacco leaf grade, and tobacco leaf part; Appearance indicators: including color, texture and thickness, oil content, and maturity; Physical indicators: tobacco leaf moisture content, tobacco leaf temperature; Chemical indicators: total sugar, nicotine, reducing sugar, total nitrogen, total potassium, total chloride, pH, starch, sugar-nitrogen ratio, sugar-alkali ratio; Sensory evaluation indicators: in accordance with GB / T 5606.4, including gloss, aroma, harmony, off-flavors, irritation, and aftertaste; Data cleaning: Data cleaning is performed according to the indicator type, including the following steps: Data deduplication and missing value handling: An index is built based on tobacco year, tobacco grade, tobacco origin, and tobacco part. Duplicate data is located by comparison and filtered. For samples with missing values, the samples are classified and filled separately. If the missing field is numerical, the mean is used to fill it. If the missing field is categorical, 5-10 similar samples with the closest key field are found, and the mode of similar samples is used to fill the missing field. Outlier removal: Outlier removal includes two methods. The first is based on prior expert knowledge, which relies on the professional experience and judgment of experts in the tobacco industry and process personnel in specific re-drying plants to set data indicator ranges and delete outlier data that exceeds the range. The second is for numerical fields, which adopt the "6σ principle". Data that deviates from the mean by more than 6 times the standard deviation is identified as outliers and replaced with the median or the mean of adjacent normal data to avoid insufficient sample size due to removing too much data.
[0013] Data standardization: To eliminate differences in units between different fields, for numerical fields, since the production and testing data for leaf re-drying are relatively evenly distributed and have no obvious extreme values, range standardization is used to normalize the data. Through data standardization, all data are converted into numerical form; the formula is as follows: , in, For characteristic indicator data, This represents the maximum value of the corresponding feature index data. This represents the maximum value of the corresponding feature index data; Feature engineering: Based on cleaned data, through feature selection, transformation, and encoding, an effective feature dataset adapted for subsequent applications is constructed; specifically, it includes the following steps: Feature selection: Redundant features are removed, and the K-Means algorithm is used to cluster the features on the standardized sample data (using the Spearman correlation coefficient between features as the distance index), and the 9-10 features with the highest importance within the cluster are selected; Feature conversion: Based on the sensory quality evaluation standard of cigarettes, a quantile conversion method is introduced. The scores of sensory evaluation indicators such as gloss, aroma, harmony, off-flavor, irritation, and aftertaste are added together and divided into 4 categories according to [100, 85], (85, 75], (75, 60], and [60, 0], with values of 1, 0.75, 0.6, and 0, respectively. These constitute the sensory quality score characteristics, which serve as one of the quality characteristics of the finished tobacco leaves. Feature encoding: For non-categorical data, target encoding is used to effectively avoid the impact of different values of discrete features in the recipe algorithm model.
[0014] Furthermore, the construction of the meta-model for predicting tobacco quality specifically includes: Using the proportion of raw tobacco leaves and the quality characteristics of tobacco leaves in historical formulas as input, and the quality characteristics of finished tobacco leaves as output, a dataset is constructed and divided into training set, test set and validation set; Formula prediction models were constructed and trained using the XGBoost algorithm and the neural network algorithm, respectively. The predicted values of the XGBoost model and the neural network model on the validation set are used as meta-features, and the true values of the validation set are used as meta-labels. An ensemble calibration meta-model is trained to form the final meta-model for predicting tobacco quality.
[0015] Furthermore, the objective function of the XGBoost algorithm in constructing and training the recipe prediction model is to minimize the sample mean square error, as detailed below: , in, For the sample size, For the first in the sample The true value of each sample data point For the sample data, the first The predicted value for each sample; the model optimization objective formula is: .
[0016] Furthermore, the specific features of the neural network algorithm in constructing and training the recipe prediction model include: A neural network consists of an input layer, an output layer, and one hidden layer. Set the loss function of the neural network to mean squared error: , in, For the sample size, For the first in the sample The true value of each sample data point For the sample data, the first The predicted value for each sample.
[0017] Furthermore, the XGBoost model and the neural network model are calibrated using an ensemble calibration method to construct a meta-model for predicting tobacco quality, specifically including: Meta-feature construction: The validation set predictions of the XGBoost model and the neural network model are used as meta-features. The true values of the validation set are used as meta tags. ,in, These are the validation set predictions for the XGBoost model. These are the predicted values for the validation set of the neural network model. Meta-model training: The function for constructing the meta-model is as follows: , in, It is a logical function; Validation set integration calibration: comparing the calibrated model The value is used to verify and confirm the calibrated meta-model in the validation set.
[0018] Furthermore, the hybrid optimization strategy specifically includes: Set optimization objectives and constraints for formula design; optimization objectives should include at least a quality objective function and a cost objective function, with the quality objective function using a nonlinear penalty function based on the target interval; constraints should include at least a total raw tobacco leaf input ratio of 1 and inventory constraints for each raw tobacco leaf. Using a linear programming solver, optimization is performed with key chemical indicators as the single objective to obtain the first set of solutions for the proportions. The second set of solutions is obtained by using a multi-objective particle swarm optimization algorithm with mass objective function and cost objective function as multiple objectives. The first and second matching solution sets are merged, and duplicate solutions are removed. The contribution of each solution is calculated according to the preset quality and cost weights, and the final ordered solution set is generated by sorting the solutions according to their contribution.
[0019] Furthermore, the optimization objectives and constraints of the formulation design specifically include: For a given set of raw tobacco leaves Set of tobacco leaf quality characteristic indicators Parameters include: weight of finished tobacco leaves The index vector for each type of raw tobacco leaf is: Then the first The first type of raw tobacco leaves The indicators are The amount of each type of raw tobacco leaf fed into the plant. The set of indicators for finished tobacco leaves is as follows: , of which The indicators are Then we have: , The proportion X of the raw tobacco leaves can be expressed as: , The characteristic index data A of the raw tobacco leaves can be expressed as: , The objective function for formula design is: , , , , in, Let be the quality objective function, where For module number The penalty function for this index, is the penalty coefficient; where Indicators The center of the target interval, Indicators The interval half-width, The module's first The degree of deviation of the indicator from the center of the target interval and Representing indicators The upper and lower bounds of the target interval; The objective function is cost. Indicates the first The cost of inputs for tobacco-like products; The relevant function constraints are as follows: , The total quantity of each type of raw tobacco leaf selected is limited to a maximum of the total quantity of finished tobacco leaves. Indicates the first The stock of various grades of tobacco, This indicates the proportion of each type of raw tobacco leaf used in the feed. Furthermore, obtaining the second matching solution set specifically includes: The optimization objective of the particle swarm optimization algorithm is set as follows: , , This indicates the smallest deviation from the specified target for finished tobacco leaves; This indicates the lowest cost; The dataset is encoded using unsigned binary integers. The population size and number of iterations are set, and the optimization is iterated continuously to output the Pareto suboptimal solution set. ; The specific steps for merging the first and second ratio solution sets are as follows: Merged solution set: The solution set calculated by the linear programming solver The solution set calculated by the multi-objective particle swarm optimization algorithm Remove duplicate solutions and construct a merged solution set. ; Multi-dimensional sorting of merged solution sets: according to The data is sorted according to the set quality and cost weights, and the contribution is calculated using the following formula: , in, For the first The contribution of each solution set. for The weight of the target for The weight of the target; Based on contribution The solution set is sorted to generate an ordered solution set. .
[0020] Using the above method, based on the constructed multidimensional tobacco leaf quality feature library, a hybrid calculation method combining linear programming solver and swarm intelligence algorithm is adopted for formula calculation. It can meet the multi-objective requirements proposed by users, such as chemical composition, inventory targets, and cost targets, and can adapt to different user needs, providing multiple formula schemes that meet the requirements.
[0021] This application also includes a digital formulation design system for threshing and re-drying tobacco leaves based on a tobacco leaf quality characteristic library, and applies a digital formulation design method for threshing and re-drying tobacco leaves based on a tobacco leaf quality characteristic library, including: The data integration module is used to integrate tobacco leaf index data from external systems and testing equipment through at least one of the following methods: system interface, file transfer, database access, and OPC unified architecture. The system basic configuration module provides basic configuration data for the data integration module, quality characteristic data management module, model management module, and intelligent formula design module, ensuring the normal operation of the modules; The tobacco leaf quality characteristic data management module, connected to the data integration module, is used to manage the quality characteristic data of raw tobacco, selected tobacco leaves, finished tobacco leaves and module trial tobacco leaves, and to build and maintain the tobacco leaf quality characteristic library based on the integrated data. The model management module, connected to the tobacco quality characteristic data management module, is used for modeling, training, and calibrating the formula prediction model, as well as constructing the multi-objective optimization formula algorithm model. The intelligent formula design module is connected to the model management module and the tobacco quality characteristic data management module, respectively. It is used to automatically perform formula calculations based on the optimization goals and parameters set by the user, call the corresponding models and data, generate and recommend multiple optimized ratio schemes.
[0022] The system described above, based on the tobacco leaf quality feature library and formulation algorithm, adopts a highly modular system framework design. In actual deployment and application, the tobacco leaf quality feature library is continuously improved through data collection and accumulation, enhancing the adaptability of formulation generation. Corresponding formulation algorithm models are generated for different formulations, improving the system's support for and generalization of different formulations.
[0023] Compared with existing technologies, the advantages of this invention are: 1. Improve the comprehensiveness of the tobacco leaf quality characteristic database; This invention constructs a tobacco leaf quality characteristic database based on the appearance indicators, physical indicators, chemical indicators, sensory evaluation indicators, and inherent indicators of tobacco leaves. It covers different types of tobacco leaves, such as raw tobacco leaves and finished tobacco leaves, and covers relevant tobacco leaf materials in the process of leaf threshing and re-drying formula design. It also covers key quality indicators in the tobacco leaf processing process, thus comprehensively improving the comprehensiveness and representativeness of the tobacco leaf quality characteristic database. 2. Improve the ability of formula design to respond to users' multi-objective needs; This invention adopts a hybrid solution method that combines linear programming solvers with swarm intelligence algorithms to solve multiple raw material tobacco leaf ratio schemes to meet users' multi-objective needs such as chemical composition, inventory targets, and cost targets, thereby satisfying users' mixed needs for quality and cost. 3. Improve the efficiency of user formula design: This invention constructs a system for formula design based on tobacco leaf quality characteristic library, which can provide process guidance and simple data entry and formula calculation functions. Users can generate raw tobacco leaf ratio schemes and view the display of finished product prediction indicators by performing simple operations on the system interface. Compared with the traditional manual method of using Excel, it effectively improves the efficiency of user formula design. The system also provides an interface to realize the digital release of formula data. Attached Figure Description
[0024] Figure 1 This is a flowchart of the method described in this application. Figure 2 A flowchart for constructing the tobacco leaf quality characteristic library of this application. Figure 3 This is a schematic diagram of the system in this application. Detailed Implementation
[0025] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0026] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0027] Please see Figure 1-3 A digital formulation design method based on a tobacco leaf quality characteristic library, such as... Figure 1 As shown, it includes the following steps: Based on multi-dimensional indicator data of tobacco raw materials and finished tobacco products, a tobacco quality characteristic database is constructed; the multi-dimensional indicator data includes at least inherent indicators, appearance indicators, physical indicators, chemical indicators and sensory evaluation indicators. Based on the raw tobacco leaf quality characteristic data and finished tobacco leaf quality characteristic data in the tobacco leaf quality characteristic database, as well as the proportion of raw tobacco leaves in historical formulas, at least two different types of prediction models are trained and integrated to construct a meta-model for tobacco leaf quality prediction. The historical formula data takes the proportion and quality characteristics of raw tobacco leaves as input and the quality characteristics of finished tobacco leaves as output. It receives multi-objective optimization requirements set by users, and calculates the formula based on the meta-model for predicting tobacco quality using a hybrid optimization strategy that combines linear programming solver and particle swarm optimization algorithm, generating and outputting an ordered set of solutions for tobacco raw material ratio schemes. Based on the meta-model for predicting tobacco quality, the finished product indicators are predicted for the blending schemes in the ordered solution set, and the prediction results are output.
[0028] like Figure 2 As shown, constructing the tobacco leaf quality characteristic library specifically includes: This application studies the index data of tobacco raw materials and finished tobacco products. The index data includes: inherent indexes, appearance indexes, physical indexes, chemical indexes, and sensory evaluation indexes. Details are as follows: Inherent indicators: tobacco leaf origin, tobacco leaf year, tobacco leaf grade, and tobacco leaf part; Appearance indicators: including color, texture and thickness, oil content, and maturity; Physical indicators: tobacco leaf moisture content, tobacco leaf temperature; Chemical indicators: total sugar, nicotine, reducing sugar, total nitrogen, total potassium, total chloride, pH, starch, sugar-nitrogen ratio, sugar-alkali ratio; Sensory evaluation indicators: in accordance with GB / T 5606.4, including gloss, aroma, harmony, off-flavors, irritation, and aftertaste; Data cleaning: Data cleaning is performed according to the indicator type, including the following steps: Data deduplication and missing value handling: An index is built based on tobacco year, tobacco grade, tobacco origin, and tobacco part. Duplicate data is located by comparison and filtered. For samples with missing values, the samples are classified and filled separately. If the missing field is numerical, the mean is used to fill it. If the missing field is categorical, 5-10 similar samples with the closest key field are found, and the mode of similar samples is used to fill the missing field. Outlier removal: Outlier removal includes two methods. The first is based on prior expert knowledge, relying on the professional experience and judgment of tobacco industry experts and specific re-drying plant process personnel to set data indicator ranges and delete outlier data exceeding those ranges. The second method, for numerical fields, uses… According to the "principle", data that deviates from the mean by more than 6 times the standard deviation is considered outlier and is replaced by the median or the mean of adjacent normal data to avoid insufficient sample size due to excessive data removal.
[0029] Data standardization: To eliminate differences in units between different fields, for numerical fields, since the production and testing data for leaf re-drying are relatively evenly distributed and have no obvious extreme values, range standardization is used to normalize the data. Through data standardization, all data are converted into numerical form; the formula is as follows: , in, For characteristic indicator data, This represents the maximum value of the corresponding feature index data. This represents the maximum value of the corresponding feature index data; Feature engineering: Based on cleaned data, through feature selection, transformation, and encoding, an effective feature dataset adapted for subsequent applications is constructed. Specifically, it includes the following steps: Feature selection: Redundant features were removed, and the standardized sample data were clustered using the K-Means algorithm (using the Spearman correlation coefficient between features as the distance indicator). The 9-10 features with the highest importance within each cluster were selected. The clustering analysis results are as follows: For raw tobacco leaves, the 10 most important characteristics include: tobacco leaf origin, tobacco leaf grade, tobacco leaf year, part of the plant, total sugar, nicotine, reducing sugar, total nitrogen, tobacco leaf moisture, and color. For finished tobacco leaves, the nine most important characteristics include: tobacco leaf grade, aroma, off-flavors, irritation, aftertaste, total sugar, nicotine, sugar-to-nicotine ratio, and sugar-to-nitrogen ratio. Feature conversion: Based on the sensory quality evaluation standard of cigarettes, a quantile conversion method is introduced. The scores of sensory evaluation indicators such as gloss, aroma, harmony, off-flavor, irritation, and aftertaste are added together and divided into 4 categories according to [100, 85], (85, 75], (75, 60], and [60, 0], with values of 1, 0.75, 0.6, and 0, respectively. These constitute the sensory quality score characteristics, which serve as one of the quality characteristics of the finished tobacco leaves. Feature encoding: For non-categorical data, target encoding is used, and the specific steps are as follows: For a non-categorical indicator, each category is mapped to a unique integer in the range [1, 1000]. For example, the origin of tobacco leaves is represented by 1 for Xuanwei and 2 for Yuxi; this effectively avoids the influence of different values of discrete features in the formulation algorithm model.
[0030] The specific steps involved in constructing a meta-model for predicting tobacco quality include: Using the proportion and quality characteristics of raw tobacco leaves in historical formulas as input and the quality characteristics of finished tobacco leaves as output, a dataset is constructed and divided into training, testing and validation sets. Formula prediction models were constructed and trained using the XGBoost algorithm and the neural network algorithm, respectively. The predicted values of the XGBoost model and the neural network model on the validation set are used as meta-features, and the true values of the validation set are used as meta-labels. An ensemble calibration meta-model is trained to form the final meta-model for predicting tobacco quality.
[0031] Construct a dataset and divide it into training, test, and validation sets: Using the proportion of raw tobacco leaves in historical formulas and the quality characteristics of raw tobacco leaves as input, and the quality characteristics of finished tobacco leaves as output, a dataset is constructed and divided into training, testing, and validation sets in an 8:1:1 ratio. Hierarchical K-fold cross-validation (K=5) is employed, dividing the dataset into K equal parts. Each part is used alternately as the validation set, and the remaining K-1 parts are used as the training set. The average of the K validation results is taken as the final generalization error of the model.
[0032] The sample data input dimensions include: tobacco leaf origin, tobacco leaf grade, tobacco leaf year, part of the plant, total sugar, nicotine, reducing sugar, total nitrogen, tobacco leaf moisture, color, and tobacco leaf proportion; The dimensions of the sample data output include: tobacco leaf grade, aroma, off-flavors, irritation, aftertaste, total sugar, nicotine, sugar-to-nicotine ratio, sugar-to-nitrogen ratio, and sensory quality score; Model building and training: Based on the data scale, XGBoost and neural networks were used to build recipe prediction models.
[0033] The XGBoost algorithm for building and training a recipe prediction model specifically includes: Configure core model parameters: The model objective function is the sample mean squared error. , in, For the sample size, For the first in the sample The true value of each sample data point For the sample data, the first The predicted value for each sample; the model optimization objective formula is: , Model training: Using a pre-split dataset, with the test set as the basis... The value is a monitoring indicator; the early shutdown cycle is set. When continuous Round Validation Set If the value does not decrease, terminate training to avoid overfitting. Parameter evaluation and tuning: After training, set... The threshold value is 0.2, and the coefficient of determination is... With a threshold of 0.9, calculate the objective function for the test set. Coefficient of determination If the data has not yet reached the threshold range, adjust parameters such as tree depth, learning rate, and regularization coefficient, and repeat the model training-parameter evaluation and tuning process until the validation set performance meets the preset threshold.
[0034] The specific steps involved in constructing and training a recipe prediction model using neural network algorithms include: Design the network structure: Set the number of neurons in the input layer Its dimension is consistent with the input data dimension of the dataset; the number of neurons in the output layer is set. Its dimensions are consistent with the output dimensions of the dataset. We set up a hidden layer with 15 neurons, using the sigmoid function as the activation function. The output layer uses a linear function, and the intermediate layers use the following activation function: , The activation function used in the output layer is: ,in, , For one in Fixed numbers; Model training and tuning: Set the loss function to mean squared error. , in, For the sample size, For the first in the sample The true value of each sample data point For the sample data, the first Predicted values for each sample; Set the initial learning rate to The number of iterations is The early stop cycle is When continuous If the accuracy on the validation set does not improve, training is terminated to avoid model overfitting. The training dataset is then input into the constructed neural network for training. The weights and biases of the network are continuously adjusted until the performance function value is less than the set value. After multiple trials, the optimal parameters are obtained, including the weight matrix and bias from the input layer to the intermediate layer and from the intermediate layer to the output layer.
[0035] The XGBoost model and the neural network model are calibrated using an ensemble calibration method to construct a meta-model for predicting tobacco quality, specifically including: Meta-feature construction: The validation set predictions of the XGBoost model and the neural network model are used as meta-features. The true values of the validation set are used as meta tags. ,in, These are the validation set predictions for the XGBoost model. These are the predicted values for the validation set of the neural network model. Meta-model training: The function for constructing the meta-model is as follows: , in, The function is a logistic function; a mapping model between meta-features and meta-labels is constructed using the logistic function, and the model is trained using the meta-features as input.
[0036] Validation set integration calibration: comparing the calibrated model The value is used to verify and confirm the calibrated meta-model in the validation set.
[0037] Hybrid optimization strategies specifically include: Set optimization objectives and constraints for formula design; optimization objectives should include at least a quality objective function and a cost objective function, with the quality objective function using a nonlinear penalty function based on the target interval; constraints should include at least a total raw tobacco leaf input ratio of 1 and inventory constraints for each raw tobacco leaf. Using a linear programming solver, optimization is performed with key chemical indicators as the single objective to obtain the first set of solutions for the proportions. The second set of solutions is obtained by using a multi-objective particle swarm optimization algorithm with mass objective function and cost objective function as multiple objectives. The first and second matching solution sets are merged, and duplicate solutions are removed. The contribution of each solution is calculated according to the preset quality and cost weights, and the final ordered solution set is generated by sorting the solutions according to their contribution.
[0038] The optimization objectives and constraints of the formulation design specifically include: For a given set of raw tobacco leaves Set of tobacco leaf quality characteristic indicators Parameters include: weight of finished tobacco leaves The index vector for each type of raw tobacco leaf is: Then the first The first type of raw tobacco leaves The indicators are The amount of each type of raw tobacco leaf fed into the plant. The set of indicators for finished tobacco leaves is as follows: , of which The indicators are Then we have: , The proportion X of the raw tobacco leaves can be expressed as: , The characteristic index data A of the raw tobacco leaves can be expressed as: , The objective function for formula design is: , , , , in, Let be the quality objective function, where For module number The penalty function for this index, is the penalty coefficient; where Indicators The center of the target interval, Indicators The interval half-width, The module's first The degree of deviation of the indicator from the center of the target interval and Representing indicators The upper and lower bounds of the target interval; The objective function is cost. Indicates the first The cost of inputs for tobacco-like products; The relevant function constraints are as follows: , The total quantity of each type of raw tobacco leaf selected is limited to a maximum of the total quantity of finished tobacco leaves. Indicates the first The stock of various grades of tobacco, This indicates the proportion of each type of raw tobacco leaf used in the feed. The optimal tobacco raw material ratio solution set is calculated using a linear programming solver.
[0039] A linear programming solver is a mathematical optimization tool that can quickly find the global optimum and rapidly solve single-objective optimization problems by setting key metrics. For nicotine, solve for the precision threshold. = Solve for the upper limit of time. =30 seconds, start the solver, and after the solution is completed, output the solution set of tobacco raw material ratio. .
[0040] The solution set for the second ratio specifically includes: The optimization objective of the particle swarm optimization algorithm is set as follows: , , This indicates the smallest deviation from the specified target for finished tobacco leaves; This indicates the lowest cost; The dataset is encoded using unsigned binary integers, and the population size is set. Number of iterations Iterative optimization is continuously performed to output a Pareto suboptimal solution set. ; The specific steps for merging the first and second ratio solution sets are as follows: Merged solution set: The solution set calculated by the linear programming solver The solution set calculated by the multi-objective particle swarm optimization algorithm Remove duplicate solutions and construct a merged solution set. ; Multi-dimensional sorting of merged solution sets: according to The data is sorted according to the set quality and cost weights, and the contribution is calculated using the following formula: , in, For the first The contribution of each solution set. for The weight of the target for The weight of the target; Based on contribution The solution set is sorted to generate an ordered solution set. .
[0041] Based on a meta-model for predicting tobacco quality, the expected finished tobacco product indicators can be predicted based on the proportions of raw tobacco materials. Details are as follows: Based on the generated ordered solution set Call the trained and calibrated tobacco leaf quality prediction meta-model, and input... The model outputs predicted characteristic data of the finished tobacco product, namely the various indicators that the finished product can achieve, based on the provided proportions of tobacco raw materials and their characteristic data.
[0042] The expected targets achieved by each scheme ratio obtained through the above method are as follows: Table 1 shows the expected performance indicators achieved by each of the proposed schemes.
[0043] This application also includes a digital formulation design system for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library, comprising: Data integration module: Integrates data from external systems and testing equipment through system interface calls, file transfer, database access, and OPC unified architecture. The external systems connected include production management system, quality management system, warehouse management system, and manufacturing execution system. The system basic configuration module provides basic configuration data for the data integration module, quality characteristic data management module, model management module, and intelligent formula design module, ensuring the normal operation of the modules; Tobacco Leaf Quality Characteristic Data Management Module: Manages raw tobacco quality characteristic data, selected tobacco leaf quality characteristic data, finished tobacco leaf quality characteristic data, and module-produced tobacco leaf quality characteristic data; Model management module: performs formula prediction modeling and multi-objective optimization formula algorithm modeling; Intelligent formula design module: Based on the selected formula prediction model and multi-objective optimization formula algorithm model, and based on the user-set optimization objectives and related parameters, it automatically completes the calculation and recommendation of multiple optimized ratio schemes, and provides a comparison of the optimization indicators of each scheme; The system's basic configuration module is connected to the data integration module, tobacco leaf quality characteristic data management module, model management module, and intelligent formula design module; the data integration module is connected to the tobacco leaf quality characteristic data management module; the tobacco leaf quality characteristic data management module is connected to the model management module and intelligent formula design module; and the model management module is connected to the intelligent formula design module.
[0044] The tobacco leaf quality characteristic data management module constructs a tobacco leaf quality characteristic library based on the integrated data synchronized by the data integration module, and outputs processed tobacco leaf quality characteristic data. The tobacco leaf quality characteristic data is transmitted to the model management module to support the construction of the formulation design algorithm model. The intelligent formulation design module automatically calculates and recommends multiple optimized ratio schemes based on the formulation design algorithm model provided by the model management module and the quality characteristic data provided by the tobacco leaf quality characteristic data management module, and provides a comparison of the optimization indicators of each scheme. The system basic configuration module provides system basic configuration data support for the operation of the data integration module, the tobacco leaf quality characteristic data management module, the model management module, and the intelligent formulation design module.
[0045] The system also boasts excellent cross-platform scalability, supports online access via browser, and at the hardware level supports x86_64 and aarch64 instruction set architectures, making it deployable on domestic processor platforms such as Intel, AMD, and Phytium. At the software level, it is compatible with mainstream Linux distributions, CentOS, Kylin, and other domestic operating systems, facilitating deployment and subsequent iterative updates.
[0046] The embodiments described above merely illustrate specific implementation methods of this application, and while the descriptions are detailed and specific, they should not be construed as limiting the scope of protection of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the technical solution of this application, and these modifications and improvements all fall within the scope of protection of this application.
Claims
1. A digital formulation design method for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library, characterized in that, Includes the following steps: Based on multi-dimensional indicator data of tobacco raw materials and finished tobacco products, a tobacco quality characteristic database is constructed; the multi-dimensional indicator data includes at least inherent indicators, appearance indicators, physical indicators, chemical indicators and sensory evaluation indicators. Based on the raw tobacco leaf quality characteristic data and finished tobacco leaf quality characteristic data in the tobacco leaf quality characteristic database, as well as the proportion of raw tobacco leaves in historical formulas, at least two different types of prediction models are trained and integrated to construct a meta-model for tobacco leaf quality prediction. The historical formula data takes the proportion and quality characteristics of raw tobacco leaves as input and the quality characteristics of finished tobacco leaves as output. It receives multi-objective optimization requirements set by users, and calculates the formula based on the meta-model for predicting tobacco quality using a hybrid optimization strategy that combines linear programming solver and particle swarm optimization algorithm, generating and outputting an ordered set of solutions for tobacco raw material ratio schemes. Based on the meta-model for predicting tobacco quality, the finished product indicators are predicted for the blending schemes in the ordered solution set, and the prediction results are output.
2. The digital formulation design method for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library as described in claim 1, characterized in that, The construction of the tobacco leaf quality characteristic library includes: Determine the data range for the indicators; handle duplicate and missing values; remove outliers; standardize the data; select features; transform features; encode features; among which: Define the scope of indicator data: study the indicator data of tobacco raw materials and finished tobacco products, including: inherent indicators, appearance indicators, physical indicators, chemical indicators, and sensory evaluation indicators; Data deduplication and missing value handling: Build an index based on tobacco year, tobacco grade, tobacco origin, and tobacco part, and filter duplicate data; use mean imputation to complete data with missing numerical fields; find the mode of 5-10 similar samples to fill data with missing categorical fields. Outlier removal: First, data index ranges are set based on prior expert knowledge, and data exceeding the index range is deleted; second, numerical data with a difference of more than 6 standard deviations above or below the mean is replaced with the median or the mean of adjacent normal data. Data standardization: The data is normalized using range standardization; the formula is as follows: , in, For characteristic indicator data, This represents the maximum value of the corresponding feature index data. This represents the maximum value of the corresponding feature index data; Feature selection: The K-Means algorithm is used to cluster the features of the standardized sample data. The Spearman correlation coefficient between features is used as the distance index, and the 9-10 features with the highest importance within the cluster are selected. Feature conversion: Based on the GB / T 5606.4 standard, a quantile conversion method is introduced. The scores of sensory evaluation indicators such as gloss, aroma, harmony, off-flavors, irritation, and aftertaste are added together and divided into 4 intervals according to [100,85], (85,75], (75,60], and [60,0], with values of 1, 0.75, 0.6, and 0, respectively, to form the sensory quality score characteristics. Feature encoding: For non-categorical data, target encoding is used.
3. The method for designing digital formulations for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library as described in claim 1, characterized in that, The construction of the meta-model for predicting tobacco quality specifically includes: Using the proportion of raw tobacco leaves and the quality characteristics of tobacco leaves in historical formulas as input, and the quality characteristics of finished tobacco leaves as output, a dataset is constructed and divided into training set, test set and validation set; Formula prediction models were constructed and trained using the XGBoost algorithm and the neural network algorithm, respectively. The predicted values of the XGBoost model and the neural network model on the validation set are used as meta-features, and the true values of the validation set are used as meta-labels. An ensemble calibration meta-model is trained to form the final meta-model for predicting tobacco quality.
4. The method for digital formulation design of threshing and re-drying tobacco leaves based on a tobacco quality characteristic library according to claim 3, characterized in that, The objective function of the XGBoost algorithm in constructing and training the recipe prediction model is to minimize the sample mean square error, as detailed below: , in, For the sample size, For the first in the sample The true value of each sample data point For the sample data, the first The predicted value for each sample; the model optimization objective formula is: 。 5. The method for designing digital formulations for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library according to claim 4, characterized in that, The specific features of the neural network algorithm used to construct and train the recipe prediction model include: A neural network consists of an input layer, an output layer, and one hidden layer. Set the loss function of the neural network to mean squared error: , in, For the sample size, For the first in the sample The true value of each sample data point For the sample data, the first The predicted value for each sample.
6. The method for designing digital formulations for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library according to claim 3, characterized in that, The XGBoost model and neural network model are calibrated using an integrated calibration method to construct a meta-model for predicting tobacco quality, specifically including: Meta-feature construction: The validation set predictions of the XGBoost model and the neural network model are used as meta-features. The true values of the validation set are used as meta tags. ,in, These are the validation set predictions for the XGBoost model. These are the predicted values for the validation set of the neural network model. Meta-model training: The function for constructing the meta-model is as follows: , in, It is a logical function; Validation set integration calibration: comparing the calibrated model The value is used to verify and confirm the calibrated meta-model in the validation set.
7. The method for digital formulation design of threshing and re-drying tobacco leaves based on a tobacco quality characteristic library according to claim 1, characterized in that, The hybrid optimization strategy specifically includes: Set optimization objectives and constraints for formula design; optimization objectives should include at least a quality objective function and a cost objective function, with the quality objective function using a nonlinear penalty function based on the target interval; constraints should include at least a total raw tobacco leaf input ratio of 1 and inventory constraints for each raw tobacco leaf. Using a linear programming solver, optimization is performed with key chemical indicators as the single objective to obtain the first set of solutions for the proportions. The second set of solutions is obtained by using a multi-objective particle swarm optimization algorithm with mass objective function and cost objective function as multiple objectives. The first and second matching solution sets are merged, and duplicate solutions are removed. The contribution of each solution is calculated according to the preset quality and cost weights, and the final ordered solution set is generated by sorting the solutions according to their contribution.
8. The method for designing digital formulations for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library according to claim 7, characterized in that, The optimization objectives and constraints of the formulation design specifically include: For a given set of raw tobacco leaves Set of tobacco leaf quality characteristic indicators Parameters include: weight of finished tobacco leaves The index vector for each type of raw tobacco leaf is: Then the first The first type of raw tobacco leaves The indicators are The amount of each type of raw tobacco leaf fed into the plant. The set of indicators for finished tobacco leaves is as follows: , of which The indicators are Then we have: , The proportion X of the raw tobacco leaves can be expressed as: , The characteristic index data A of the raw tobacco leaves can be expressed as: , The objective function for formula design is: , , , , in, Let be the quality objective function, where For module number The penalty function for this index, is the penalty coefficient; where Indicators The center of the target interval, Indicators The interval half-width, The module's first The degree of deviation of the indicator from the center of the target interval and Representing indicators The upper and lower bounds of the target interval; The objective function is cost. Indicates the first The cost of inputs for tobacco-like products; The relevant function constraints are as follows: , The total quantity of each type of raw tobacco leaf selected is limited to a maximum of the total quantity of finished tobacco leaves. Indicates the first The stock of various grades of tobacco, This indicates the proportion of each type of raw tobacco leaf used in the feed.
9. The method for designing digital formulations for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library according to claim 8, characterized in that, The specific steps to obtain the second matching solution set include: The optimization objective of the particle swarm optimization algorithm is set as follows: , , This indicates the smallest deviation from the specified target for finished tobacco leaves; This indicates the lowest cost; The dataset is encoded using unsigned binary integers. The population size and number of iterations are set, and the optimization is iterated continuously to output the Pareto suboptimal solution set. ; The specific steps for merging the first and second ratio solution sets are as follows: Merged solution set: The solution set calculated by the linear programming solver The solution set calculated by the multi-objective particle swarm optimization algorithm Remove duplicate solutions and construct a merged solution set. ; Multi-dimensional sorting of merged solution sets: according to The data is sorted according to the set quality and cost weights, and the contribution is calculated using the following formula: , in, For the first The contribution of each solution set. for The weight of the target for The weight of the target; Based on contribution The solution set is sorted to generate an ordered solution set. .
10. A digital formulation design system for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library, characterized in that, The application of the digital formulation design system for threshing and re-drying tobacco leaves based on a tobacco quality characteristic library as described in any one of claims 1-8 includes: The data integration module is used to integrate tobacco leaf index data from external systems and testing equipment through at least one of the following methods: system interface, file transfer, database access, and OPC unified architecture. The system basic configuration module provides basic configuration data for the data integration module, quality characteristic data management module, model management module, and intelligent formula design module, ensuring the normal operation of the modules; The tobacco leaf quality characteristic data management module, connected to the data integration module, is used to manage the quality characteristic data of raw tobacco, selected tobacco leaves, finished tobacco leaves and module trial tobacco leaves, and to build and maintain the tobacco leaf quality characteristic library based on the integrated data. The model management module, connected to the tobacco quality characteristic data management module, is used for modeling, training, and calibrating the formula prediction model, as well as constructing the multi-objective optimization formula algorithm model. The intelligent formula design module is connected to the model management module and the tobacco quality characteristic data management module, respectively. It is used to automatically perform formula calculations based on the optimization goals and parameters set by the user, call the corresponding models and data, generate and recommend multiple optimized ratio schemes.
Citation Information
Cited By
A method and system for formulation optimization of rubber-construction waste mix
CN122157908A