Data-driven optimization method for process parameters of traditional chinese medicine manufacturing process

By optimizing process parameters using the maximum information coefficient and particle swarm optimization algorithm, the problem of quantifying process parameters and quality indicators in the manufacturing process of traditional Chinese medicine was solved. This improved the accuracy of the quality prediction model for traditional Chinese medicine products and optimized the process parameters, thereby enhancing the rationality of decision-making in the production process.

CN116414095BActive Publication Date: 2025-12-23SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310413626.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-12-23
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

In the current manufacturing process of traditional Chinese medicine, it is difficult to determine the quantitative relationship between process parameters and quality indicators. Manual experience methods are highly subjective, and data-driven methods are difficult to screen key parameters and lack real data verification, which makes it difficult to optimize process parameters.

Method used

The maximum information coefficient is used to measure the linear and nonlinear relationship between process parameters and quality. A quality prediction model is constructed, and the key process parameters are optimized by combining the particle swarm optimization algorithm. The optimal solution is found in the process parameter search space through data preprocessing and particle swarm optimization algorithm.

Benefits of technology

This improved the accuracy of the quality prediction model for traditional Chinese medicine products, avoided the blindness of human experience, optimized process parameters, and enhanced the rationality and reliability of decision-making in the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116414095B_ABST
    Figure CN116414095B_ABST
Patent Text Reader

Abstract

The application discloses a data-driven traditional Chinese medicine manufacturing process parameter optimization method, which comprises the following steps: calculating the maximum mutual information coefficient (MIC) between the historical process parameters and product quality of a traditional Chinese medicine product production process, constructing a quality prediction model (PM-AdaBoost), calculating an adaptive function according to the maximum mutual information coefficient and the mean square error of the quality prediction model, initializing a particle swarm, calculating the adaptive function, and updating the speed and position of the particle in the process parameter search space through multiple iterations to obtain the key process parameters in the traditional Chinese medicine product production process and the optimized quality prediction model, and further taking the mean square error of the optimized quality prediction model as the adaptive function and utilizing a particle swarm optimization algorithm to obtain the optimized key process parameters. The application measures the linear and nonlinear relationship between variables through the maximum information coefficient, selects the maximum information coefficient and the quality prediction mean square error as the standard for screening the key process parameters and constructing the quality prediction model, and has an absolute advantage in accuracy. Based on the quality prediction model, the key is optimized through the particle swarm algorithm, and the particle swarm has an absolute advantage in convergence speed as the process parameter optimization algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of Chinese medicine manufacturing process information processing, and particularly relates to a data-driven Chinese medicine manufacturing process parameter optimization method. BACKGROUND

[0002] The existing Chinese medicine manufacturing process parameter optimization methods mainly include an artificial experience method and a data-driven method. The parameter values obtained by relying on artificial experience are only applicable to specific Chinese medicine products, and when the production conditions change, experience needs to be accumulated again, which wastes a large amount of human and time resources. The data-driven process parameter optimization relies on collected production data and machine learning, deep learning and other technologies, and before the process parameters are optimized, key process parameters affecting product quality need to be preselected, a quality prediction model is established by relying on the selected parameters, and finally, the process parameters are optimized based on the prediction model by using an optimization algorithm.

[0003] At present, the difficulties of the data-driven Chinese medicine product manufacturing process parameter optimization mainly include the following points: 1) it is difficult to quantify the relationship between the process parameters and the quality indicators in the manufacturing process; 2) the production of Chinese medicine products experiences multiple stages, and it is difficult to determine the key process parameters affecting the product quality in each stage, and if all the process parameters are used, the quality prediction model accuracy will be affected and the difficulty of subsequent process parameter optimization will be increased. The existing Chinese medicine manufacturing process parameter selection methods still have many deficiencies: 1) the method based on artificial experience is only applicable to specific products, and the artificial experience has certain subjectivity and blindness, when a new important product appears, there is no reference experience and manual, only accumulated data, and it is difficult to quickly optimize the process parameters. 2) the data-driven process parameter optimization method is difficult to preselect the key parameters affecting the product quality from the numerous parameters, establish an accurate quality prediction model, and optimize the process parameters. 3) it is difficult to quantify the correlation between the process parameters and the quality indicators in the manufacturing process. 4) when the process parameters are optimized, the actual production status is not considered, and there is a lack of real data verification. SUMMARY

[0004] The present application aims at the data collected in the existing traditional Chinese medicine product manufacturing process, which often has the characteristics of data missing, abnormal value, multiple collinearity and the like, leading to the difficulty in directly utilizing the production data for analysis, and the difficulty in the numerous factors affecting the quality of traditional Chinese medicine products in the manufacturing process of traditional Chinese medicine products, and proposes a data-driven traditional Chinese medicine manufacturing process process parameter optimization method, which measures the linear and nonlinear relationship between variables through the maximum information coefficient, and selects the maximum information coefficient and the quality prediction mean square error as the standard for screening the key process parameters and constructing the quality prediction model, which has an absolute advantage in accuracy; based on the quality prediction model, the key is optimized through the particle swarm algorithm, and the particle swarm algorithm has an absolute advantage in convergence speed as the process parameter optimization algorithm.

[0005] The present application is realized by the following technical solutions:

[0006] The present application relates to a data-driven traditional Chinese medicine manufacturing process process parameter optimization method, which calculates the maximum mutual information coefficient (MIC) between the historical process parameters and product quality in the production process of traditional Chinese medicine products, constructs a quality prediction model (PM-AdaBoost), calculates the fitness function according to the maximum mutual information coefficient and the mean square error of the quality prediction model, initializes the particle swarm, calculates the fitness function, and updates the speed and position of the particle in the process parameter search space through multiple iterations, obtains the key process parameters in the production process of traditional Chinese medicine products and the optimized quality prediction model, and further uses the mean square error of the optimized quality prediction model as the fitness function, and uses the particle swarm optimization algorithm to obtain the optimized key process parameters.

[0007] The historical process parameters of the traditional Chinese medicine product production process include: traditional Chinese medicine production process parameters X and traditional Chinese medicine product quality Y, x1, x2, …, x i, …, x n ∈X, which represents the traditional Chinese medicine product production process parameters, wherein represents the vector composed of the samples collected by the traditional Chinese medicine production process parameters, n represents the number of process parameters, N represents the number of samples, Y=(Y (1) , Y (2) , …Y (N) ) T represents the vector composed of the samples of the traditional Chinese medicine product quality.

[0008] The historical process parameters of the traditional Chinese medicine product production process are preferably pre-processed, specifically including: missing value processing, abnormal value processing and data standardization processing.

[0009] The missing value processing, i.e. single filling method, is to reduce the influence of incomplete data on the modeling of the production process, and the mean value of the sample data of the same type of parameters is used for filling.

[0010] The abnormal value processing refers to: adopting 3σ method to eliminate abnormal data, due to various disturbances in the collection, test instrument response drift, equipment failure, data recording deviation, analysis personnel operation error and other reasons, the obtained data is not reliable, and the quality of the data is reduced. The abnormal value refers to the abnormal data deviating from other data in the data set, and the deviation of such data is usually not caused by the normal fluctuation of the parameter. The abnormal value will greatly interfere with the data modeling, resulting in low model accuracy and decreased generalization ability, therefore, the 3σ method is adopted to eliminate abnormal data.

[0011] The data standardization processing refers to: adopting maximum and minimum value method for data standardization processing, various stirring speed, vacuum degree, temperature, viscosity value, holding time and other different dimension parameters are involved in the production process of traditional Chinese medicine products, the horizontal difference between different indicators is large, and the dimensions are different. In order to reduce the difference between variables and improve the convergence speed and training accuracy in the training of prediction model, the maximum and minimum value method is adopted to standardize the data.

[0012] The maximum mutual information coefficient is used to measure the correlation between the process parameters and the product quality in the production process of traditional Chinese medicine, which is obtained by the following method: the process parameters x i and the product quality Y are discrete in two-dimensional space, and are represented by scatter plot, the current two-dimensional space is divided into a certain number of intervals in X and Y directions respectively, then the current scatter is viewed in each grid, that is, MIC = MAX XY<B {I(x i , Y) / (log2(min(x i , Y)))}, wherein: I(x i , Y) represents the mutual information of process parameters x i and product quality Y, wherein, I(x i , Y) = -∑∑(P(x i , Y)log(P(x i , Y) / (P(x i )P(Y))), P(Y) is the probability density function of quality Y = Y j , P(x i ) is the probability density function of process parameters .

[0013] The fitness function calculated according to the maximum mutual information coefficient and the mean square error of the quality prediction model refers to: Wherein: f i is the i th process parameter, Y is the quality index parameter, n sThe number of co-selected parameters is the mean square error between the predicted value of the quality prediction model and the true value of the quality.

[0014] The initialization of the particle swarm refers to: when solving the key process parameter selection and quality prediction model construction problem by using the particle swarm optimization algorithm, first, the number of particles m is set in the process parameter search space of the particles, and the position and speed of each particle in the search space are randomly assigned.

[0015] The velocity and position of the particle in the process parameter search space are updated, specifically: Wherein: is the weight, the velocity v of the particle swarm i =(v i1 , v i2 , …, v im ) T , i = 1, 2, …, m, represents the position vector of particle i in the j dimension in the t iteration, the position z of the particle in the space i =(z i1 , z i2 , …, z im ) T , i = 1, 2, …, m, the optimal position p searched by the particle i =(p i1 , p i2 , …, p im ) T , i = 1, 2, …, m, the optimal position p searched by the group g =(p g1 , p g2 , …, p gm ) T , i = 1, 2, …, m, r1r2 is a random number between 0 and 1, and c1c2 are the particle individual learning factor and the group learning factor respectively; the fitness function is calculated again after updating, and the iteration is stopped when the number of iterations is reached and the current Chinese medicine production process parameter set and the optimized quality prediction model PM-AdaBoost are output.

[0016] The optimized key process parameter obtained by using the particle swarm optimization algorithm refers to: taking the mean square error between the true value of the quality of traditional Chinese medicine and the predicted value of the optimized quality prediction model as the fitness function, initializing the particle swarm population, i.e. setting the number of particles in the key process parameter search space of the particles and randomly assigning the position and speed of each particle in the search space, and after multiple iterations of the particles, i.e. updating the velocity and position of the particle in the process parameter search space, wherein: is a weight, the velocity v' of the particle swarm i = (v' i1 , v' i2 , …, v' im ) T , i = 1, 2, …, m, represents the position vector of particle i in the j dimension in the t iteration, the position z' of the particle in the space i = (z' i1 , z' i2 , …, z' im ) T , i = 1, 2, …, m, represents the position of the particle in the space, the optimal position p' searched by the particle i = (p' i1 , p' i2 , …, p' im ) T , i = 1, 2, …, m', the optimal position p' searched by the swarm g = (p' g1 , p' g2 , …, p' gm ) T , i = 1, 2, …, m', r'1r'2 are random numbers between 0 and 1, and c'1c'2 are the individual learning factor and the swarm learning factor of the particle, respectively; the fitness function is calculated again after updating, and the iteration is stopped when the number of iterations is reached and the current traditional Chinese medicine production process parameter optimization value, i.e. the optimized key process parameter, is output.

[0017] The present application relates to a system for implementing the above method, comprising: a data acquisition and preprocessing module, a key variable selection and quality prediction model construction module, and a traditional Chinese medicine product key process parameter optimization module, wherein: the data acquisition and preprocessing module acquires historical data obtained from industrial field instruments, and obtains a historical data set after missing value processing, outlier detection processing and data standardization; the key variable selection and quality prediction module selects a variable set closely related to the quality variable according to the historical data set and constructs a quality prediction model, thereby eliminating redundant information, reducing the difficulty of quality prediction modeling and the complexity of the model, learning the nonlinear function relationship between the key parameters and the quality to achieve accurate quality prediction; the traditional Chinese medicine product key process parameter optimization module optimizes the key process parameters according to the quality prediction model, with the mean square error of the quality prediction as the fitness function, to achieve the best process parameter combination in the production process of traditional Chinese medicine products and ensure the production of high-quality traditional Chinese medicine products.

[0018] Technical effects

[0019] The application integrates key parameter selection and quality prediction model construction of traditional Chinese medicine product production through key variable selection and quality prediction model construction module. Compared with the prior art, the application realizes data-driven key parameter selection in traditional Chinese medicine industrial production, avoids the blindness of subjective selection of workers, pre-selects key parameters affecting product quality from numerous parameters, establishes an accurate quality prediction model, optimizes process parameters, and obviously improves the accuracy of the prediction model. The particle swarm optimization of key process parameters avoids the subjectivity and blindness of artificial experience setting. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 A PM-AdaBoost model construction flowchart of the application is shown in the figure.

[0021] Figure 2 An AdaBoost algorithm flowchart is shown in the figure.

[0022] Figure 3 An experimental comparison chart of quality prediction results of the embodiment is shown in the figure.

[0023] Figure 4 An iteration curve of process parameter optimization of the embodiment is shown in the figure. DETAILED DESCRIPTION

[0024] As shown in the figure, a data-driven traditional Chinese medicine manufacturing process parameter optimization method related to the embodiment comprises: Figure 1 Step A: collect data and perform data preprocessing. In the embodiment, 13 kinds of traditional Chinese medicine manufacturing process parameters including feeding amount, wall rotating speed, homogenizing rotating speed, vacuum degree 1, vacuum degree 2, stirring rotating speed, temperature rising temperature, viscosity value, holding temperature, standing time, spraying body temperature, left and right glue storage box temperature, and rolling mold revolution are collected, and one traditional Chinese medicine product quality is collected. In order to be suitable for subsequent calculation and analysis, the collected industrial data set is supplemented with missing values, outliers are removed, and standardized processing is performed.

[0025] Step B: initialize the particle swarm population. When solving the key process parameter selection and quality prediction model construction problem by using the particle swarm optimization algorithm, first determine the process parameter search space, set the number of particles, and initialize the position and speed of the particles in the search space.

[0026] Step C: update the speed and position of the particles in the process parameter search space each time, calculate the fitness function again after updating, judge whether the iteration number is reached, if the iteration number is reached, stop iteration, otherwise continue iteration, and output the current traditional Chinese medicine production process parameter set and quality prediction model PM-AdaBoost when the iteration is ended.

[0027]

[0028] ​Step D: Based on the key process parameters and the quality prediction model, the process parameters are optimized by using the particle swarm optimization algorithm with the mean square error of quality prediction as the fitness function, the process parameter search space is determined, the number of particles is set, and the initial position and velocity of the particle swarm are initialized.

[0029] Step E: Update the velocity and position of the particle in the process parameter search space each time, calculate the fitness function MSE again after updating, judge whether the iteration number is reached, if the iteration number is reached, stop iteration, otherwise continue iteration, and output the optimization result of the key process parameters of traditional Chinese medicine production process when the iteration is ended.

[0030] As shown in Figure 3 , the MSE calculated by the model of the present embodiment is 0.01738, which is reduced by 0.14412 compared with the MSE of 0.1615 without process parameter selection.

[0031] Based on the key process parameters and the quality prediction model, the process parameters are optimized by using the particle swarm optimization algorithm with the mean square error of quality prediction as the fitness function, and the optimization result of the key process parameters is output.

[0032] As shown in Figure 4 and Table 1, the optimization result of the final optimization of each key process parameter is obtained by particle swarm iteration optimization.

[0033] Table 1

[0034]

[0035] Through specific actual experiments, the computer is configured as Intel(R) Core(TM) i7-8700 CPU @ 3.20 GHz 32.00 G RAM and runs in Python 3.7. Based on 13 process parameters and traditional Chinese medicine product quality samples provided by a traditional Chinese medicine production factory in Beijing, the data is shown in Table 2.

[0036] Table 2

[0037]

[0038]

[0039] Compared with the prior art, the application as a whole realizes data-driven selection of key parameters in traditional Chinese medicine industrial production, uses the maximum mutual information coefficient between traditional Chinese medicine manufacturing process parameters and quality variables and the mean square error of a quality prediction model as selection criteria, avoids the blindness of worker subjective selection, selects key parameters which have important influence on quality prediction from numerous parameters, and constructs a quality prediction model with high prediction accuracy, intelligently optimizes the traditional Chinese medicine manufacturing process parameters based on the constructed quality prediction model, obtains the optimal setting value of the key process parameters, and improves the rationality and reliability of the traditional Chinese medicine product production process decision.

[0040] The above specific embodiments can be adjusted in different ways by those skilled in the art without departing from the principles and purposes of the application, the protection scope of the application is subject to the claims and is not limited by the above specific embodiments, and each implementation scheme within the scope is subject to the constraints of the application.

Claims

1. A data-driven method for optimizing process parameters in the manufacturing process of traditional Chinese medicine, characterized in that, By calculating the maximum mutual information coefficient between historical process parameters and product quality in the production process of traditional Chinese medicine products, and constructing a quality prediction model, the fitness function is calculated based on the maximum mutual information coefficient and the mean square error of the quality prediction model. Through particle swarm optimization, calculating the fitness function, and iteratively updating the velocity and position of particles in the process parameter search space, the key process parameters and optimized quality prediction model in the production process of traditional Chinese medicine products are obtained. Furthermore, the mean square error of the optimized quality prediction model is used as the fitness function, and the optimized key process parameters are obtained using the particle swarm optimization algorithm. The historical process parameters of the production process of the traditional Chinese medicine products have been preprocessed, specifically including: missing value processing, outlier processing and data standardization. The maximum mutual information coefficient, which measures the correlation between process parameters and product quality in the production of traditional Chinese medicine, is obtained as follows: (The process parameters are then...) and product quality Discretized in a two-dimensional space and represented using a scatter plot, the current two-dimensional space is... , The directions are divided into a certain number of intervals, and then the distribution of the current scatter points within each square is examined to determine the maximum mutual information coefficient. ,in: Indicate process parameters and the quality of traditional Chinese medicine products Mutual information, among which, , It is mass Y= The probability density function, Process parameters The probability density function; The historical process parameters of the traditional Chinese medicine product manufacturing process include: process parameters of the traditional Chinese medicine production process. and the quality of traditional Chinese medicine products , This indicates the process parameters in the production of traditional Chinese medicine products, among which... , represents a vector composed of samples collected from a certain parameter in the production process of traditional Chinese medicine. Indicates the number of process parameters. Indicates the number of samples. , representing the vector formed by the quality samples of traditional Chinese medicine products.

2. The data-driven method for optimizing process parameters in the manufacturing process of traditional Chinese medicine according to claim 1, characterized in that, The missing value handling mentioned above, namely the single imputation method, refers to using the mean of sample data with the same type of parameters to impute missing values ​​in order to reduce the impact of incomplete data on production process modeling. The outlier handling mentioned above refers to: using the 3σ method to remove outlier data; The data standardization process mentioned above refers to the use of the maximum and minimum value method for data standardization.

3. The data-driven method for optimizing process parameters in the manufacturing process of traditional Chinese medicine according to claim 1, characterized in that, The initialization of the particle swarm refers to the following: when using the particle swarm optimization algorithm to solve the problem of selecting key process parameters and constructing a quality prediction model, firstly, in the process parameter search space of the particles, the number of particles m is set, and the position and velocity of each particle in the search space are randomly assigned.

4. The data-driven method for optimizing process parameters in the manufacturing process of traditional Chinese medicine according to claim 1, characterized in that, The velocity and position of the updated particles in the process parameter search space are specifically as follows: , ,in: As weights, the velocity of the particle swarm , Represents particles In the t-th iteration A 3D position vector, representing the particle's position in space. The optimal position found by the particle Optimal position found by the group search r1 and r2 are random numbers between 0 and 1, and c1 and c2 are the individual particle learning factor and the group learning factor, respectively. After updating, the fitness function is recalculated. When the number of iterations is reached, the iteration stops and the current set of process parameters for the TCM production process and the optimized quality prediction model PM-AdaBoost are output.

5. The data-driven method for optimizing process parameters in the manufacturing process of traditional Chinese medicine according to claim 1, characterized in that, The aforementioned use of particle swarm optimization (PSO) to obtain optimized key process parameters refers to: using the mean squared error between the actual quality value of traditional Chinese medicine and the predicted value of the optimized quality prediction model as the fitness function, initializing the particle swarm population, that is, in the search space of key process parameters, setting the number of particles and randomly assigning values ​​to the position and velocity of each particle in the search space, and after multiple iterations of the particles, that is, updating the velocity and position of the particles in the process parameter search space each time, specifically: , ,in: As weights, the velocity of the particle swarm , Represents particles In the t-th iteration A 3D position vector, representing the particle's position in space. Optimal position found by the particle Optimal position found by the group search r'1 and r'2 are random numbers between 0 and 1, and c'1 and c'2 are the individual particle learning factor and the group learning factor, respectively. After updating, the fitness function is calculated again. When the number of iterations is reached, the iteration stops and the optimized values ​​of the current process parameters of the TCM production process are output, that is, the optimized key process parameters.

6. A system for implementing the data-driven process parameter optimization method for traditional Chinese medicine manufacturing as described in any one of claims 1-5, characterized in that, include: The system comprises three modules: a data acquisition and preprocessing module, a key variable selection and quality prediction model construction module, and a key process parameter optimization module for traditional Chinese medicine (TCM) products. Specifically: the data acquisition and preprocessing module collects historical data from industrial field instruments, processes it for missing values ​​and outliers, and standardizes the data to obtain a historical dataset; the key variable selection and quality prediction module selects a set of variables closely related to quality variables based on the historical dataset and constructs a quality prediction model, thereby eliminating redundant information, reducing the difficulty and complexity of quality prediction modeling, and learning the nonlinear functional relationship between key parameters and quality to achieve accurate quality prediction; the key process parameter optimization module for TCM product production optimizes key process parameters based on the quality prediction model, using the mean square error of quality prediction as the fitness function, to achieve the optimal combination of process parameters in the production process of TCM products, ensuring high-quality TCM product production.