A process optimization method based on traditional Chinese medicine production data mining

By obtaining the optimal data points for traditional Chinese medicine production through data sampling, evaluation and iterative calculations, the difficulty of determining parameter regions in complex data space is solved, and the optimization and precise control of traditional Chinese medicine production processes are achieved.

CN115525697BActive Publication Date: 2025-10-03GUANGDONG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211273092.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2025-10-03
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

In the production process of traditional Chinese medicine, existing technologies find it difficult to determine the effective parameter area in the complex high-dimensional data space, which makes it difficult to optimize the production process. In addition, the clustering algorithm converges slowly in large-scale data, making it difficult to select the K value, which easily leads to local optimality.

Method used

A data sampling function is used to obtain the initial data set, and the data evaluation model is used to evaluate the effectiveness. A prediction evaluation model is established by combining the support vector machine and kernel ridge regression model. The adaptive optimization function is used to iteratively calculate in the data space of interest to obtain the optimal data points, construct a high-quality data set, and train the convolutional neural network model for precise control.

Benefits of technology

It enables precise exploration of data areas of interest during the production process of traditional Chinese medicine, improves the accuracy and efficiency of data mining, optimizes the production process of traditional Chinese medicine, and enhances the control accuracy of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115525697B_ABST
    Figure CN115525697B_ABST
Patent Text Reader

Abstract

The present invention discloses a process optimization method based on traditional Chinese medicine production data mining, comprising: first, performing data sampling on collected traditional Chinese medicine production data through a data sampling function to obtain an initial data set; second, using a constructed data evaluation model to perform validity evaluation on the sampled data points in the initial data set to obtain an evaluation data set; and, introducing a kernel model in machine learning to perform evaluation training on the evaluation data set to establish a prediction evaluation model; based on the prediction evaluation model, using an adaptive optimization function to perform iterative operations in a data space of interest to obtain optimal data points, and forming a corresponding high-quality data set with a series of obtained optimal data points; finally, training a traditional Chinese medicine production optimization decision model through data in the high-quality data set, and realizing precise control of the traditional Chinese medicine production process based on the traditional Chinese medicine production optimization decision model to achieve optimization of the traditional Chinese medicine production process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of big data analysis and traditional Chinese medicine production technology, and in particular to a process optimization method based on traditional Chinese medicine production data mining. Background Art

[0002] Process parameter analysis is a crucial step in the production of Traditional Chinese Medicine (TCM). During the TCM manufacturing process, process data points generated by each process unit accumulate over the course of each production batch. With the continuous advancement of big data technology, leveraging big data to mine these massive amounts of accumulated data points, extracting underlying patterns and implicit relationships within the process, and better guiding TCM production processes and making informed decisions, has become a crucial research topic. Data mining is a key step in big data analysis, specifically exploring initial data to uncover valuable insights and prepare data for subsequent big data analysis steps, such as model building and data analysis.

[0003] In the research on big data analysis of traditional Chinese medicine production, most of the current research work is to find a model relationship between process parameters and the quality of traditional Chinese medicine production, so as to conduct data analysis, obtain the optimal process parameters, and improve the quality of traditional Chinese medicine production. Chinese invention patent application CN111888788A "A recurrent neural network control method and system suitable for traditional Chinese medicine extraction and concentration" (Zhang Zirui; He Yan; Chen Xuesong; Cai Shuting; Xiong Xiaoming; Zhang Riwei; Teng Xiao; Cao Fan; Xu Yaotong) replaces the traditional PID control system in the traditional Chinese medicine extraction and concentration link with an improved recurrent neural network control system. The collected data samples are used for training to obtain a neural network model, and the network model is used to output the optimal parameters in the traditional Chinese medicine extraction and concentration link. It does not perform data mining analysis on the data samples used for training, and the quality of the data used for model training still has a lot of room for improvement. Chinese invention patent application CN112348360A, "A Traditional Chinese Medicine Production Process Parameter Analysis System Based on Big Data Technology" (Xie Zhijian, Zhang Jinghai, Wang Zhenyu, Zhao Feifei, and Zhang He), utilizes traditional Chinese medicine production data through data mining to establish a BP neural network model to predict traditional Chinese medicine quality indicators using production process parameter data, thereby optimizing the production process. The system uses a clustering algorithm to perform data mining on the collected production process parameter data, removing invalid data from the production process parameter dataset based on the clustering results to improve the dataset quality. However, clustering algorithms converge slowly for large-scale data, and their K value (the number of clusters) is sensitive and difficult to select, especially for multi-category, multi-dimensional datasets, which can easily lead to local optimality.

[0004] For a highly complex chemical process like traditional Chinese medicine production, the characteristic parameter categories in its production data are diverse, and the data space in which the data resides is high-dimensional, making it extremely difficult to identify effective and feasible parameter regions. Using data mining methods to extract valid data from traditional Chinese medicine production data has become a key issue in big data analysis in traditional Chinese medicine production. Summary of the Invention

[0005] The purpose of the present invention is to provide a process optimization method based on traditional Chinese medicine production data mining, to solve the difficulty of the existing technology in determining the feasible area for exploration in complex data space, to achieve precise control of the traditional Chinese medicine production process, and thus to achieve the purpose of optimizing the production process.

[0006] In order to achieve the above tasks, the present invention adopts the following technical solutions:

[0007] A process optimization method based on traditional Chinese medicine production data mining, comprising:

[0008] First, the collected traditional Chinese medicine production data are sampled through the data sampling function to obtain the initial data set; secondly, the constructed data evaluation model is used to evaluate the effectiveness of the sampled data points in the initial data set to obtain the evaluation data set; and the kernel model in machine learning is introduced to evaluate and train the evaluation data set to establish a predictive evaluation model; based on the predictive evaluation model, the adaptive optimization function is used to perform iterative operations in the data space of interest to obtain the optimal data points, and the obtained series of optimal data points constitute the corresponding high-quality data set; finally, the traditional Chinese medicine production optimization decision model is trained with the data in the high-quality data set, and based on the traditional Chinese medicine production optimization decision model, precise control of the traditional Chinese medicine production process is achieved, and optimization of the traditional Chinese medicine production process is achieved.

[0009] Furthermore, the data sampling function is used to sample the collected traditional Chinese medicine production data to obtain an initial data set, including:

[0010] The collected traditional Chinese medicine production data is a multi-dimensional data space Building a data sampling model Among them, s i is the i-th p-dimensional sub-data space x i The sampling step size, is the sampling grid; using the data sampling model For sampling of Chinese medicine production data, the sampling step size s in the data sampling model is set i To determine the sampling grid The size of , thereby determining the amount of sampled data.

[0011] Furthermore, the constructed data evaluation model is used to perform validity evaluation on the sampled data points in the initial data set to obtain an evaluation data set, including:

[0012] Defining a mapping

[0013]

[0014] in, Represents the number field, p represents the dimension of the field; mapping From p-dimensional data space It is represented by the mapping of the sampled data points in to the solution space S;

[0015] The solution space S is defined as:

[0016]

[0017] The solution space S is composed of the tensor product of the classification space η and the target space τ, where:

[0018] The classification space η is used to describe whether the sampled data point is a valid data point or an invalid data point. The classification space η is represented by a set of valid data (valid) and invalid data (invalid). Its data evaluation category is y, which is the sampling data point x through the above mapping The obtained The target space τ is defined as Represents a number field;

[0019] Each sampled data point is defined as an independent target t in the target space τ, t∈τ; the value of the target t reflects the effectiveness of the sampled data point; through the data evaluation model, D init By executing the mapping, we can get the data point evaluation result d(x). The data point evaluation result d(x) is determined by the sampled data point x, the evaluation category y corresponding to the sampled data point x, and the target t. d(x)≡(x,y,t);

[0020] Collect the results of evaluating all sampled data points, which is defined as the evaluation dataset D expl .

[0021] Furthermore, the establishment of the prediction and evaluation model includes:

[0022] An evaluation prediction model ε is constructed based on two independent kernel model estimators, support vector machine (SVM) and kernel ridge regression (KRR). The evaluation prediction model is used to evaluate and predict unevaluated data points in the data space.

[0023] Define the evaluation prediction model ε as the evaluation dataset Dexpl and data space The tensor product is mapped to the solution space S and the prediction result probability space η p The tensor product of is specifically expressed as:

[0024]

[0025] Furthermore, in the prediction and evaluation model, the solution space S contains the prediction and evaluation categories and prediction targets Among them, SVM is used for the results Classification, KRR is used to predict the optimal target parameters

[0026] Based on the kernel model, it is assumed that there is a Mapping φ to feature space F:

[0027]

[0028] Its inner product <φ(x),φ(x')> represents the data space The Gaussian kernel k of the two sampled data points x and x';

[0029] The mapping of the two feature spaces of SVM and KRR is:

[0030]

[0031] and:

[0032]

[0033] φ SVM Represents data space To the classification feature space F S The mapping of φ KRR Represents data space To the target regression feature space F R The mapping of Gaussian kernel k SVM (x,x') and k KRR The hyperparameters in (x,x') are γ SVM and γ KRR ;

[0034] Prediction result probability space η p Contains the effective prediction evaluation category probability obtained under the prediction evaluation condition of the sampled data point x Specifically expressed as:

[0035]

[0036] In the above formula, f(x) represents the prediction result output by the prediction evaluation model, and a and b are the probability model parameters.

[0037] Furthermore, the adaptive optimization function is used to perform iterative operations in the data space of interest to obtain the optimal data point, including:

[0038] Construct an adaptive optimization function U(u,w) to obtain the optimal data point x new ; The adaptive optimization function U(u,w) consists of the optimization vector u and the weight vector w, specifically:

[0039]

[0040] Where ||w||1 represents the l1 norm of w, and the superscript T represents the transpose;

[0041] That is, the optimal data point x new Expressed as:

[0042]

[0043] Furthermore, the optimization vector u consists of three parts, which are specifically defined as:

[0044]

[0045] Among them, U S The predicted probability of the data point class is based on the effective prediction evaluation class probability The Shannon information entropy S is used to represent the reliability of the predicted data point evaluation category, which is specifically expressed as:

[0046]

[0047] in, represents the probability of effective prediction evaluation category;

[0048] U o Predict the credibility of the target, according to the predicted target Establish a target prediction credibility evaluation mechanism with the existing target t, which is specifically expressed as:

[0049]

[0050] In the above formula, the initial data set D is init The maximum and minimum targets obtained in the evaluation are t max and t min , the prediction target obtained in the evaluation prediction model ε Re-divide and normalize it, that is

[0051] Ur is the classification feature space distance, which is specifically expressed as:

[0052]

[0053] Among them, ||·||2 represents the l2 norm, e is the base of natural logarithm, γ C represents hyperparameters;

[0054] The weight vector w is defined as:

[0055]

[0056] The components of the weight vector s≥0, o≥0 and r≥0 respectively determine U S 、U o and U r The degree of influence on data point optimization.

[0057] Furthermore, by constructing the adaptive optimization function U(u,w), the optimal data point x is obtained. new It can be specifically expressed as:

[0058]

[0059] Use the established mapping relationship That is, using the data evaluation model to evaluate x new Evaluate and return x new The evaluation result is d(x new ) to high-quality dataset D expl middle.

[0060] Furthermore, the method of training an accurate and efficient TCM production optimization decision model based on the data in the high-quality data set and achieving precise control of the TCM production process based on the TCM production optimization decision model includes:

[0061] The data in the high-quality data set is the effective production data mined, and these data are used to train the convolutional neural network model; after the network model is trained, the real-time traditional Chinese medicine production data is input into the convolutional neural network model, and the convolutional neural network model outputs the traditional Chinese medicine production data value at the next moment through calculation, that is, the predicted production data; at the same time, the predicted production data is compared with the pre-set ideal production data range for interval judgment and corresponding production decisions are made. If it exceeds the interval range, the corresponding process parameter value is adjusted in the traditional Chinese medicine production process according to the specific exceeding of the ideal interval range, thereby accurately controlling the traditional Chinese medicine production process.

[0062] Compared with the prior art, the present invention has the following technical features:

[0063] The present invention extracts effective data in the traditional Chinese medicine production process through the method of data mining. Aiming at the difficulty of determining the feasible area of ​​data exploration in the unknown data space, the target item is introduced into the data mining algorithm model, and the data mining focus is guided to the data area of ​​interest. An adaptive optimization function is constructed based on the kernel model in machine learning. By training a relatively small number of data sets, data prediction and evaluation are realized in the data subspace of interest, and the optimal data point is obtained through sequential iterative mining. Finally, a series of data points obtained by data mining constitute the corresponding data set. The database composed of these data sets can provide a good data analysis basis for the production optimization decision model established subsequently, thereby accurately controlling the production process and optimizing the traditional Chinese medicine production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 Schematic diagram of the process of establishing a dataset for evaluation;

[0065] Figure 2 This is a comparison diagram of the effects of an embodiment of the method of the present invention. DETAILED DESCRIPTION

[0066] The present invention proposes a process optimization method based on Chinese medicine production data mining. The method performs validity evaluation and screening of Chinese medicine production data in a sequential iterative manner. In order to solve the difficulty of determining a feasible area for data exploration in a complex data space, the present invention constructs a data mining algorithm model to evaluate data quality and mine the optimal data points. First, the collected Chinese medicine production data is sampled by a data sampling function to obtain an initial data set; second, the constructed data evaluation model is used to evaluate the validity of the sampled data points in the initial data set to obtain an evaluation data set; and a kernel model in machine learning is introduced to evaluate and train the evaluation data set to establish a prediction evaluation model; based on the prediction evaluation model, an adaptive optimization function is used to perform iterative operations in the data space of interest to obtain the optimal data points, and the obtained series of optimal data points constitute the corresponding high-quality data set; finally, an accurate and efficient Chinese medicine production optimization decision model is trained based on the data in the high-quality data set, and the Chinese medicine production process is precisely controlled based on the Chinese medicine production optimization decision model, thereby optimizing the Chinese medicine production process.

[0067] Referring to the accompanying drawings, a process optimization method based on traditional Chinese medicine production data mining of the present invention comprises the following specific steps:

[0068] Step 1: Use the data sampling function to sample the collected traditional Chinese medicine production data to obtain the initial data set.

[0069] In order to obtain the initial dataset D init, it is necessary to sample the massive amount of TCM production data (process data) collected. The huge amount of TCM production data is a multi-dimensional data space. (its dimension is set to p), which can be regarded as a one-dimensional sub-data space x i , i=1,2,…,p, the data set x is specifically expressed as:

[0070]

[0071] Here, p represents the dimension of the data space, that is, the number of one-dimensional sub-data spaces.

[0072] In the present invention, the sampling method of Chinese medicine production data is to establish a data sampling model Among them, s i is the i-th p-dimensional sub-data space x i The sampling step size, is the sampling grid; using the data sampling model Sampling the production data of traditional Chinese medicine, that is, in the p-dimensional data space Sampling data points in the data sampling model; by setting the sampling step s in the data sampling model i To determine the sampling grid The size of , thereby determining the amount of sampled data.

[0073] Through data sampling model Data sampling is performed on the traditional Chinese medicine production data (i.e., data set x), and the n data points obtained by sampling are defined as the initial data set D init , specifically expressed as:

[0074] D init =(x1,x2,x3,...,x n )

[0075] Among them, x n is the nth sample data in the initial data set.

[0076] In order to more clearly illustrate the above data sampling mechanism, an example is given to illustrate it in detail.

[0077] Assume that the obtained traditional Chinese medicine production data is a two-dimensional data set x', which is specifically expressed as follows:

[0078]

[0079] And the data space where the data is located It consists of one-dimensional sub-data spaces x'1 = (-3, 3) and x'2 = (-3, 3), which can be expressed as:

[0080]

[0081] In the data space In the data sampling model, In the data sampling model, set the sub-data space x' of each dimension i The sampling step size s i ' are all 1, that is, to build a 6×6 sampling grid Realize the data space A sampling of 49 data points.

[0082] Step 2: Use the constructed data evaluation model to evaluate the effectiveness of the sampled data points in the initial data set to obtain an evaluation data set.

[0083] Establish a data evaluation model and evaluate the initial data set D init The sampled data points x1,x2,x3,...,x n The evaluation is performed and the dataset after evaluation is defined as D expl .

[0084] The data evaluation model is constructed as follows:

[0085] Defining a mapping

[0086]

[0087] in, Represents the number field, p represents the dimension of the field; mapping From p-dimensional data space It is represented by the mapping of the sampled data points x in to the solution space S.

[0088] The solution space S is defined as:

[0089]

[0090] The solution space S is composed of the tensor product of the classification space η and the target space τ, where:

[0091] The classification space η is used to describe whether the sampled data points are valid data points or invalid data points. The classification space η is defined as:

[0092] η={valid,invalid}

[0093] The classification space η is represented by a set of valid data (valid) and invalid data (invalid), and its data evaluation category is y, because y is the sampling data point x through the above mapping relationship So it can be expressed as The specific definitions are as follows:

[0094]

[0095] Among them, ≡ represents identity, and η represents the classification space.

[0096] The target space τ is defined as:

[0097]

[0098] in, Represents a number field.

[0099] Each sampled data point is defined as an independent target t in the target space τ, t∈τ; the size of the target t value reflects the effectiveness of the sampled data point; one of the prerequisites for obtaining the target interval range of the valid data point is to solve the maximum value of the target t under the set extreme value conditions.

[0100] Through data evaluation model, D init By executing the mapping, we can get the data point evaluation result d(x). The data point evaluation result d(x) is determined by the sampled data point x, the evaluation category y corresponding to the sampled data point x, and the target t. It can be expressed in the identity form:

[0101] d(x)≡(x,y,t)

[0102] Collect the results of evaluating all sampled data points, which is defined as the evaluation dataset D expl :

[0103] D expl ={d(x1),d(x2),d(x3)...,d(x n )}

[0104] Step 3: Introduce the kernel model in machine learning to establish a predictive evaluation model by evaluating and training the evaluation data set; then, based on the predictive evaluation model, use the adaptive optimization function to perform iterative operations in the data space of interest to obtain the optimal data points, and form the corresponding high-quality data set with a series of optimal data points.

[0105] Step 3.1, evaluate the dataset D expl Conduct evaluation training and build an evaluation prediction model ε.

[0106] An evaluation prediction model ε is constructed based on two independent kernel model estimators, support vector machine (SVM) and kernel ridge regression (KRR); the evaluation prediction model is used to evaluate and predict unevaluated data points in the data space.

[0107] Define the evaluation prediction model ε as the evaluation dataset D expl and data space The tensor product is mapped to the solution space S and the prediction result probability space η p The tensor product of is specifically expressed as:

[0108]

[0109] In order to further demonstrate the mechanism of the above model, the following analysis and explanation are given:

[0110] The solution space S contains the prediction evaluation categories ( η is the classification space) and the prediction target (

[0111] τ is the target space); SVM is used for the results Classification, KRR is used to predict the optimal target parameters

[0112] Based on the kernel model, it is assumed that there is a Mapping φ to feature space F:

[0113]

[0114] Its inner product <φ(x),φ(x')> represents the data space The Gaussian kernel k of the two sampled data points x and x' is specifically expressed as:

[0115] <φ(x),φ(x′)>=k(x,x′)

[0116]

[0117] Among them, γ is a hyperparameter of the feature space metric.

[0118] Since the feature space of SVM is not necessarily the same as the feature space of KRR, based on the above model derivation, there is a mapping between the two feature spaces of SVM and KRR, namely:

[0119]

[0120] and:

[0121]

[0122] φ SVM Represents data space To the classification feature space F S The mapping of φ KRR Represents data space To the target regression feature space F R The Gaussian kernel k SVM(x,x') and k KRR The hyperparameters in (x,x') are γ SVM and γ KRR .

[0123] Prediction result probability space η p Contains the effective prediction evaluation category probability obtained under the prediction evaluation condition of the data point x Specifically expressed as:

[0124]

[0125] In the above formula, f(x) represents the prediction result output by the prediction evaluation model, and a and b are the probability model parameters.

[0126] In short, the evaluation prediction model ε is based on the previous evaluation data set D expl The evaluation results of the predicted data points and their associated probabilities in the data space of interest.

[0127] Step 3.2: Use the adaptive optimization function to iterate in the data space of interest to obtain the optimal data point x new .

[0128] Construct an adaptive optimization function U(u,w) to obtain the optimal data point x new The adaptive optimization function U(u,w) constructed by the present invention is composed of the optimization vector u and the weight vector w, and its specific expression is:

[0129]

[0130] Here, ||w||1 represents the l1 norm of w, and the superscript T represents the transpose.

[0131] That is, the optimal data point x new It can be expressed as:

[0132]

[0133] In order to further demonstrate the mechanism of the above model, the following analysis and explanation are given:

[0134] The optimization vector u consists of three parts, which are specifically defined as:

[0135]

[0136] Among them, U S The predicted probability of the data point class is based on the effective prediction evaluation class probability The Shannon information entropy S is used to represent the reliability of the predicted data point evaluation category, which is specifically expressed as:

[0137]

[0138] U o Predict the credibility of the target, according to the predicted target Establish a target prediction credibility evaluation mechanism with the existing target t, which is specifically expressed as:

[0139]

[0140] In the above formula, the maximum target and minimum target t obtained under extreme conditions are used max and t min , the prediction target obtained in the evaluation prediction model ε Re-divide and normalize it, that is Prediction target Normalization is shown below:

[0141]

[0142] Among them, the maximum target and the minimum target t max and t min The initial data set D in step 2 init The constraints obtained in the evaluation are that the data evaluation category is valid, that is, sty = valid, the maximum target and the minimum target t max and t min Specifically expressed as:

[0143]

[0144] U r is the classification feature space distance, which is specifically expressed as:

[0145]

[0146] Among them, ||·||2 represents the l2 norm, e is the base of natural logarithm, γ C Represents a hyperparameter of the kernel estimator.

[0147] By definition, the sampled data point x and D expl The more similar the nearest neighbors in U are, the r Therefore, U r Can ensure that the space region of interest, exploring new data points x new .

[0148] The weight vector w is defined as:

[0149]

[0150] The components of the weight vector s≥0, o≥0 and r≥0 respectively determine U S 、U o and U r The impact on data point optimization, that is, w determines the degree of data mining exploration.

[0151] In summary, an adaptive optimization function U is constructed based on the optimization vector u and the weight vector w. The optimization vector u consists of three components with different meanings. These components are weighted by the weights s, o, and r in the weight vector w. Therefore, the degree of data mining can be controlled by adjusting the weights.

[0152] Step 3.3, evaluate the optimal data point x new , and obtain high-quality data sets.

[0153] By constructing the adaptive optimization function U(u,w), the optimal data point x is obtained new It can be specifically expressed as:

[0154]

[0155] Use the mapping established in step 2 That is, using the data evaluation model to evaluate x new Evaluate and return x new The evaluation result is d(x new ) to high-quality dataset D expl middle.

[0156] Step 3.4: In each successive iteration from step 3.1 to step 3.3, an optimal data x is obtained. new , and d(x new )Return D expl , therefore, the results collected from n iterations make it possible to update the dataset D expl When the total number of evaluation data points reaches the preset value N, the iteration ends, the data mining is completed, and a high-quality data set D is obtained. expl .

[0157] Step 4: Train an accurate and efficient TCM production optimization decision model using data from high-quality data sets, and achieve precise control of the TCM production process and optimization of the TCM production process based on the TCM production optimization decision model.

[0158] High-quality dataset D explThe data in the data is the effective production data mined, and these data are used to train the convolutional neural network model; after the network model is trained, the real-time traditional Chinese medicine production data is input into the convolutional neural network model, and the convolutional neural network model outputs the traditional Chinese medicine production data value at the next moment through calculation, that is, the predicted production data; at the same time, the predicted production data is compared with the pre-set ideal production data range for interval judgment and corresponding production decisions are made. If it exceeds the interval range, the corresponding process parameter value is adjusted in the traditional Chinese medicine production process according to the specific exceeding of the ideal interval range, so as to achieve precise control of the traditional Chinese medicine production process and ultimately achieve the purpose of optimizing the production process.

[0159] Example:

[0160] Taking the common alcohol precipitation process in Chinese medicine production as an example, adding alcohol flow rate, ethanol concentration, ethanol temperature, and alcohol precipitation supernatant volume in the alcohol precipitation process is a key process parameter. Using the data mining model proposed by the present invention for application in Chinese medicine production, data mining exploration is carried out on the process parameters in the alcohol precipitation process, a production decision-making method is established, and accurate control of the alcohol precipitation process is achieved, thereby optimizing the alcohol precipitation process. In conjunction with Figure 1, the implementation method of the present invention is described.

[0161] 1) First, input the data points of the key parameters of the alcohol precipitation process, and its data space is determined by the alcohol addition flow rate F V ∈[a,b], ethanol concentration C E ∈[c,d], ethanol temperature T E ∈[e,f], alcohol precipitation supernatant volume V SL ∈[g,h] is a four-dimensional data space.

[0162] Using parameter vector x:

[0163]

[0164] The data space can therefore be written as

[0165]

[0166] Using the data sampling function established in step 1 In the parameter subspace Set the grid sampling window in Data Space The input data points in are sampled to obtain the initial data set D init =D(x1,x2,x3,...,x n ).

[0167] In addition, in order to better conduct data mining and exploration of the process parameters in the alcohol precipitation process, the ratio of the concentrate and ethanol fully mixed during the alcohol precipitation process is taken as the target t.

[0168]

[0169] Where m0 is the amount of concentrate used in the alcohol precipitation; m1 is the amount of ethanol used; m2 is the total mass of the resulting supernatant; S0 is the total solids content of the concentrate; S2 is the total solids content of the supernatant. H is the retention rate of the index component.

[0170] 2) Data evaluation

[0171] The obtained initial data set D of the alcohol precipitation process init Evaluate the data points in step 2, evaluate the model using the data in step 2, and perform the mapping Get the data point evaluation result d(x)

[0172] For mapping The data evaluation category y and target t are expressed as:

[0173]

[0174] The optimal data point is obtained by solving the maximum value of the target t. The evaluated data set D is obtained. expl .

[0175] 3) Exploration and optimization

[0176] According to step 3, an evaluation prediction model ε is constructed based on two independent kernel model estimators, support vector machine (SVM) and kernel ridge regression (KRR). Based on the prediction results obtained by the evaluation prediction model ε, the optimal data point x in the alcohol precipitation process is obtained by constructing an adaptive optimization function U(u,w) new .

[0177] 4) Evaluate x new

[0178] Use the mapping relationship established in step 2 That is, using the data evaluation model to evaluate x new Evaluate and return the evaluation result d(x new ).

[0179] 5) Iterative calculation

[0180] D expl ∪d(x new )→D expl

[0181] Update D expl ;

[0182] 6) When n=N, end data mining and output D expl , update the data mining library.

[0183] 7) Establish an optimization decision model for the alcohol precipitation process based on data mining

[0184] The mined valid data from the alcohol precipitation process is used to train a convolutional neural network model. Once trained, the model calculates and outputs the next production value for the alcohol precipitation process, known as the predicted production data. Simultaneously, the predicted production data is compared with a pre-set ideal production data range and the corresponding alcohol precipitation production decision is made. If the predicted production data falls outside the ideal range, the relevant process parameters in the alcohol precipitation process are adjusted based on the specific reasons for exceeding the ideal range: alcohol addition flow rate, ethanol concentration, ethanol temperature, and alcohol precipitation supernatant volume. This achieves precise control of the alcohol precipitation process and optimizes it.

[0185] In order to prove the effectiveness of the process optimization method based on Chinese medicine production data mining proposed in the present invention, for the above-mentioned alcohol precipitation process, the method proposed in the present invention and the traditional method (i.e., using the original production data of the alcohol precipitation process) were used to establish an alcohol precipitation process optimization decision model, and comparative tests were carried out. The experimental results are as follows: Figure 2 As shown, under stable model convergence conditions, the optimization decision accuracy of the proposed method is approximately 98%, while that of the traditional method is approximately 83%. The proposed method surpasses the traditional method in both model convergence and optimization decision accuracy. Experimental results demonstrate that the proposed process optimization method based on Traditional Chinese Medicine production data mining has excellent accuracy and efficiency.

Claims

1. A process optimization method based on traditional Chinese medicine production data mining, characterized in that: include: First, the data sampling function is used to sample the collected traditional Chinese medicine production data to obtain the initial data set, including: Input the data points of the key parameters of the alcohol precipitation process, and its data space is determined by the alcohol addition flow rate F V ∈[a,b], ethanol concentration C E ∈[c,d], ethanol temperature T E ∈[e,f], alcohol precipitation supernatant volume V SL ∈[g,h] constitutes a four-dimensional data space; Using parameter vector x: Data space is written Using data sampling functions In the parameter subspace Set the grid sampling window in Data Space The input data points in are sampled to obtain the initial data set D init =D(x1,x2,x3,...,x n ); In order to better explore the process parameters in the alcohol precipitation process through data mining, the ratio of the concentrate and ethanol fully mixed in the alcohol precipitation process is taken as the target t, that is, Wherein, m0 is the amount of concentrate used in alcohol precipitation; m1 is the amount of ethanol used; m2 is the total mass of the obtained supernatant; S0 is the total solid mass content of the concentrate; S2 is the total solid content in the supernatant; H is the retention rate of the index component; Secondly, the constructed data evaluation model is used to evaluate the effectiveness of the sampled data points in the initial data set to obtain an evaluation data set. In addition, the kernel model in machine learning is introduced to evaluate and train the evaluation data set to establish an evaluation prediction model. The evaluation prediction model is constructed based on two independent kernel model estimators: support vector machine and kernel ridge regression. Based on the prediction results obtained by the evaluation prediction model, the adaptive optimization function U(u,w) constructed by the optimization vector u and the weight vector w is used to obtain the optimal data point x in the alcohol precipitation process. new , and the obtained series of optimal data points constitute the corresponding high-quality data set; Finally, an alcohol precipitation process optimization decision model is trained using data from high-quality data sets. After the model is trained, the convolutional neural network model calculates and outputs the production data value of the alcohol precipitation process at the next moment, that is, the predicted production data. At the same time, the predicted production data is compared with the pre-set ideal production data range to make an interval judgment and make a corresponding alcohol precipitation production decision. If the interval range is exceeded, the relevant process parameter values ​​in the alcohol precipitation process of traditional Chinese medicine are adjusted according to the specific range that exceeds the ideal interval: alcohol addition flow rate, ethanol concentration, ethanol temperature, and volume of alcohol precipitation supernatant. This achieves precise control of the alcohol precipitation process of traditional Chinese medicine and optimizes the alcohol precipitation process of traditional Chinese medicine.

2. The process optimization method based on traditional Chinese medicine production data mining according to claim 1, characterized in that, The constructed data evaluation model is used to perform validity evaluation on the sampled data points in the initial data set to obtain an evaluation data set, including: Defining a mapping in, Represents the number field, p represents the dimension of the field; mapping From p-dimensional data space It is represented by the mapping of the sampled data points in to the solution space S; The solution space S is defined as: The solution space S is composed of the tensor product of the classification space η and the target space τ, where: The classification space η is used to describe whether the sampled data point is a valid data point or an invalid data point. The classification space η is represented by a set of valid data and invalid data. Its data evaluation category is y, which is the sampling data point x through the above mapping The obtained The target space τ is defined as Represents a number field; Each sampled data point is defined as an independent target t in the target space τ, t∈τ; the value of the target t reflects the effectiveness of the sampled data point; through the data evaluation model, D init By executing the mapping, we can get the data point evaluation result d(x). The data point evaluation result d(x) is determined by the sampled data point x, the evaluation category y corresponding to the sampled data point x, and the target t. d(x)≡(x,y,t); Collect the results of evaluating all sampled data points, which is defined as the evaluation dataset D expl .

3. The process optimization method based on traditional Chinese medicine production data mining according to claim 2, characterized in that, The establishment of the evaluation prediction model comprises: An evaluation prediction model ε is constructed based on two independent kernel model estimators, support vector machine (SVM) and kernel ridge regression (KRR). The evaluation prediction model is used to evaluate and predict unevaluated data points in the data space. Define the evaluation prediction model ε as the evaluation dataset D expl and data space The tensor product is mapped to the solution space S and the prediction result probability space η p The tensor product of is specifically expressed as:

4. The process optimization method based on traditional Chinese medicine production data mining according to claim 3, characterized in that, In the evaluation prediction model, the solution space S contains the prediction evaluation categories and prediction targets Among them, SVM is used for the results Classification, KRR is used to predict the optimal target parameters Based on the kernel model, it is assumed that there is a Mapping φ to feature space F: Its inner product <φ(x),φ(x')> represents the data space The Gaussian kernel k of the two sampled data points x and x'; The mapping of the two feature spaces of SVM and KRR is: and: φ SVM Represents data space To the classification feature space F S The mapping of φ KRR Represents data space To the target regression feature space F R The mapping of Gaussian kernel k SVM (x,x') and k KRR The hyperparameters in (x,x') are γ SVM and γ KRR ; Prediction result probability space η p Contains the effective prediction evaluation category probability obtained under the prediction evaluation condition of the sampled data point x Specifically expressed as: In the above formula, f(x) represents the prediction result output by the evaluation prediction model, and a and b are the probability model parameters.

5. The process optimization method based on traditional Chinese medicine production data mining according to claim 1, characterized in that, Use the adaptive optimization function to perform iterative operations in the data space of interest to obtain the optimal data point, including: Construct an adaptive optimization function U(u,w) to obtain the optimal data point x new ; The adaptive optimization function U(u,w) consists of the optimization vector u and the weight vector w, specifically: Where ||w||1 represents the l1 norm of w, and the superscript T represents the transpose; That is, the optimal data point x new Expressed as:

6. The process optimization method based on traditional Chinese medicine production data mining according to claim 1, characterized in that, The optimization vector u consists of three parts, which are specifically defined as: Among them, U S The predicted probability of the data point class is based on the effective prediction evaluation class probability The Shannon information entropy S is used to represent the reliability of the predicted data point evaluation category, which is specifically expressed as: in, represents the probability of effective prediction evaluation category; U o Predict the credibility of the target, according to the predicted target Establish a target prediction credibility evaluation mechanism with the existing target t, which is specifically expressed as: In the above formula, the initial data set D is init The maximum and minimum targets obtained in the evaluation are t max and t min , the prediction target obtained in the evaluation prediction model ε Re-divide and normalize it, that is U r is the classification feature space distance, which is specifically expressed as: Among them, ||·||2 represents the l2 norm, e is the base of natural logarithm, γ C represents hyperparameters; The weight vector w is defined as: The components of the weight vector s≥0, o≥0 and r≥0 respectively determine U S 、U o and U r The degree of influence on data point optimization.

7. The process optimization method based on traditional Chinese medicine production data mining according to claim 1, characterized in that: By constructing the adaptive optimization function U(u,w), the optimal data point x is obtained new Specifically expressed as: Use the established mapping relationship That is, using the data evaluation model to evaluate x new Evaluate and return x new The evaluation result is d(x new ) to high-quality dataset D expl middle.

Citation Information

Patent Citations

  • Recurrent neural network control method and system suitable for traditional Chinese medicine extraction and concentration

    CN111888788A

  • Traditional Chinese medicine production process parameter analysis system based on big data technology

    CN112348360A

  • LS-SVM model-based ultrasonic extraction process optimization method for effective components of salvia miltiorrhiza

    CN110988153A

  • Clock tree comprehensive optimal strategy prediction method, system and application

    CN113505562A