A batch production process monitoring method based on neighborhood difference feature analysis and extraction
The reference data is obtained through the time neighborhood window and the differential characteristics are analyzed, and the problems of inequality and time-varying of penicillin production batch data are solved, real-time monitoring and abnormal detection of penicillin production batches are realized.
Patent Information
- Application Number
- CN202210650166.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-30
AI Technical Summary
The inequality and time-varying characteristics of penicillin production batch data make it difficult for the prior art to achieve accurate process monitoring and abnormal detection.
Reference data is obtained through the time neighborhood window, and the difference characteristics of new production batch data are analyzed in real time, and the difference characteristics directly used for process monitoring are extracted.
Real-time monitoring of penicillin production batches is achieved, abnormalities or failures can be detected in a timely manner, and challenges of data inequality and time-varying characteristics.
Smart Images

Figure CN114967624B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for monitoring the operating state of an intermittent process, and particularly to a batch production process monitoring method based on the analysis and extraction of neighborhood difference features. Background Art
[0002] With the rapid development of biotechnology, the intermittent process has become an important production mode in modern industry, and the penicillin fermentation process belongs to a typical intermittent process, with characteristics such as multi-stage, time-varying, and multi-batch. Penicillin fermentation is a process completed intermittently according to production batches, and its operation is mainly divided into two stages. The first stage is to cultivate a large number of penicillin-producing bacteria in a culture tank. When the reproduction of the bacteria reaches a certain level, it enters the second stage where the bacteria produce penicillin. In the second stage, glucose material needs to be continuously added to the reactor to promote the activity of the bacteria. For the entire production process of penicillin, many factors such as the tank temperature, pH value in the tank, stirring power, and ventilation volume have important impacts on the processing process and product quality.
[0003] In order to improve the quality of penicillin products, it is necessary to monitor the entire production process of penicillin and timely detect abnormal or faulty penicillin production batches. However, for such a complex process as penicillin production, it is difficult to establish an accurate mechanism model. Therefore, in the past few years, the penicillin batch production process monitoring has been implemented through a sampling data-driven approach. This is mainly due to the wide application of the Distributed Control System (DCS) in penicillin production, making the real-time measurement and data transmission of sampling data easier. Since the penicillin production specifically includes two stages, the different operations in the two stages result in obvious differences in the sampling data. Moreover, the difference in the production time length of each batch of penicillin leads to unequal numbers of sampling data for each batch, that is, the problem of unequal data length. Therefore, there are relatively large technical implementation difficulties in conducting data-driven process monitoring for penicillin production.
[0004] In addition, considering the multi-stage time-varying characteristics of penicillin production batch data, using a fixed-point feature extraction mechanism or model will gradually reduce the abnormal detection sensitivity of the corresponding method. It is worth emphasizing that from the perspective of process monitoring, performing feature analysis and extraction on multi-batch data is a necessary means to achieve process monitoring, but the real purpose is to detect faults or abnormalities in the process monitoring. In order to address the problem of unequal data length of penicillin production batches and at the same time consider the time-varying characteristics of batch data, the feature analysis and extraction centered on process monitoring should have self-adaptive characteristics, that is, it can be carried out immediately for newly sampled online data instead of being fixed. Summary of the Invention
[0005] The main technical problem to be solved by the present invention is: in view of the variable-length and time-varying problems of penicillin production batch data, how to obtain reference data through a time neighborhood window and instantaneously analyze and extract differential features directly used for process monitoring. Specifically speaking, the method of the present invention first obtains reference data within a similar time neighborhood from the reference batch matrix under multiple normal batches through a time neighborhood window, and instantaneously analyzes the difference between the data vector of the newly produced batch and the reference data vector, so as to directly monitor the change range of the differential feature to realize real-time monitoring of whether there are abnormalities or faults in the penicillin production batch.
[0006] The technical solution adopted by the method of the present invention to solve the above problems is: a batch production process monitoring method based on neighborhood differential feature analysis and extraction, including the following steps:
[0007] Step (1): Obtain a sample data set of J penicillin normal production batches from the historical database of penicillin production batches, and respectively form corresponding batch matrices X1, X2,..., X j ; where, the batch matrix corresponding to the j-th penicillin normal production batch is specifically composed of N j 10×1-dimensional data vectors, j ∈ {1, 2,..., J}, represents a 10×N j dimensional real number matrix, R represents the set of real numbers, and the first column vector in X j is the data vector at the first sampling moment of the j-th penicillin normal production batch, and the N j -th column vector in X j is the data vector at the N j -th sampling moment of the j-th penicillin normal production batch.
[0008] It should be noted that the arrangement order of the 10 data in the data vector at each sampling moment is: ventilation rate, stirring power, glucose feeding temperature, glucose feeding rate, coolant flow rate, acid-base flow rate, reactor temperature, pH value, glucose concentration, and penicillin concentration.
[0009] Step (2): Set the reference length of the time neighborhood window to be equal to L, start the production of the latest penicillin batch, and obtain the sample data at each sampling moment; where, L is a positive integer.
[0010] Step (3): Construct a 10×1-dimensional data vector x i from the 10 sample data at the latest sampling moment, and respectively obtain reference data vectors from the batch matrices X1, X2,..., X j , so as to merge and form a reference data matrix X t, specifically as shown in steps (3.1) to (3.4).
[0011] Step (3.1): After recording the current sampling moment as the ζ-th sampling moment of the latest production batch of penicillin, initialize j = 2.
[0012] Step (3.2): Determine whether ζ is greater than L; if not, form a reference data matrix X by combining the column vectors from the 1st column to the (ζ + L)-th column in X1 t ; if so, form a reference data matrix X by combining the column vectors from the (ζ - L)-th column to the min{ζ + L, N1}-th column in X1 t ; where min{ζ + L, N1} represents selecting the minimum value between ζ + L and N1.
[0013] Step (3.3): Determine whether ζ is greater than L; if not, record the column vectors from the 1st column to the (ζ + L)-th column in X j as reference data vectors v1, v2,..., v ζ+L in sequence, and then update the reference data matrix X according to the formula X t = [X t , v1, v2,..., v ζ+L ; if so, record the column vectors from the (ζ - L)-th column to the t -th column in X j as reference data vectors in sequence, and then update the reference data matrix X according to the formula ; where represents selecting the minimum value between ζ + L and N t ; where represents selecting the minimum value between ζ + L and N j 1.
[0014] Step (3.4) Determine whether j is less than J; if so, set j = j + 1 and then return to step (3.3); if not, obtain the final reference data matrix X t .
[0015] Step (4): Calculate the mean vector μ t and the standard deviation vector δ t of all column vectors in the reference data matrix X t , and then perform standardization processing on the data vector x according to the formula t to obtain the online data vector Then perform standardization processing on each column vector in the reference data matrix X t in the same way to obtain the neighborhood matrix where means dividing the elements at the same positions in the two vectors on the left and right of the symbol.
[0016] Step (5): Use the online data vector and the neighborhood matrix to perform immediate extraction of differential features, thereby obtaining the immediate transformation vector w t , and the specific implementation process is shown in Steps (5.1) to (5.2).
[0017] Step (5.1): Calculate the distance between each column vector in and . Then, mark the C column vectors in with the smallest distance to as u1, u2,..., u C ; where C is an integer less than M. The distance between any column vector z in and is calculated according to the formula .
[0018] Step (5.2): Calculate the immediate coefficient vector β ∈R t according to the formula C×1 . Then, solve the eigenvalue problem to obtain the eigenvector p corresponding to the largest eigenvalue λ, and calculate the immediate transformation vector w according to the formula t ; where U = [u1, u2,..., u C .
[0019] It should be noted that the implementation process of the above Steps (5.1) to (5.2) is actually to perform a transformation on the online data vector t through the immediate transformation vector w , so as to maximize the reconstructed error of the converted neighbors while minimizing the fluctuation change of the corresponding features of each reference data vector within the similar time neighborhood in the time neighborhood window matrix , that is:
[0020]
[0021] By constructing the Lagrangian function , the solution of the above formula ④ can be realized: First, calculate the partial derivative of φ with respect to w t :
[0022]
[0023] Take the extreme value when the above formula ② is equal to zero, and thus obtain the generalized eigenvalue problem defined by . In addition, since the 10 sample data selected for the penicillin production process are not significantly correlated with each other, is an invertible symmetric matrix. In other words, multiplying both sides of the equation in the generalized eigenvalue problem by on the left simultaneously, we obtain the eigenvalue problem
[0024] To avoid the problem of using different technical terms for the same symbol, the eigenvector p is used for intermediate transition in the above eigenvalue problem, that is, in step (5.2) In addition, since the calculation result of is a J×1 dimensional column vector, and the rank of is equal to 1, so there is only one non-zero eigenvalue in the eigenvalue problem in step (5.2), which is the largest eigenvalue.
[0025] Step (6): According to the formula the instantaneous difference feature y is calculated. t After that, according to the formula the reference eigenvector ξ is calculated, and then the largest element and the smallest element in ξ are recorded as ξ max and ξm in respectively.
[0026] Step (7): Determine whether the condition ξ min ≤y t ≤ξ max is satisfied; if so, the penicillin production batch runs normally, and step (8) is executed; if not, step (9) is executed to determine whether the penicillin production batch runs normally.
[0027] Step (8): Determine whether the penicillin production of this batch is over; if not, return to step (3) to continue implementing batch production process monitoring using the sample data at the next sampling moment; if so, the penicillin production of this batch runs normally, and the data vectors at all sampling moments of this production batch are formed into a batch matrix X J+1 in the order of sampling time. After that, set J = J + 1, then clean the penicillin production equipment and return to step (2).
[0028] Step (9): Return to step (3) to continue implementing batch production process monitoring using the sample data at the next sampling moment. If the instantaneous difference features corresponding to consecutive A sampling moments do not satisfy the judgment condition in step (7), the penicillin production batch runs abnormally, and the penicillin production of this batch is immediately stopped.
[0029] Through the above implementation steps, the advantages of the method of the present invention are introduced as follows.
[0030] The main advantages of the method of the present invention are as follows: First, it can handle the problem of unequal batch lengths in penicillin production batch data. The solution is to obtain the reference data vectors within the adjacent time neighborhood from the normal batch matrix through the time neighborhood window. Second, it can adaptively analyze and extract the immediate difference features for anomaly detection. The specific implementation method is the immediate difference feature analysis involved in the method of the present invention. Third, it can handle the time-varying characteristics of penicillin production batch data by adding new penicillin production batch data in normal operation into the batch matrix. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the implementation process of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0032] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0033] The present invention discloses a batch production process monitoring method based on neighborhood difference feature analysis and extraction. The following describes the specific implementation manner of the method of the present invention with reference to the Figure 1 schematic diagram of the implementation process shown below.
[0034] Step (1): From the historical database of penicillin production batches, obtain a sample data set of J normal penicillin production batches, and respectively form corresponding batch matrices X1, X2,..., X J .
[0035] In the actual penicillin production process in operation, a corresponding distributed control system (abbreviation: DCS) is equipped. The DCS can obtain and store 10 sample data collected at each sampling moment in real time. For the sake of consistency, these 10 sample data can be arranged in the following order to form a 10×1-dimensional data vector: ventilation rate, stirring power, glucose feeding temperature, glucose feeding rate, coolant flow rate, acid-base flow rate, reactor temperature, pH value, glucose concentration, and penicillin concentration.
[0036] Step (2): Set the reference length of the time neighborhood window to be equal to L, start the production of the latest penicillin batch, and obtain the sample data at each sampling moment.
[0037] Step (3): Form a 10×1-dimensional data vector x t with the 10 sample data at the latest sampling moment, and respectively obtain the reference data vectors from the batch matrices X1, X2,..., X J , so as to merge and form a reference data matrix X t , specifically as shown in steps (3.1) to (3.4);
[0038] Step (4): Calculate the mean vector μ t of all column vectors in t the reference data matrix X and the standard deviation vector δ t . After that, according to the formula , perform standardization processing on the data vector x t to obtain the online data vector . Then, perform standardization processing on each column vector in the reference data matrix X t in the same way to obtain the neighborhood matrix
[0039] Step (5): Perform instant extraction of differential features on the online data vector and the neighborhood matrix to obtain the instant transformation vector w t . The specific implementation process is shown in Steps (5.1) to (5.2).
[0040] Step (6): Calculate the instant differential feature y according to the formula. After that, calculate the reference feature vector ξ according to the formula t . Then, record the largest element and the smallest element in ξ as ξ max and ξ min .
[0041] Step (7): Determine whether the condition ξ min ≤y t ≤ξ max is satisfied. If so, the penicillin production batch runs normally and Step (8) is executed; if not, Step (9) is executed to determine whether the penicillin production batch is abnormal.
[0042] Step (8): Determine whether the penicillin production of this batch is over. If not, return to Step (3) to continue monitoring the batch production process using the sample data at the next sampling moment. If so, the penicillin production of this batch runs normally. After forming the batch matrix X by arranging the data vectors at all sampling moments of this production batch in the order of sampling time J+1 , set J = J + 1, then clean the penicillin production equipment and return to Step (2).
[0043] Step (9): Return to Step (3) to continue monitoring the batch production process using the sample data at the next sampling moment. If the instant differential features corresponding to consecutive A sampling moments do not satisfy the judgment condition in Step (7), the penicillin production batch runs abnormally and the penicillin production of this batch is immediately stopped.
Claims
1. A batch production process monitoring method based on neighborhood difference feature analysis, characterized in that Specifically, it includes the following steps: Step (1): From the historical database of penicillin production batches, obtain a sample data set of J normal penicillin production batches, and respectively form the corresponding batch matrices X1, X2, …, X J ; where the batch matrix corresponding to the j-th normal penicillin production batch is specifically composed of N j 10×1 data vectors, j ∈ {1, 2, …, J}, denotes a 10×N j dimensional real matrix, R represents the set of real numbers, the first column vector in X j is the data vector at the first sampling moment of the j-th normal penicillin production batch, and the N j -th column vector in X j is the data vector at the N j -th sampling moment of the j-th normal penicillin production batch; Step (2): Set the reference length of the time neighborhood window to be equal to L, start the production of the latest penicillin production batch, and obtain the sample data at each sampling moment; where L is a positive integer; Step (3): Assemble the 10 sample data at the latest sampling moment into a 10×1 data vector x t , and obtain reference data vectors from the batch matrices X1, X2, …, X J respectively, so as to merge and form a reference data matrix X t , as specifically shown in Steps (3.1) to (3.4); Step (3.1): After recording the current sampling moment as the ζ-th sampling moment of the latest penicillin production batch, initialize j = 2; Step (3.2): Determine whether ζ is greater than L; if not, then form a reference data matrix X from the column vectors of the 1st to the (ζ + L)-th columns in X1 t ; if so, then form a reference data matrix X from the column vectors of the (ζ - L)-th to the min{ζ + L, N1}-th columns in X1 t ; where min{ζ + L, N1} represents the selection of the minimum value between ζ + L and N1; Step (3.3): Determine whether ζ is greater than L; if not, then record the column vectors of the 1st to (ζ + L)th columns in X j as reference data vectors v1, v2,..., v ζ+L in sequence. After that, update the reference data matrix X according to the formula X t = [X t , v1, v2,..., v ζ+L ; if so, then record the column vectors of the (ζ - L)th to t th columns in X j as reference data vectors in sequence. After that, update the reference data matrix X according to the formula ; where, represents the minimum value of ζ + L and N t ; ; j ; Step (3.4): Determine whether j is less than J; if so, after setting j = j + 1, return to Step (3.3); if not, obtain the final reference data matrix X t ; Step (4): Calculate the average vector μ t of all column vectors in t and the standard deviation vector δ t in the reference data matrix X. After that, according to the formula perform standardization processing on the data vector x t to obtain the online data vector Then, perform standardization processing on each column vector in the reference data matrix X t in the same way to obtain the neighborhood matrix where the symbol means dividing the elements at the same positions in the left and right vectors; Step (5): Use the online data vector and the neighborhood matrix to perform immediate extraction of differential features, thereby obtaining the immediate conversion vector w t , and the specific implementation process is shown in steps (5.1) to (5.2); Step (5.1): Calculate the distance between each column vector in and Then, mark the C column vectors in with the smallest distance to C as u1, u2,..., u Step (5.2): According to the formula calculate the instantaneous coefficient vector β t ∈R C×1 After that, solve the eigenvalue problem for the eigenvector p corresponding to the largest eigenvalue λ, and calculate the instantaneous conversion vector w according to the formula ; where U = [u1, u2,..., u t ; C Step (6): According to the formula calculate to obtain the instant difference feature y t After that, according to the formula calculate to obtain the reference feature vector ξ, and then record the largest element and the smallest element in ξ as ξ max and ξ min ; Step (7): Determine whether the condition ξ is satisfied min ≤ y t ≤ ξ max ; if so, the penicillin production batch runs normally and step (8) is executed; if not, step (9) is executed to determine whether the penicillin production batch runs normally; Step (8): Determine whether the production of this batch of penicillin is completed; if not, return to step (3) and continue to implement batch production process monitoring using the sample data at the next sampling moment; if so, the production of this batch of penicillin is operating normally, and the data vectors at all sampling moments of this production batch are composed into a batch matrix X in the order of sampling time J+1 After that, set J = J + 1, then clean the penicillin production equipment and return to step (2); Step (9): Return to Step (3) and continue to implement batch production process monitoring using the sample data of the next sampling moment. If the immediate difference features corresponding to consecutive A sampling moments do not meet the judgment conditions in Step (7), the penicillin production batch runs abnormally, and immediately stop the penicillin production of this batch.
2. The batch production process monitoring method based on neighborhood difference feature analysis extraction according to claim 1, characterized in that In the data vectors at each sampling moment, the arrangement order of the 10 data is as follows: ventilation rate, stirring power, glucose feeding temperature, glucose feeding rate, coolant flow rate, acid-base flow rate, reactor temperature, pH value, glucose concentration, and penicillin concentration.
Citation Information
Patent Citations
kernel learning monitoring method for penicillin production process under unequal-length batch conditions
CN103901855A
Online and real-time batch process monitoring method based on k nearest neighbor
CN104808648A