Intermittent production process soft measurement method based on CDPC-JGCN
By employing the CDPC-JGCN method, which utilizes transfer entropy to select causal auxiliary variables and causal density peak clustering, a graph convolutional network model is constructed and calibrated online. This addresses the issue of insufficient accuracy in soft measurement models during intermittent production, enabling efficient online measurement of key variables.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-07
AI Technical Summary
In intermittent production processes, key variables are difficult to measure online in real time using physical sensors. Existing data-driven soft measurement methods are limited by process multimodalities and batch differences, resulting in insufficient accuracy of soft measurement models.
The CDPC-JGCN method is adopted to select causal auxiliary variables by means of transfer entropy, divide modes by causal density peak clustering, construct a graph convolutional network model, and introduce real-time learning for online calibration to establish a multimodal soft measurement model.
It improves the accuracy of soft measurement of key variables in intermittent production processes, adapts to process changes, and enables online real-time measurement.
Smart Images

Figure CN121808265A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intermittent production process monitoring technology, and particularly relates to a soft measurement method for intermittent production processes based on CDPC-JGCN (CausalDensity Peak Clustering - Just-in-time Graph Convolutional Networks). Background Technology
[0003] Batch production processes, which produce the same product through repeated batch operations, are suitable for multi-variety, small-batch production and have been widely used in fine chemicals, food processing, microelectronics, and other fields. Real-time measurement of key variables in batch production processes is crucial for intelligent monitoring and ensuring product quality. However, due to limitations in measurement technology and the on-site operating environment, some key variables in batch production processes are difficult to measure online in real time using physical sensors, directly affecting the effectiveness of process monitoring. Soft measurement technology provides an effective solution to this problem. Unlike physical sensors, soft measurement selects measurable auxiliary variables that are correlated with the key variables (dominant variables), establishes a soft measurement model between the auxiliary and dominant variables, and achieves online real-time measurement of the key variables (dominant variables). This technology has been widely applied in actual batch production processes.
[0004] In existing soft measurement methods for batch production processes, data-driven methods are used because they do not rely on complete prior process knowledge. However, these data-driven methods select auxiliary variables based on the correlation between variables, which limits the improvement of the accuracy of the soft measurement model. The inherent characteristics of batch production processes result in multiple operational stages or multiple modes, with significant differences in process characteristics between different modes. Furthermore, raw material fluctuations mean that batch production processes are not strictly repetitive, leading to differences in batch distribution and affecting the accuracy of the constructed soft measurement model. Therefore, this paper proposes a soft measurement method for batch production processes based on CDPC-JGCN. This method selects auxiliary variables based on the transfer entropy between measurable process variables and dominant variables to obtain a set of causal auxiliary variables for soft measurement of batch production processes. Then, it uses a causal density peak clustering algorithm to divide the batch production process into multiple modes and constructs a multimodal soft measurement model based on graph convolutional networks. Finally, it introduces real-time learning to construct an online calibration strategy for the soft measurement model, enabling online measurement of key variables (dominant variables) in the batch production process and improving the accuracy of soft measurement in batch production processes. Summary of the Invention
[0005] This invention aims to improve the accuracy of soft measurement of key variables (dominant variables) in batch production processes. It proposes a soft measurement method for batch production processes based on CDPC-JGCN, comprising the following steps:
[0006] Step 1: Collect multiple batches of intermittent production process data, standardize the data, calculate the transfer entropy between the measurable variable and the dominant variable, select causal auxiliary variables for soft measurement based on the magnitude of the transfer entropy, and construct a set of causal auxiliary variables for soft measurement of intermittent production process;
[0007] Step 2: Combining the causal density peak clustering algorithm, the intermittent production process is divided into multiple modes to achieve mode division of the intermittent production process, obtain the process dataset of each mode, and use graph convolutional networks to construct soft measurement models for each mode;
[0008] Step 3: Introduce real-time learning, use kernel functions to calculate the similarity between online and offline samples of process data, select a similar sample set of process data based on the similarity, update the soft measurement model parameters online, and realize the online calibration of the soft measurement model.
[0009] Step one specifically includes:
[0010] Collect measurable variable data from multiple batches of intermittent production processes. ,in The number of process variables, For the number of batch process data, The number of process data sampling points; according to equation (1), along the sampling time direction, Standardized to
[0011]
[0012] In the formula, This refers to standardized intermittent production process data; Data for intermittent production processes The mean along the sampling time direction; Data for intermittent production processes Standard deviation along the sampling time direction.
[0013] Then calculate the measurable variables of the intermittent production process. and dominant variables Intermittent transfer entropy, given measurable variables in an intermittent production process. , and Inter-transfer entropy The calculation formula is
[0014]
[0015] In the formula, Copula entropy; These are time-domain parameters; It is the dominant variable in intermittent production processes. The Middle Values of variables at any given time; and Before the dominant variable and the measurable variable of the intermittent production process, respectively. Value of the variable at time.
[0016] To determine measurable variables in a batch production process and dominant variables The direction of causal relationship between them, defining the causal effect index. for
[0017]
[0018] In the formula, for and The transfer entropy between them.
[0019] When the measurable variables and the dominant variables in a batch production process When the value is greater than 0, the measurable variable is the dependent variable of the dominant variable, and in this case, the measurable variable is selected as a causal auxiliary variable for soft sensor modeling. This is based on the causal effect index. Select causal auxiliary variables and construct a set of causal auxiliary variables for soft measurement of intermittent production processes. ,in This represents the number of causal auxiliary variables.
[0020] Step two specifically includes:
[0021] The intermittent production process is divided into multiple modes using the causal density peak algorithm. First, the calculation is performed. Data from each process Local density of each sample point The calculation formula is
[0022]
[0023] In the formula, A set of causal auxiliary variables for intermittent production processes The Middle Time sample points and the Time sample points The distance between; To cut off the distance.
[0024] A long short-term memory network process model was trained using intermittent production process data. The resulting process model was:
[0025]
[0026] In the formula, Indicates the first The model prediction at time; The length of the time window; Representation process model; These are model parameters; for The Middle Time to Data samples at any given time.
[0027] Since there is a causal relationship between the modality center and the characteristics of the intermittent production process, the selected modality center should be as close as possible to the behavior predicted by the model, and the degree of closeness should be... The calculation formula is
[0028]
[0029] In the formula, For the first Time interval over sample The degree of similarity to the model's predicted behavior. The larger, the more it indicates It more closely resembles the predictive behavior of the model; This represents the dimension of the intermittent production process sample.
[0030] Then calculate the relative distance between the data sample points of each process.
[0031]
[0032] In the formula, and These represent the process model at the 1st... Time and the The prediction error at any given time; and They represent and The value after z-score standardization.
[0033] Due to the complexity of the intermittent production process's operating state, the local density of process data samples is low in regions with frequent mode switching. When the intermittent production process is stable, the corresponding process data samples in this region have a high local density. To avoid interference from non-modal centroids in high-density regions on the selection of modal centroids in low-density regions, decision values for the process data samples are calculated. for
[0034]
[0035] In the formula, ; for The mean; for The covariance matrix; This is a matrix transpose operation.
[0036] The number of modal centers and the allocation of remaining process data samples can be determined in the following two steps:
[0037] (1) Decision values of all process data sample points Sort them in descending order and calculate the mean of all decision values. and standard deviation Decision value greater than + Process data sample points are used as candidate mode centers. The remaining process data sample points are assigned to the mode center with the smallest relative distance. If no process data sample points are assigned to the candidate mode center... ,but It should not be selected as a modal center;
[0038] (2) Divide the intermittent production process into (1) A modality, if Similar modes are fused based on the inter-modal similarity index. Modal calculation... and modality The similarity index is
[0039]
[0040] In the formula, , and These represent the modes of the intermittent production process. and modality The temporal center separation, process data sample difference, and local density similarity are calculated as follows:
[0041]
[0042]
[0043]
[0044] In the formula, and These are the intermittent production process modes. and modality The number of samples; To take the absolute value; and They represent The Middle Time and the Data samples at any given time; This is for calculating Euclidean distance.
[0045] If the difference between the two modes with the highest similarity index and the two modes with the lowest similarity index exceeds a threshold... Then the two modalities with the greatest similarity will be fused, and (2) will be repeated until... Or all modes may fail to merge. Ultimately, the intermittent production process will be divided into... Each modality has a right endpoint. and datasets for each modality ,in .
[0046] Using modal datasets of intermittent production processes Training a soft measurement model based on a graph convolutional network yields a multimodal soft measurement model, which can be represented as follows:
[0047]
[0048] In the formula, Indicates the first A soft measurement model for each modality; Indicates the first The predicted value of the time-matter model.
[0049] Step three specifically includes:
[0050] Given online sample of intermittent production process data , It is the sampling time, first based on the sampling time. Determine online samples of process data The modality belongs to which the process data is calculated, and then the online samples and historical samples of the corresponding modality are compared according to equation (14). distance for:
[0051]
[0052] In the formula, The kernel width of the kernel function; It is an exponential function.
[0053] Will Arrange them from smallest to largest, and select the first one. Historical data samples from each process are used as training samples for the soft measurement model to obtain a training set for soft measurement model calibration. .
[0054] When the Online samples of time process data Upon arrival, the soft measurement model corresponding to the mode of the sample is selected, and then the selected similar samples are used. Fine-tuning the parameters of the modal model yields a soft measurement model for predicting key variables at that moment. , No. Prediction results of key variables in time-intermittent production processes for
[0055]
[0056] Advantages of this invention: This invention fully considers the impact of batch differences in intermittent production processes on soft sensor models, introduces transfer entropy to calculate the causal relationship between auxiliary and dominant variables, and selects causal auxiliary variables based on the magnitude of causal effect values; then, it uses causal density peak clustering to achieve modal segmentation of intermittent production processes, obtains multimodal datasets, and uses GCN to construct soft sensor models for each modality; finally, it uses online calibration of soft sensor models based on real-time learning, realizing the construction of soft sensor models based on CDPC-JGCN, thus improving the accuracy of predictions by soft sensor models for intermittent production processes. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of a soft measurement method for intermittent production processes based on CDPC-JGCN, as described in this invention. Figure 2 These are the predicted results of penicillin concentrations in the first batch using different soft measurement methods. Detailed Implementation
[0058] The present invention will be further described below with reference to examples and accompanying drawings. It should be noted that the embodiments do not limit the scope of protection claimed by the present invention.
[0059] Example
[0060] IndPensim is a simulation platform for industrial fed-batch penicillin fermentation, and its data characteristics closely resemble those of actual industrial penicillin fermentation. The sampling interval was 0.2 hours, with 1150 sampling points per batch, for a total of 10 batches of data collected. Six batches were used as training data, two as validation data, and two as test batches. The variables for the industrial fed-batch penicillin fermentation process are shown in Table 1, with variable 15 (penicillin concentration) serving as the key variable (dominant variable) in the soft sensing.
[0061] Table 1. Variables in the industrial penicillin fermentation process
[0062] Serial Number variable unit Serial Number variable unit 1 time h 9 temperature K 2 Basic feed rate L / h 10 Heat production kJ 3 Increase the flow rate of cold water L / h 11 Percentage of carbon dioxide in exhaust gas % 4 Heating water flow rate L / h 12 Oxygen uptake g / min 5 Dissolved oxygen concentration mg / L 13 Oxygen percentage in exhaust gas % 6 reactor volume L 14 carbon dioxide release rate g / h 7 reactor quality kg 15 penicillin concentration g / L 8 pH -
[0063] The algorithm flow of this invention applied to the industrial penicillin fermentation process is as follows: Figure 1 As shown, the specific steps are as follows:
[0064] Step 1: Standardize the process data from 6 batches using Z-score, then calculate the propagation entropy between the measurable variables and the dominant variable for each process, and select the causal auxiliary variable set based on the magnitude of the propagation entropy. .
[0065] Step 2: Calculate the local density of each sample in the industrial penicillin fermentation process according to equation (4). Data from industrial penicillin fermentation process The long short-term memory network process model was obtained through training, and the model prediction error of samples at each time step in the industrial penicillin fermentation process was obtained according to equations (5) and (6). Then, the relative distance between each sample is calculated according to equation (7). Based on equation (8), the decision values of the data samples for each process in the industrial penicillin fermentation process are obtained. Based on equations (9) to (12), candidate modal center samples for the industrial penicillin fermentation process were obtained. According to the relative distance between each process data sample and the modal center sample, they were assigned to the corresponding modal, thus dividing the industrial penicillin fermentation process into 3 modalities and obtaining a multimodal process dataset. The right endpoint times of each mode are 385h, 652h and 1150h respectively; finally, the graph convolutional network soft measurement model of each mode is established according to equation (13).
[0066] Step 3: When the online process data samples arrive, calculate the kernel distance similarity between the online process data samples and the offline samples according to formula (14), sort the similarity from largest to smallest, select the top 10% of offline samples to construct a similar dataset of online samples, adjust the parameters of the offline soft measurement model online, and realize the online prediction of industrial penicillin concentration.
[0067] Figure 2The prediction results of different soft sensing methods on batch 1 are shown. Figure 2 It is evident that, considering the multimodal characteristics of the industrial penicillin fermentation process, the soft sensor method constructing individual modal models has higher prediction accuracy than the global modeling method. Among the multimodal soft sensor modeling methods, the DPC-GCN method exhibits significant errors due to neglecting the temporal characteristics of the penicillin fermentation process data during modality segmentation. In contrast, IDPC-GCN and WSDPC-GCN overcome this limitation, significantly improving the performance of the soft sensor model. Compared to these methods, CDPC-JGCN employs CDPC in modality segmentation, constructs a GCN soft sensor model for each modality, and uses JITL for model calibration, resulting in smaller prediction errors. Experimental results demonstrate that the CDPC-JGCN soft sensor model constructed in this invention has high prediction accuracy, validating the effectiveness of the soft sensor method described in this invention.
Claims
1. A soft measurement method for intermittent production processes based on CDPC-JGCN, characterized in that: The method includes the following steps: Step 1: Collect multiple batches of intermittent production process data, standardize the data, calculate the transfer entropy between the measurable variables and the dominant variables, select causal auxiliary variables for soft measurement based on the magnitude of the transfer entropy, and construct a set of causal auxiliary variables for soft measurement of the intermittent production process; the intermittent production process data are industrial penicillin fermentation process variables, specifically including time, basic feed flow rate, cooling water flow rate, heating water flow rate, dissolved oxygen concentration, reactor volume, reactor mass, pH, temperature, heat production, tail gas carbon dioxide percentage, oxygen uptake rate, tail gas oxygen percentage, carbon dioxide release rate, and penicillin concentration; Step 2: Combining the causal density peak clustering algorithm, the intermittent production process is divided into multiple modes to achieve mode division of the intermittent production process, obtain the process dataset of each mode, and use graph convolutional networks to construct soft measurement models for each mode; Step 3: Introduce real-time learning, use kernel functions to calculate the similarity between online and offline samples of process data, select a similar sample set of process data based on the similarity, update the soft measurement model parameters online, and realize the online calibration of the soft measurement model.
2. The soft measurement method for intermittent production processes based on CDPC-JGCN according to claim 1, characterized in that: Step one specifically includes: Collect measurable variable data from multiple batches of intermittent production processes. ,in The number of process variables, For the number of batch process data, The number of process data sampling points; according to equation (1), along the sampling time direction, Standardized to ; In the formula, This refers to standardized intermittent production process data; Data for intermittent production processes The mean along the sampling time direction; Data for intermittent production processes Standard deviation along the sampling time direction; Then calculate the measurable variables of the intermittent production process. and dominant variables Intermittent transfer entropy, given measurable variables in an intermittent production process. , and Inter-transfer entropy The formula for calculation is: ; In the formula, Copula entropy; These are time-domain parameters; It is the dominant variable in intermittent production processes. The Middle Values of variables at any given time; and Before the dominant variable and the measurable variable of the intermittent production process, respectively. Values of variables at any given time; To determine measurable variables in a batch production process and dominant variables The direction of causal relationship between them, defining the causal effect index. for ; In the formula, for and Transfer entropy between; When the measurable variables and the dominant variables in a batch production process When the value is greater than 0, the measurable variable is the dependent variable of the dominant variable, and this measurable variable is selected as the causal auxiliary variable for soft measurement modeling; according to the causal effect index Select causal auxiliary variables and construct a set of causal auxiliary variables for soft measurement of intermittent production processes. ,in This represents the number of causal auxiliary variables.
3. The soft measurement method for intermittent production processes based on CDPC-JGCN according to claim 2, characterized in that: Step two specifically includes: The intermittent production process is divided into multiple modes using the causal density peak algorithm. First, the calculation is performed. Data from each process Local density of each sample point The calculation formula is ; In the formula, A set of causal auxiliary variables for intermittent production processes The Middle Time sample points and the Time sample points The distance between; To cut off the distance; A long short-term memory network process model was trained using intermittent production process data. The resulting process model was: ; In the formula, Indicates the first The model prediction at time 1; The length of the time window; Representation process model; These are model parameters; for The Middle Time to Data samples at any given time; Since there is a causal relationship between the modality center and the characteristics of the intermittent production process, the selected modality center should be as close as possible to the behavior predicted by the model, and the degree of closeness should be... The calculation formula is ; In the formula, For the first Time interval over sample The degree of similarity to the model's predicted behavior. The larger, the more it indicates It more closely resembles the predictive behavior of the model; Dimensions for samples of intermittent production processes; Then calculate the relative distance between the data sample points of each process. ; In the formula, and These represent the process model at the 1st... Time and the The prediction error at any given time; and They represent and The value after z-score standardization; Due to the complexity of the intermittent production process's operating state, the local density of process data samples is low in regions with frequent mode switching. When the intermittent production process is stable, the corresponding process data samples in this region have a high local density. To avoid interference from non-modal centroids in high-density regions on the selection of modal centroids in low-density regions, decision values for the process data samples are calculated. for ; In the formula, ; for The mean; for The covariance matrix; This is a matrix transpose operation; The number of modal centers and the allocation of remaining process data samples can be determined in the following two steps: (1) Decision values of all process data sample points Sort them in descending order and calculate the mean of all decision values. and standard deviation Decision value greater than + Process data sample points are used as candidate mode centers. The remaining process data sample points are assigned to the mode center with the smallest relative distance. If no process data sample points are assigned to the candidate mode center, then... ,but It should not be selected as a modal center; (2) Divide the intermittent production process into (1) A modality, if Similar modes are fused based on the inter-modal similarity index; modal similarity is calculated. and modality The similarity index is ; In the formula, , and These represent the modes of the intermittent production process. and modality The temporal center separation, process data sample difference, and local density similarity are calculated as follows: ; ; ; In the formula, and These are the intermittent production process modes. and modality The number of samples; To take the absolute value; and They represent The Middle Time and the Data samples at any given time; For calculating Euclidean distance; If the difference between the two modes with the highest similarity index and the two modes with the lowest similarity index exceeds a threshold... Then the two modalities with the greatest similarity will be fused, and (2) will be repeated until... Or all modes may fail to merge; ultimately, the intermittent production process will be divided into Each modality has a right endpoint. and datasets for each modality ,in ; Using modal datasets of intermittent production processes Training a soft measurement model based on a graph convolutional network yields a multimodal soft measurement model, denoted as: ; In the formula, Indicates the first A soft measurement model for each modality; Indicates the first The predicted value of the time-matter model.
4. The soft measurement method for intermittent production processes based on CDPC-JGCN according to claim 3, characterized in that: Step three specifically includes: Given online sample of intermittent production process data , It is the sampling time, first based on the sampling time. Determine online samples of process data The modality belongs to which the process data is calculated, and then the online samples and historical samples of the corresponding modality are compared according to equation (14). distance for: ; In the formula, The kernel width of the kernel function; It is an exponential function; Will Arrange them from smallest to largest, and select the first one. Historical data samples from each process are used as training samples for the soft measurement model to obtain a training set for soft measurement model calibration. ; When the Online samples of time process data Upon arrival, the soft measurement model corresponding to the mode of the sample is selected, and then the selected similar samples are used. Fine-tuning the parameters of the modal model yields a soft measurement model for predicting key variables at that moment. , No. Prediction results of key variables in time-intermittent production processes for 。