Multi-mode process virtual sample generation method with multi-view feature fusion

CN120705711BActive Publication Date: 2026-09-11KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510893715.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2026-09-11
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供多视角图特征融合的多模式过程虚拟样本生成方法,以解决上述背景技术中提出由于积累的数据量不够,导致构建的软测量模型难以准确捕捉生产系统的动态特性,进而影响对目标质量变量的实时估计的问题

Benefits of technology

[0026]与现有技术相比,本发明的有益效果是:该多视角图特征融合的多模式过程虚拟样本生成方法:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705711B_ABST
    Figure CN120705711B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of industrial process control, and particularly discloses a multi-mode process virtual sample generation method based on multi-view graph feature fusion, which comprises the following steps: S1, collecting sensor data in an industrial process by means of offline detection and distributed control, and establishing an industrial process database; S2, performing normalization processing on the collected sensor data based on a Z-Score method to obtain a normalized data set. The multi-mode process virtual sample generation method based on multi-view graph feature fusion realizes mode discrimination by means of a Gaussian mixture model and combines a metric learning method of mode preserving embedding, so as to solve the problem of data loss caused by insufficient initial data of a new process, sensor failure or working condition switching in an industrial scene, accurately reflect the feature distribution of different process modes, and enable a soft measurement model to capture more comprehensive dynamic characteristics of an industrial process after the original samples are combined and trained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial process control technology, specifically to a method for generating multi-mode process virtual samples through multi-view graph feature fusion. Background Technology

[0002] With the rapid advancement of Industry 4.0 and intelligent manufacturing, process industries are accelerating their transformation towards full-process digitalization, networking, and intelligence. Against this backdrop, the demand for real-time data from industrial automation systems is growing exponentially. However, the real-time measurement of key quality variables such as component concentration and reaction efficiency often faces significant bottlenecks: high-precision sensor deployment costs are high, and offline measurement has large latency, making it difficult to meet the needs of real-time process monitoring. This has become a key obstacle to realizing the self-sensing, self-decision-making, and self-adaptive capabilities of production systems. Soft measurement technology provides a solution to these problems. This technology uses easily collectable auxiliary variables as input and target quality variables as output. It achieves real-time estimation of difficult-to-measure quality variables by constructing predictive models. Its core modeling methods are divided into first-principles models (FPMs) and data-driven models. FPMs rely on precise mechanisms and are suitable for simple systems; data-driven models rely on artificial intelligence and industrial big data and are more suitable for high-dimensional nonlinear scenarios.

[0003] As a key supporting technology for the intelligent transformation of process industries, data-driven soft measurement methods have shown enormous application potential. However, due to the high complexity of production systems and the rapid changes in external demands, data acquisition limitations are common in real-world industrial scenarios, leading to severe small sample problems and significantly reducing modeling accuracy. For example, insufficient data in the early stages of new processes makes it difficult for models to capture dynamic characteristics; redundant data under normal operating conditions reduces the effectiveness of feature extraction; sensor failures and operating condition switching cause data gaps and incomplete sample coverage; and the superimposed high-dimensionality, nonlinearity, and noise interference exacerbate the modeling difficulty under limited data. Overcoming the small sample bottleneck is a key challenge to improving the accuracy of soft measurement. Summary of the Invention

[0004] The purpose of this invention is to provide a method for generating virtual samples of multi-mode processes through multi-view graph feature fusion, in order to solve the problem mentioned in the background art that the constructed soft measurement model is difficult to accurately capture the dynamic characteristics of the production system due to insufficient accumulated data, thereby affecting the real-time estimation of target quality variables.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a multi-mode process virtual sample generation method for multi-view graph feature fusion, comprising the following steps: S1. Collect sensor data from industrial processes using offline detection and distributed control methods to establish an industrial process database; determine auxiliary variables through mechanistic analysis of industrial processes. and predictor variables Auxiliary variables Input variables, predictor variables As output variables, the dataset in the industrial process database includes auxiliary variables. and predictor variables ; S2. The collected sensor data is normalized using the Z-Score method to obtain a normalized dataset. ,in Indicates the number of samples; Indicates the inclusion of auxiliary variables and predictor variables The total number of variables, This indicates the number of samples in the training set. This indicates the number of samples in the test set; the dataset... Divided into training set and test set ,in ; S3. Train the Gaussian Mixture Model (GMM) on the training set. Pattern discrimination is performed to obtain the probability of each sample belonging to each pattern. The pattern with the highest probability is taken as the predicted pattern for the sample, and this pattern information is integrated into the original dataset to finally generate a dataset with pattern annotations. ; S4. By employing a metric learning method based on pattern-preserving embeddings, sample balancing is performed between patterns, thereby constructing a balanced training dataset. ; S5. Based on the balanced sample set The training process incorporates generative networks that integrate process pattern information and graph networks, and utilizes the trained model to generate enhanced samples. ; S6. Enhance the sample Compared with the original sample Merge and retrain the soft sensor models to improve the model's prediction accuracy;

[0006] Preferred, The process of predicting patterns in the original data in S3 includes the construction of the GMM model and pattern prediction. The construction process of the GMM model is as follows: First, given labeled training data ; in , for Dimensional input variables; , Output as a single variable; Assuming the observation data is from The mixture is composed of a weighted blend of Gaussian distributions, with each blend component corresponding to a latent pattern. The probability density function with strong expressive power can be constructed by weighted summation and can be described by the following formula: ; in, The observed data is represented by the feature vector, and the dimension of the feature vector can be any direction. The number of components in the mixture; It is the first A Gaussian-distributed mixture of weights, satisfying and , Representing the The probability density function of a Gaussian distribution is defined as follows: ;

[0007] in, Indicates the dimension of the data; It is the first The mean vector of Gaussian components; No. The covariance matrix of Gaussian components, This determines the shape and scale of the distribution; The determinant of the covariance matrix. Influence on the normalization factor; The GMM model parameters The specific optimization steps, estimated using the Expectation-Maximization (EM) algorithm, can be described as follows: First, the values ​​of each parameter are initially selected randomly; Step E: Calculate each data point Belongs to the The posterior probability of each Gaussian component is calculated using the following formula: ; M-step: Update parameters, calculated using the following formula: ; ; ; in, Indicates belonging to the first The expected number of samples for each component; Then, repeat the E-step and M-step until the GMM model converges; The pattern prediction process is as follows: Based on the trained GMM, the probability assignment for each sample pattern can be described as follows: ; in This represents a well-trained GMM model; For pattern label vectors, Represented as ,in Indicates the first The sample belongs to the first The probability of each pattern is calculated, and the pattern with the highest probability is selected as the data pattern. After adding pattern information to the original dataset, a new dataset is obtained. .

[0008] By adopting the above technical solution, and by identifying the GMM implementation pattern, combined with the metric learning method of pattern-preserving embedding, the problem of insufficient data in the early stage of new processes, data loss caused by sensor failure or switching of operating conditions in industrial scenarios can be solved.

[0009] Preferred, The pattern embedding-preserving metric learning method in S4 employs an autoencoder (AE), which includes an encoder and a decoder. The encoder consists of a two-layer feedforward network, as shown in the following formula: ; in, For the original sample, For encoder weights, For encoder bias, As an optional activation function, the sigmoid function is selected. The latent features of the output; The decoder uses the same architecture as the encoder, as shown in the following formula: ; in, As a latent feature, For decoder weights, For encoder bias, For the reconstructed samples of the output, To characterize the pattern information of the data, a mode-preserving mapping metric learning method is incorporated into the loss function of the self-decoder. This brings the representations of the same patterns closer together and widens the representations of different patterns further apart. The loss function is expressed as follows: ; in, Indicates the reconstruction loss; The triplet represents the learning loss.

[0010] By adopting the above technical solution, the balanced training dataset enables the soft measurement model to capture the dynamic characteristics of industrial processes more comprehensively, especially improving the modeling ability for rare operating conditions and reducing prediction bias under small sample conditions.

[0011] Preferred, The reconstruction loss and triplet metric learning loss in the overall loss of the pattern-preserving embedding-based metric learning method are described as follows: ; in, It is the number of samples. It is the first One input sample, These are reconstructed samples; The pattern-preserving embedding-based metric learning objective constructs a discriminative latent space structure, causing similar samples to cluster and dissimilar samples to move away, and designs a triplet metric learning loss. The loss function The formula for representing is as follows: ; ; ; ; in, The model represents the embedding vector of the samples; This represents the anchor sample. For from the anchor point Positive samples randomly selected from the same pattern Indicates from the anchor point Negative samples randomly selected from different patterns;

[0012] express and The Euclidean distance between them; , , These are the probability vectors of anchor points, positive samples, and negative samples belonging to each pattern, respectively. This represents the probability that the anchor point and the positive sample belong to the same pattern. This represents the probability that the anchor point and the negative sample belong to the same pattern.

[0013] By employing the above technical solution, the reconstruction loss can ensure that the generated samples retain the physical characteristics of the original data, thus avoiding information loss.

[0014] Preferably, the generative network in S5 adopts a generative adversarial network architecture. The generative adversarial network includes a generator and a discriminator. The generator takes the probability value of the pattern to which the sample belongs as an input conditional variable to guide the generation of pattern-related data. The framework of the generative model introduces a regressor under pattern conditions and a sample classifier with pattern as the output target through a regularization term to enhance the generative model's ability to perceive pattern information. A co-training method is used to train the regression model and the classification model. The discriminator introduces a multi-view graph convolutional feature enhancement mechanism and constructs a feature relationship topology graph through a feature correlation matrix to achieve explicit modeling of the constraint relationship between feature nodes.

[0015] By adopting the above technical solution and through the collaborative training mechanism that complements regression and classification tasks, the generator's ability to generate data for rare working conditions can be enhanced.

[0016] Preferred, The generator is described as follows: The pattern probabilities predicted by the GMM are used as conditional variables input to the generator to guide the generation of pattern-related data. A regressor under pattern conditions and a sample classifier with pattern as the output target are introduced to enhance the generative model's ability to perceive pattern information. The mapping function of the regressor is described as follows: ; in, This represents the mapping relationship in the regression model; The mapping function of the pattern classifier is described as follows: ; in This represents the mapping relationship of the classifiers. It is the first One input sample; During the generator training process, the mapping relationship between process variables and quality variables under different modes, as well as the modeling performance of classification relationships, are backpropagated to the generator to guide its learning. The generator's objective function is described as follows: ; in, and This indicates that two weighting coefficients are used to weigh the two loss components; It represents the number of virtual samples generated during the network update process; It is the first generated by the generator. Quality variables of each sample; The first one predicted by the regressor Individual sample quality variables; It is the generated first The data belongs to the first The probability of each pattern, The classifier predicts the first... The pattern probability of each data point.

[0017] By adopting the above technical solution, the pattern attribution accuracy of the generated data can be improved through pattern probability condition input and classifier constraints.

[0018] Preferred, The collaborative training method for training the regression model and the classification model is as follows: Assume the original generated data is represented as... , For the generated data quality variables corrected by the regressor, The generated data pattern probabilities are corrected by the classifier. For data adjusted based on regressors The predicted pattern attribution probability when updating the classifier is achieved through knowledge transfer via a collaborative training framework of regression and classification models. The interaction optimization process is as follows: The regressor optimizes the mapping relationship of quality variables and outputs an adjusted data representation. ; The classifier corrects the pattern probability distribution and outputs the corrected data representation. ; Through a cross-training mechanism, using adjusted data Train the classifier to predict probabilities Approximating the true probability of membership The same applies to training the regressor. The optimization objective functions of the regression model and the classification model are described as follows: ; ; in, For the quality variables predicted by the regression model during collaborative training; For the classification model's prediction during collaborative training... The sample belongs to the first The probability of each pattern; It represents the number of virtual samples generated in a single iteration.

[0019] By adopting the above technical solution, collaborative training can make the quality variables and pattern probabilities of the generated data mutually constrain each other, avoiding the bias caused by single task optimization.

[0020] Preferred, The multi-view graph convolutional feature enhancement mechanism first establishes a dual-channel feature correlation analysis mechanism: The Pearson correlation coefficient is used to characterize the linear correlation properties, and the obtained linear correlation matrix is ​​used as the edge connection strength of the nodes to ensure that the graph structure can represent the linear dependence between features. The Spearman rank correlation coefficient is used to capture the nonlinear correlation between features, and the obtained nonlinear correlation matrix is ​​used as the edge connection strength of the nodes to ensure that the graph structure can represent the nonlinear dependency structure between features. Based on this, two learnable weight parameters are introduced. The hidden features under two correlation metrics are adaptively concatenated with the original data features to form a multi-view map feature extraction structure. The fused features are then used as new data features to distinguish between true and false data. The objective function of the discriminator is: ; in This indicates the expected calculation. Representative Discriminator Represents generator, Used to identify the authenticity of generated samples. It is a gradient penalty term. For real samples, To generate samples.

[0021] By adopting the above technical solution, and using Pearson and Spearman correlation coefficients to characterize linear and nonlinear dependencies respectively, the problem of difficulty in modeling complex relationships between features in industrial data can be solved.

[0022] Preferred, The Pearson correlation coefficient and Spearman correlation coefficient were calculated as follows: For the Pearson correlation coefficient: assuming the sample sets of two paired labeled continuous variables are represented as follows: ,in For the sample size, the covariance between the two variables is... It can be calculated as follows: ; in, and The first one sample peacekeeping 3D eigenvalues; and It is the first peacekeeping The mean of the dimensional features; The Pearson correlation coefficient between the variables It can be calculated as: ; in and These are the standard deviations of two continuous variables; Regarding the Spearman correlation coefficient: the Spearman correlation coefficient is used to capture the nonlinear correlation between two continuous variables, and its calculation formula is as follows: ; ; in, Indicates the number of samples; and express and The rank of the sorted items from smallest to largest; This indicates the difference in rank between them.

[0023] The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 8 is characterized in that: The gradient penalty is as follows: ;

[0024] in, It's gradient calculation. It belongs to the first Real samples of each pattern With generated data Linear interpolation between them.

[0025] By adopting the above technical solution, gradient penalty constraints can make the transition between generated samples and real samples in the feature space more natural.

[0026] Compared with the prior art, the beneficial effects of the present invention are: the multi-view image feature fusion multi-mode process virtual sample generation method: 1. This invention achieves pattern discrimination through Gaussian mixture model and combines it with a metric learning method that preserves pattern embedding, thereby solving the problem of data loss in industrial scenarios caused by insufficient data in the early stages of new processes, sensor failures, or changes in operating conditions. It can accurately reflect the feature distribution of different process modes. After training by merging it with the original samples, the soft measurement model can capture more comprehensive dynamic characteristics of industrial processes, improve the modeling ability for rare operating conditions, and significantly improve the prediction bias problem caused by incomplete data coverage under small sample conditions. 2. In this invention, a pattern-condition regressor and classifier are introduced into the generative adversarial network framework. Combined with a multi-view graph convolution feature enhancement mechanism, the linear dependence of features is characterized by Pearson correlation coefficient and nonlinear correlation is captured by Spearman rank correlation analysis. A feature relationship topology graph is constructed to achieve explicit modeling of complex constraint relationships between industrial process variables. This ensures that the generated virtual samples closely resemble the distribution of real industrial data. After co-training with the original data, the soft measurement model's feature extraction capability for high-dimensional, nonlinear, and noisy industrial data can be enhanced, significantly improving the prediction accuracy and robustness of the model under complex working conditions. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the process flow of the method of the present invention; Figure 2 This is a schematic diagram of the structural framework of the present invention; Figure 3 This is a schematic diagram of the CTC process flow structure of the present invention; Figure 4 This is a schematic diagram of the t-SNE dimensionality reduction distribution structure of the original data in this invention; Figure 5 This is a schematic diagram of the t-SNE dimensionality reduction distribution structure of the fused data after adding balanced samples in this invention; Figure 6 This is a schematic diagram of the matrix concentration prediction and actual value curves during the CTC process after data enhancement using the method of this invention and other methods. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] Please see Figures 1-6 This invention provides a technical solution: a method for generating virtual samples in a multi-mode process by fusing multi-view graph features.

[0030] The first step involves collecting sensor data from industrial processes using offline detection and distributed control methods to establish an industrial process database. Based on this, the mechanisms of the industrial processes are analyzed in depth to determine auxiliary variables. and predictor variables Auxiliary variables Input variables, predictor variables As output variables, the data stored in the database constitutes a labeled dataset. This dataset Includes auxiliary variables and predictor variables The second step will involve processing the dataset. It is divided into training set and test set, where It represents the number of samples.

[0031] The second step is to normalize the collected sensor data using the Z-Score method to obtain a normalized dataset. ,in Indicates the number of samples; Indicates the inclusion of auxiliary variables and predictor variables The total number of variables; This indicates the number of samples in the training set. This indicates the number of samples in the test set; [The dataset is then divided into several parts]. Divided into training set and test set ,in .

[0032] The third step is to train a Gaussian Mixture Model (GMM) on the training set. Pattern discrimination is performed to obtain the probability of each sample belonging to each pattern. The pattern with the highest probability is taken as the predicted pattern for the sample, and this pattern information is integrated into the original dataset to finally generate a dataset with pattern annotations. ;

[0033] The Gaussian Mixture Model (GMM) is established as follows: The construction process of the GMM model is as follows: The core idea of ​​the Gaussian Mixture Model (GMM) used is to maximize the joint probability distribution of the observed data under the model. Specifically, the labeled training data is given in the second step. ; in , for Dimensional input variables; , Output as a single variable; Assuming the observation data is from The mixture is composed of a weighted blend of Gaussian distributions, with each blend component corresponding to a latent pattern. The probability density function with strong expressive power can be constructed by weighted summation and can be described by the following formula: ; in, The observed data is represented by the feature vector, and the dimension of the feature vector can be any direction. The number of components in the mixture; It is the first A Gaussian-distributed mixture of weights, satisfying and , Representing the The probability density function of a Gaussian distribution is defined as follows: ; in, Indicates the dimension of the data; It is the first The mean vector of Gaussian components; No. The covariance matrix of Gaussian components, This determines the shape and scale of the distribution; The determinant of the covariance matrix. Influence on the normalization factor; GMM model parameters The estimation process using the Expectation-Maximization (EM) algorithm includes the following iterative steps: First, the values ​​of each parameter are initially selected randomly; Step E: Calculate each data point Belongs to the The posterior probability of each Gaussian component is calculated using the following formula: ; M-step: Update parameters, calculated using the following formula: ; ; ; in, Indicates belonging to the first The expected number of samples for each component; Then, repeat the E-step and M-step until the GMM model converges; The pattern prediction process is as follows: After the above iterations converge, the trained GMM model can be obtained, and the probability assignment for each sample pattern is as follows: ; in This represents a well-trained GMM model; For pattern label vectors, Represented as ,in Indicates the first The sample belongs to the first The probability of each pattern is calculated, and the pattern with the highest probability is selected as the data pattern. After adding pattern information to the original dataset, a new dataset is obtained. .

[0034] The fourth step involves employing a pattern-preserving embedding-based metric learning method to balance the samples among patterns, thereby constructing a balanced training dataset. ; Specifically, the metric learning method based on pattern embedding preservation employs an autoencoder (AE), which consists of an encoder and a decoder. The encoder is composed of a two-layer feedforward network, as shown in the following formula: ; in, For the original sample, For encoder weights, For encoder bias, As an optional activation function, the sigmoid function is selected. The latent features of the output; The decoder uses the same architecture as the encoder, as shown in the following formula: ; in, As a latent feature, For decoder weights, For encoder bias, For the reconstructed samples of the output, To characterize the pattern information of the data, a mode-preserving mapping metric learning method is incorporated into the loss function of the self-decoder. This brings the representations of the same patterns closer together and widens the representations of different patterns further apart. The loss function is expressed as follows: ; in, Indicates the reconstruction loss; The triplet represents the learning loss. The reconstruction loss and triplet metric learning loss in the overall loss of the pattern-preserving embedding-based metric learning method are described as follows:

[0035] ; in, It is the number of samples. It is the first One input sample, These are reconstructed samples; In metric learning, to improve the separability of sample features in the embedding space, relative distance optimization is usually performed by constructing positive-negative sample pairs or triplet sample pairs. However, in actual training, mode collapse is prone to occur, where multiple samples of different categories or modalities excessively overlap in the embedding space, leading to a loss of representational ability. Therefore, this invention proposes a pattern-aware smoothing mechanism. This mechanism adjusts the geometric relationship between features by identifying the similarity between samples, thereby enhancing the diversity and discriminativeness among embedded features. The loss function is as follows: ; ; ; ; in, The model represents the embedding vector of the samples; This represents the anchor sample. For from the anchor point Positive samples randomly selected from the same pattern Indicates from the anchor point Negative samples randomly selected from different patterns; express and The Euclidean distance between them; , , These are the probability vectors of anchor points, positive samples, and negative samples belonging to each pattern, respectively. This represents the probability that the anchor point and the positive sample belong to the same pattern. This represents the probability that the anchor point and the negative sample belong to the same pattern. use After training the pattern-aware metric learning model constructed above and achieving convergence, use this model to encode the training data to obtain hidden features. Two features belonging to the same pattern are randomly selected, interpolated, and decoded to obtain new samples under the corresponding pattern as pattern balancing samples. .

[0036] Fifth step, based on the balanced sample set The training process incorporates generative networks that integrate process pattern information and graph networks, and utilizes the trained model to generate enhanced samples. ; The specific steps for building and training a generative model are as follows: (1) The discriminator based on graph network feature extraction is constructed by first establishing a dual-channel feature association analysis mechanism: the linear association characteristics are characterized by Pearson correlation coefficient, and the obtained linear correlation matrix is ​​used as the edge connection strength of the node to ensure that the graph structure can represent the linear dependency relationship between features; the nonlinear association characteristics between features are captured by Spearman rank correlation analysis, and the obtained nonlinear correlation matrix is ​​used as the edge connection strength of the node to ensure that the graph structure can represent the nonlinear dependency structure between features; on this basis, two learnable weight parameters are introduced. The hidden features under the two correlation metrics are adaptively concatenated with the original data features to form a multi-view graph feature extraction structure. The two correlation coefficients are described as follows: For the Pearson correlation coefficient: assuming the sample sets of two paired labeled continuous variables are represented as follows... ,in For the sample size, the covariance between the two variables is... It can be calculated as follows: ; in, and The first one sample peacekeeping 3D eigenvalues; and It is the first peacekeeping The mean of the dimensional features; Pearson correlation coefficient between variables It can be calculated as: ; in and These are the standard deviations of two continuous variables; For the Spearman correlation coefficient: The Spearman correlation coefficient is used to capture the non-linear correlation between two continuous variables, and its calculation formula is as follows: ; ; in, Indicates the number of samples; and express and The rank of the sorted items from smallest to largest; This indicates the difference in rank between them; The fused features are used as new data features to distinguish between true and false data. The objective function of the discriminator is: ; in, This indicates the expected calculation. Representative Discriminator Represents generator, Used to identify the authenticity of generated samples. It is a gradient penalty term. For real samples, To generate samples, The gradient penalty is as follows: ; in It belongs to the first Real samples of each pattern With generated data Linear interpolation between them.

[0037] (2) Construction of the pattern-related generator The generative network adopts a generative adversarial network (GAN) architecture, which includes a generator and a discriminator. The generator takes the probability value of the pattern to which the sample belongs as a conditional variable as input, guiding the generation of pattern-related data. The framework of the generative model introduces a pattern-conditional regressor and a sample classifier with the pattern as the output target through a regularization term, so as to enhance the generative model's ability to perceive pattern information. The pattern probabilities predicted by GMM are used as conditional variables input to the generator to guide the generation of pattern-related data. A regressor under pattern conditions and a sample classifier with pattern as the output target are introduced to enhance the generative model's ability to perceive pattern information. The relationship between process variables and quality variables under modal information conditions is as follows: ; in, This represents the mapping relationship in the regression model; The mapping function of the pattern classifier is described as follows: ; in This represents the mapping relationship between classifiers; During generator training, the mapping relationship between process variables and quality variables under different modes, as well as the modeling performance of classification relationships, are backpropagated to the generator to guide its learning. The generator's objective function is described as follows: ; in, and This indicates that two weighting coefficients are used to weigh the two loss components; It represents the number of virtual samples generated during the network update process; It is the first generated by the generator. Quality variables of each sample; The first one predicted by the regressor Individual sample quality variables; It is the generated first The data belongs to the first The probability of each pattern, The classifier predicts the first... The pattern probability of each data point.

[0038] (3) Construction of collaborative training A collaborative training method is used to train the regression and classification models. The discriminator incorporates a multi-view graph convolutional feature enhancement mechanism, constructing a feature relationship topology graph through a feature correlation matrix to explicitly model the constraints between feature nodes. After one iteration of the generation model, the generated data is input into both models to improve their generalization ability. The co-training process for training regression and classification models is as follows: Assume the original generated data is represented as... , For the generated data quality variables corrected by the regressor, The generated data pattern probabilities are corrected by the classifier. For data adjusted based on regressors The predicted pattern attribution probability when updating the classifier is achieved through knowledge transfer via a collaborative training framework of regression and classification models. The interaction optimization process is as follows: The regressor optimizes the mapping relationship of quality variables and outputs an adjusted data representation. ; The classifier corrects the pattern probability distribution and outputs the corrected data representation. ; Through a cross-training mechanism, using adjusted data Train the classifier to predict probabilities Approximating the true probability of membership The same applies to training the regressor. The optimization objective functions for regression and classification models are described below: ; ; in, For the quality variables predicted by the regression model during collaborative training; For the classification model's prediction during collaborative training... The sample belongs to the first The probability of each pattern; It represents the number of virtual samples generated in a single iteration. This ensures that regression and classification models can extract reliable supervisory information during the iteration process and pass it to the generator, thereby helping the generator learn regression knowledge and classification information from the input data, improving the quality of generated data and the generalization ability of the prediction model.

[0039] Step 6: Use the results obtained in step 4. The dataset is used to train and iterate the generator network until a Nash equilibrium is reached.

[0040] Step 7: Using the trained generative model, generate several virtual sample data sets. These virtual data sets have similar statistical characteristics and process correlations to the original data. Merge these generated virtual samples with the original training data to construct an enhanced training dataset. Based on this enhanced training dataset, retrain the soft sensor model and finally establish the quality variables. Regarding process variables The mapping relationship is expressed as: ; in, This is the soft measurement model function obtained by training the above augmented dataset.

[0041] The embodiments of this invention use root mean square error (RMSE) and coefficient of determination. To validate the prediction results, the smaller the RMSE value... A larger value indicates a smaller prediction error, meaning better prediction performance of the soft sensor modeling method. The calculation formula is as follows: ; ; in, Represents the number of samples. Indicates the predicted output. The table shows the actual output values, and The RMSE represents the mean of the sample output values; the lower the RMSE value, the better. The higher the value, the higher the accuracy of the model's predictions.

[0042] The following industrial example of a specific chlortetracycline fermentation process illustrates the performance of the method proposed in this invention. In the chlortetracycline fermentation process, the concentration of the chlortetracycline matrix is ​​a crucial indicator in the feedback fermentation control process. However, currently, the concentration of the chlortetracycline matrix cannot be detected online. To improve the control level of chlortetracycline fermentation, soft-sensor modeling of the chlortetracycline matrix concentration is necessary. Table 1 shows the nine auxiliary variables selected for the key predictor variable, chlortetracycline matrix concentration.

[0043] Table 1. Explanation of Auxiliary Variables

[0044] This experiment collected 13 batches of data from a fermenter. The first 8 batches, comprising 200 data points, were used as the training set; the remaining 5 batches, comprising 127 data points, were used as the test set. Through repeated experiments, the network structure and initialization parameters of different generative models were determined. To verify the effectiveness of the method proposed in this invention, four other data generation methods were compared: (1) GPR: Gaussian process regression; where a parent covariance function with a noise term is used; (2)SVAE: Sample generation based on supervised variational autoencoder; (3) WGAN_gp: WGAN with gradient penalty is used for sample generation; (4) VA_WGAN: A generative model that combines VAE and WGAN-gp; (5)MR_GAN: A correlation generative adversarial network based on regression modeling (MR-GAN) for data augmentation-based soft sensor modeling; (6)MGGAN_gp: The method proposed in this invention.

[0045] Based on the experimental results, the prediction performance reached its optimal level when the data augmentation amount was 160 in this case. The experimental results are as follows: Table 2. Impact of different data augmentation methods on the accuracy of soft sensor modeling in CTC chemical processes. GPR 0.4374 0.9085 SVAE 0.3701 0.9182 VA_WGAN 0.3519 0.9343 WGAN_gp 0.3696 0.9275 MR_GAN 0.3542 0.9334 MGGAN_gp 0.3288 0.9468

[0046] Table 2 compares the improvements of five sample generation methods on the performance of the soft sensor prediction model when generating 160 samples. Analysis of Table 1 shows that these existing models can improve the model's prediction performance to some extent. Specifically, SVAE, VA_WGAN, WGAN_gp, and MR_GAN improved prediction accuracy by 13.56%, 19.55%, 15.73%, and 19.02%, respectively. The proposed MGGAN_gp reduced the root mean square error from 0.4374 to 0.3288, an improvement of 24.83%, demonstrating the effectiveness of the proposed MGGAN_gp sample generation method in the chlortetracycline CTC chemical process. This case study demonstrates that the present invention has good sample generation capabilities.

[0047] Working principle: Industrial sensor data is collected through offline detection and distributed control to establish a database. The data is normalized based on the Z-Score method, and training and test sets are divided. The training set is used to perform pattern discrimination through Gaussian Mixture Model (GMM) to generate a dataset with pattern annotation. By adopting a metric learning method based on pattern-preserving embedding, sample balancing is performed between patterns to construct a balanced training dataset. A generative network that integrates process pattern information and graph network is trained based on the balanced sample set. Augmented samples are generated using the trained model. The augmented samples are merged with the original samples to retrain the soft measurement model to improve the model's prediction accuracy.

[0048] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

Claims

1. A method for generating virtual samples in a multi-mode process through multi-view graph feature fusion, characterized in that: Includes the following steps: S1. Collect sensor data from industrial processes using offline detection and distributed control methods to establish an industrial process database; determine auxiliary variables through mechanistic analysis of industrial processes. and predictor variables Auxiliary variables Input variables, predictor variables As output variables, the dataset in the industrial process database includes auxiliary variables. and predictor variables ; S2. The collected sensor data is normalized using the Z-Score method to obtain a normalized dataset. ,in Indicates the number of samples; Indicates the inclusion of auxiliary variables and predictor variables The total number of variables, This indicates the number of samples in the training set. This indicates the number of samples in the test set; the dataset... Divided into training set and test set ,in ; S3. Train the Gaussian Mixture Model (GMM) on the training set. Pattern discrimination is performed to obtain the probability of each sample belonging to each pattern. The pattern with the highest probability is taken as the predicted pattern for the sample, and this pattern information is integrated into the original dataset to finally generate a dataset with pattern annotations. ; S4. By employing a metric learning method based on pattern-preserving embeddings, sample balancing is performed between patterns, thereby constructing a balanced training dataset. ; S5. Based on the balanced sample set The training process incorporates generative networks that integrate process pattern information and graph networks, and utilizes the trained model to generate enhanced samples. ; S6. Enhance the sample Compared with the original sample Merge and retrain the soft sensor model to improve the model's prediction accuracy; The generative network in S5 adopts a generative adversarial network (GAN) architecture, which includes a generator and a discriminator. The generator takes the probability value of the pattern to which the sample belongs as an input conditional variable to guide the generation of pattern-related data. The generator framework introduces a pattern-conditional regressor and a sample classifier with the pattern as the output target through a regularization term to enhance the generator's ability to perceive pattern information. A co-training method is used to train the regression model and the classification model. The discriminator introduces a multi-view graph convolution feature enhancement mechanism, which constructs a feature relationship topology graph through a feature correlation matrix to achieve explicit modeling of the constraint relationship between feature nodes. The generator is described as follows: The pattern probability predicted by the GMM is used as a conditional variable input to the generator to guide the generation of pattern-related data. A regressor under pattern conditions and a sample classifier with pattern as the output target are introduced to enhance the generator's ability to perceive pattern information. The mapping function of the regressor is described as follows: ; in, This represents the mapping relationship in the regression model. It represents the probability that the i-th sample belongs to each pattern. It is the quality variable of the sample predicted by the regressor; The mapping function of the pattern classifier is described as follows: ; in This represents the mapping relationship of the classifiers. It is the first One input sample; During the generator training process, the mapping relationship between process variables and quality variables under different modes, as well as the modeling performance of classification relationships, are backpropagated to the generator to guide its learning. The generator's objective function is described as follows: ; in, This indicates the expected calculation. Representative discriminator, Represents generator, Used to identify the authenticity of generated samples. It represents the number of virtual samples generated during the network update process. and This indicates that two weighting coefficients are used to weigh the two loss components; It is the first generated by the generator. Quality variables of each sample; The first one predicted by the regressor Individual sample quality variables; It is the generated first The data belongs to the first The probability of each pattern, The classifier predicts the first... The pattern probability of each data point; The collaborative training method for training the regression model and the classification model is as follows: Assume the original generated data is represented as... , For the generated data quality variables corrected by the Regressor, The generated data pattern probabilities are corrected by the classifier. For data adjusted based on regressors The predicted pattern attribution probability when updating the classifier is achieved through knowledge transfer via a collaborative training framework of regression and classification models. The interaction optimization process is as follows: The regressor optimizes the mapping relationship of quality variables and outputs an adjusted data representation. ; The classifier corrects the pattern probability distribution and outputs the corrected data representation. ; Through a cross-training mechanism, using adjusted data Train the classifier to predict probabilities Approximating the true probability of membership The same applies to training the regressor. The optimization objective functions of the regression model and the classification model are described as follows: ; ; in, For the quality variables predicted by the regression model during collaborative training; For the classification model's prediction during collaborative training... The sample belongs to the first The probability of each pattern; It represents the number of virtual samples generated in a single iteration.

2. The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 1, characterized in that: The process of predicting patterns in the original data in S3 includes the construction of the GMM model and pattern prediction. The construction process of the GMM model is as follows: First, given labeled training data ; in , for Dimensional input variables; , Output as a single variable; Assume the observation data is from The mixture is composed of a weighted blend of Gaussian distributions, with each blend component corresponding to a latent pattern. The probability density function with strong expressive power can be constructed by weighted summation and can be described by the following formula: ; in, The observed data is represented by the feature vector, and the dimension of the feature vector can be any direction. The number of components in the mixture; It is the first A Gaussian-distributed mixture of weights, satisfying and , Representing the The probability density function of a Gaussian distribution is defined as follows: ; in, Indicates the dimension of the data; It is the first The mean vector of Gaussian components; No. The covariance matrix of Gaussian components, This determines the shape and scale of the distribution; The determinant of the covariance matrix. Influence on the normalization factor; The GMM model parameters The specific optimization steps for estimating using the EM algorithm by maximizing expectation can be described as follows: First, the values ​​of each parameter are initially selected randomly; Step E: Calculate each data point Belongs to the The posterior probability of each Gaussian component is calculated using the following formula: ; M-step: Update parameters, calculated using the following formula: ; ; ; in, Indicates belonging to the first The expected number of samples for each component; Then, repeat the E-step and M-step until the GMM model converges; The pattern prediction process is as follows: Based on the trained GMM, the probability assignment for each sample pattern can be described as follows: ; in This represents a well-trained GMM model; For pattern label vectors, Represented as ,in Indicates the first The sample belongs to the first The probability of each pattern is calculated, and the pattern with the highest probability is selected as the data pattern. After adding pattern information to the original dataset, a new dataset is obtained. .

3. The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 1, characterized in that: The mode-preserving embedding-based metric learning method in S4 employs an autoencoder (AE), which includes an encoder and a decoder. The encoder consists of a two-layer feedforward network, as shown in the following formula: ; in, For the original sample, For encoder weights, For encoder bias, As an optional activation function, the sigmoid function is selected. The output is the latent feature; The decoder uses the same architecture as the encoder, as shown in the following formula: ; in, As a latent feature, For decoder weights, For decoder bias, For the reconstructed samples of the output, To characterize the pattern information in the data, a pattern-preserving embedding metric learning method is incorporated into the self-decoder loss function. This brings the representations of the same patterns closer together and widens the representations of different patterns further apart. The loss function is expressed as follows: ; in, Indicates the reconstruction loss; The triplet represents the learning loss.

4. The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 3, characterized in that: The reconstruction loss in the overall loss of the pattern-preserving embedding-based metric learning method is described as follows: ; in, It is the number of samples. It is the first One input sample, These are reconstructed samples; The pattern-preserving embedding-based metric learning objective constructs a discriminative latent space structure, causing similar samples to cluster and dissimilar samples to move away, and designs a triplet metric learning loss. The loss function The formula for representing is as follows: ; ; ; ; in, The model represents the embedding vector of the samples; Anchor represents the anchor point sample. For from the anchor point Positive samples randomly selected from the same pattern Indicates from the anchor point Negative samples randomly selected from different patterns; express and The Euclidean distance between them; , , These are the probability vectors of anchor points, positive samples, and negative samples belonging to each pattern, respectively. This represents the probability that the anchor point and the positive sample belong to the same pattern. This represents the probability that the anchor point and the negative sample belong to the same pattern.

5. The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 1, characterized in that: The multi-view graph convolutional feature enhancement mechanism first establishes a dual-channel feature correlation analysis mechanism: The Pearson correlation coefficient is used to characterize the linear correlation properties, and the resulting linear correlation matrix is ​​used as the edge connection strength of the nodes to ensure that the graph structure can represent the linear dependence between features. The Spearman correlation coefficient is used to capture the nonlinear correlation between features, and the obtained nonlinear correlation matrix is ​​used as the edge connection strength of the nodes to ensure that the graph structure can represent the nonlinear dependency structure between features. Based on this, two learnable weight parameters are introduced. The hidden features under two correlation metrics are adaptively concatenated with the original data features to form a multi-view map feature extraction structure. The fused features are then used as new data features to distinguish between true and false data. The objective function of the discriminator is: ; in This indicates the expected calculation. Representative discriminator, Represents generator, Used to identify the authenticity of generated samples. It is a gradient penalty term. For real samples, To generate samples.

6. The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 5, characterized in that: The Pearson correlation coefficient and Spearman correlation coefficient were calculated as follows: For the Pearson correlation coefficient: assuming the sample sets of two paired labeled continuous variables are represented as follows: ,in For the sample size, the covariance between the two variables is... It can be calculated as: ; in, and The first one sample peacekeeping 3D eigenvalues; and No. peacekeeping The mean of the dimensional features; The Pearson correlation coefficient between the variables It can be calculated as: ; in and These are the standard deviations of two continuous variables; Regarding the Spearman correlation coefficient: the Spearman correlation coefficient is used to capture the nonlinear correlation between two continuous variables, and its calculation formula is as follows: ; ; in, Indicates the number of samples; and express and The rank of the sorted items from smallest to largest; This indicates the difference in rank between them.

7. The multi-view graph feature fusion method for generating virtual samples in a multi-mode process according to claim 6, characterized in that: The gradient penalty term is as follows: ; in, It's gradient calculation. It belongs to the first Real samples of each pattern With generated data Linear interpolation between them.

Citation Information

Patent Citations

  • Multi-working-condition process adaptive soft measurement modeling method based on local double-weighted probability hidden variable regression model

    CN114239400A

  • Bearing small sample data expansion method and system based on variational auto-encoder

    CN120180084A