Process parameter prediction method and system based on adversarial domain self-adaption

By aligning the process parameter feature distribution under multiple operating conditions by adversarial domain adaptive method, the problem of poor generalization performance of the process parameter prediction model in multiple operating conditions is solved, and a higher accuracy and stable prediction effect is achieved.

CN120336827APending Publication Date: 2025-07-18TIANJIN DEV ZONE JINGNUOHANHAI DATA TECH CO LTD +1

Patent Information

Application Number
CN202510519563.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Under multi-condition conditions, the generalization performance of the existing process parameter prediction model is poor, resulting in poor accuracy of prediction results. Traditional working condition recognition and domain adaptation methods cannot effectively solve the problem of data distribution differences under different working conditions.

Method used

Adversarial domain adaptation is adopted to conduct adversarial training on features through discriminators and gradient flip layers to align the feature distribution of the source domain and the target domain, combine the similarity calculation of cross-domain nearest neighbor sample pairs, and generate a prediction model to achieve feature alignment under different operating conditions.

Benefits of technology

It improves the accuracy and robustness of process parameter prediction, can adapt to data distribution changes under multiple operating conditions, and improves the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336827A_ABST
    Figure CN120336827A_ABST
Patent Text Reader

Abstract

The invention discloses a process parameter prediction method and system based on antagonism domain self-adaption, and the method comprises the following steps: obtaining process parameter time sequence data of a source working condition domain and a target working condition domain, the data comprising a plurality of process variables and corresponding target quality parameter labels; extracting features of data of the source domain and the target domain, and performing adversarial training on the features by using a discriminator in combination with a gradient overturning layer, so that feature distribution of the source domain and the target domain is aligned; calculating the similarity between a source domain sample and a target domain sample, selecting a cross-domain nearest neighbor sample pair, and enhancing the inter-domain consistency by minimizing the feature representation distance of the similar sample pair; meanwhile, training the data to realize feature alignment under different working conditions; performing model training by using the data after feature alignment to generate a prediction model; and inputting the real-time process parameters of the target working condition into the prediction model, and outputting a target quality parameter prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of process parameter prediction, and specifically relates to a process parameter prediction method and system based on adversarial domain adaptation. Background Art

[0002] In industrial processes and technological processes, in order to improve product quality, it is necessary to predict future industrial parameters through historical data. Currently, there are methods in the prior art that use deep learning for prediction to optimize and change the technological process. However, due to the existence of multiple working conditions in the technological process, when there are multiple working conditions, the data distributions of different working conditions vary greatly, resulting in limited generalization performance of the model, and thus poor accuracy of the final prediction result.

[0003] Referring to the publication number CN119578653A, the present invention discloses a sizing process parameter recommendation method, device, computer device, and storage medium. The method includes: obtaining production order data to be predicted; preprocessing the production order data to be predicted; inputting the preprocessing result into a prediction model for process parameter prediction; outputting a prediction result; the prediction model is obtained by preprocessing and tagging production order data as a sample set, and training a multi-weight main model in an iterative training manner. Specifically, by dynamically adjusting the quality scores of each preprocessed and tagged data and integrating them into the loss function for training; the multi-weight main model includes several base models and a meta-model, and different weights are assigned according to the prediction performance of the base models, and the results predicted by the weighted base models are integrated through the meta-model. By implementing the method of the present invention, the role of multi-model enhanced stacking can be exerted, reducing the dependence on manual experience, improving the accuracy of process parameters and the first-pass production success rate, and reducing product rework and production costs.

[0004] The prior art has a certain impact on the process parameter prediction performance under different working conditions. In order to improve the performance of the model under multi-working condition data, the existing methods are mainly divided into working condition recognition and domain adaptation. The method of working condition recognition divides the existing data into working conditions, then analyzes and measures a certain amount of process data under different working conditions, combines it with its index variables to form a historical working condition database, and establishes a multi-working condition prediction model to improve the overall prediction accuracy. However, it is difficult for the historical database to cover all working conditions, and this method cannot accurately identify new working conditions generated during the production process. Therefore, the method based on working condition recognition is limited in solving the problem of process parameter prediction under multiple working conditions. The domain adaptation method is committed to reducing the data distribution differences between different working conditions. Among them, the TCA method and the BDA method have been applied to the process field, but the TCA and BDA methods will perform feature mapping on the original process data, and the mapped process data is not readable, that is, it is impossible to split a single variable from the original data, and it is impossible to implement a data augmentation method. Summary of the Invention

[0005] To achieve the above object and other related objects, the present invention discloses a process parameter prediction method based on adversarial domain adaptation, comprising the following steps: S1: Obtain the process parameter time series data of the source working condition domain and the target working condition domain, and the data includes process variables and corresponding target quality parameter labels; S2: Extract the features of the source domain and target domain data, and use a discriminator combined with a gradient reversal layer to perform adversarial training on the features to align the feature distributions of the source domain and the target domain; S3: Calculate the similarity between the source domain and target domain samples, select cross-domain nearest neighbor sample pairs, and enhance the inter-domain consistency by minimizing the feature representation distance of the similar sample pairs; S4: Combine the data in step S3 to train step S2 to achieve feature alignment under different working conditions; S5: Use the data after feature alignment to train the model to generate a prediction model; S6: Input the real-time process parameters of the target working condition into the prediction model and output the predicted value of the target quality parameter.

[0006] Further, step S2 includes: The discriminator distinguishes the source domain and target domain features through a classification loss function; The gradient reversal layer reversely propagates the domain classification loss gradient during backpropagation to drive the feature extraction module to generate domain-invariant features.

[0007] Further, the method for selecting cross-domain nearest neighbor sample pairs includes: Randomly select data in the source domain; Measure the sample similarity based on the Euclidean distance in the feature space; Use the data in the target domain corresponding to the lowest sample similarity and the corresponding data in the source domain as cross-domain nearest neighbor sample pairs.

[0008] Further, measuring the sample similarity based on the Euclidean distance in the feature space includes: First, randomly select samples in the source domain ; ; Use the Euclidean distance as a metric to measure the distance between the label of and the process data of all samples in the target domain to find the most similar sample ; Let be the data of the i-th random sample in be the process data, be the label data; The calculation formula of the Euclidean distance is as follows: ; Among them, time_step refers to the number of process data time steps available, is the j-th label data of the i-th sample in the source domain data, is the j-th label data of the k-th sample in the target domain.

[0009] Furthermore, when minimizing the feature representations of similar sample pairs, the loss function is: ; where sim() represents the cosine similarity function, and the model captures cross-domain consistent representations by minimizing the cosine similarity between the nearest neighbor data, where is 's feature representation, is 's feature representation.

[0010] Furthermore, step S2 includes: Extract the features of the source domain data and the target domain data respectively through the feature extractor; Reverse the gradient of the feature extractor and use the discriminator for classification; The feature extractor re-extracts features according to the classification results of the discriminator, and the discriminator conducts classification. Repeat the above steps for adversarial training to complete the alignment of the feature distributions of the source domain and the target domain.

[0011] Furthermore, the total loss in step S2 is: ; Among them, is the total amount of data in the source domain , is the non-linear regression head, is 's random sample, is the feature extractor, is the total amount of data in the target domain , R is the gradient reversal layer, is the discriminator, is the true label of the true working condition of the data sample, is the loss function for the prediction stage, is the classification score of the discriminator for the i-th sample in the source domain, is the classification score of the discriminator for the i-th sample in the target domain.

[0012] Furthermore, in step S4, combining the data in step S3, train step S2, and the loss function is: ; Among them, and are adjustable hyperparameters, is the loss function in step S2, is the loss function when minimizing the feature representations of similar sample pairs in step 3.

[0013] In a second aspect, the present invention provides a process parameter prediction system based on adversarial domain adaptation, including: A data acquisition module for acquiring multi-condition process parameter time-series data; An adversarial cross-domain feature alignment module for achieving cross-domain feature distribution alignment through adversarial training; A similarity sample cross-domain alignment module for generating contrastive learning constraints for cross-domain similar sample pairs; A joint training module for optimizing the loss functions of the adversarial cross-domain feature alignment module and the similarity sample cross-domain alignment module to generate a cross-condition domain adaptation prediction model; A prediction output module for real-time output of the quality parameter prediction result based on the trained model.

[0014] Furthermore, the adversarial cross-domain feature alignment module includes: A feature extractor , a non-linear regression head and a discriminator , among which the feature extractor and the discriminator introduce the gradient reversal technique to form adversarial learning, promoting the feature extractor to align cross-domain features; The feature extractor and the non-linear regression head use labels for regression training.

[0015] Furthermore, the discriminator uses cross-entropy as the loss function of the discriminator.

[0016] Furthermore, optimizing the loss functions of the adversarial cross-domain feature alignment module and the similarity sample cross-domain alignment module includes: The total loss is: ; Among them, and are adjustable hyperparameters, is the loss function of the adversarial cross-domain feature alignment module, is the loss function of the similarity sample cross-domain alignment module; Among them, is: ; where sim() represents the cosine similarity function, and the model captures cross-domain consistent representations by minimizing the cosine similarity between nearest neighbor data, is the data domain randomly selected samples in feature representation of, is the data domain in the sample the most similar sample feature representation of.

[0017] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the above method.

[0018] In a fourth aspect, the present invention discloses a computer-readable storage medium, which stores a computer program executable by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of the method.

[0019] By adopting the above technical solutions, for multi-condition scenarios, a fermentation product concentration prediction model ADA-SSR based on adversarial domain adaptation is proposed to solve the problem of poor model generalization caused by multi-condition data. The proposed ADA-SSR consists of two modules. One module is an adversarial cross-domain feature alignment module, which sets a discriminator and constructs adversarial training between the encoder and the discriminator to align cross-domain data feature representations. The other module is a similar sample cross-domain alignment module, which finds the nearest neighbor data by calculating the Euclidean distance between cross-condition data and captures cross-domain consistent representations by minimizing the cosine similarity of the nearest neighbor data. Based on the domain adaptation prediction method, by adaptively reducing the data distribution differences under different conditions, the problem of poor model generalization performance is solved, thereby improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not limit the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 is a flowchart of an embodiment of the present invention; Figure 2 is a framework diagram of the ADA-SSR method; Figure 3 is a schematic diagram of the adversarial cross-domain feature alignment module of the present invention; Figure 4Distribution diagrams of data under different working conditions; Figure 5 Comparison diagrams of each algorithm when working condition 1 is the source domain and different working conditions are the target domain ; Figure 6 Comparison diagrams of each algorithm when working condition 2 is the source domain and different working conditions are the target domain ; Figure 7 Comparison diagrams of each algorithm when working condition 3 is the source domain and different working conditions are the target domain ; Figure 8 Distribution diagrams of data under different working conditions after adaptation; Figure 9 Prediction results of penicillin concentration in the test set of each working condition when working condition 1 is the source domain, working condition 2 is the target domain, and the number of labels is 2; Figure 10 Prediction results of penicillin concentration in the test set of each working condition when working condition 1 is the source domain, working condition 2 is the target domain, and the number of labels is 16. Specific implementation manners

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0022] Referring to Figure 1 , the embodiments of the present invention provide a process parameter prediction method based on adversarial domain adaptation, including the following steps: S1: Obtain the time series data of process parameters in the source working condition domain and the target working condition domain, and the data includes multiple process variables and corresponding target quality parameter labels; S2: Extract the features of the source domain and target domain data, and use the discriminator combined with the gradient reversal layer to perform adversarial training on the features to align the feature distributions of the source domain and the target domain; S3: Calculate the similarity between the source domain and target domain samples, select the cross-domain nearest neighbor sample pairs, and enhance the inter-domain consistency by minimizing the feature representation distance of the similar sample pairs; S4: Train the data in steps S2 and S3 simultaneously to achieve feature alignment under different working conditions; S5: Use the aligned data to train the model to generate a prediction model; S6: Input the real-time process parameters of the target working condition into the prediction model and output the predicted value of the target quality parameter.

[0023] Further, in step S2: The discriminator distinguishes source domain and target domain features through a classification loss function; The gradient reversal layer reversely propagates the domain classification loss gradient during backpropagation to drive the feature extraction module to generate domain-invariant features.

[0024] Further, the method for selecting cross-domain nearest neighbor sample pairs includes: Randomly select data in the source domain; Measure the sample similarity based on the Euclidean distance in the feature space; Use the data in the target domain corresponding to the lowest sample similarity and the corresponding data in the source domain as cross-domain nearest neighbor sample pairs.

[0025] Further, measuring the sample similarity based on the Euclidean distance in the feature space includes: First, in the source domain randomly select samples ; Use the Euclidean distance as the measurement standard to measure the distance between the label of and the process data of all samples in the target domain to find the most similar sample For is the data of the i-th random sample in is the process data, is the label data; The calculation formula of the Euclidean distance is as follows: ; where time_step refers to the number of available process data time steps, is the j-th label data of the i-th sample in the source domain data, is the j-th label data of the k-th sample in the target domain.

[0026] Further, when minimizing the feature representation of similar sample pairs, the loss function is: ; where sim() represents the cosine similarity function, and the model captures cross-domain consistent representations by minimizing the cosine similarity between nearest neighbor data.

[0027] Further, step S2 includes: Extract the features of the source domain data and the target domain data respectively through the feature extractor; Reverse the gradient of the feature extractor and use the discriminator for classification; The feature extractor re - extracts features according to the classification results of the discriminator, and the discriminator conducts classification. The above steps are repeatedly executed for adversarial training to complete the alignment of the feature distributions between the source domain and the target domain.

[0028] Furthermore, the total loss in step S2 is: ; where, is the total amount of data in the source domain , is the non - linear regression head, is a random sample in , is the feature extractor, is the total amount of data in the target domain, R is the gradient reversal layer, is the discriminator, is the true label of the true working condition of the data sample, is the loss function for the prediction stage, is the classification score of the discriminator for the i - th sample in the source domain, is the classification score of the discriminator for the i - th sample in the target domain.

[0029] Furthermore, in step S4, step S2 and step S3 are trained simultaneously, and the loss function is: ; where, and are adjustable hyperparameters, is the loss function in step S2, is the loss function in step S3.

[0030] On the other hand, the present invention provides a process parameter prediction system based on adversarial domain adaptation, including: A data acquisition module for obtaining multi - condition process parameter time - series data; An adversarial cross - domain feature alignment module for achieving cross - domain feature distribution alignment through adversarial training; A similarity sample cross - domain alignment module for generating contrastive learning constraints for cross - domain similar sample pairs; A joint training module for optimizing the loss functions of the adversarial cross - domain feature alignment module and the similarity sample cross - domain alignment module to generate a cross - condition domain - adaptive prediction model; A prediction output module for real - time output of the quality parameter prediction result based on the trained model.

[0031] Furthermore, the adversarial cross - domain feature alignment module includes: A feature extractor , non - linear regression head and discriminator , where the feature extractor and the discriminator introduce the gradient reversal technique to form adversarial learning, promoting the feature extractor to align cross - domain features; The feature extractor and the non - linear regression head use labels for regression training.

[0032] Furthermore, the discriminator uses cross - entropy as the loss function of the discriminator.

[0033] Furthermore, optimizing the loss functions of the adversarial cross - domain feature alignment module and the similarity sample cross - domain alignment module includes: The total loss is: ; where, and are adjustable hyperparameters, is the loss function of the adversarial cross - domain feature alignment module, is the loss function of the similarity sample cross - domain alignment module; where, is: ; where sim() represents the cosine similarity function, and the model captures cross - domain consistent representations by minimizing the cosine similarity between nearest - neighbor data, is the feature representation of randomly selected samples in the data domain , is the data domain in which the sample the most similar sample feature representation.

[0034] The following details the model proposed by the present invention. This model consists of two core sub - modules: the adversarial cross - domain feature alignment module and the similarity sample cross - domain alignment module, and these two modules can enable the model to adaptively align features in different data domains.

[0035] Referring to Figure 2 , the adversarial cross - domain feature alignment module is based on the principle of adversarial learning and aims to guide the model to adaptively align the data distributions of different working conditions. This module consists of a feature extractor , non - linear regression head and discriminator composed, where and introduce the gradient reversal technique to form adversarial learning, promoting Align cross - domain features; With Use labels for regression training.

[0036] For ease of explanation, define the process data under working condition 1 as , and define the process data under working condition 2 as . The discriminator consists of a single fully - connected layer and a softmax function. By comparing the feature representations of data samples from different working conditions, the discriminator can provide a discrimination score, which reflects the degree of the model's recognition of the inter - domain differences. Specifically, in the data domain , randomly select data , calculate the feature representation ; similarly, randomly select in the data domain , and calculate its feature representation . Then, use the discriminator to score these feature representations, , and at the same time hope that for and the scoring gap is as large as possible to indicate that the model can effectively distinguish data under different working conditions. Therefore, choose cross - entropy as the loss function of the discriminator, take the score output by the discriminator as the predicted label, and take the true working condition of the data sample as the true label , for example, assign label 1 to domain , and label 0 to domain . By minimizing , it can prompt the discriminator to better recognize and distinguish the data distributions under different working conditions. But to achieve the purpose of aligning the features between cross - domain samples, that is, to make generate similar representations for samples from different data domains, this module introduces the gradient reversal technique.

[0037] Gradient reversal is an adversarial learning technique used to reduce the distribution difference between the source domain and the target domain. In gradient reversal, the feature extractor is used to extract features from the input data, while the discriminator is used to classify or discriminate these features. The key idea of gradient reversal is that during the training process, by reversing the gradient of the feature extractor to deceive the classifier so that it is difficult to distinguish data from the source domain and the target domain. Specifically, when the classifier tries to distinguish data from two domains, gradient reversal will make the features learned by the feature extractor "useless" to the classifier, that is, the classifier cannot distinguish different domains through these features. In this way, the feature extractor is forced to learn domain - independent general features, thereby reducing the distribution difference between the two domains. Therefore, during the training process, the gradient reversal mechanism will reverse The gradient of to counteract the training process, so that the obtained features are useless for the discriminator, thereby reducing the feature distribution difference between the two working condition data sets. Define The gradient reversal layer between is ; Therefore, the total loss of this module can be expressed as: ; In the similar sample cross-domain alignment module, first randomly select a sample in the data domain , and calculate its feature representation . Secondly, use the Euclidean distance as the metric to measure the label of and the process data of all samples in the data domain , find the most similar sample and calculate the representation ; where time_step refers to the number of available process data time steps. Finally, by minimizing the and distance between, to optimize the model parameters. For the most similar samples and , the loss function is: ; where sim() represents the cosine similarity function, and the model captures cross-domain consistent representations by minimizing the cosine similarity between the nearest neighbor data. Therefore, the total loss of ADA-SSR is: ; where, and are adjustable hyperparameters.

[0038] To reflect the technical effect of this technical solution, the penicillin fermentation concentration is used as the process data in this embodiment for demonstration.

[0039] Penicillin is an antibacterial drug produced by fermentation using specific strains of Penicillium. The penicillin fermentation process generally consists of three stages: the first stage is the exponential growth stage of Penicillium, during which a large number of Penicillium are produced; the second stage is the process of Penicillium producing penicillin. During this stage, precise control of pH value, dissolved oxygen, nutrient supply, etc. can produce penicillin for a long time; the third stage is the apoptosis of Penicillium. During this process, the concentration of penicillin cannot be measured in real time, resulting in a lag in process control. Therefore, it is necessary to predict the penicillin concentration in real time to assist the operator in making decisions. During the penicillin fermentation process, the initial conditions, feeding changes, environmental changes, etc. of different fermentation batches will all cause changes in the working conditions. For example, a slight change in temperature may affect the metabolic rate of Penicillium and the synthesis efficiency of penicillin; a change in pH value will affect the physiological state and metabolic pathways of cells; insufficient or excessive oxygen supply will affect the yield and quality of penicillin. That is to say, the distribution of process data generated under different working conditions also shows differences.

[0040] Therefore, in order to simulate the changes in working conditions, 30 batches of experimental data were generated in the experiments of this paper. Specifically, in the simulation, other initial conditions were fixed as default values. Under the conditions of initial substrate concentrations of 5, 10, and 15, three groups of different process data were generated, with 10 batches of data in each group, which were respectively defined as working condition 1, working condition 2, and working condition 3. Figure 4 Figure 5 is a scatter plot of the data of 3 working conditions after being reduced to three dimensions by principal component analysis, indicating that there are certain distribution differences in the data under different working conditions.

[0041] To illustrate the effect of ADA-SSR, in this embodiment, a comparative experiment was carried out under 16-tag conditions and two key issues were explored: (1) Whether the change in working conditions will reduce the model accuracy? (2) Can the proposed adversarial domain adaptation method in this paper replace the traditional transfer learning method? For comparison, four benchmark models were selected for the comparative experiment: TCA-GMM-PLS; SR; TCA-SSR and BDA-SSR. Among them, TCA-GMM-PLS uses TCA to perform distribution adaptation on the source domain and the target domain, and then combines GMM and PLS to train the model. SR uses a transformer and directly trains without performing distribution adaptation; TCA-SSR and BDA-SSR are based on a transformer and use TCA and BDA respectively for distribution adaptation, and then train the model.

[0042] Table 1 shows the prediction results of 5 models with operating condition 1 as the source domain and operating condition 2 as the target domain. As can be seen from Table 1, compared with SR and other models, the performance of SR is the worst. This is because SR does not consider the change in data distribution between different operating conditions, resulting in the model trained under operating condition 1 being inapplicable under operating condition 2. Comparing TCA-SSR, BDA-SSR and ADA-SSR, it is found that the performance of the three models is close, indicating that the domain adaptation method proposed in this chapter can be used as an alternative to traditional transfer learning methods, and then on the basis of integrating the methods in Chapter 3, solve the problems brought by multiple operating conditions.

[0043] To further illustrate the model effect, Figures 5 - 7 The effects of each algorithm in predicting penicillin concentration with different source domains and different target domains are shown. "1->2" means operating condition 1 is used as the source domain and operating condition 2 is used as the target domain, and so on. Observing the effects of each model, it can be seen that the performance of the prediction model will decline when the operating condition changes. Among them, the performance of the SR model drops the most severely. This is because SR does not set a domain adaptation strategy, resulting in its inability to align the data distributions of different operating conditions, resulting in poor performance across operating conditions. Although the effects of each model decline across operating conditions, in most cases, the performance of ADA-SSR drops the least. In a few cases, the effects of ADA-SSR are the same as those of TCA-SSR and BDA-SSR, indicating that ADA-SSR has good stability and can establish a prediction model with strong robustness. Figure 8 The results of the first three dimensions selected after the PCA dimensionality reduction of the three operating condition data after ADA-SSR adapts to the data of each operating condition are shown. Comparing Figure 4 it is found that the differences in the data distributions of different operating conditions are significantly reduced, further indicating that ADA-SSR can align the data distributions under different operating conditions and can replace traditional transfer learning methods.

[0044] Table 1 Prediction results of algorithms under multiple operating conditions

[0045] To evaluate the effectiveness of each module of ADA-SSR, ablation experiments are carried out in this embodiment under the condition of 16 labels, trying to illustrate two problems: (1) Is the proposed adversarial cross-domain alignment module effective? (2) Is the proposed similar sample cross-domain alignment module effective? Therefore, two baseline models are designed in this embodiment for ablation experiments: DA-SSR-D and DA-SSR-N. Among them, the DA-SSR-D model does not set the adversarial cross-domain alignment module, and the DA-SSR-N model does not set the similar sample cross-domain alignment module.

[0046] Table 2 compares the test sets of the three models with operating condition 1 as the source domain and operating condition 2 as the target domain under the condition of 16 labels 。It can be seen that the effect of the model in this paper is better than that of other models. The results of ADA-SSR-N are poor because ADA-SSR-N only considers the most similar samples between different data domains and does not consider aligning the distribution differences of other data, thus leading to a decrease in accuracy. At the same time, it also shows that the similar sample cross-domain alignment module plays a major role in the method of this chapter. Comparing ADA-SSR-D and SR, the effect is better after setting up the adversarial cross-domain alignment module, which shows the effectiveness of this module.

[0047] Table 2 Results of ablation experiments

[0048] Experiments were carried out under the conditions of 2 labels and 16 labels respectively to try to illustrate two problems: (1) Can RL-SSR improve the effect of the model under the condition of label sparsity? (2) Can RL-SSR improve the effect of the model under multi-condition data? (3) Is RL-SSR better than the existing models? For this reason, five benchmark models were used in this embodiment for experiments: SR, TCA-GMM-PLS, ANN, CL-SSR, and ADA-SSR. It should be specially noted that among them, ANN used a special training method, that is, first supervised training was carried out in the source domain, then a new hidden layer was added, and finally the parameters of the new hidden layer were retrained using the target domain data to achieve the migration from the source domain to the target domain. The remaining methods were trained using the source domain data, and the target domain data was only used for testing.

[0049] Table 3 compares the average of the test sets of RL-SSR and other models under 2 labels and 16 labels , where "1->2 (Test on 1)" means that working condition 1 is used as the source domain, working condition 2 is used as the target domain, and the test data of working condition 1 is used for testing, and so on. Figure 9 The prediction result curves of each model for penicillin concentration under 2 labels are shown, Figure 10 The prediction result curves of each model for penicillin concentration under 16 labels are shown.

[0050] From Table 3, by comparing SR, CL-SSR, ADA-SSR, and RL-SSR, it can be found that RL-SSR achieved the best accuracy, which indicates that the model addressed the issues of label sparsity and multi-condition data. Observing the results of ANN, it can be found that ANN has good performance in the target domain but performs poorly in the source domain. This is because when changing the network topology and continuing training, the adaptability of the model to the source domain is damaged. This can show good performance under the condition of small samples in the target domain, but it cannot take into account the predictions of both the source domain and the target domain at the same time. Further comparing TCA-GMM-PLS, ANN, and RL-SSR, it can be found that RL-SSR has the best effect, proving that the model proposed in this paper is better than the existing methods.

[0051] Table 3 Comparison experiment results

[0052] In this embodiment, for multi - working - condition scenarios, a fermentation product concentration prediction model ADA - SSR based on adversarial domain adaptation is proposed to solve the problem of poor model generalization caused by multi - working - condition data. The proposed ADA - SSR consists of two modules. One module is the adversarial cross - domain feature alignment module, which sets a discriminator to construct adversarial training between the encoder and the discriminator to align the cross - domain data feature representations. The other module is the similar sample cross - domain alignment module, which calculates the Euclidean distance between cross - working - condition data to find the nearest neighbor data and captures the cross - domain consistent representation by minimizing the cosine similarity of the nearest neighbor data. Experiments are carried out on the Indpensim dataset. Through comparative experiments and ablation experiments, it is proved that ADA - SSR can replace traditional domain adaptation methods to solve the multi - working - condition data problem. Through the feature visualization method, it is proved that ADA - SSR aligns the scattered multi - working - condition data distributions into close data distributions, eliminating the feature shift between different working conditions. Through the fusion experiment of RL - SSR, it is proved that the combination of CL - SSR and ADA - SSR can solve both the label sparsity problem and the multi - working - condition data problem, improving the prediction accuracy of the fermentation product concentration.

[0053] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined.

[0054] For method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0055] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A process parameter prediction method based on adversarial domain adaptation, characterized in that It includes the following steps: S1: Obtain the process parameter time series data of the source operating condition domain and the target operating condition domain, where the data includes process variables and corresponding target quality parameter tags; S2: Extract the features of the source domain and target domain data, and use the discriminator combined with the gradient reversal layer to perform adversarial training on the features to align the feature distributions of the source domain and the target domain; S3: Calculate the similarity between the source domain and target domain samples, select the cross-domain nearest neighbor sample pairs, and enhance the inter-domain consistency by minimizing the feature representation distance of the similar sample pairs; S4: Combine the data in step S3 to train step S2 to achieve feature alignment under different operating conditions; S5: Use the data after feature alignment to train the model to generate a prediction model; S6: Input the real-time process parameters of the target operating condition into the prediction model and output the predicted value of the target quality parameter.

2. The method according to claim 1, wherein Step S2 includes: The discriminator distinguishes the source domain and target domain features through the classification loss function; The gradient reversal layer reverses the domain classification loss gradient during backpropagation to drive the feature extraction module to generate domain-invariant features.

3. The method according to claim 1, characterized in that The method for selecting cross-domain nearest neighbor sample pairs includes: Randomly select data in the source domain; Measure the sample similarity based on the Euclidean distance in the feature space; Use the data in the target domain corresponding to the lowest sample similarity and the corresponding data in the source domain as the cross-domain nearest neighbor sample pairs.

4. The method according to claim 3, wherein Measuring the sample similarity based on the Euclidean distance in the feature space includes: First, in the source domain randomly select samples ; Use the Euclidean distance as a metric to measure the label of and find the distance between the process data of all samples in the target domain to find the most similar sample ; Let be the data of the i-th random sample in, be the process data, and be the label data; The calculation formula of the Euclidean distance is as follows: ; where time_step refers to the number of process data time steps available, is the j-th label data of the i-th sample in the source domain data, is the j-th label data of the k-th sample in the target domain.

5. The method according to claim 4, wherein When minimizing the feature representations of similar sample pairs, the loss function is as follows: ; Where sim() represents the cosine similarity function, which enables the model to capture cross-domain consistent representations by minimizing the cosine similarity between nearest neighbor data, where is the feature representation of , and is the feature representation of .

6. The method according to claim 1, wherein Step S2 includes: Extract the features of the source domain data and target domain data respectively through the feature extractor; Reverse the gradient of the feature extractor and use the discriminator for classification; The feature extractor re-extracts features according to the classification result of the discriminator, and the discriminator performs classification. Repeat the above steps for adversarial training to complete the alignment of the feature distributions of the source domain and the target domain.

7. The method according to claim 1, characterized in that, The total loss in step S2 is: ; Among them, is the total amount of data in the source domain , is the non-linear regression head is a random sample in is the feature extractor is the target domain is the total amount of data, R is the gradient reversal layer is the discriminator is the true label of the true working condition of the data sample is the loss function for the prediction stage is the classification score of the discriminator for the i-th sample in the source domain is the classification score of the discriminator for the i-th sample in the target domain 8. The method according to claim 1, wherein In step S4, combine the data in step S3 to train step S2, and the loss function is: ; Among them, and are adjustable hyperparameters, is the loss function in step S2, is the loss function when minimizing the feature representations of similar sample pairs in step 3.

9. A process parameter prediction system based on adversarial domain adaptation, characterized in that, It includes: A data acquisition module for obtaining multi-condition process parameter time series data; An adversarial cross-domain feature alignment module that realizes cross-domain feature distribution alignment through adversarial training; A similarity sample cross-domain alignment module for generating contrast learning constraints for cross-domain similar sample pairs; A joint training module for optimizing the loss functions of the adversarial cross-domain feature alignment module and the similarity sample cross-domain alignment module to generate a cross-operating condition domain adaptive prediction model; A prediction output module that outputs the predicted result of the quality parameter in real time based on the trained model.

10. The system according to claim 9, wherein The adversarial cross-domain feature alignment module includes: Feature extractor , non-linear regression head and discriminator , where the feature extractor and the discriminator introduce the gradient reversal technique to form adversarial learning, promoting the feature extractor to align cross-domain features; Feature extractor and the non-linear regression head are trained for regression using the labels.

11. The system according to claim 10, characterized in that The discriminator uses cross-entropy as the loss function of the discriminator.

12. The system according to claim 9, characterized in that, Optimizing the loss functions of the adversarial cross-domain feature alignment module and the similarity sample cross-domain alignment module includes: The total loss is: ; Among them, and are adjustable hyperparameters, is the loss function of the adversarial cross-domain feature alignment module, is the loss function of the similarity sample cross-domain alignment module; Among them, is: ; where sim() represents the cosine similarity function, and the model captures cross-domain consistent representations by minimizing the cosine similarity between the nearest neighbor data, is the data domain randomly selected samples from the feature representation of, is the data domain in the sample the most similar sample the feature representation of.

13. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

14. A computer-readable storage medium, characterized in that, It stores a computer program executable by an electronic device, and when the computer program runs on the electronic device, the electronic device executes the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Shaping process parameter recommendation method and device, computer equipment and storage medium

    CN119578653A

Cited By

  • Multi-objective collaborative optimization pipe network automatic transmission and distribution control method

    CN121364636A