Multi-production-line Time Series Prediction Method for Performance Indexes in the Esterification Stage of Polyester Fiber

By introducing transfer learning and multi-source domain deep migration networks in the esterification stage of polyester fibers, the difficulty in building a time series prediction model caused by insufficient data on a single production line is solved, and a more accurate and generalized prediction effect is achieved.

CN115238962BActive Publication Date: 2025-06-20DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210722212.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-06-20
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

Under the multi-production line conditions, the historical label data obtained by a single production line is limited, making it difficult to build an accurate and generalized time series prediction model.

Method used

Transfer learning strategies are introduced to enhance dataset variability using historical sensor data from other production lines, thereby helping the target domain to build time series prediction models. The specific solution includes segmenting the source domain time series based on the maximum entropy theory, forming multiple sub-source domains, and using the multi-source domain deep migration network (DA-MDTN) for training to realize multi-source domain migration.

Benefits of technology

Through the combination of transfer learning and multi-source domain deep migration network, the problem of insufficient data is effectively alleviated, the time series prediction performance of performance indicators in the target domain is improved, and the generalization ability of the model is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115238962B_ABST
    Figure CN115238962B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning. A production line with stable operating conditions and continuous production for more than two months is selected as the source domain, and a production line lacking a historical data set and requiring the construction of a time series prediction model is selected as the target domain. A multi-production-line time series prediction model for performance indicators in the esterification stage of polyester fibers based on transfer learning is used to transfer available knowledge from the historical data set in the esterification stage of the source domain to predict the true values of the performance indicators in the esterification stage of the target domain for a period of time in the future. In view of the problem that it is difficult to establish an accurate and strongly generalized time series prediction model due to limited historical label data obtained from a single production line, the present invention introduces a transfer learning strategy, uses historical sensor data from other production lines to increase the data set difference, and fills the data gap caused by the lack of samples, thereby helping the target domain to construct a time series prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the construction of a time series prediction model for a polyester fiber polymerization process under multi-production line conditions, aiming at the problem that it is difficult to construct a time series prediction model due to insufficient time series data obtained for a single production line. Background Art

[0002] Polyester fiber, also known as "polyester", is one of the three major synthetic fibers (polyester, nylon, acrylic). Due to a series of excellent properties such as high strength, high modulus, easy stretching and shrinking, heat resistance and light resistance, it is widely used in production fields such as clothing, textiles, and construction. Currently, it has developed into the largest variety of synthetic fibers in the world, accounting for up to 80% in the fiber market. Polyester fiber refers to a virgin fiber (i.e., polyester fiber) formed by taking terephthalic acid (PTA) and ethylene glycol (EG) as raw materials through a polymerization process to form polyethylene terephthalate (PET) melt, and then through a melt conveying process and a spinning process. The polymerization process is the first process and also the most important link, which includes three major stages: esterification, pre-polycondensation, and final polycondensation, and is a complex process industry process with complex chemical processes and diverse production equipment. Due to the complexity, non-linearity, coupling, and time-variation of the polymerization process, real-time state monitoring and data analysis of the system process are the key to ensuring production efficiency.

[0003] During the production process, a large amount of sensor data obtained based on the monitoring and data acquisition system is time series data, which contains rich historical information of the industrial process. This information can reflect the specific operating conditions of the industrial process and the trend of dynamic changes of indicators over time. Time series prediction is to predict some key variables of the industrial process based on the historical time series data collected by industrial sensors, so as to understand the process trend, correctly evaluate the current state of the system, perceive the abnormal state of the system in advance, and realize the monitoring, control, etc. of process indicators.

[0004] The large-scale deployment of sensors has made great progress in data-driven industrial time series prediction technologies. These technologies rely on historical data collected by sensors to infer the working conditions and health status of future production equipment. Although research shows that these methods have achieved certain effects, they require a large amount (usually a long time span) of historical data to train learning models to achieve accurate prediction. For newly built production lines or production equipment with newly installed sensors, it is often difficult to collect a large enough historical data set in time, and it is often difficult to construct a time series prediction model. To solve the above problems, transfer learning can transfer existing knowledge from a source domain (different production lines, different production conditions) with the same task but different data distributions and sufficient training data to the target domain, thereby helping the target domain establish a time series prediction model.

[0005] Traditional machine learning methods all follow the assumption that the training and test data satisfy the same data distribution. Transfer learning breaks this assumption and can transfer knowledge from related domains (or called source domains) to improve the learning performance of the model in the target domain or minimize the labeled data required to train the learner in the target domain. Summary of the Invention

[0006] The present invention aims to solve the above problems existing in the prior art and provides a multi-production-line time series prediction method for the performance indicators in the esterification stage of polyester fibers. Based on the esterification stage of the polyester fiber polymerization process, in view of the problem that it is difficult to establish an accurate and strong generalization ability time series prediction model due to limited historical label data obtained from a single production line, a transfer learning strategy is introduced to increase the dataset diversity using the historical sensor data of other production lines and fill the data gap caused by the lack of samples, thereby helping to construct a time series prediction model for the target domain.

[0007] The modeling solution adopted by the present invention is as follows: First, the source domain time series is segmented based on the maximum entropy theory to make the data distribution in each sub-source domain approximately invariant. After segmenting the source domain, the multi-sub-source domain and target domain data are input into a domain adversarial-based multi-source deep transfer network (DA-MDTN) for training. The target distribution is regarded as a weighted combination of multiple sub-source domain distributions, and multi-source domain transfer is achieved through the way of domain adversarial. The transfer route of DA-MDTN is: minimizing the distribution difference between the target domain and each sub-source domain through a multi-way adversarial strategy, using the domain confusion score to represent the possibility that the target domain samples belong to different sub-source domains under deep features as the weight coefficient of each prediction; through a training strategy including pre-training, multi-way adversarial learning, and predictor adaptation, making the most of the labeled data in the target domain, improving the prediction performance of the model in the target domain and ensuring the generalization ability of the model.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A multi-production-line time series prediction method for the performance indicators in the esterification stage of polyester fibers based on transfer learning, select a production line with stable working conditions (the polyester product index requirements do not change) and continuous production for more than two months as the source domain, and a production line lacking a historical dataset (only collecting historical data for no more than three days) and requiring to construct a time series prediction model as the target domain. Adopt a multi-production-line time series prediction model for the performance indicators in the esterification stage of polyester fibers based on transfer learning, transfer available knowledge from the historical dataset in the esterification stage of the source domain, combine with the historical dataset in the esterification stage of the target domain, and predict the true values of the performance indicators in the esterification stage of the target domain for a period of time in the future;

[0010] The modeling steps of the multi-production line time series prediction model for the performance indicators in the esterification stage of polyester fibers based on transfer learning are as follows:

[0011] (1) Divide the source domain time series into multiple segments, regarded as multiple sub-source domains, and convert the target domain time series and all sub-source domain time series into time series datasets that can be input into the multi-source domain deep transfer network, namely the target domain time series dataset and multiple sub-source domain time series datasets;

[0012] (2) Input the target domain time series dataset and multiple sub-source domain time series datasets into the multi-source domain deep transfer network (Domain Adversarial-based Multi-source domain Deep Transfer Network, DA-MDTN) for training to obtain a trained network model, and the trained network model can be directly used for the time series prediction task in the esterification stage of the target domain;

[0013] After the modeling is completed, input the process feature time series data in the esterification stage of the target domain into the trained network model to obtain the predicted values of the performance indicators in the esterification stage for multiple future time steps.

[0014] As a preferred technical solution:

[0015] For the multi-production line time series prediction method for the performance indicators in the esterification stage of polyester fibers based on transfer learning as described above, the segmentation of the source domain time series adopts the segmentation method based on maximum entropy. The purpose is to avoid the time series distribution drift problem caused by the large time span of the source domain time series to the greatest extent. The segmentation method based on maximum entropy makes the sample distribution within each segment approximately consistent by finding the most dissimilar time segments after segmentation, that is, the segmentation objective is as follows:

[0016]

[0017] Among them,

[0018] K0 is a predefined parameter to avoid excessive segmentation of the time series;

[0019] K is the number of sub-source domains;

[0020] d is the distance metric function between data distributions;

[0021] D i,j represents a sub-interval, |D i,j | is the interval length;

[0022] n is the length of the source domain time series;

[0023] Δ1 and Δ2 are predefined parameters to avoid the interval being too large or too small, which makes it difficult to capture the distribution information of the interval.

[0024] The learning objective of the above optimization problem is to maximize the average distribution distance by searching for the boundaries of K and the corresponding time segments, so that the distributions of different time segments are as different as possible. The sample distribution within the same time zone can be regarded as unchanged, which can better perform domain adaptation. Consider using the greedy algorithm to solve this optimization problem. The time series is evenly divided into 10 segments, where each segment is the smallest unit period that cannot be further divided, and the value of K is searched in {2, 3, 5, 7, 10}. Given K, let A and B represent the start and end points of the time series. First, consider K = 2 and traverse among the nine candidate segmentation points to find the segmentation point C that maximizes d(D AC , D CB ); after determining point C, consider K = 3 and use the same method to determine the segmentation point D, and sequentially determine K - 1 segmentation points.

[0025] For the multi-production-line time series prediction method of the performance index in the polyester fiber esterification stage based on transfer learning as described above, the process of converting the time series into a time series data set is as follows:

[0026] Multivariate time series data where L is the length of the observation sequence and m is the sequence dimension. (The superscript T represents transpose) is the observation point of X at time j; the specified prediction sequence (i.e., the single sequence to be predicted) in X is where is the observation point of x p at time j; the remaining sequences in X except the specified prediction sequence are non-prediction sequences, and the non-prediction sequences are where is the observation point of x np at time j;

[0027] The multivariate time series data X is converted into a time series data set by the sliding window method where N is the size of the data set; where l is the window size, represents the predicted value of the h-th time step in the future of the prediction sequence, and h is the prediction step. The constructed three-dimensional time series data set H can be directly input into the network for training by the batch processing method.

[0028] For the multi-production-line time series prediction method of the performance index in the polyester fiber esterification stage based on transfer learning as described above, the multi-source domain deep transfer network includes four components: a feature extractor F, a domain discriminator D, a predictor P, and a parameter-free target domain predictor P tBetween the feature extractor F and the domain discriminator D is the Gradient Reversal Layer (GRL); the feature extractor F is shared by the target domain and all sub-source domains, each sub-source domain corresponds to a domain discriminator D and a predictor P, and the final result is output by the target domain predictor P t Output.

[0029] For the multi-production-line time series prediction method of polyester fiber esterification stage performance indicators based on transfer learning as described above, the feature extractor F is based on a recurrent neural network, including two layers of gated recurrent unit (GRU) networks and two layers of fully connected networks. The feature extractor F maps the time series features of the target domain and K sub-source domains into a common feature space, and adapts to the feature distributions of the target domain and each sub-source domain to the greatest extent. During the training process, an adversarial learning method is adopted to obtain the best mapping of F, which can learn the specific relationships between domains while retaining the time series information of the target domain and K sub-source domains to obtain domain-invariant features.

[0030] For the multi-production-line time series prediction method of polyester fiber esterification stage performance indicators based on transfer learning as described above, each sub-source domain corresponds to a domain discriminator D, and the multi-way domain discriminator is denoted as For the multi-way adversarial strategy; for a sample x from the target domain or the j-th sub-source domain, after the features are extracted by the feature extractor F, they are given to the domain discriminator D, and D classifies whether F(x) comes from sub-source domain j or the target domain. The domain label of the source domain is 0, and the domain label of the target domain is 1. Samples from sub-source domain j will only pass through the specific domain discriminator to obtain the discrimination result However, for a sample x from the target domain t will pass through the multi-way domain discriminator to generate K domain discrimination results which are used to update the multi-way domain discriminator

[0031] The inter-domain difference can be measured by the classification loss of the domain discriminator. The smaller the inter-domain difference, the greater the classification loss. Therefore, the multi-way domain discriminator is used to provide the domain confusion score of the target domain sample x t to K predictors The calculation formula is as follows:

[0032]

[0033] Among them,

[0034]

[0035] Among them,

[0036] is the domain discriminator corresponding to the j-th sub-source domain;

[0037] is the equilibrium constant;

[0038] is obtained by taking the average of the losses on the domain discriminator When the probability t that x comes from the sub-source domain j increases, the domain confusion score S c increases, indicating that F(x t ) is similar to the features from the sub-source domain j.

[0039] Each sub-source domain corresponds to a predictor P, and the predictors for multiple sub-source domains (multi-way predictors) are represented as For the time series data from the sub-source domain j, only the predictor is activated and the backpropagation of the gradient is performed; for the time series data x t from the target domain, it will pass through the multi-way predictor and output K prediction results The predictor is first pre-trained with the labeled data of the sub-source domain j, and then trained with the labeled data of the source domain and the target domain after multi-way adversarial learning to adapt to the target domain.

[0040] For the multi-production-line time series prediction method of the performance index of the polyester fiber esterification stage based on transfer learning as described above, the target domain predictor P t finally outputs the prediction results of the target domain samples. It has no parameters. For each target feature F(x t ), the target domain predictor P t takes the domain confusion scores corresponding to each sub-source domain and re-weights the prediction results of the corresponding predictors and then accumulates to obtain the final predicted value P t (F(x t )) The specific calculation formula is as follows:

[0041]

[0042] For the multi-production-line time series prediction method of the performance index of the polyester fiber esterification stage based on transfer learning as described above, the multi-source domain deep transfer network is trained by a three-stage method. The training process includes pre-training, multi-way adversarial learning, and predictor adaptation. First, the feature extractor F and the multi-way predictor are pre-trained with the source domain data Then, the multi-way predictor is fixed and the feature extractor F and the multi-way domain discriminator are updated according to the multi-way adversarial strategy using the labeled source domain and the unlabeled target domain training sets Finally, update the feature extractor F and the multi-way predictor with the labeled data in the source domain and the target domain. And calculate the target domain predictor P. t . The specific process is as follows:

[0043] (1) Pre-training: Jointly train the feature extractor F and the multi-way predictor with the labeled time series datasets of K sub-source domains. The training process in this stage can be expressed as:

[0044]

[0045] Among them,

[0046]

[0047] Among them,

[0048] L ps (F, P) is the prediction loss of the K sub-source domain datasets, calculated by the mean absolute error (MAE);

[0049] y is the true value label corresponding to the sample x;

[0050] After pre-training, F can successfully extract the time series features of the source domain samples, and each sub-source domain corresponding to can also better predict the label values of the corresponding sub-source domain samples. However, the performance of the target domain samples on the pre-trained model is very poor because the target domain distribution has not been adapted to the distributions of each sub-source domain at this time. In the next step, domain adaptation is performed through an adversarial strategy to enable F to extract domain-invariant features;

[0051] (2) Multi-way adversarial training: Fix the multi-way predictor Use the unlabeled time series dataset in the target domain and the labeled time series datasets in the K sub-source domains to update the feature extractor F and the multi-way domain discriminator according to the multi-way adversarial strategy.

[0052] In the adversarial training of each line in the multi-way adversarial strategy, seek the feature extractor F that can maximize the domain discriminator loss, and at the same time seek the domain discriminator D that can minimize the classification loss; at the same time, in order to enable F to retain the ability to extract time series features, add a prediction regression loss to the optimization objective. Then the optimization objective of the multi-way adversarial strategy is as follows:

[0053]

[0054] Among them,

[0055] L dis (F, D) is the domain discriminator classification loss;

[0056] For the prediction loss of the K sub-source domain datasets with the predictor fixed;

[0057]

[0058]

[0059] Among them,

[0060] K is the number of sub-source domains;

[0061] E represents calculating the mean;

[0062] represents the multi-way predictor with its parameters fixed and not participating in parameter update during the backpropagation process, represents the predictor with its parameters fixed;

[0063] During the training process, the domain discriminator tries to correctly distinguish the domain labels of the deep features of the samples (source domain: 0, target domain: 1). At the same time, the feature extractor tries to deceive the domain discriminator so that it cannot correctly classify. F and D learn against each other, and finally the domain discriminator cannot distinguish whether the extracted features come from the sub-source domain or the target domain, achieving the purpose of making F extract domain-invariant features.

[0064] (3) Predictor adaptation, using the labeled time-series datasets in the K sub-source domains and the labeled time-series dataset in the target domain to update the feature extractor F and the multi-way predictor and calculating to obtain the predictor P for the target domain t ; After multi-way adversarial learning, DA-MDTN can already obtain good domain-invariant features, but it may not necessarily achieve good prediction results in the prediction task of the target domain. If you want to apply the source domain predictor in the target domain and make the model not lose its generalization ability, it is required that the predictor performs well in both domains;

[0065] To finally obtain an ideal predictor, fine-tune F and by combining the labeled samples in the source domain and the target domain. Then the training process in the predictor adaptation stage is as follows:

[0066]

[0067] Among them,

[0068] L pst (F, P) is the predictor loss for the target domain dataset and the K sub-source domain datasets:

[0069]

[0070] Among them,

[0071] y is the label corresponding to the source domain sample x;

[0072] y t is the target domain sample x t corresponding label.

[0073] The target domain predictor is obtained by integrating multiple predictors. After the DA-MDTN training is completed, when the unlabeled samples in the target domain arrive, first calculate the domain confusion scores of the corresponding multiple domain discriminators, and then weight the prediction results of each predictor for output.

[0074] The multi-production line time series prediction method for the performance index of the polyester fiber esterification stage based on transfer learning as described above. The historical data set of the esterification stage in the source domain contains sensor data collected every 1 minute for two consecutive months. The historical data set of the esterification stage in the target domain contains sensor data collected every 1 minute for 2 consecutive days. The sensors in the production line serving as the source domain and the production line serving as the target domain are the same. The sensor data includes a prediction sequence and a non-prediction sequence, which is a 14-dimensional multivariate time series. The non-prediction sequence includes: esterification kettle pressure, esterification kettle liquid level, esterification mixer current, process tower vapor pipeline temperature, esterification kettle temperature, esterification separation tower top pressure, esterification separation tower middle pressure, separation tower top temperature, separation tower temperature, separation tower liquid level, esterification separated water reflux flow rate, water reflux pump frequency, and bottom tower EG pump reflux flow rate. The prediction sequence is the density index of the product - oligomer in the esterification stage.

[0075] Beneficial effects

[0076] (1) The present invention adopts a source domain segmentation method based on the maximum entropy theory, and effectively alleviates the problem of temporal distribution drift caused by the large amount of temporal data and long time span in the source domain by segmenting the source time series into multiple segments and then performing migration work.

[0077] (2) The multi-source domain migration network in the present invention learns domain-invariant and transferable feature representations through multi-way adversarial training, and at the same time uses the domain confusion score to represent the probability that a target domain sample belongs to a certain sub-source domain, realizing the construction of a fine-grained integrated prediction model.

[0078] (3) The network training in the present invention adopts a three-stage training method to effectively improve the training efficiency, and adds a predictor adaptation stage after multi-way adversarial training, ensuring the prediction performance of the model in the target domain while not losing the generalization ability of the model. Description of the drawings

[0079] Figure 1 is a multi-production line time series prediction model modeling method for the performance index of the polyester fiber esterification stage based on transfer learning;

[0080] Figure 2 is the DA-MDTN network structure;

[0081] Figure 3 is the DA-MDTN training and validation process;

[0082] Figure 4 is the single-step prediction result of two multi-source domain transfer models;

[0083] Figure 5 is the prediction error box plot. Detailed implementation manners

[0084] The present invention will be further described below in conjunction with the detailed implementation manners. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0085] As Figure 1 shown, a multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning specifically comprises the following steps:

[0086] (1) Acquisition of the target domain time series dataset and multiple sub-source domain time series datasets

[0087] (1.1) Select a production line with stable working conditions (the polyester product index requirements do not change) and continuous production for more than two months as the source domain, and take the sensor data collected every 1 minute for two consecutive months of this production line as the original source domain time series;

[0088] (1.2) Select a production line lacking a historical dataset (only collecting historical data for no more than three days) and requiring the construction of a time series prediction model as the target domain, and take the sensor data collected every 1 minute for 2 consecutive days of this production line as the original target domain time series;

[0089] Among them, the sensors in the production line serving as the target domain are the same as those in the production line serving as the source domain; the sensor data is a 14-dimensional multi-source time series, including a prediction sequence and a non-prediction sequence. The non-prediction sequence includes: esterification kettle pressure, esterification kettle liquid level, esterification agitator current, process tower vapor pipeline temperature, esterification kettle temperature, esterification separation tower top pressure, esterification separation tower middle pressure, separation tower top temperature, separation tower temperature, separation tower liquid level, esterification separated water reflux flow rate, water reflux pump frequency, and tower bottom EG pump reflux flow rate. The prediction sequence is the density index of the product - oligomer in the esterification stage;

[0090] (1.3) Preprocess the data of the original source domain time series and the original target domain time series to obtain the source domain time series and the target domain time series;

[0091] Among them, the preprocessing includes outlier detection, data smoothing, and normalization;

[0092] During outlier detection, outliers are generally defined as observed values that deviate too much from the normal working conditions due to uncontrollable factors such as problems with measuring instruments or abnormal sensor working states. Generally, a time window ω is defined, and within this time window, the 3δ rule is used to detect outliers and replace them with the average of the previous time step and the next time step. Specifically, for all feature dimensions of the time series m is the dimension, and calculate the standard deviation δ of X within the time window ω j of j When the observed value j within X meets the condition:

[0093]

[0094] where is the mean of the sequence X j , it can be regarded as an outlier, and then is replaced with Outlier detection is performed on all time windows in turn. In the present invention, the size ω of the time window is selected as the time series length of one day;

[0095] During data smoothing, the moving average method is used to minimize the influence of noise. The moving average method constructs a new time series by successively taking the average value within the original sequence window. Specifically:

[0096]

[0097] where is the data at the new time T obtained by the moving average; x T is the data at time T of the original sequence; w is the size of the moving window. In the present invention, the size w of the moving window is selected as 5;

[0098] The z-score method is used to normalize the feature data. For all feature dimensions of the time series The normalization calculation formula is as follows:

[0099]

[0100] where is the original data after normalization; is the mean of the sequence X j ; δ j is the standard deviation of the sequence X j ;

[0101] (1.4) Adopt the segmentation method based on maximum entropy to segment the source domain time series into K sub - intervals, making the sample distribution within each interval approximately the same, and regarding them as K sub - source domains;

[0102] The objectives of segmentation are as follows:

[0103]

[0104] Among them, K0 is a predefined parameter to avoid over - segmentation of the time series; K is the number of sub - source domains; d is the distance metric function between data distributions. In this invention, d is selected as the maximum mean discrepancy (MMD) function; D i,j represents a sub - interval, and |D i,j | is the interval length; n is the length of the source domain time series; Δ1 and Δ2 are predefined parameters to avoid it being difficult to capture the distribution information of the interval due to the interval being too large or too small;

[0105] Maximize the average distribution distance by searching for K and the boundaries of the corresponding time segments, making the distributions of different time segments as different as possible, and the sample distribution within the same time zone can be regarded as unchanged, which can better perform domain adaptation; Consider using the greedy algorithm to solve the segmentation problem. Divide the time series evenly into 10 segments, where each segment is the smallest unit period that cannot be further segmented, and search for the value of K in {2, 3, 5, 7, 10}; Given K, use A and B to represent the start and end points of the time series. First, consider K = 2, traverse among the nine candidate segmentation points to find the segmentation point C that maximizes d(D AC , D CB ); After determining point C, consider K = 3, and use the same method to determine the segmentation point D, and successively determine K - 1 segmentation points;

[0106] (1.5) Convert the target domain time series and all sub - source domain time series into time - series data sets, obtaining the target domain time - series data set and multiple sub - source domain time - series data sets;

[0107] Among them, the process of converting the time series into a time - series data set is as follows:

[0108] Multivariate time - series data where L is the length of the observation sequence and m is the sequence dimension, is the observation point of X at the j - th moment; The specified prediction sequence (i.e., the single sequence to be predicted) in X is where is the observation point of x p at the j - th moment; The remaining sequences in X except the specified prediction sequence are non - prediction sequences, and the non - prediction sequences are where is the observation point of x np at the j - th moment;

[0109] The multivariate time series data X is converted into a time series dataset by the sliding window method where N is the size of the dataset; where l is the window size, represents the predicted value of the h-th time step in the future of the prediction sequence, and h is the prediction step size; the constructed three-dimensional time series dataset H can be directly input into the network for training through the batch processing method;

[0110] (2) Training of the network model

[0111] Input the target domain time series dataset and multiple sub-source domain time series datasets obtained in step (1) into the Domain Adversarial-based Multi-source domain Deep Transfer Network (DA-MDTN) for training, and then obtain the trained network model according to the three-stage training method. The trained network model can be directly used for the time series prediction task in the esterification stage of the target domain;

[0112] Among them, the multi-source domain deep transfer network structure is as Figure 2 shown. The multi-source domain deep transfer network contains four components: a feature extractor F, a domain discriminator D, a predictor P, and a parameter-free target domain predictor P t , and there is a Gradient Reversal Layer (GRL) between the feature extractor F and the domain discriminator D. The target domain and all sub-source domains share the feature extractor F. Each sub-source domain corresponds to a domain discriminator D and a predictor P, and the final result is output by the target domain predictor P t ;

[0113] The feature extractor F is based on a recurrent neural network and contains two layers of gated recurrent unit (GRU) networks and two layers of fully connected networks. The feature extractor F maps the time series features of the target domain and K sub-source domains into a common feature space, adapting to the feature distributions of the target domain and each sub-source domain to the greatest extent; during the training process, an adversarial learning method is adopted to obtain the best mapping of F, which can learn the specific relationships between domains while retaining the time series information of the target domain and K sub-source domains to obtain domain-invariant features;

[0114] Each sub-source domain corresponds to a domain discriminator D, so the domain discriminators of multiple sub-source domains (multi-way domain discriminators) are represented as for the multi-way adversarial strategy; for a sample x from the target domain or the j-th sub-source domain, after extracting features through the feature extractor F, it is handed over to the domain discriminator D. D classifies whether F(x) comes from sub-source domain j or the target domain. The domain label of the source domain is 0, and the domain label of the target domain is 1; samples from sub-source domain j will only pass through the specific domain discriminator to obtain the discrimination result However, the sample x from the target domain t will pass through the multi-way domain discriminator to generate K domain discrimination results which are used to update the multi-way domain discriminator The inter-domain difference can be measured by the classification loss of the domain discriminator. The smaller the inter-domain difference, the larger the classification loss. Therefore, the multi-way domain discriminator is used to provide the domain confusion score of the target domain sample x to K predictors t The calculation formula of which is as follows: The calculation formula of which is as follows:

[0115]

[0116] where

[0117]

[0118] where is the domain discriminator corresponding to the j-th sub-source domain; is the balance constant; is the number of samples in the j-th sub-source domain; is obtained by taking the average of the losses on the domain discriminator When the probability that x t comes from the sub-source domain j increases, the domain confusion score S c increases, indicating that F(x t ) is similar to the features from the sub-source domain j;

[0119] Each sub-source domain corresponds to a predictor P, and the multi-way predictor is represented as For the time series data from the sub-source domain j, only the predictor is activated and the gradient is backpropagated; for the time series data x t from the target domain, it will pass through the multi-way predictor to output K prediction results The predictor is first pre-trained with the labeled data of the sub-source domain j, and then trained with the labeled data of the source domain and the target domain after multi-way adversarial learning to adapt to the target domain;

[0120] The target domain predictor P t finally outputs the prediction result of the target domain sample, which has no parameters. For each target feature F(x t ), the target domain predictor P t takes the domain confusion score corresponding to each sub-source domain to re-weight the prediction results of the corresponding predictors and then accumulates them to obtain the final predicted value P t(F(x t )),The specific calculation formula is as follows:

[0121]

[0122] The specific parameters of the multi-source domain deep transfer network are shown in Table 1;

[0123] Table 1 Specific parameters of the multi-source domain deep transfer network

[0124]

[0125]

[0126] The training process of the multi-source domain deep transfer network includes pre-training, multi-way adversarial training, and predictor adaptation;

[0127] In the pre-training stage, the feature extractor F and the multi-way predictor are jointly trained using the labeled time-series datasets in K sub-source domains The number of training rounds is set to 30; the training process of this stage can be expressed as:

[0128]

[0129] Among them,

[0130]

[0131] Among them, L ps (F, P) is the prediction loss of the K sub-source domain datasets, calculated by the mean absolute error (MAE); y is the true value label corresponding to the sample x;

[0132] In the multi-way adversarial training stage, the multi-way predictor is fixed The feature extractor F and the multi-way domain discriminator are updated respectively according to the multi-way adversarial strategy using the unlabeled time-series dataset in the target domain and the labeled time-series datasets in K sub-source domains In the multi-way adversarial strategy, the adversarial training of each line seeks the feature extractor F that can maximize the domain discriminator loss, and at the same time seeks the domain discriminator D that can minimize the classification loss; at the same time, in order to enable F to retain the ability to extract time-series features, the prediction regression loss is added to the optimization objective, then the optimization objective of the multi-way adversarial strategy is as follows:

[0133]

[0134] Among them, L dis (F, D) is the domain discriminator classification loss; is the prediction loss of the K sub-source domain datasets under the fixed predictor;

[0135]

[0136]

[0137] Among them, K is the number of sub-source domains; E represents the calculation of the mean; represents a multi-way predictor whose parameters are fixed and do not participate in parameter updates during the backpropagation process, represents a predictor whose parameters are fixed;

[0138] During the training process, the domain discriminator tries to correctly distinguish the domain labels of the deep features of the samples (source domain: 0, target domain: 1). At the same time, the feature extractor tries to deceive the domain discriminator so that it cannot correctly classify. F and D learn adversarially with each other, and finally the domain discriminator cannot distinguish whether the extracted features come from the sub-source domain or the target domain, achieving the purpose of making F extract domain-invariant features;

[0139] During the predictor adaptation stage, the feature extractor F and the multi-way predictor are updated respectively using the labeled time-series datasets in the K sub-source domains and the labeled time-series datasets in the target domain The number of training rounds is set to 30, and then the target domain predictor P is calculated t ;

[0140] Among them, the target domain predictor is obtained by integrating the multi-way predictor. The training process in the predictor adaptation stage is as follows:

[0141]

[0142] Among them, L pst (F, P) is the prediction loss of the target domain dataset and the K sub-source domain datasets:

[0143]

[0144] Among them, y is the label corresponding to the source domain sample x, and y t is the label corresponding to the target domain sample x t ;

[0145] (3) Prediction of the performance index of the esterification stage for multiple future time steps

[0146] Input the process feature time-series data of the esterification stage in the target domain into the trained network model. When the unlabeled samples in the target domain arrive, the prediction results of the predictor are output. First, calculate the domain confusion scores of the corresponding multiple domain discriminators, and then weight the predicted values of each path for multiple future time steps;

[0147] Among them, the training and validation process of DA-MDTN is as Figure 3 shown.

[0148] Introduce the setting criteria for two comparative experiments: 1) The single best sub-source domain: In the multi-source domain, take out the sub-source domain with the best migration effect for the comparative experiment; 2) Source combine: Aggregate multiple sub-source domains into a single source domain (in our experiment, it means not splitting the source domain) for the comparative experiment; 3) Multi-source: Do not process the multi-source domain; The first setting criterion can verify whether introducing other source domains is valuable for improving single-source domain migration, and the second setting criterion can verify whether the multi-source domain migration structure is valuable. The migration experiment results of different setting criteria are shown in Table 2;

[0149] Table 2 Migration experiment results of different models

[0150]

[0151] In Table 2, TCA-LSTM is the combination of the traditional migration method Transfer Component Analysis and the LSTM network, DTN (Deep Transfer Network) is the deep transfer network, SDS-DTM (SourceDomain Segmentation-based Deep Transfer Model) is the deep transfer model based on source domain segmentation, the Only source domain method means training the model only with source domain data and then directly using it for the prediction task of the target domain, and DANN (Domain Adversarial Neural Network) is the single-source domain migration method based on domain adversarial; DANN achieved the best migration effect under the first two setting criteria. Compared with DTN that uses the distribution distance as the training loss regularization term, DANN can better extract the domain-invariant features of samples through adversarial learning; Under the second setting criterion, all sub-source domains are combined into one source domain, which is equivalent to not splitting the source domain. Due to the influence of the time series distribution drift, the prediction accuracy of all methods is much worse than that under the Single best criterion; Under the third multi-source setting criterion, the prediction errors of the two models SDS-DTM and DA-MDTN for the time series data of the target domain are smaller than the prediction error of DANN, which has the best effect under the Single best criterion, indicating that introducing other source domains can improve the single-source domain migration effect; At the same time, the prediction effect of DANN, which has the best effect under the Source combine criterion, is far inferior to that of SDS-DTM and DA-MDTN, indicating that the multi-source domain migration structure has research value for the migration task between polyester esterification stage production lines.

[0152] Figure 4The single-step prediction curves of SDS-DTM and DA-MDTN are given to more intuitively show the prediction effect; as shown in the figure, the prediction curve of DA-MDTN fits better with the true value curve, and it can capture the change in the sequence trend more timely and make adjustments near the peak of the sequence.

[0153] Figure 5 The comparison of the absolute prediction errors of multi-source domain migration and different single-source domain migrations is given. It can be seen from the error box plot that the multi-source domain migration model of DA-MDTN has a better prediction effect than the single-source domain migration model of DANN, indicating that DA-MDTN can combine the useful knowledge of each sub-source domain for migration, verifying the effectiveness of the multi-source domain migration model.

Claims

1. A multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning, characterized in that: Select a production line with stable operating conditions and continuous production for more than two months as the source domain, and a production line lacking historical data sets as the target domain. Adopt a multi-production line time series prediction model for the performance indicators in the esterification stage of polyester fibers based on transfer learning to transfer available knowledge from the historical data set in the esterification stage of the source domain and predict the true values of the performance indicators in the esterification stage of the target domain for a period of time in the future; The modeling steps of the multi-production line time series prediction model for the performance indicators in the esterification stage of polyester fibers based on transfer learning are as follows: (1) Divide the source domain time series into multiple segments, regarded as multiple sub-source domains, and convert the target domain time series and all sub-source domain time series into time series data sets; (2) Input the target domain time series data set and multiple sub-source domain time series data sets into a multi-source domain deep transfer network for training to obtain a trained network model.

2. The multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 1, characterized in that, The segmentation of the source domain time series adopts a segmentation method based on maximum entropy, and the segmentation objectives are as follows: Among them, K0 is a predefined parameter to avoid excessive segmentation of the time series; K is the number of sub-source domains; d is a distance metric function between data distributions; Indicates a sub-interval, is the interval length; n is the length of the source domain time series; Δ1 and Δ2 are predefined parameters to avoid being difficult to capture the distribution information of the interval due to the interval being too large or too small.

3. The multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 2, characterized in that, The process of converting the time series into a time series data set is as follows: Multivariate time series data where L is the length of the observation sequence and m is the sequence dimension, is the observation point of X at time j; the specified prediction sequence in X is where is the observation point of x p at time j; the remaining sequences in X other than the specified prediction sequence are non-prediction sequences, and the non-prediction sequences are where is the observation point of x np at time j; The multivariate time series data X is converted into a time series dataset by the sliding window method where N is the size of the dataset; where l is the window size, represents the predicted value of the h-th time step in the future of the prediction sequence, and h is the prediction step size.

4. The multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 3, characterized in that, The multi-source domain deep transfer network includes four components: a feature extractor F, a domain discriminator D, a predictor P, and a parameter-free target domain predictor P t , and there is a gradient reversal layer between the feature extractor F and the domain discriminator D; the feature extractor F is shared by the target domain and all sub-source domains, each sub-source domain corresponds to a domain discriminator D and a predictor P, and the final result is output by the target domain predictor P t Output.

5. The multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 4, characterized in that, The feature extractor F takes a recurrent neural network as the backbone, including two layers of gated recurrent unit networks and two layers of fully connected networks. The feature extractor F maps the time series features of the target domain and K sub-source domains into a common feature space to best adapt to the feature distributions of the target domain and each sub-source domain.

6. The multi-production-line time series prediction method for performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 5, characterized in that, The multi-way domain discriminator is used to provide the target domain sample x to K predictors t with the domain confusion score The calculation formula is as follows: Among them, Among them, is the domain discriminator corresponding to the j-th sub-source domain; is the equilibrium constant; is the number of samples in the j-th sub-source domain; is obtained by averaging the loss on the domain discriminator . When the probability t that x comes from the j-th sub-source domain increases, the domain confusion score S c increases, indicating that F(x t ) is similar to the features from the j-th sub-source domain.

7. The multi-production-line time series prediction method for the performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 6, wherein, For each target feature F(x t ), the target domain predictor P t takes the domain confusion scores corresponding to each sub-source domain to re-weight the prediction results of the corresponding predictors and then accumulates to obtain the final predicted value P t (F(x t )) with the specific calculation formula as follows:

8. The multi-production-line time series prediction method for the performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 7, wherein, The multi-source domain deep transfer network is trained by a three-stage method: (1) Pre-training: Jointly train the feature extractor F and the multi-way predictor with the labeled time-series datasets in K sub-source domains (2) Multi-channel adversarial training: Fix the multi-channel predictor Use the unlabeled time-series dataset in the target domain and the labeled time-series datasets in K sub-source domains to update the feature extractor F and the multi-channel domain discriminator respectively according to the multi-channel adversarial strategy In the multi-way adversarial strategy, the adversarial training of each line seeks a feature extractor F that can maximize the loss of the domain discriminator, and at the same time seeks a domain discriminator D that can minimize the classification loss; the optimization objectives of the multi-way adversarial strategy are as follows: Among them, is the classification loss of the domain discriminator; To fix the average prediction loss of the K sub-source domain datasets for the predictor; Among them, Indicates calculating the mean value; Indicates that the parameters of the predictor are fixed and do not participate in parameter updates during backpropagation, Indicates that the parameters of the predictor are fixed; (3) Predictor adaptation: Update the feature extractor F and the multi-way predictor using the labeled time-series datasets in the K sub-source domains and the labeled time-series dataset in the target domain respectively, and calculate the target domain predictor P. And calculate the target domain predictor P. t .

9. The multi-production-line time series prediction method for the performance indicators in the esterification stage of polyester fibers based on transfer learning according to claim 8, wherein, The historical data set in the esterification stage of the source domain contains sensor data collected every 1 minute for two consecutive months. The historical data set in the esterification stage of the target domain contains sensor data collected every 1 minute for 2 consecutive days; the sensors in the production line used as the source domain and the production line used as the target domain are the same. The sensor data includes a prediction sequence and a non-prediction sequence, which is a 14-dimensional multivariate time series. The non-prediction sequence includes: esterification kettle pressure, esterification kettle liquid level, esterification agitator current, process tower vapor pipeline temperature, esterification kettle temperature, esterification separation tower top pressure, esterification separation tower middle pressure, separation tower top temperature, separation tower temperature, separation tower liquid level, esterification separation water reflux flow rate, water reflux pump frequency, and bottom tower EG pump reflux flow rate. The prediction sequence is the density index of the product - oligomer in the esterification stage.