Printing and dyeing process sewage discharge prediction method based on improved Informer model

By pre-processing the sewage discharge data of printing and dyeing enterprises and using the improved Informer model to predict, the problem of large error in sewage discharge prediction in the prior art is solved, and high-accurate sewage discharge prediction is achieved.

CN120069173APending Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510077160.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has large errors when predicting the wastewater discharge of printing and dyeing processes, and fails to effectively consider the interaction of complex nonlinear relationships and multivariate data.

Method used

By preprocessing the daily sewage discharge data of printing and dyeing enterprises, a training data set is generated, and training is performed using the improved Informer model to predict future sewage discharges.

Benefits of technology

It has achieved high accuracy in sewage emission forecasting, helping printing and dyeing companies to reduce sewage emissions more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069173A_ABST
    Figure CN120069173A_ABST
Patent Text Reader

Abstract

The invention discloses a printing and dyeing process sewage discharge prediction method based on an improved Informer model. The method comprises the following steps: step 1, collecting a sewage discharge data table in a daily production process of a printing and dyeing enterprise; 2, performing data cleaning on the data table, and performing data preprocessing operation by using methods of missing value processing, abnormal value processing, data standardization and the like to obtain a training data set; step 3, using the training data set to train a sewage discharge prediction model based on the improved Informer printing and dyeing process; step 4, carrying out back propagation on the training error, updating a model weight parameter, and storing the model after training is finished; and step 5, using the trained model to predict sewage discharge conditions in the next few days. According to the method, the sewage discharge amount of the printing and dyeing process is predicted by using the improved Informer model, and the method has relatively high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting the sewage discharge volume of a printing and dyeing process based on an improved Informer model. Technical Background

[0002] The printing and dyeing process is an important process in the textile industry, mainly using various dyes, auxiliaries and chemicals to dye and print fabrics. In the printing and dyeing process, water is an indispensable element, used in processes such as dyeing, printing, washing and shaping. The wastewater generated in these processes contains various harmful substances, such as dye residues, auxiliaries, heavy metals and salts, which pose a potential threat to the environment. With the rapid development of the textile industry, the supervision and emission reduction of the sewage discharge volume of enterprises' printing and dyeing have become more urgent. The prediction of the sewage discharge volume of printing and dyeing is of great significance for enterprise production planning, environmental protection, compliance testing and resource planning.

[0003] Traditional prediction methods are mainly based on statistical features and mathematical models, but do not consider complex non-linear relationships and the interaction of multivariate data, often resulting in large errors in prediction results. At present, the printing and dyeing industry generally adopts measures such as optimizing the process flow, workshop scheduling, and updating equipment to achieve a certain degree of energy conservation and emission reduction, but the effect is relatively limited. With the development of big data and Internet of Things technologies, establishing a prediction model based on deep learning to analyze sewage discharge volume data can help printing and dyeing enterprises predict the sewage discharge volume of the printing and dyeing process in their future production processes, so as to achieve better sewage emission reduction effects. Summary of the Invention

[0004] In order to overcome the limitations of existing methods in predicting the sewage discharge volume of the printing and dyeing process, the present invention preprocesses the data related to the daily sewage discharge volume of printing and dyeing enterprises and uses it to train an improved Informer model, and uses this model to predict the sewage discharge volume of the printing and dyeing process, with high accuracy.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] A method for predicting the sewage discharge volume of a printing and dyeing process based on an improved Informer model, the method for predicting the sewage discharge volume of the printing and dyeing process includes the following steps:

[0007] Step 1, collect the sewage discharge volume data table Data in the daily production process of a printing and dyeing enterprise, including data for N days with the following 9 fields: date t, water consumption w 1 , water consumption for printing process w 2 , water consumption for dyeing process w 3 , water saving for printing process w 4 , hot water saving for printing process w 5, domestic water w 6 , printing sewage discharge w 7 and dyeing sewage discharge w 8 ;

[0008] Step 2: Clean the data in the data table Data, and perform data preprocessing operations using methods for handling missing values, handling outliers, and data standardization to obtain the training dataset S = {S i |i = 1,.., N}, where S i represents the relevant data for the i-th day, and S i is expressed as follows:

[0009] S i = {t, w 1 , w 2 , w 3 , w 4 , w 5 , w 6 , w 7 , w 8}(i ∈ [1, N]);

[0010] Step 3: Use the training dataset S to train a printing and dyeing process sewage discharge prediction model based on the improved Informer;

[0011] Step 4: Backpropagate the training error Loss kernel-mse to update the model weight parameters. Each time training is performed, the sliding window is slid back by 1 day. After the sliding window slides to the last batch of data, one iteration ends. After the model is trained and iterated E times, save the trained model Model;

[0012] Step 5: Use the trained model Model to predict the printing and dyeing sewage discharge data for the next m days. The user inputs the date T, and the model predicts the printing and dyeing sewage discharge data for m days after date T.

[0013] Furthermore, in the above Step 2, the data preprocessing process is as follows:

[0014] (2.1) Handling of missing data: If more than 60% of the feature data in Data is missing, i.e., the field value is empty, then delete the discharge record; otherwise, fill the missing feature values with the mean;

[0015] (2.2) Handling of abnormal data: If there is a situation where the feature data in Data differs from the mean by several times, or the feature data is negative, then replace the discharge record with the mean;

[0016] (2.3) Perform standardization operations with a mean of 0 and a variance of 1 on the processed data.

[0017] Furthermore, in step 3, the detailed process of model training is as follows:

[0018] (3.1) Each time of training, data is taken from the training set S based on the sliding window mechanism to generate the encoder input X, decoder input Y, and label output Y of the model. label , initialize the sliding window size as n, that is, the printing and dyeing sewage discharge data for n days. The model predicts the printing and dyeing sewage discharge data for m days after n days, where m < n; take a sliding window, that is, n days of data in S, and use the date, water consumption, water consumption for printing process, water consumption for dyeing process, water saving for printing process, hot water saving for printing process, and domestic water as the encoder input X = {X i | i = 1, 2,..., n}, where X i represents the data of the i-th day:

[0019] X i = {t, w 1 , w 2 , w 3 , w 4 , w 5 , w 6}(i ∈ [1, n])

[0020] Take the data of the last m days in the sliding window and the data of m days filled with 0, a total of 2m days of data. Use the date t, printing sewage discharge w 7 , and dyeing sewage discharge w 8 as the decoder input Y = {Y i | i = 1, 2,..., 2m}, where Y i represents the data of the i-th day:

[0021]

[0022] Take the data of the last m days after the sliding window as the label output Y label , that is, the real data of m days after n days that the model needs to predict;

[0023] (3.2) Input X into the encoder and Y into the decoder, perform the same feature embedding EMB with dimension d on the input X and Y, and obtain X emb and Y emb respectively. The feature embedding EMB includes scalar embedding U, position embedding PE, and date embedding SE, and the formula is as follows:

[0024] EMB = U + PE + SE

[0025] Among them, U represents one-dimensional convolution on the data features in the input data except for the date t, and PE represents calculating the positional embedding of the input data in this batch of data. The positional encoding corresponding to the position p is represented in the odd and even positions, and the formula is as follows:

[0026]

[0027] SE represents calculating the date embedding for the date t in the input data, and the date embedding is represented as follows:

[0028] SE = [month, week, day]

[0029] Among them, month represents the month part of the date t, week represents the day of the week, and day represents the day part of the date t;

[0030] (3.3) The encoder calculates the probability sparse self-attention for the feature-embedded X emb The probability sparse self-attention formula is as follows: attention

[0031]

[0032] Among them, Q, K, and V represent the parameter matrices Query, Key, and Value obtained by performing different linear transformations on the input X emb Since the different dot products q*k in Query and Key have different effects on self-attention, fewer dot products contribute the vast majority of the self-attention scores, that is, the probability distribution of the self-attention mechanism is sparse; the Top-K queries with the largest contribution are found by calculating the KL divergence between the attention probability distribution of each query in the Query matrix and the uniform distribution to form The formula for calculating the KL divergence is as follows:

[0033]

[0034] Among them, q i is the value of the Q matrix, is the value of the K matrix, is the different dot product, d is the dimension of the feature embedding of X emb The smaller , the higher the similarity between the query q i and the uniform distribution, and the lower the contribution to the self-attention score. The Top-K with the lowest similarity are found to form

[0035] (3.4) Using the attention distillation operation to assign higher weights to the dominant attention-based dominant features to obtain the main feature X feature The attention distillation formula is as follows:​

[0036] X feature = MaxPool(Elu(Conv1d(X attention )))

[0037] Where Conv1d represents a one-dimensional convolutional operation on the time series, uses the Elu activation function, and finally performs a max pooling operation;

[0038] (3.5) The model takes X emb Sequences of length 1 / 2 and 1 / 4 as copies and inputs them into the encoder to obtain the corresponding copy features. Finally, the obtained copy features are fused with the main features to obtain the output features X feature map ;

[0039] (3.6) The decoder calculates the probability sparse self-attention for the feature-embedded Y emb to obtain Y attention and then calculates the attention scores with the output X feature map of the encoder. Finally, a prediction output Y prediction is obtained through a fully connected layer;

[0040] (3.7) Use the Kemel-MSE loss function to calculate the Loss prediction between Y label and Y kernel-mse , and the Kemel-MSE loss function is expressed as follows:

[0041]

[0042] Furthermore, in step 5, the prediction process is as follows:

[0043] (5.1) The user inputs a date T, obtains the printing and dyeing sewage discharge data S1 for n days before date T, and processes S1 into the encoder input X;

[0044] (5.2) Obtain the decoder input Y consisting of the data S2 for m days before date T and the data filled with 0 for m days;

[0045] (5.3) Input X and Y into the model Model to predict and obtain Y prediction ;

[0046] (5.4) After inverse normalizing Y prediction , obtain the prediction results of the printing and dyeing sewage discharge for the next m days after date T.

[0047] The technical concept of the present invention is as follows: based on the daily sewage discharge data provided by printing and dyeing enterprises, data preprocessing is carried out to generate a training dataset, and then the improved Informer model is trained to obtain a printing and dyeing sewage discharge prediction model, which is used to predict the sewage discharge situation of printing and dyeing enterprises in the printing and dyeing process link in the future for a period of time.

[0048] The beneficial effect of the present invention is: it has high accuracy. Brief Description of the Drawings

[0049] Figure 1 It is the overall flowchart for implementing the printing and dyeing process sewage discharge prediction method based on the improved Informer model of the present invention.

[0050] Figure 2 It is the flowchart of the data preprocessing stage.

[0051] Figure 3 It is the network structure diagram of the Informer model.

[0052] Figure 4 It is the training flowchart of the printing and dyeing sewage discharge prediction model.

[0053] Figure 5 It is the prediction flowchart of the printing and dyeing sewage discharge prediction model. Detailed Embodiment

[0054] The following further describes the present invention with reference to the drawings.

[0055] Refer to Figures 1 to 5 , a printing and dyeing process sewage discharge prediction method based on an improved Informer model: according to the daily sewage discharge data of the printing and dyeing process provided by printing and dyeing enterprises, predict the sewage discharge situation of orders in the printing and dyeing link of printing and dyeing enterprises in the future for a period of time. The printing and dyeing setting sewage discharge prediction method includes the following steps:

[0056] Step 1: Collect the sewage discharge data table Data during the daily production process of printing and dyeing enterprises;

[0057] Table 1 is the description of the basic information of the printing and dyeing sewage discharge data table of printing and dyeing enterprises:

[0058]

[0059] Table 1

[0060] Step 2: Clean the data table Data, and perform preprocessing operations such as dealing with missing values, dealing with outliers, and data standardization to obtain the preprocessed training dataset S;

[0061] The processing process of the data preprocessing is as follows:

[0062] (2.1) Handling of missing data: Table 2 shows a partial data sample of the printing and dyeing wastewater discharge data table. It can be seen that in the first and second sample data, both are null values, with more than 60% of the information missing in the data, so they are deleted; in the third, fifth, and sixth data samples, some data are missing. At this time, the method of mean imputation can be used to handle the missing values.

[0063]

[0064] Table 2

[0065] (2.2) Handling of abnormal data: If the characteristic data of printing and dyeing has a situation where it differs from the mean by several times, or the wastewater discharge is negative, that is, there are abnormal data in the characteristic values, then the discharge record will be replaced with the mean. As shown in Table 2, in the data of 2022-03-01, the discharge of printed polluted water is 13362.0, which is obviously several times the mean, and the data is abnormal; in the data of 2022-02-28, 2022-03-02, and 2022-03-03, the discharge of printed wastewater is negative, and the data is abnormal. For abnormal data, the mean is also used to replace the abnormal data items.

[0066] (2.3) Standardization operation with a mean of 0 and a variance of 1 for the data: Table 3 shows a partial data sample of the training dataset S obtained after the standardization operation on the wastewater discharge data table:

[0067]

[0068] Table 3

[0069] Step 3: Use the training dataset S to train the printing and dyeing process wastewater discharge prediction model based on the improved Informer; the process is as follows:

[0070] (3.1) First, set the sliding window size n = 32, that is, 32 days of data, and predict the data for the next m = 5 days. The process of one training is described below. Table 4 is the encoder input X generated by the sliding window. Table 5 is the decoder input Y generated by the sliding window, and Table 6 is the label output Y label , that is, the real data for the next 5 days predicted by the model.

[0071]

[0072]

[0073] Table 4

[0074] 2022-01-28 -0.717199 1.023858 ... ... ... 2022-02-01 -0.472890 1.325916 2022-02-02 0 0 ... ... ... 2022-02-06 0 0

[0075] Table 5

[0076] 2022-02-02 -0.640048 0.555931 ... ... ... 2022-02-06 -0.627190 -0.005593

[0077] Table 6

[0078] (3.2) Input X into the encoder and input Y into the decoder. Perform the same feature embedding EMB on the input X and Y respectively to obtain X emb and X emb

[0079] (3.3) The encoder calculates the probability sparse self-attention for the feature-embedded X emb to obtain X attention .

[0080] (3.4) Use the attention distillation operation to assign higher weights to the dominant features with dominant attention to obtain the main feature X feature .

[0081] (3.5) Take 1 / 2 and 1 / 4 of the sequence length of X emb as copies and input them into the encoder to obtain the corresponding copy features. Finally, fuse the obtained copy features with the main feature to obtain the output feature X feature map of the final encoder.

[0082] (3.6) The decoder calculates the probability sparse self-attention for the feature-embedded Y emb to obtain T attention and then calculates the attention score with the output X feature map of the encoder. Finally, obtain the predicted output Y prediction through a fully connected layer.

[0083] (3.7) Calculate the Loss prediction between Y label and Y kernel-mse .

[0084] Step 4: Backpropagate the error to update the model weight parameters. After each training, slide the sliding window backward by 1 day. When the sliding window slides to the last batch of data in the training dataset S, one round of iteration ends. After the model training iterates E = 100 times, save the trained model Model.

[0085] Step 5: Use the trained prediction model Model to predict the sewage discharge situation in the next 5 days.

[0086] The user inputs the date 2023-02-24, and the model predicts the printing and dyeing sewage discharge volume in the next 5 days. The model obtains the printing and dyeing sewage discharge volume data for 32 days before 2023-02-24 and forms the encoder input X as shown in Table 7:

[0087] 2023-01-24 -0.098866 -0.539368 -0.807541 -0.913076 -0.888833 -0.025125 2022-01-25 -0.476504 -0.522405 -0.142381 -0.849057 -0.539026 -0.059980 ... ... ... ... ... ... ... 2022-02-24 0.616985 0.520807 0.969454 0.502447 -1.051243 -0.071599

[0088] Table 7

[0089] Obtain the data for the 5 days from February 20, 2023 to February 24, 2023 and the data for 5 days filled with 0 to form the decoder input Y as shown in Table 8.

[0090]

[0091]

[0092] Table 8

[0093] Input X and Y into the model, and the model calculates the prediction result Y prediction , as shown in Table 9:

[0094] 0.1017 -0.3292 -0.0492 0.1306 0.0718 0.1319 0.3830 0.0563 0.7958 0.3003

[0095] Table 9

[0096] The prediction result Y prediction After inverse normalization, the prediction result of the printing and dyeing wastewater discharge is obtained, as shown in Table 10:

[0097] Date Print Pollution Water Dye Pollution Water 2023-02-25 991.7856 1620.2565 2023-02-26 954.3465 2218.2712 2023-02-27 984.3653 2220.0054 2023-02-28 1061.5824 2121.5662 2023-03-01 1164.0411 2439.0039

[0098] Table 10

[0099] Those of ordinary skill in the art in this technical field should recognize that the above content is only used to illustrate the present invention, rather than to limit the present invention. As long as it is within the scope of the spirit of the present invention, changes and modifications to the above examples will fall within the scope of the claims of the present invention.

Claims

1. A method for predicting the discharge of wastewater from printing and dyeing processes based on an improved Informer model, characterized in that: The method for predicting the discharge amount of printing and dyeing process wastewater comprises the following steps: Step 1, collect the sewage discharge data table Data in the daily production process of the printing and dyeing enterprises, including the following 9 fields with data for N days: date t, water consumption w1, printing process water w2, dyeing process water w3, printing process water saving w4, printing process hot water saving w5, domestic water w6, printing sewage discharge w7 and dyeing sewage discharge w8; Step 2: Clean the data table Data, use the missing value processing, outlier processing, and data standardization methods to perform data preprocessing operations to obtain the training data set S = {S i |i=1,..,N},S i represents the relevant data of the i-th day, S i It is expressed as follows: <h2 style=";text-align:left;direction:ltr">S<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (t, w1, w2, w3, w4, w5, w6, w7, w8) (i∈[1, N]) Step 3: Use the training data set S to train a printing and dyeing process wastewater discharge prediction model based on the improved Informer; Step 4: Set the training error Loss kernel-mse Back propagation, update the model weight parameters, slide the sliding window back 1 day for each training, and the iteration ends after the sliding window slides to the last batch of data. After the model training iteration E times, save the trained model Model; Step 5: Use the trained model Model to predict the printing and dyeing wastewater discharge data for the next m days. The user inputs the date T, and the model predicts the printing and dyeing wastewater discharge data m days after the date T.

2. The method for predicting the discharge of wastewater from printing and dyeing processes based on the improved Informer model according to claim 1, characterized in that: In step 2, the data preprocessing process is as follows: (2.1) Processing of missing data: If more than 60% of the characteristic data of Data is missing, that is, the field value is empty, the emission record is deleted, otherwise the mean is used to fill the missing characteristic values; (2.2) Processing of abnormal data: If the characteristic data of Data differs from the mean by several times, or the characteristic data is negative, the emission record is replaced by the mean; (2.3) The processed data is standardized to a mean of 0 and a variance of 1.

3. The method for predicting the discharge of wastewater from printing and dyeing processes based on the improved Informer model according to claim 1 or 2, characterized in that: In step 3, the detailed process of model training is as follows: (3.1) Each training takes data from the training set S based on the sliding window mechanism to generate the encoder input X, decoder input Y, and label output Y of the model. label , initialize the sliding window size as n, that is, the printing and dyeing wastewater discharge data for n days. The model predicts the printing and dyeing wastewater discharge data for m days after n days, where m < n; take a sliding window of n days of data from S, and use the date, water consumption, water consumption for printing process, water consumption for dyeing process, water saving for printing process, hot water saving for printing process, and domestic water as the encoder input X = {X i | i = 1, 2,..., n}, where X i represents the data of the i-th day: <h2 style=";text-align:left;direction:ltr">X<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (t,w1,w2,w3,w4,w5,w6) (i∈[1,n]) Take the last m days of data in the sliding window and the m days of data filled with 0, a total of 2m days of data, and use the date t, printing wastewater discharge w7, and dyeing wastewater discharge w8 as the input of the decoder Y = {Y i |i=1,2,...,2m}, where Y i Represents the data for the i-th day: Take the data of m days after the sliding window as the label output Y label , that is, the model needs to predict the real data m days after n days; (3.2) Input X into the encoder and input Y into the decoder. The input X and Y are embedded with the same features of dimension d using EMB to obtain X and Y. emb and Y emb , feature embedding EMB includes scalar embedding U, position embedding PE and date embedding SE, the formula is as follows: EMB=U+PE+SE Among them, U represents the one-dimensional convolution of the data features in the input data except date t, and PE represents the calculation of the position embedding of the input data in this batch of data, where the position corresponding to the position p is encoded in the representation of odd and even bits. The formula is as follows: SE means calculating the date embedding of the date t in the input data. The date embedding is expressed as follows: SE = [month, week, day] Among them, month represents the month part of date t, week represents the day of the week, and day represents the day part of date t; (3.3) Encoder X after feature embedding emb Compute the probability sparse self-attention X attention , the probabilistic sparse self-attention formula is as follows: Where Q, K, and V represent the input X emb The parameter matrices Query, Key, and Value obtained by different linear transformations. Since different dot products q*k in Query and Key have different effects on self-attention, fewer dot products contribute to most of the self-attention scores, that is, the probability distribution of the self-attention mechanism is sparse; By calculating the KL divergence between the attention probability distribution and the uniform distribution of each query in the Query matrix, we find the Top-K query components with the greatest contribution. Matrix, the formula for calculating KL divergence is as follows: where q i is the value of the Q matrix, is the value of the K matrix, That is, different dot products, d is X emb The dimension of feature embedding, The smaller the query value, the smaller the query value. i The higher the similarity with the uniform distribution, the lower the contribution to the self-attention score. Find the Top-K with the lowest similarity. composition (3.4) Use the attention distillation operation to give higher weights to the dominant features with dominant attention, and obtain the main feature X feature , the attention distillation formula is as follows: X feature =MapPool(Elu(Conv1d(X attention ))) Conv1d represents a one-dimensional convolution operation on the time series, using the Elu activation function, and finally performing a maximum pooling operation; (3.5) The model will be X emb 1 / 2 and 1 / 4 of the sequence length are input as copies to the encoder to obtain the corresponding copy features. Finally, the obtained copy features are fused with the main features to obtain the output features X of the final encoder. featuremap ; (3.6) The decoder performs feature embedding on Y emb Calculate the probability sparse self-attention to get Y attention After the encoder output X featuremap Calculate the attention score and finally get the predicted output Y through a fully connected layer prediction ; (3.7) Use Kernel-MSE loss function to calculate Y prediction With Y label Loss kernel-mse , the Kernel-MSE loss function is expressed as follows:

4. The method for predicting the discharge of wastewater from printing and dyeing processes based on the improved Informer model according to claim 1 or 2, characterized in that: In step 5, the prediction process is as follows: (5.1) The user inputs a date T, obtains the printing and dyeing wastewater discharge data S1 n days before date T, and processes S1 into the encoder input X; (5.2) Obtain the data S2 of m days before date T and the data of m days filled with 0 to form the decoder input Y; (5.3) Input X and Y into the model and predict Y prediction ; (5.4) prediction After de-standardization, the prediction results of printing and dyeing wastewater discharge in the next m days after date T are obtained.