Method for predicting content of free calcium oxide in cement clinker

Through the combination of dynamic adversarial domain adaptation and local attention Transformer, the stability and accuracy of the prediction of free calcium oxide content in cross-domain cement clinker production are solved, and rapid adaptation and efficient prediction of new factory data are achieved.

CN120356583APending Publication Date: 2025-07-22HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433347.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the production of cross-domain cement clinker, traditional models have deteriorated the predictive performance of free calcium oxide content due to differences in kiln equipment and fluctuations in raw material components, and the existing domain adaptation methods have problems of premature alignment convergence and insufficient timing modeling capabilities.

Method used

The method of fusion of dynamic adversarial domain adaptation and local attention Transformer is adopted to generate feature extraction networks to conduct domain adversarial training, dynamically adjust the loss weight, and combine local attention mechanism and residual gating unit to achieve the prediction of the free calcium oxide content of cross-domain cement clinker.

Benefits of technology

It improves the stability and accuracy of cross-domain prediction, can quickly adapt to new factory data, and solves the stability and accuracy problems of traditional methods in cross-domain prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356583A_ABST
    Figure CN120356583A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting the content of free calcium oxide in cement clinker, which comprises the following steps of: 1, respectively acquiring cement firing process variable data and free calcium oxide content data of two cement production lines in continuous time, respectively forming a source domain data set A and a target domain data set B, and preprocessing; 2, generating a feature extraction network, inputting the source domain data set A and the target domain data set B which are preprocessed in the step 1 into the feature extraction network for field adversarial training, and obtaining fusion features through dynamic domain adaptation in the adversarial training; and step 3, generating an improved Transform model, training the improved Transform model by using the fusion features obtained in the step 2, and outputting a predicted value and a performance evaluation index by the trained improved Transform model. The method is beneficial to solving the problem of predicting the content of free calcium oxide in the cross-domain cement clinker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cement clinker production, and specifically to a method for predicting the content of free calcium oxide in cement clinker. Background Art

[0002] In the production of cement clinker, the content of free calcium oxide is a core index for measuring the quality of the calcination process. Existing prediction methods mainly rely on historical data of a single factory to establish regression models, such as linear regression based on partial least squares (PLS), support vector machine (SVM), or conventional neural networks. However, when the model is applied to a new factory, due to differences in kiln equipment, fluctuations in raw material composition, and shifts in the distribution of operating parameters, traditional models often exhibit significant performance degradation. Existing domain adaptation methods such as maximum mean discrepancy (MMD) and domain adversarial training (DANN) can alleviate the distribution differences, but have the following defects: (1) The static alignment strategy ignores the dynamic evolution of the domain discriminator, resulting in premature convergence of the alignment; (2) The ability to model time series is insufficient, and traditional recurrent neural networks (RNNs) are difficult to capture long-range process parameter dependencies. Summary of the Invention

[0003] The present invention provides a method for predicting the content of free calcium oxide in cement clinker by fusing dynamic adversarial domain adaptation and local attention Transformer, so as to solve the problem that it is difficult to achieve rapid on-site analysis of the measurement of the content of free calcium oxide in the cross-domain cement clinker production process in the prior art.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] A method for predicting the content of free calcium oxide in cement clinker, comprising the following steps:

[0006] Step 1: Respectively collect the data of the cement burning process variables within a continuous time of two cement production lines, and the data of the content of free calcium oxide in the products produced by the two cement production lines within the continuous time obtained from laboratory tests, wherein the collection frequencies of the data of the cement burning process variables of the two cement production lines are different;

[0007] Use the data of the cement burning process variables and the content of free calcium oxide of one of the cement production lines to form a source domain dataset A, and use the data of the cement burning process variables and the content of free calcium oxide of the other cement production line to form a target domain dataset B;

[0008] Preprocess the data in the source domain dataset A and the target domain dataset B respectively;

[0009] Step 2: Generate a feature extraction network, which includes an input layer, multiple fully connected layers, a gradient reversal layer, a domain discriminator, and an output layer;

[0010] Input the pre - processed source - domain dataset A and target - domain dataset B in step 1 into the feature extraction network for domain - adversarial training. During the adversarial training, the feature extraction network extracts the high - dimensional features of the data in the source - domain dataset A and the data in the target - domain dataset B through the forward propagation of the gradient reversal layer, and reverses the gradient direction when the gradient reversal layer back - propagates to confuse the domain discriminator;

[0011] During the adversarial training, according to the extracted high - dimensional features, calculate the regression loss of the source - domain prediction of the feature extraction network, the adversarial loss of the domain discriminator, and the distribution difference loss of the cross - domain features. Calculate the total loss during training from the weighted sum of the regression loss of the source - domain prediction, the adversarial loss of the domain discriminator, and the distribution difference loss of the cross - domain features, and dynamically adjust the weight ratios of the adversarial loss of the domain discriminator and the distribution difference loss of the cross - domain features in the total loss according to the domain discrimination accuracy of the domain discriminator, thereby dynamically balancing the optimization objectives of domain discrimination and distribution alignment, completing the dynamic domain adaptation during the adversarial training, and finally obtaining the fused features through the dynamic domain adaptation in the adversarial training;

[0012] Step 3: Generate an improved Transformer model. The improved Transformer model includes an input layer, a Transformer encoder, and an output layer. The Transformer encoder introduces a local attention mechanism to divide sliding windows to extract local temporal features, and fuses features at different levels through a residual gating unit to retain long - range dependency information;

[0013] Use the fused features obtained in step 2 to train the improved Transformer model, and the trained improved Transformer model outputs the predicted value of the free calcium oxide content and the performance evaluation index.

[0014] Further, the pre - processing in step 1 includes mean filtering, outlier removal, and normalization processing.

[0015] Further, in the feature extraction network described in step 2, the input layer receives the process variable data in each dataset after the pre - processing in step 1, and gradually extracts high - dimensional abstract features through multiple fully - connected layers; the gradient reversal layer is embedded in the feature extraction path, retains the feature values during forward propagation, and reverses the direction and adjusts the amplitude of the gradient signal of the domain discriminator during back - propagation; the domain discriminator is composed of stacked fully - connected layers, and the output layer outputs the domain classification result through the Sigmoid activation function, so that the feature extraction network generates fused features that are indistinguishable from the factory source; the free calcium oxide content data in the source - domain dataset A is used as a supervision signal and input to the output layer for calculating the regression loss of the source - domain prediction.

[0016] Further, in step 2, the dynamic adjustment strategy for the weight ratios of the adversarial loss of the domain discriminator and the distribution difference loss of cross-domain features in the total loss is as follows:

[0017] At the initial state, the weights of the regression loss, the adversarial loss, and the distribution difference loss are evenly distributed; when the domain discrimination accuracy is significantly higher than the preset threshold, the weight of the adversarial loss is increased to strengthen the feature confusion effect, and at the same time, the weight of the distribution difference loss is decreased; when the domain discrimination accuracy is lower than the critical value, the weight of the distribution difference loss is increased; thus, through a continuously gradual weight adjustment method, the stability of model training is ensured.

[0018] Further, in the improved Transformer model generated in step 3, the local attention mechanism of the Transformer encoder divides the fused features obtained in step 2 into overlapping sliding windows, and calculates the feature correlation weights within each window; the residual gating unit dynamically selects the ratio of the current layer features to the skip connection features through a gating signal to achieve multi-scale feature fusion; the encoder levels aggregate global temporal patterns through cross-layer connections.

[0019] Further, the performance evaluation metrics in step 3 include MSE, MAE, and MMD.

[0020] Compared with the prior art, the advantages of the present invention are as follows:

[0021] 1. Compared with the traditional domain adaptation method that adopts a fixed-weight adversarial training strategy, the present invention designs a dynamic adversarial domain adaptation mechanism, automatically adjusts the weight ratio of the adversarial loss and the distribution difference loss based on the real-time accuracy of the domain discriminator, that is, when the discrimination accuracy is higher than the threshold, the adversarial weight is increased to strengthen the feature confusion, and vice versa, the kernel space distribution matching is enhanced, avoiding premature convergence of feature alignment and helping to improve the stability of cross-domain prediction.

[0022] 2. Aiming at the long-range dependence and local feature correlation of industrial data in the new dry-process cement burning process, which are difficult to balance, the traditional global attention mechanism has a high computational complexity, and the deep network is prone to losing detailed information, the present invention proposes a local sliding window attention mechanism, divides the input sequence into overlapping windows for local feature correlation calculation, and designs a residual gating unit to dynamically fuse multi-layer features, retaining multi-scale temporal patterns, which helps to solve the prediction accuracy problem.

[0023] 3. Considering the problem that the offline training model is difficult to adapt to the input of new factory data, the present invention realizes the fast domain adaptation and incremental prediction of new data by fine-tuning the feature extractor and the dynamic loss weight, and at the same time sets a gradient clipping threshold to prevent overfitting, which helps to solve the problem of cross-domain prediction of the free calcium oxide content in cement clinker. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is the schematic diagram of the method in the embodiment of the present invention. Specific implementation manners

[0025] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0026] As Figure 1 shown, this embodiment discloses a method for predicting the content of free calcium oxide in cement clinker, including the following steps:

[0027] Step 1: Respectively collect the data of the cement burning process variables within a continuous time for two cement production lines, and the data of the content of free calcium oxide in the products produced by the two cement production lines during the continuous time obtained through laboratory tests, wherein the collection frequencies of the data of the cement burning process variables for the two cement production lines are different.

[0028] In this embodiment, the data of the cement burning process variables collected for each cement production line includes a total of 10-dimensional process variable data, namely the temperature at the outlet of the precalciner, the coal feeding amount of the precalciner, the raw material feeding amount, the temperature of the tertiary air, the temperature of the flue gas chamber at the kiln tail, the coal feeding amount at the kiln head, the temperature of the kiln head hood, the current of the kiln main machine, the negative pressure of the kiln head hood, and the pressure under the grate of the F2 fan.

[0029] Use the data of the cement burning process variables and the content of free calcium oxide in one of the cement production lines to form the source domain dataset A, and use the data of the cement burning process variables and the content of free calcium oxide in the other cement production line to form the target domain dataset B. In this embodiment, it is assumed that the source domain dataset A contains nA samples, and the target domain dataset samples, where and are respectively the process variable data of the i-th sample in the source domain dataset A and the process variable data of the j-th sample in the target domain dataset B, and are respectively the content of free calcium oxide in the i-th sample in the source domain dataset A and the content of free calcium oxide in the j-th sample in the target domain dataset B.

[0030] Preprocess the data in the source domain dataset A and the target domain dataset B respectively. The preprocessing includes mean filtering, outlier removal, and normalization.

[0031] The specific mean filtering process is as follows: perform sliding window smoothing on the time series variable. Let the size of the mean filtering window be w, and the filtered value at time t is where w is the size of the mean filtering window, t is the time point, is the filtered value at time point t, and x t-k is the original data value at time point t - k. High-frequency noise is eliminated through mean filtering.

[0032] In this embodiment, the process of removing outliers is as follows: using the 3σ principle, if |x ik -μ k |>3σ k , then it is determined as an outlier and the entire sample is deleted, where μ k , σ k are the mean and standard deviation of the k-th feature respectively, and x ik is the value of the k-th feature of the i-th sample.

[0033] In this embodiment, the process of standardization is as follows: the data is mapped to zero mean and unit variance through Z-score standardization, as shown in the following formula:

[0034]

[0035] where μ k and σ k are the mean and standard deviation of the k-th feature respectively, x ik is the value of the k-th feature of the i-th sample, and x′ ik is the standardized value of the k-th feature of the i-th sample.

[0036] In this embodiment, the mean filter window w = 5. After removing outliers and standardization, the process variable data X A in the source domain dataset A and the process variable data X B in the target domain dataset B have a feature mean approaching 0 and a standard deviation of 1, eliminating the dimension difference.

[0037] Finally, the free calcium oxide content as the target variable and the 10-dimensional process variables as the auxiliary variables are separated to obtain the process variable data X A in the source domain dataset A after preprocessing, the free calcium oxide data y A in the source domain dataset A, the process variable data X B in the target domain dataset B, and the free calcium oxide data y B in the target domain dataset B.

[0038] Step 2: Generate a feature extraction network. The feature extraction network includes an input layer, a feature extractor, a gradient reversal layer, a domain discriminator, a predictor, and an output layer. The dataset processed in Step 1 is input into the feature extractor, and high-dimensional features of the source domain and the target domain are extracted through forward propagation. When backpropagating, the gradient direction is reversed to confuse the domain discriminator. The predictor receives the features output by the feature extractor and outputs the predicted value of the free calcium oxide content through a fully connected layer.

[0039] Specifically, the input layer receives the cement firing process variable data in each dataset after preprocessing in step 1, and gradually extracts high-dimensional abstract features through the multi-layer perceptron (MLP) in the feature extractor; the gradient reversal layer is embedded in the feature extraction path, retaining the feature values during forward propagation and reversing the direction and adjusting the amplitude of the gradient signal of the domain discriminator during backpropagation. The domain discriminator is composed of stacked fully connected layers; the output layer outputs the domain classification result through the Sigmoid activation function, so that the feature extraction network generates fusion features that are indistinguishable from the factory source; the predictor receives the features output by the feature extractor and outputs the predicted value of the free calcium oxide content through the fully connected layer.

[0040] The feature extraction network realizes domain adversarial training through the gradient reversal layer (GRL). The forward propagation of GRL is an identity mapping, and the gradient is multiplied by the negative coefficient -λ during backpropagation.

[0041] The forward process is shown as follows:

[0042] F(x) = x

[0043] where F(x) is the forward propagation function of the gradient reversal layer, x is the input data, that is, the high-dimensional data extracted through the fully connected layer.

[0044] The backward process is shown as follows:

[0045]

[0046] where -λ is the negative coefficient of the gradient reversal layer, is the backward propagation gradient of the gradient reversal layer, and I is the identity matrix.

[0047] Input the source domain dataset A and the target domain dataset B after preprocessing in step 1 into the feature extraction network for domain adversarial training. During the adversarial training process, the feature extraction network extracts the high-dimensional features of the data in the source domain dataset A and the data in the target domain dataset B through the forward propagation of the gradient reversal layer, and reverses the gradient direction during the backpropagation of the gradient reversal layer to confuse the domain discriminator.

[0048] The network includes a feature extractor F θ , a domain discriminator D φ and a predictor P ψ , and the objective function is:

[0049]

[0050] where: is the regression loss of source domain prediction; is the adversarial loss of the domain discriminator; is the distribution difference loss of cross-domain features.

[0051] During adversarial training, based on the extracted high-dimensional features, calculate the regression loss of the source domain prediction of the feature extraction network, the adversarial loss of the domain discriminator, and the distribution difference loss (MMD) of cross-domain features. Calculate the total loss during training from the weighted sum of the regression loss of the source domain prediction, the adversarial loss of the domain discriminator, and the distribution difference loss of cross-domain features. And according to the domain discrimination accuracy of the domain discriminator, dynamically adjust the weight ratio of the adversarial loss of the domain discriminator and the distribution difference loss of cross-domain features in the total loss, thereby dynamically balancing the optimization objectives of domain discrimination and distribution alignment, completing the dynamic domain adaptation during adversarial training, and finally obtaining the fused features through the dynamic domain adaptation in adversarial training;

[0052] In this example, the feature extractor Fθ is a two-layer multi-layer perceptron MLP (input is 10-dimensional, hidden layer is 64-dimensional, ReLU activation), and the domain discriminator D φ contains a GRL layer and a Sigmoid output. During training, the source domain samples X A in the source domain dataset A and the target domain samples X B in the target domain dataset B are θ extracted into 64-dimensional features F A and F B . D φ receives F A and F B and outputs the domain label (0 represents the source domain, 1 represents the target domain), and forces F θ to generate domain-invariant features through the adversarial loss.

[0053] In this embodiment, the dynamic domain adaptation process is as follows:

[0054] S1. Calculate the adversarial loss of the domain discriminator, using binary cross-entropy, as shown in the following formula:

[0055]

[0056] where, is the adversarial loss of the domain discriminator, n A and n B are the number of samples in the source domain dataset A and the target domain dataset B respectively, x i is the feature vector of the i-th sample, l i is the domain label of the i-th sample, l i =0 represents the source domain, l i =1 represents the target domain, and D φ (F θ (x i ) is the output of the domain discriminator for the i-th sample.

[0057] S2. Calculate the distribution difference loss MMD of cross - domain features, that is, measure the feature distribution difference through multi - kernel maximum mean difference, as shown in the following formula:

[0058]

[0059] Among them, is the distribution difference loss of cross - domain features, and are the feature vectors of the i - th sample in the source - domain dataset A and the j - th sample in the target - domain dataset B respectively, and k( x , y) is the kernel function used to calculate the similarity between two samples. a is the bandwidth parameter of the Gaussian kernel, and ‖x - y‖ is the Euclidean distance between two samples.

[0060] S3. Dynamically adjust the weight ratio of the adversarial loss of the domain discriminator and the distribution difference loss of cross - domain features. The dynamic adjustment strategy is as follows:

[0061] At the initial state, evenly distribute the weights of the regression loss, adversarial loss, and distribution difference loss; when the domain discrimination accuracy is significantly higher than the preset threshold, increase the weight of the adversarial loss to strengthen the feature confusion effect, and at the same time reduce the weight of the distribution difference loss; when the domain discrimination accuracy is lower than the critical value, increase the weight of the distribution difference loss; thus, through a continuously gradual weight adjustment method, ensure the stability of model training. That is, adjust the weight α according to the domain discrimination accuracy η, as shown in the following formula:

[0062]

[0063] The final joint optimization objective is

[0064] In this example, 32 samples are randomly sampled in each batch during training to calculate D φ for F A and F B to obtain the discrimination accuracy η. If η > 0.6, then α = 0.7, emphasizing adversarial training; if η < 0.6, then α = 0.3, emphasizing MMD alignment. The dynamic weight balances domain discrimination and distribution alignment, improving the cross - domain generalization ability.

[0065] Step 3: Generate an improved Transformer model. The improved Transformer model includes an input layer, a Transformer encoder, and an output layer. The Transformer encoder in the improved Transformer model is composed of local attention and a residual gated unit (RGU). The Transformer encoder introduces a local attention mechanism to divide the sliding window to extract local temporal features, and fuses features at different levels through the residual gated unit to retain long-range dependency information.

[0066] Specifically, the Transformer encoder introduces a local attention mechanism, divides the sequence into windows of length w, and calculates the multi-head attention within the window as shown in the following formula:

[0067]

[0068] where Attention(Q, K, V) is the calculation function of the multi-head attention mechanism, Q is the query vector, K is the key vector, V is the value vector, W i Q , W i K , W i V is the projection matrix of the i-th attention head, dk is the dimension of the key vector, and Softmax converts the dot product result into a probability distribution.

[0069] The Transformer encoder introduces a residual gated unit. The output x of the current layer is fused with the residual r through a gating mechanism as shown in the following formula:

[0070] g = σ(W g [x; r] + b g )

[0071] where g is the gating weight, σ is the Sigmoid function, W g and b g are the gating weight matrix and the gating bias term, respectively.

[0072] The prediction target of the improved Transformer model T ω is shown in the following formula:

[0073]

[0074] where, is the fused feature of the i-th sample, and y (i) is the free calcium oxide content corresponding to the i-th sample.

[0075] After training is completed, the improved Transformer model can receive a new sample x new , and after being θ feature-extracted by F, the predicted value of the output is input

[0076] Using the fused feature F obtained in step 2 fused to train the improved Transformer model, and project it linearly to d model = 64 dimensions. The encoder has 2 layers, with 4 heads of attention in each layer, and the window size w = 10. The RGU dynamically fuses the local attention and the output of the feed-forward network through the gating weight g, and retains the long-range dependence. The fused sample data is input into the improved Transformer for 100 rounds of training, with a batch size of 32 and a learning rate of 0.01. The final output layer maps the sequence to the predicted value of the free calcium oxide content and outputs it.

[0077] The performance evaluation metrics of the prediction results of the improved Transformer model in this embodiment include the mean square error MSE, the mean absolute error MAE, and the maximum mean discrepancy MMD distance. Among them:

[0078] The calculation formula of MSE is as follows:

[0079]

[0080] where n is the number of samples, y i is the actual value, is the predicted value.

[0081] The calculation formula of MAE is as follows:

[0082]

[0083] where n is the number of samples, y i is the actual value, is the predicted value.

[0084] The calculation formula of the MMD distance is as follows:

[0085]

[0086] where P and Q are two distributions, x i and y j are samples drawn from these two distributions respectively, is the mapping function that maps to a high-dimensional space through the kernel function, m and n are the sizes of P and Q respectively, and H is the feature space.

[0087] Using ablation experiments, first input the data of groups A and B after dynamic adversarial domain adaptation into the improved Transformer model described above, then input the data of group B into the improved Transformer model described above, and finally input the data of group A into the improved Transformer model described above. If the MMD distance decreases, it proves that the alignment of feature distributions is effective; if the prediction errors MSE and MAE decrease, it indicates that the feature extractor successfully confuses the domain discriminator and can achieve cross-domain generalization.

[0088] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. The embodiments described in the present invention are only descriptions of the preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Among the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without contradiction. As long as such a combination does not violate the idea of the present invention, it should also be regarded as the content disclosed in this disclosure. To avoid unnecessary repetition, the present invention does not separately describe various possible combination methods.

[0089] The present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention and without departing from the design idea of the present invention, various modifications and improvements made by those skilled in the art to the technical solutions of the present invention should fall within the protection scope of the present invention. The technical content claimed by the present invention has been fully recorded in the claims.

Claims

1. A method for predicting the free calcium oxide content in cement clinker, characterized in that, It includes the following steps: Step 1: Collect the cement burning process variable data of two cement production lines within a continuous time respectively, and the free calcium oxide content data in the products produced by the two cement production lines in the continuous time obtained from laboratory tests, wherein the data collection frequencies of the cement burning process variable data of the two cement production lines are different; Taking the cement burning process variable data and free calcium oxide content data of one cement production line as the source domain dataset A and taking the cement burning process variable data and free calcium oxide content data of another cement production line as the target domain dataset B ; Preprocess the data in the source domain dataset A and the target domain dataset B separately; Step 2: Generate a feature extraction network, which includes an input layer, multiple fully connected layers, a gradient reversal layer, a domain discriminator, and an output layer; The source domain dataset after preprocessing in step 1 A , the target domain dataset B are input into the feature extraction network for domain adversarial training. During the adversarial training process, the feature extraction network extracts the high-dimensional features of the data in the source domain dataset A and the data in the target domain dataset B through the forward propagation of the gradient reversal layer, and reverses the gradient direction during the backpropagation of the gradient reversal layer to confuse the domain discriminator; During adversarial training, according to the extracted high-dimensional features, calculate the regression loss of the source domain prediction of the feature extraction network, the adversarial loss of the domain discriminator, and the distribution difference loss of cross-domain features. Calculate the total loss during training from the weighted sum of the regression loss of the source domain prediction, the adversarial loss of the domain discriminator, and the distribution difference loss of cross-domain features, and dynamically adjust the weight ratios of the adversarial loss of the domain discriminator and the distribution difference loss of cross-domain features in the total loss according to the domain discrimination accuracy of the domain discriminator, thereby dynamically balancing the optimization objectives of domain discrimination and distribution alignment, completing the dynamic domain adaptation during adversarial training, and finally obtaining fused features through the dynamic domain adaptation in adversarial training; Step 3: Generate an improved Transformer model, which includes an input layer, a Transformer encoder, and an output layer, wherein the Transformer encoder introduces a local attention mechanism to divide sliding windows to extract local temporal features, and fuses features at different levels through a residual gating unit to retain long-range dependence information; Use the fused features obtained in Step 2 to train the improved Transformer model, and output the predicted value of the free calcium oxide content and the performance evaluation index by the trained improved Transformer model.

2. The method for predicting the content of free calcium oxide in cement clinker according to claim 1, wherein The preprocessing in Step 1 includes mean filtering, outlier removal, and normalization processing.

3. The method for predicting the free calcium oxide content in cement clinker according to claim 1, wherein, In the feature extraction network described in Step 2, the input layer receives the process variable data in each dataset after preprocessing in Step 1, and gradually extracts high-dimensional abstract features through multiple fully connected layers; the gradient reversal layer is embedded in the feature extraction path, retains the feature values during forward propagation, and reverses the direction and adjusts the amplitude of the gradient signal of the domain discriminator during backpropagation; The domain discriminator is composed of stacked fully connected layers, and the output layer outputs the domain classification result through the Sigmoid activation function, so that the feature extraction network generates fused features that are indistinguishable from the factory sources; The free calcium oxide content data in the source domain dataset A is used as a supervision signal and input to the output layer to calculate the regression loss of the source domain prediction.

4. The method for predicting the free calcium oxide content in cement clinker according to claim 1, wherein In Step 2, the dynamic adjustment strategy for the weight ratios of the adversarial loss of the domain discriminator and the distribution difference loss of cross-domain features in the total loss is: At the initial state, the weights of the balanced distribution regression loss, adversarial loss, and distribution difference loss are allocated; when the domain discrimination accuracy rate is significantly higher than the preset threshold, the weight of the adversarial loss is increased to strengthen the feature confusion effect, and at the same time, the weight of the distribution difference loss is decreased; when the domain discrimination accuracy rate is lower than the critical value, the weight of the distribution difference loss is increased; thus, through a continuously gradual weight adjustment method, the stability of model training is ensured.

5. The method for predicting the free calcium oxide content in cement clinker according to claim 1, characterized in that In the improved Transformer model generated in step 3, the local attention mechanism of the Transformer encoder divides the fused features obtained in step 2 into overlapping sliding windows, and calculates the feature correlation weights within each window; the residual gating unit dynamically selects the ratio of the current layer features to the skip connection features through the gating signal to achieve multi-scale feature fusion; the global temporal patterns are aggregated through cross-layer connections between the encoder levels.

6. The method for predicting the content of free calcium oxide in cement clinker according to claim 1, characterized in that, The performance evaluation metrics in step 3 include MSE, MAE, and MMD.