A Prediction Method for Superheat Degree of Aluminum Electrolysis Mixed-Frequency Data Based on Attention Mechanism

The method addresses the challenge of mixed-frequency data by using long short-term memory networks and attention mechanisms to enhance prediction accuracy in fields like industrial aluminum electrolysis, finance, and transportation.

CN114444811BActive Publication Date: 2025-07-15XIAN GUANGLIN HUIZHI ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210132540.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-07-15
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the data quantity differences and timing dependence problems caused by different sampling frequencies in mixed data, resulting in low prediction accuracy.

Method used

Using an attention mechanism-based method, aluminum electrolytic overheating prediction is carried out through sliding window processing, long and short-term memory network coding, convolutional neural network feature extraction and timing attribute attention mechanism, mixed frequency data characteristics of different frequencies are fused to predict aluminum electrolytic overheat.

Benefits of technology

It improves the accuracy of mixed data prediction, effectively utilizes the characteristics of the original data, and improves the accuracy of the prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444811B_ABST
    Figure CN114444811B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data mining, and specifically relates to a method for predicting the superheat degree of aluminum electrolysis mixed-frequency data based on an attention mechanism. The method includes: obtaining the original mixed-frequency data; processing the original mixed-frequency data by using a sliding window to obtain the processed mixed-frequency data; encoding the processed mixed-frequency data by using a long short-term memory network to obtain an encoding matrix and an embedding vector; extracting and fusing the time series features from the encoding matrix by using a convolutional neural network to obtain a fused encoding matrix; learning an attribute feature context vector and a time series feature context vector from the fused encoding matrix by using a time series attribute attention mechanism; and obtaining the prediction result of the target prediction index according to the embedding vector, the attribute feature context vector, and the time series feature context vector. The present invention effectively utilizes the characteristics of the original mixed-frequency data itself to mine the important information related to the prediction target, and improves the accuracy of the prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining, and particularly relates to a method for predicting the superheat degree of aluminum electrolysis mixed-frequency data based on an attention mechanism. Background Art

[0002] In the context of industrial and commercial decision-making in the real world, time series prediction is a key component. However, in some practical applications, due to cost and technological limitations, it is not practical to collect these time series at a uniform frequency. This problem of using time series sampled at different frequencies to predict key decision-making indicators in real decision-making scenarios is called the mixed-frequency data prediction problem. In the context of the big data era, the mixed-frequency data prediction problem is becoming increasingly common. For example, in the field of industrial electrolytic aluminum, high-frequency sensor data is combined with low-frequency experimental data to predict the superheat degree during the industrial aluminum electrolysis process. In the financial field, a combination of quarterly, monthly, and daily data is used to predict the gross domestic product. In the transportation field, data collected by sensors with different sampling frequencies is used to predict the number of traffic collisions. Therefore, how to efficiently utilize data with different collection frequencies for data mining has become an urgent problem to be solved.

[0003] Mixed-frequency data truly exists in the objective world and has corresponding application scenarios in various fields. However, due to its special data form, mixed-frequency data is difficult to process, resulting in relatively few research works in machine learning at present. Generally speaking, there are the following difficulties in processing mixed-frequency data: First, mixed-frequency data comes from different data sources, and each data source is sampled at a different sampling frequency. The number of samples between different frequencies is not equal, and it is impossible to directly perform information fusion at the data layer. Second, in most scenarios, mixed-frequency data is time series data, so it is difficult to model the time series dependence of different frequencies and the time series dependence across frequencies. Finally, the characteristics and characteristic dimensions of different data sources of mixed-frequency data vary greatly. However, to jointly describe a unified decision-making system, the dimensionality difference of different data sources is an obstacle that must be overcome when fusing information sources into a unified representation.

[0004] Every field in reality is a complex system with various possible non - linear relationships. To explore these non - linear relationships, scholars in various fields have proposed some solutions. For example: vector autoregressive model, piece - wise linear model, polynomial pattern, non - parametric model, decision tree, support vector machine, artificial neural network and other methods. The above models are all based on data sampled at equal frequencies for regression analysis. When using such models to process mixed - frequency data, usually, the data collected at different frequencies will be aligned to the frequency of the target variable according to the sampling frequency of the target variable. Traditional methods are generally divided into two types. One is to convert high - frequency data into co - frequency data that matches the sampling frequency of low - frequency data by summing or averaging, which ignores the diversity and volatility in high - frequency variables and also causes information loss. The other is to convert low - frequency data into high - frequency co - frequency data by methods such as interpolation filling, which makes the artificially specified data distribution of low - frequency data may bring secondary errors to the model, ultimately resulting in a poor prediction effect. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention proposes a method for predicting the superheat degree of aluminum electrolysis mixed - frequency data based on the attention mechanism. The method includes:

[0006] S1: Obtain different index data related to the superheat degree of aluminum electrolysis at different sampling frequencies as the original mixed - frequency data; among them, the indexes related to the superheat degree of aluminum electrolysis include: molecular ratio, alumina concentration, current density, voltage and electrolysis temperature;

[0007] S2: Process the original mixed - frequency data using a sliding window to obtain the mixed - frequency data divided by the sliding window;

[0008] S3: Encode the mixed - frequency data divided by the sliding window using a long short - term memory network to obtain an encoded matrix and an embedding vector;

[0009] S4: Extract temporal features from the encoded matrix using a convolutional neural network, and fuse the extracted temporal features to obtain a fused encoded matrix;

[0010] S5: Use a temporal attribute attention mechanism to learn from the fused encoded matrix to obtain an attribute feature context vector and a temporal feature context vector;

[0011] S6: According to the embedding vector, the attribute feature context vector and the temporal feature context vector, use a fully - connected neural network to obtain the prediction result of the superheat degree of aluminum electrolysis; adjust the parameters of the electrolytic cell according to the prediction result; among them, the parameters of the electrolytic cell include: current, voltage and blanking time interval.

[0012] Preferably, the process of processing the original mixed-frequency data using a sliding window includes: determining the sliding window size according to the sampling time interval and the optimal lag order of the sampling frequency time series data of the target prediction index; sliding the sliding window according to the time interval of the maximum frequency in the mixed-frequency data to obtain the mixed-frequency data divided by the sliding window.

[0013] Preferably, the mixed-frequency data divided by the sliding window is:

[0014]

[0015] Wherein, represents the mixed-frequency data matrix at time t, represents the time series data of the k-th frequency, n k represents the data dimension of the k-th frequency, l k represents the data volume of the k-th frequency data, and K represents the number of frequency types of the mixed-frequency data.

[0016] Preferably, the formula for encoding the mixed-frequency data divided by the sliding window is:

[0017]

[0018] Wherein, represents the encoding matrix at time t, represents the preprocessed mixed-frequency data, K represents the number of frequency types of the mixed-frequency data, represents the encoding matrix element, F k (·) represents the long short-term memory network, m k represents the data dimension after encoding the data of the k-th frequency by the long short-term memory network, l k represents the data volume of the k-th frequency data.

[0019] Preferably, the embedding vector is:

[0020]

[0021] Wherein, h t represents the embedding vector of the mixed-frequency data at time t, represents the embedding vector of the k-th frequency data, K represents the number of frequency types of the mixed-frequency data, and m represents the dimension after splicing the embedding vectors of each frequency data.

[0022] Preferably, the fusion encoding matrix is:

[0023]

[0024]

[0025] Wherein, Denote the fusion coding matrix at time t, and * denotes the convolution calculation. Denote the feature matrix of the k-th frequency data. Denote the elements of the coding matrix. Denote the c-th convolution kernel for processing the k-th frequency data, C denotes the total number of convolution kernels, and K denotes the number of frequency types of the mixed-frequency data.

[0026] Preferably, the formula for learning the fusion coding matrix using the temporal attribute attention mechanism is:

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033] Among them, Denote the weight of the i-th attribute feature vector in the fusion coding matrix. Denote the weight of the j-th temporal feature vector in the fusion coding matrix, Energy() represents a multi-layer perceptron. Denote the attribute feature vector. Denote the temporal feature vector, h t Denote the embedding vector of the mixed-frequency data at time t, W af Denote the parameters of the multi-layer perceptron, W tf Denote the parameters of the multi-layer perceptron. Denote the attribute feature attention vector. Denote the temporal feature attention vector. Denote the weight of the m-th attribute feature vector in the fusion coding matrix. Denote the weight of the C-th temporal feature vector in the fusion coding matrix, Normalize() represents normalization. Denote the attribute feature context vector. Denote the temporal feature context vector, m represents the attribute feature dimension, and C represents the total number of convolution kernels.

[0034] Preferably, the formula for obtaining the prediction target of the event is:

[0035]

[0036] Among them, denotes the predicted target value at time t, Δ is the prediction step size, and W t denotes the first trainable parameter, denotes the second training parameter, denotes the third training parameter, denotes the attribute feature context vector, denotes the time series feature context vector, W h denotes the fourth trainable parameter, h t denotes the embedding vector of the mixed-frequency data at time t.

[0037] The beneficial effects of the present invention are as follows: The present invention regards mixed-frequency data as a special time series multi-source data, uses a convolutional neural network to fuse the features of each data source at the feature layer, and solves the problem of data volume differences brought by different sampling frequencies in mixed-frequency data; combines a time series attribute attention mechanism to identify important information for the prediction target from the fused features, thereby improving the prediction accuracy; the present invention effectively utilizes the characteristics of the original mixed-frequency data itself to mine important information related to the prediction target and improves the accuracy of the prediction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of the overheat prediction method for aluminum electrolysis mixed-frequency data based on the attention mechanism in the present invention;

[0039] Figure 2 is a schematic diagram of the framework of the overheat prediction method for aluminum electrolysis mixed-frequency data based on the attention mechanism in the present invention;

[0040] Figure 3 is a schematic diagram of the abstract description of the mixed-frequency data organization form in the present invention;

[0041] Figure 4 is a framework diagram of the long short-term memory network in the present invention;

[0042] Figure 5 is a schematic diagram of the detailed calculation method of the convolutional network in the present invention;

[0043] Figure 6 is a schematic diagram of the detailed calculation method of the time series attribute attention mechanism in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] The present invention proposes a method for predicting mixed-frequency data based on an attention mechanism, as Figure 1 shown, the method comprising:

[0046] S1: Obtain different index data related to the superheat degree of aluminum electrolysis at different sampling frequencies as the original mixed-frequency data; wherein, the indexes related to the superheat degree of aluminum electrolysis include: molecular ratio, alumina concentration, current density, voltage, and electrolysis temperature;

[0047] S2: Process the original mixed-frequency data using a sliding window to obtain the mixed-frequency data divided by the sliding window;

[0048] S3: Encode the mixed-frequency data divided by the sliding window using a long short-term memory network to obtain an encoded matrix and an embedding vector;

[0049] S4: Extract the temporal features of the encoded matrix using a convolutional neural network, and fuse the extracted temporal features to obtain a fused encoded matrix;

[0050] S5: Learn the fused encoded matrix using a temporal attribute attention mechanism to obtain an attribute feature context vector and a temporal feature context vector;

[0051] S6: According to the embedding vector, the attribute feature context vector, and the temporal feature context vector, use a fully connected neural network to obtain the prediction result of the superheat degree of aluminum electrolysis; adjust the various indexes of the electrolytic cell according to the prediction result; wherein, the parameters of the electrolytic cell include: current, voltage, and blanking time interval.

[0052] A preferred embodiment of the present invention is as follows:

[0053] As Figure 2 shown, and h t respectively represent the input data at time t, the encoded matrix of the mixed-frequency data, the fused encoded matrix, and the embedding vector; the sliding window selects the input mixed-frequency data from the mixed-frequency data, denoted as Then, allocate a long short-term memory network model for each frequency of data to encode the mixed-frequency data to obtain and h t ; thereafter, use K×C one-dimensional convolutional kernels of length l k to calculate and obtain wherein each frequency is allocated C convolutional kernels to extract or expand the temporal features; and respectively represent the attribute feature context vector and the temporal feature context vector, is the prediction result.

[0054] The mixed-frequency data comes from different data sources, and each data source is collected at a different sampling frequency; the organization of the mixed-frequency data is as Figure 3 shown, and the blank spaces in the figure represent sampling gaps. The original mixed-frequency data can be expressed as:

[0055]

[0056]

[0057] where T = max{T 1 ,..., T K}, n k and T k represent the dimension and length of the sequence respectively.

[0058] The process of processing the original mixed-frequency data using a sliding window includes: determining the sliding window size according to the sampling time interval of the time-series data of the prediction target sampling frequency and the optimal lag order p, where p is calculated by the grid search method; sliding the sliding window according to the time interval of the maximum frequency in the mixed-frequency data to obtain the mixed-frequency data divided by the sliding window; the mixed-frequency data divided by the sliding window is:

[0059]

[0060] where represents the mixed-frequency data matrix at time t, represents the time-series data of the k-th frequency, n k represents the data dimension of the k-th frequency, that is, there are n k attributes sampled at the k-th frequency, l k represents the data volume of the k-th frequency data, and K represents the number of frequency types of the mixed-frequency data.

[0061] As Figure 4 shown, encoding the mixed-frequency data divided by the sliding window using a group of long short-term memory networks includes: each frequency is assigned a long short-term memory network to learn the time-series characteristics of a specific frequency, the time series of each frequency is independently encoded, and the encoding matrix and the embedding vector h t are obtained; specifically, the formula for encoding the mixed-frequency data divided by the sliding window is:

[0062]

[0063]

[0064] where represents the encoding matrix at time t, represents the preprocessed mixed-frequency data, and K represents the number of frequency types of the mixed-frequency data. represents the encoding matrix element, F k (·) represents the long short-term memory network, m k represents the data dimension after encoding the data of the k-th frequency by the long short-term memory network, l k represents the data volume of the data of the k-th frequency. represents the final state after encoding the data of the k-th frequency by the long short-term memory network.

[0065] The embedding vector representation of the mixed-frequency data at time t is:

[0066]

[0067] where, h t represents the embedding vector of the mixed-frequency data at time t. represents the embedding vector of the data of the k-th frequency, K represents the number of frequency types of the mixed-frequency data, and m is the dimension after concatenating the embedding vectors of each frequency data, that is

[0068] Traditional prediction models cannot handle mixed-frequency data largely because the data volumes of various attributes in the mixed-frequency data are different. The advantage of the convolutional neural network lies in its strong ability to extract features. As Figure 5 shown, a convolutional neural network with a set of convolutional kernels with the same number is used to extract or expand the time series features, and a convolutional neural network with the same number of convolutional kernels is assigned to the data of each frequency; the convolutional neural network can adaptively align the mixed-frequency data, enabling the mixed-frequency data to achieve information fusion at the feature layer and obtaining a fusion encoding matrix; compared with the traditional alignment method, the present invention has fewer steps of artificially intervening in the data, which is beneficial to improving the prediction accuracy by utilizing the characteristics of the mixed-frequency data itself; specifically, assuming that each convolutional neural network model has C convolutional kernels, the set of convolutional kernels on the data of the k-th frequency is:

[0069]

[0070] The output of each convolutional neural network is a set of time series attributes and feature attributes, which can be expressed as:

[0071]

[0072]

[0073] where, represents the feature matrix of the data of the k-th frequency, * represents the convolution calculation, represents the element of the encoding matrix. It represents the c-th convolution kernel for processing the k-th frequency data. C represents the total number of convolution kernels, and its size is also the dimension of the temporal feature.

[0074] The temporal features in [[ ]] can be represented in the form of a column vector:

[0075]

[0076] The attribute features in [[ ]] can be represented in the form of a row vector:

[0077]

[0078] Fuse the temporal features obtained by convolution at each frequency to obtain a fused encoding matrix; the fused encoding matrix is:

[0079]

[0080] The task of temporal data prediction is different from natural language processing. The traditional attention mechanism focuses on the importance of a certain moment for the prediction result, while in temporal data prediction, the importance of different relevant variables for the prediction target is different. For example Figure 6 As shown in [[ ]], in order to better learn the context-related vector from the fused encoding matrix, the present invention adds attention in the attribute dimension on the basis of the traditional attention, aiming to meet the need to distinguish the influence degree of different time series on the prediction target in the time series prediction task. This strategy identifies important information from both the temporal and attribute dimensions, not only can effectively identify the fluctuations of data in the temporal dimension, but also can assign reasonable weights to different relevant indicators. The setting of this attention mechanism better meets the actual scenario requirements of temporal data prediction; the formula for learning the context-related vector from the fused encoding matrix by using the temporal-attribute attention mechanism is:

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087] Among them, represents the weight of the i-th attribute feature vector in the fused encoding matrix, represents the weight of the j-th temporal feature vector in the fused encoding matrix, Energy() represents a multi-layer perceptron, represents the attribute feature vector, represents the time series feature vector, h t represents the embedding vector of the mixed-frequency data at time t, represents the parameters of the multi-layer perceptron, represents the parameters of the multi-layer perceptron, represents the attribute feature attention vector, represents the time series feature attention vector, represents the weight of the m-th attribute feature vector in the fusion coding matrix, represents the weight of the C-th time series feature vector in the fusion coding matrix, Normalize() represents normalization, represents the attribute feature context vector, represents the time series feature context vector, m represents the attribute feature dimension, C represents the total number of convolution kernels, and its size is also the time series feature dimension.

[0088] According to the embedding vector, the attribute feature context vector, and the time series feature context vector, using a fully connected neural network, the prediction target of the event is obtained, and the formula is:

[0089]

[0090] where, represents the predicted target value at time t, Δ is the prediction step, W t represents the first trainable parameter, represents the second trainable parameter, represents the third trainable parameter, represents the attribute feature context vector, represents the time series feature context vector, W h represents the fourth trained parameter, h t represents the embedding vector of the mixed-frequency data at time t.

[0091] In industrial electrolytic aluminum, it is usually required that the superheat degree be maintained within a certain range. If the predicted value is not within this range, technicians can adopt some technical means to control it so that the prediction target meets the ideal range; applying the present invention to predict the superheat degree in industrial electrolytic aluminum, at this time, the index data related to the superheat degree of aluminum electrolysis is used as the mixed-frequency data, and the prediction result of the superheat degree can be obtained by using the present invention. According to the prediction result, the parameters of the electrolytic cell are adjusted, and the quality of aluminum production can be improved.

[0092] The present invention is not limited to the application in the field of electrolytic aluminum prediction. Optionally, the present invention can also be applied to many different fields, for example: the present invention is applied to energy consumption prediction to predict the energy consumption value of the next month. If the predicted value is larger than that of the previous month, the relevant energy department can increase energy output to meet market demand; the present invention is applied to the financial field, and users can combine daily data such as stock volatility, monthly data such as consumer price index and industrial production to predict the GDP of the next quarter. Reasonable prediction of GDP can enable experts to judge the macroeconomic operation status and formulate correct macroeconomic policies. In addition, due to the diversity of wearable devices, the acquisition frequency of sensors is also diverse. The present invention can be applied to the field of human activity recognition. According to the readings of the three-axis acceleration sensor and the heart rate sensor, human activities can be identified, which can effectively carry out daily needs such as health management.

[0093] The present invention regards mixed data as a special time series multi-source data, and uses convolutional neural networks to fuse the features of each data source at the feature layer to solve the problem of data volume differences caused by different sampling frequencies in mixed data; combines the time series attribute attention mechanism to identify important information for the prediction target from the fused features, thereby improving the prediction accuracy; the present invention effectively utilizes the characteristics of the original mixed data itself to mine important information related to the prediction target, thereby improving the accuracy of the prediction results.

[0094] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A superheat prediction method for aluminum electrolysis mixed-frequency data based on an attention mechanism, characterized in that Including: S1: Obtain different index data related to the superheat degree of aluminum electrolysis at different sampling frequencies as the original mixed-frequency data; among them, the indexes related to the superheat degree of aluminum electrolysis include: molecular ratio, alumina concentration, current density, voltage, and electrolysis temperature; S2: Process the original mixed-frequency data using a sliding window to obtain the mixed-frequency data divided by the sliding window; S3: Encode the mixed-frequency data divided by the sliding window using a long short-term memory network to obtain an encoded matrix and an embedding vector; S4: Extract the temporal features of the encoded matrix using a convolutional neural network, and fuse the extracted temporal features to obtain a fused encoded matrix; S5: Learn the fused encoded matrix using a temporal attribute attention mechanism to obtain an attribute feature context vector and a temporal feature context vector; the formula for learning the fused encoded matrix using a temporal attribute attention mechanism is: Among them, represents the weight of the i-th attribute feature vector in the fusion coding matrix, represents the weight of the j-th temporal feature vector in the fusion coding matrix, Energy() represents a multi-layer perceptron, r i represents the attribute feature vector, z j represents the temporal feature vector, h t represents the embedding vector of the mixed-frequency data at time t, W af represents the parameter of the multi-layer perceptron, W tf represents the parameter of the multi-layer perceptron, represents the attribute feature attention vector, represents the temporal feature attention vector, represents the weight of the m-th attribute feature vector in the fusion coding matrix, represents the weight of the C-th temporal feature vector in the fusion coding matrix, Normalize() represents normalization, represents the attribute feature context vector, represents the temporal feature context vector, m represents the attribute feature dimension, C represents the total number of convolutional kernels; S6: According to the embedding vector, the attribute feature context vector, and the temporal feature context vector, use a fully connected neural network to obtain the prediction result of the superheat degree; adjust the parameters of the electrolytic cell according to the prediction result; among them, the parameters of the electrolytic cell include: current, voltage, and blanking time interval.

2. The overheat degree prediction method for aluminum electrolysis mixed-frequency data based on the attention mechanism according to claim 1, wherein The process of processing the original mixed-frequency data using a sliding window includes: determining the size of the sliding window according to the sampling time interval and the optimal lag order of the sampling frequency time series data of the target prediction index; sliding the sliding window according to the time interval of the maximum frequency in the mixed-frequency data to obtain the mixed-frequency data divided by the sliding window.

3. The overheat degree prediction method for aluminum electrolysis mixing frequency data based on the attention mechanism according to claim 2, wherein The mixed-frequency data divided by the sliding window is: Among them, represents the mixing data matrix at time t, represents the time-series data of the k-th frequency, n k represents the dimension of the k-th frequency data, l k represents the data volume of the k-th frequency data, and K represents the number of frequency types of the mixing data.

4. A method for predicting the superheat degree of aluminum electrolysis mixing frequency data based on the attention mechanism according to claim 1, characterized in that The formula for encoding the mixed-frequency data divided by the sliding window is: Among them, represents the mixing data coding matrix at time t, represents the preprocessed mixing data, and K represents the number of frequency types of the mixing data, represents the element of the coding matrix, F k (·) represents the long short-term memory network, m k represents the data dimension after the data of the k-th frequency is encoded by the long short-term memory network, l k represents the data volume of the data of the k-th frequency.

5. The superheat prediction method for aluminum electrolysis mixed-frequency data based on the attention mechanism according to claim 1, wherein The embedding vector is: Among them, h t represents the embedding vector of the mixed-frequency data at time t, represents the embedding vector of the k-th frequency data, K represents the number of frequency types of the mixed-frequency data, and m represents the dimension after splicing the embedding vectors of each frequency data.

6. The overheat prediction method for aluminum electrolysis mixing frequency data based on the attention mechanism according to claim 1, wherein The fused encoded matrix is: Among them, represents the fusion coding matrix at time t, and * represents the convolution calculation. represents the feature matrix of the k-th frequency data. represents the element of the coding matrix. represents the c-th convolution kernel for processing the k-th frequency data. C represents the total number of convolution kernels, and K represents the number of frequency types of the mixed-frequency data.

7. A method for predicting the superheat degree of aluminum electrolysis mixed-frequency data based on an attention mechanism according to claim 1, characterized in that The formula for obtaining the prediction target of the event is: Among them, represents the predicted target value at time t, Δ is the prediction step, and W t represents the first trainable parameter, represents the second trainable parameter, represents the third trainable parameter, represents the attribute feature context vector, represents the time series feature context vector, and W h represents the fourth trainable parameter, and h t represents the embedding vector of the mixed-frequency data at time t.