Power load time sequence prediction method based on dynamic decomposition and feature learning

Through the combination of SSAE and supervised autoencoder, the redundancy characteristics of the power load time series are compressed, and the end-to-end prediction model is built, which solves the problems of high-frequency redundancy and noise interference in the existing technology, and improves the accuracy and efficiency of power load prediction.

CN120373141APending Publication Date: 2025-07-25NORTHEASTERN UNIV CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510640210.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing decomposition-based power load time series prediction method has problems such as many high-frequency redundant components, high feature dimensions, complex models and susceptible to noise interference, which affects the prediction accuracy and efficiency.

Method used

The feature selection method SSAE is adopted based on deep learning, combining sliding empirical wavelet transformation and supervised autoencoder, compressing redundant features, and constructing an end-to-end power load time series prediction model to achieve the fusion of feature extraction and prediction.

Benefits of technology

It effectively reduces the complexity of the model, improves the prediction accuracy and robustness, and improves the efficiency and adaptability of power load time series prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373141A_ABST
    Figure CN120373141A_ABST
Patent Text Reader

Abstract

The invention discloses a power load time sequence prediction method based on dynamic decomposition and feature learning, and relates to the field of power load prediction. The core thought of the method is that power load data in a certain time range are collected, and an original power load time sequence is preprocessed; in order to fully excavate a change mode in the power load time sequence, a sliding decomposition method is introduced to decompose the preprocessed time sequence, and local features of the time sequence are extracted; in consideration of redundant information in the sub-sequence obtained through decomposition, a feature extraction mechanism is further introduced to compress and screen the sub-sequence, key features are reserved to be used for constructing a power load time sequence prediction model, and accurate prediction of future power loads is achieved. According to the method, the expression capability of the key features of the power load time sequence is improved, the integrated process of the power load time sequence from decomposition to prediction is realized, the model complexity is effectively reduced, noise interference is suppressed, and the prediction precision and efficiency of the power load time sequence are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric load forecasting, and particularly relates to a method for forecasting electric load time series based on dynamic decomposition and feature learning. Background Art

[0002] Electric load is an important indicator for measuring electric power demand in the power system, reflecting the power consumption demand of the power grid at different time periods. With the acceleration of industrialization and urbanization, the volatility and complexity of electric load are increasing day by day. Accurately forecasting electric load is of great significance for power grid dispatching, power generation planning, and energy optimization. In order to effectively forecast electric load, electric load time series analysis methods are usually adopted. The electric load time series records the process of electric power demand changing with time, including various factors such as trends, periodic fluctuations, and seasonal variations. The complexity of these data makes the forecasting process full of challenges. Electric load time series forecasting aims to predict the future change of electric power demand by analyzing historical data, so as to help the power system make decisions in advance and ensure the stability and efficiency of the power grid operation. With the continuous increase in the proportion of new energy access and the increasingly complex operation mode of the power system, higher requirements are put forward for the accuracy and response speed of electric load time series forecasting. To address these challenges, the decomposition forecasting method has become an effective and widely used research direction in electric load time series forecasting. Decomposition forecasting improves the forecasting accuracy and efficiency by decomposing the electric load time series into multiple different components and modeling them separately. Commonly used decomposition methods include empirical mode decomposition, variational mode decomposition, and wavelet transform, etc. However, after decomposing the original time series, the existing decomposition-based electric load forecasting methods often generate a large number of high-frequency redundant components, resulting in a higher model dimension, increased computational complexity, and being more sensitive to noise, which affects the forecasting accuracy. How to effectively compress redundant features while decomposing and improve the forecasting performance of electric load time series remains a key problem to be solved urgently. Summary of the Invention

[0003] In view of the above deficiencies of the prior art, the present invention provides a method for forecasting electric load time series based on dynamic decomposition and feature learning. This method improves the expression ability of the key features of the electric load time series by compressing the features generated after decomposing the original electric load time series, and on this basis, realizes an integrated process from decomposition to forecasting of the electric load time series, thereby effectively reducing the model complexity, suppressing noise interference, and improving the accuracy and efficiency of electric load time series forecasting.

[0004] The technical solution of the present invention is as follows:

[0005] A method for forecasting electric load time series based on dynamic decomposition and feature learning, the method comprising the following steps:

[0006] Step 1: Determine the acquisition frequency and acquire the power data according to the determined acquisition frequency to obtain the original power load time series;

[0007] Step 2: Preprocess the original power load time series obtained in Step 1: fill in the missing values and perform normalization after filling in the missing values; divide the preprocessed power load time series into training sequences;

[0008] Step 3: Decompose the training sequences obtained in Step 2 to obtain the corresponding subsequences;

[0009] Step 4: Concatenate the training sequences with their subsequences to obtain the training set;

[0010] Step 5: Use the training set obtained in Step 4 to train the SSAE to obtain the power load time series prediction model;

[0011] Step 6: Apply the power load time series prediction model obtained in Step 5 to actual applications for real-time power load time series prediction.

[0012] Further, according to the power load time series prediction method, in Step 2, the interpolation method is used to fill in the missing values of the original power load time series obtained in Step 1, ensuring that the deviation rate of the data statistical characteristics of the complete power load time series obtained after filling from the data statistical characteristics of the original power load time series is less than 5%.

[0013] Further, according to the power load time series prediction method, in Step 3, the sliding empirical wavelet transform method is used to decompose the training sequences obtained in Step 2 to obtain the corresponding subsequences.

[0014] Further, according to the power load time series prediction method, the steps include the following steps:

[0015] Step 3.1: Define the sliding window of the power load time series and set the sliding window length w and the number k of subsequences to be decomposed;

[0016] Step 3.2: According to the sliding window length w and the number k of subsequences to be decomposed determined in Step 3.1, use the empirical wavelet transform method to decompose the training sequences obtained in Step 2 respectively to obtain the corresponding subsequences.

[0017] Further, according to the power load time series prediction method, in Step 3.1, w is set to an integer multiple of the period T of the power load time series.

[0018] Further, according to the power load time series prediction method, in Step 3.1, the number k of subsequences is set to 2, 3 or 4.

[0019] Further, according to the power load time series prediction method, when performing the splicing operation on the training sequence and its subsequence in step 4: at each time t, only the time step samples in the previous cycle are extracted, and these samples are spliced into the input vector z t ; The cycle length is T.

[0020] Further, according to the power load time series prediction method, step 5 includes the following steps:

[0021] Step 5.1: Use the first supervised autoencoder SAE1 to provide the input vector z of the training set obtained in step 4 t Perform feature compression and obtain z t The potential low-dimensional representation of And the encoder parameters of SAE1 The d 1 express Dimensions, and satisfy d 1 <T(k+1),其中T(k+1)表示z t Dimensions;

[0022] Step 5.2: Use the second supervised autoencoder SAE2 to transform the latent low-dimensional representation obtained in step 5.1 Further compression to obtain a more compact deep feature representation And the encoder parameters of SAE2 where d 2 express Dimensions, and satisfy d 2 <d 1 ;

[0023] Step 5.3: Based on the encoder part of the two supervised autoencoders SAE pre-trained in steps 5.1 and 5.2, cascade and then connect the output layer to build a network framework, using the input vector z provided by the training set t and supervision label y t The network is fine-tuned end-to-end to obtain a power load time series prediction model; the supervision label in Represents the normalized sample value at time t.

[0024] Further, according to the power load time series prediction method, the encoder parameters trained in steps 5.1 and 5.2 are used. and As the initialization weights of the encoder part of the first two layers in SSAE, namely the two supervised autoencoders SAE.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] Aiming at the problems existing in the existing decomposition-based power load time series prediction method, such as many high-frequency redundant components, high feature dimension, complex model and susceptibility to noise interference, this method is improved from two aspects:

[0027] (1) To compress redundant features, reduce dimensions and suppress noise interference, a feature selection method SSAE (Stacked Supervised Autoencoder) based on deep learning is used to screen and compress the features obtained by decomposing the original power load time series, so as to effectively extract key features and improve the robustness and prediction accuracy of the model;

[0028] (2) To avoid fragmentation of each link, improve the overall prediction efficiency and effect, integrate decomposition, feature extraction and prediction modeling, construct an end-to-end prediction process, realize automatic processing, and improve the prediction accuracy and system adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flowchart of the power load time series prediction method based on dynamic decomposition and feature learning in this embodiment;

[0030] Figure 2 It is an hourly curve graph of the New England power load in 2022;

[0031] Figure 3 It is a structural diagram of the supervised stacked autoencoder model in this embodiment, where (a) is the structural diagram of the first supervised autoencoder SAE1; (b) is the structural diagram of the second supervised autoencoder SAE2; (c) is the structural diagram of the power load time series prediction model;

[0032] Figure 4 It is a fitting effect diagram of the true value and the predicted value obtained by using the method of the present invention under different prediction steps in this embodiment, where Figure (a) is a comparison diagram of one-step prediction; Figure (b) is a comparison diagram of three-step prediction; Figure (c) is a comparison diagram of five-step prediction. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments. The specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0034] The core idea of the present invention is as follows: 1. Collect power load data within a certain time range, preprocess the original power load time series, and construct training and test sequences; 2. To fully extract the change patterns in the power load time series, introduce a sliding decomposition method to decompose the time series and extract its local features; 3. Considering the redundant information in the decomposed subsequences, further introduce a feature extraction mechanism to compress and screen them, retain key features for constructing a power load time series prediction model, and achieve accurate prediction of future power loads.

[0035] Figure 1 is a flowchart of the power load time series prediction method based on dynamic decomposition and feature learning in this embodiment. As Figure 1 shown, the power load time series prediction method based on dynamic decomposition and feature learning includes the following steps.

[0036] Step 1: Determine the collection frequency and collect power data according to the determined collection frequency to obtain the original power load time series;

[0037] The first step in power load prediction is data collection. Usually, power data is collected in real time through smart meters or related monitoring devices. To ensure that the collected data covers different seasons and time periods, power data can be collected once per minute or per hour to obtain a power load time series that reflects the fluctuation characteristics of the load over time. The power load hourly time series in the area of the Independent System Operator of New England (ISO New England, ISONE) in 2022 shown in Figure 2 is adopted in this embodiment.

[0038] Step 2: Preprocess the original power load time series obtained in Step 1 to obtain the preprocessed power load time series and divide it into a training sequence and a test sequence;

[0039] Step 2.1: Fill in the missing values in the original power load time series obtained in Step 1 based on the principle of being able to reflect the fluctuation changes of the power load to obtain a complete power load time series;

[0040] Due to reasons such as equipment failures, there are cases of missing data collection in the original power load time series obtained in Step 1, which cannot accurately reflect the fluctuation changes of the power load. Therefore, to ensure the comprehensiveness of the data, this embodiment uses the interpolation method to fill in the missing values in the original power load time series obtained in Step 1, ensuring that the deviation rate between the data statistical characteristics of the obtained complete power load time series and the data statistical characteristics of the original power load time series is less than 5%.

[0041] Step 2.2: normalize the complete power load time series obtained in step 1 to obtain a preprocessed power load time series and divide it into a training sequence and a test sequence;

[0042] This embodiment uses the maximum and minimum normalization shown in formula (1) to scale the complete power load time series obtained in step 1 to between 0 and 1, so as to accelerate the convergence speed of the subsequent power load time series prediction model.

[0043]

[0044] in is the normalized sample value at time t; is the original sample value at time t; and are the maximum and minimum sample values of the training sequence respectively.

[0045] Step 3: Use the sliding empirical wavelet transform method to decompose the training sequence and test sequence obtained in step 2 respectively to obtain corresponding subsequences;

[0046] The power load time series usually contains various information such as power consumption trends, cycles, small-scale human influences, and noise. These information are relatively independent in the spectrum and have different frequencies. Therefore, empirical wavelet decomposition can effectively separate various types of information in the power load time series and mine the potential detailed features. This implementation introduces a sliding empirical wavelet decomposition method, in which the decomposition window slides every time time moves forward by one point to dynamically decompose the training sequence and test sequence obtained in step 2. The specific steps are as follows:

[0047] Step 3.1: Define the sliding window of the power load time series and determine the sliding window length w and the number of subsequences k to be decomposed;

[0048] First, define a sliding window of length w. At each time t, the sliding window contains the data of w time steps before time t. In order to better capture the periodicity and trend characteristics in the power load time series, especially in the scenario with a significant daily cycle, w is usually set to an integer multiple of the period T of the power load time series, such as 24, 48 or 72 (in hours in this embodiment).

[0049] Subsequently, in order to adapt to the information such as trend, cycle, disturbance and noise in the power load time series, the number of subsequences k required to be decomposed is usually set to 2, 3 or 4 for the data in each window.

[0050] Step 3.2: According to the sliding window length w determined in Step 3.1 and the number k of subsequences to be decomposed, the training sequence and the test sequence obtained in Step 2 are respectively decomposed using the empirical wavelet transform method to obtain the corresponding subsequences.

[0051] After determining the sliding window length w and the number k of subsequences to be decomposed, at time t, for the sliding window of the power load time series with length w perform the decomposition operation to obtain k subsequences, denoted as:

[0052]

[0053] where represents the sample value of the first subsequence at time t - w; represents the sample of the first subsequence at time t - 1; represents the sample of the k-th subsequence at time t - w; represents the sample of the k-th subsequence at time t - 1.

[0054] At time t + 1, the sliding window of the power load time series slides one time step to the right and is updated to and the above decomposition operation is repeated. For both the training sequence and the test sequence, the same processing method of the sliding window of the power load time series is adopted, and the window always slides to the last sample in the current sequence.

[0055] Step 4: Concatenate the training sequence with its subsequences to obtain the training set; concatenate the test sequence with its subsequences to obtain the test set;

[0056] To improve the computational efficiency, in this implementation, the training / test sequence obtained in Step 2 is concatenated with the corresponding subsequences generated in Step 3: at each time t, only the time step samples within its previous period (length T) are extracted, and these samples are concatenated into a one-dimensional input vector z t .

[0057]

[0058] where, x norm represents the normalized original sequence, and x (i) is the i-th component sequence generated by the empirical wavelet decomposition of the sliding window of the power load time series at time t After concatenation, an input feature with dimension T(k + 1) is formed.

[0059] The corresponding supervised label is:

[0060]

[0061] where Represents the normalized sample value at time t.

[0062] Step 5: Use the training set obtained in Step 4 to train the SSAE to obtain a power load time series prediction model.

[0063] After Step 4, the input dimension of the power load time series prediction model increases to k + 1 times the original. On the one hand, it brings more features. On the other hand, high-frequency components usually contain noise in the training sequence, and directly inputting them into the power load time series prediction model may weaken the prediction performance. To further improve the generalization ability of the power load time series prediction model, it is necessary to perform deep extraction on the input. The autoencoder is a typical unsupervised learning structure that can achieve effective information extraction and dimensionality reduction by learning the compressed representation of the input data. On this basis, this embodiment introduces a Supervised Autoencoder (SAE) as a feature extraction module, and combines the supervised labels of the power load time series prediction to effectively learn the deep features of the input of the power load time series prediction model. The specific training strategy is as follows:

[0064] Step 5.1: Use the first supervised autoencoder SAE1 to perform feature compression on the input vector z provided by the training set obtained in Step 4 t to obtain its latent low-dimensional representation (where d 1 represents the dimension of 1 and satisfies d

[0065] such as Figure 3 (a) shows that the first supervised autoencoder SAE1 is used to extract low-dimensional latent features from the input vector provided by the training set obtained in Step 4 . First, the input vector z t is mapped through the encoder function to obtain the hidden feature representation where are the parameters of the SAE1 encoder. Then, through the decoder function the input is reconstructed to obtain the reconstructed input vector where are the parameters of the SAE1 decoder. At the same time, the prediction branch outputs the power load time series prediction result through the function where are the parameters of the SAE1 prediction branch. To achieve effective compression of the input information and full expression of the prediction target, SAE1 introduces a joint loss function during the training process, and its form is as follows:

[0066]

[0067] Among them, α 1 ∈[0,1] is a weighting coefficient used to balance the input reconstruction error and prediction error. represents the input reconstruction error, which measures the difference between the original input vector and the reconstructed vector. Represents the forecast error, which measures the difference between the actual load and the predicted load.

[0068] Step 5.2: Use the second supervised autoencoder SAE2 to transform the latent low-dimensional representation obtained in step 5.1 Further compression to obtain a more compact deep feature representation (where d 2 express The dimension satisfies d 2 <d 1 ) and extract its encoder parameters

[0069] like Figure 3 As shown in (b), the second supervised autoencoder SAE2 is dedicated to the potential low-dimensional representation obtained in step 5.1. Perform deeper feature extraction and information refinement, further compress its dimensions to mine more discriminative hidden features. Specifically, the input vector Encoder Function Mapping to a new low-dimensional representation in are the parameters of the SAE2 encoder. Subsequently, through the decoder function Reconstruct input and output reconstruction result in is the parameter of the SAE2 decoder. At the same time, the prediction branch is passed through the function Output the corresponding load forecast results in is the parameter of the SAE2 prediction branch. In order to achieve a balance between effective representation learning and prediction performance, SAE2 also uses a joint loss function, which is defined as follows:

[0070]

[0071] Among them, α 2 ∈[0,1] is a weighting coefficient used to balance the input reconstruction error and prediction error. represents the input reconstruction error, Represents the prediction error.

[0072] Step 5.3: Based on the encoder parts of the two supervised autoencoders SAE pre-trained in Steps 5.1 and 5.2, cascade them and then connect the output layer to construct a network framework, and use the input vector z t and the supervised label y t to perform end-to-end fine-tuning on the network to obtain a power load time series prediction model.

[0073] As Figure 3 (c) shows, cascade the encoder functions obtained from training in Steps 5.1 and 5.2 and as the feature extraction module of the overall model, and on this basis, connect the final prediction branch to establish the framework of the power load time series prediction model. To make full use of the learned feature representation ability, in this stage, use the encoder parameters obtained from training in Steps 5.1 and 5.2 and as the initial weights to improve the training efficiency and prediction performance of the overall model.

[0074] Specifically, the original input vector z t is sequentially mapped through two initialized encoders to obtain the final deep features

[0075]

[0076] Then, the deep features are input into the final prediction branch f final , and the prediction result is output:

[0077]

[0078] where θ final represents the parameters of the final prediction branch;

[0079] In this stage, the encoder and the prediction branch are combined into a unified end-to-end model, and overall fine-tuning training is performed to optimize the parameters to minimize the prediction error:

[0080]

[0081] Through this end-to-end training strategy, the model can not only inherit the pre-training ability of the encoder but also further jointly optimize to improve the overall performance.

[0082] Step 6: Use the power load time series prediction model obtained in Step 5 for real-time power load time series prediction in actual applications.

[0083] In practical applications, the deployed power load time series prediction model is loaded into the target device. The system receives the power load data stream in real time and performs missing value filling and normalization processing on the data according to Step 2. Samples of the previous continuous k time steps before the current moment are extracted in the form of a power load time series sliding window and decomposed using the empirical wavelet transform method, and the splicing operation is performed in the manner of Step 4 to obtain the model input. Subsequently, it is input into the power load time series prediction model to obtain the output, and the output result is de-normalized to obtain the power load prediction value at the current moment:

[0084]

[0085] Among them, and are the maximum sample value and the minimum sample value of the training sequence, respectively.

[0086] To evaluate the overall performance of the method of the present invention, it is compared with traditional prediction methods on the ISONE power load data set. The comparison objects include: Support Vector Regression (SVR), Multilayer Perceptron (MLP), Long Short-Term Memory (LSTM), Empirical Mode Decomposition-MLP (EMD-MLP), and Empirical Wavelet Transform-edRVFL (EWT-edRVFL). The evaluation of the prediction performance uses the Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE) as performance indicators, and the relevant expressions are shown as follows:

[0087]

[0088] Among them, is the true load value, is the predicted value generated by the power load time series prediction model. To evaluate the performance of this technology in power load prediction, experimental comparisons are carried out under one-step, three-step, and five-step prediction tasks. The prediction performance is shown in Table 1, Figure 4 shows the fitting situation between the predicted value and the actual value in a typical period, intuitively reflecting the prediction accuracy.

[0089] Comparison of Prediction Indicators in Table 1

[0090]

[0091] From Figure 4 and Table 1, it can be seen that the method of the present invention performs better than other models at all prediction steps, with the smallest error, demonstrating its advantage in prediction accuracy. Especially in the three-step and five-step predictions, its excellent ability in short-term prediction is reflected, and it can maintain a high prediction accuracy within a longer prediction range.

Claims

1. A power load time series prediction method based on dynamic decomposition and feature learning, characterized in that The method includes the following steps: Step 1: Determine the acquisition frequency and acquire power data according to the determined acquisition frequency to obtain the original power load time series; Step 2: Preprocess the original power load time series obtained in Step 1: fill in the missing values and perform normalization after filling in the missing values; Divide the training sequence from the preprocessed power load time series; Step 3: Decompose the training sequence obtained in Step 2 to obtain the corresponding subsequences; Step 4: Perform a splicing operation on the training sequence and its subsequences to obtain a training set; Step 5: Use the training set obtained in Step 4 to train the SSAE to obtain a power load time series prediction model; Step 6: Use the power load time series prediction model obtained in Step 5 in practical applications to perform real-time power load time series prediction.

2. The power load time series prediction method according to claim 1, wherein In Step 2, the interpolation method is used to fill in the missing values of the original power load time series obtained in Step 1, ensuring that the deviation rate of the data statistical characteristics of the complete power load time series obtained after filling from the data statistical characteristics of the original power load time series is less than 5%.

3. The power load time series prediction method according to claim 1, characterized in that, In Step 3, the sliding empirical wavelet transform method is used to decompose the training sequence obtained in Step 2 to obtain the corresponding subsequences.

4. The power load time series prediction method according to claim 3, characterized in that Step 3 includes the following steps: Step 3.1: Define the sliding window of the power load time series and set the sliding window length w and the number k of subsequences to be decomposed; Step 3.2: According to the sliding window length w and the number k of subsequences to be decomposed determined in Step 3.1, use the empirical wavelet transform method to decompose the training sequence obtained in Step 2 respectively to obtain the corresponding subsequences.

5. The power load time series prediction method according to claim 4, wherein In Step 3.1, w is set to an integer multiple of the period T of the power load time series.

6. The power load time series prediction method according to claim 4, characterized in that In Step 3.1, the number k of subsequences is set to 2, 3, or 4.

7. The power load time series prediction method according to claim 1, characterized in that When performing the concatenation operation on the training sequence and its subsequence in step 4: at each time step t, only the time step samples within its previous period are extracted, and these samples are concatenated into the input vector z t ; the period length is T.

8. The power load time series prediction method according to claim 7, characterized in that Step 5 includes the following steps: Step 5.1: Use the first supervised autoencoder SAE1 to perform feature compression on the input vector z provided by the training set obtained in Step 4, and obtain the latent low-dimensional representation of z t and the encoder parameters of SAE1 t where d represents the dimension and satisfies d <T(k + 1), where T(k + 1) represents the dimension of z 1 ; <T(k + 1), where T(k + 1) represents the dimension of z 1 ; t ​ Step 5.2: Further compress the latent low-dimensional representation obtained in Step 5.1 using the second supervised autoencoder SAE2 to obtain a more compact deep feature representation as well as the encoder parameters of SAE2 where d represents 2 the dimension of and satisfies d 2 < d 1 ; Step 5.3: Based on the encoder parts of the two supervised autoencoders SAE1 and SAE2 pre-trained in Steps 5.1 and 5.2, cascade them and then connect the output layer to construct a network framework, and use the input vector z t and the supervised label y t to perform end-to-end fine-tuning on the network to obtain a power load time series prediction model; the supervised label where represents the normalized sample value at time t.

9. The power load time series prediction method according to claim 8, wherein Use the encoder parameters obtained from the training in Steps 5.1 and 5.2 and as the initial weights for the encoder parts of the first two layers in the SSAE, namely the two supervised autoencoders SAE1 and SAE2.