Stock price prediction method based on Transform-KAN model

By combining Transformer-KAN model built with Transformer and KAN networks, the problem of difficult long-term dependencies in traditional stock price prediction methods is solved, and a higher-precision stock price prediction is achieved.

CN120387891APending Publication Date: 2025-07-29JILIN INST OF CHEM TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510416689.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional stock price prediction methods are difficult to effectively capture the long-term dependence relationship and complex feature interactions of stock prices, and there is a problem of gradient disappearance, resulting in low prediction accuracy.

Method used

Combining Transformer's global dependency modeling capability and KAN's powerful nonlinear fitting capability, by replacing Transformer's fully connected layer as a KAN network, building a Transformer-KAN model, using the KAN network as the output layer, deleting the Decoder part, adding a linear layer to reduce the computational complexity, and capturing nonlinear relationships.

Benefits of technology

It improves the accuracy and generalization ability of stock price predictions, can better grasp the long-term dependency relationship and short-term trading characteristics, reduce calculation complexity, and improve prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387891A_ABST
    Figure CN120387891A_ABST
Patent Text Reader

Abstract

The invention relates to the field of financial stock price markets, in particular to a stock price prediction method based on a Transform-KAN model, which comprises the following steps: S1, selecting a proper data set on a Kaggle platform, processing missing values and abnormal values of the stock price data set, dividing the data set into a training set and a test set, carrying out maximum-minimum normalization processing on the training set and the test set; s2, generating a time window feature by using a time sequence feature extraction method, and converting the time window feature into an input format suitable for a Transform structure; s3, a Transform-KAN model is constructed for stock price prediction, a long-term dependency relationship of a time sequence is extracted by using a Transform, and an MLP layer of the Transform is replaced by using KAN (Kolmogorov-Arnold Networks), so that the nonlinear fitting capability of the model is enhanced; s4, defining a loss function and an optimizer, and setting hyper-parameters such as a learning rate and a batch size; s5, inputting the data into the network, training the model by using the training set, and storing the trained model; and S6, loading the stored model, predicting the test set, and evaluating by using various evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial stock price markets, and in particular to a stock price prediction method based on the Transformer-KAN model. Background Art

[0002] Stock price prediction has always been an important research direction in the field of financial markets. The stock market is affected by many factors such as economic policies, market sentiment, and emergencies. It does not follow a simple linear relationship but shows a highly non-linear relationship. Due to various situations such as policy adjustments and wars, the changes in the stock price market are random, making it difficult to accurately predict the stock price trend. The human factors, market manipulation, and short-term fluctuations in stock market transactions increase the noise in the data set, making it easy to fit too much invalid information during the prediction process. Traditional statistical methods and machine learning methods have limitations in dealing with long-term dependencies and complex feature interactions.

[0003] The rise of deep learning has provided new ideas for stock price prediction. Deep learning models have excellent non-linear modeling capabilities, are good at capturing long-term dependencies, are suitable for fusing multi-source data, and can perform end-to-end learning to reduce manual features. With the improvement of technology computing power, the complexity of stock market data has also increased. Traditional RNN-based models such as LSTM are prone to the problem of gradient disappearance when predicting the stock market and are difficult to capture long-term dependencies. Transformer solves certain long-term dependency problems through its self-attention mechanism, but has a high computational complexity and is prone to ignoring local patterns. Therefore, it has become a difficult point to develop a deep learning model with global dependency modeling capabilities and strong non-linear fitting capabilities. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides a stock price prediction method based on the Transformer-KAN model. This method combines the global dependency modeling ability of Transformer and the strong non-linear fitting ability of KAN (Kolmogorov-Arnold Networks) to more accurately predict the stock price trend. The use of this method includes the following steps:

[0005] S1. Perform data preprocessing. The input variables are the opening price Open, closing price Close, highest price High, lowest price Low, and trading volume Volume of the stock market, and the output variable is the closing price Close. First, read the data set and perform format conversion to ensure the correct parsing and standardization of the date field. After sorting the data, the data set needs to be detected for missing values. After filling in the missing values, it is necessary to check again to ensure its integrity and consistency.

[0006] S2. After obtaining the complete dataset, it is necessary to normalize the complete dataset to improve the training effect of the model and accelerate the convergence speed. First, divide the dataset into a training set and a test set in a ratio of 8:2 to ensure that the model can learn the

[0007] latent laws of the data. Secondly, perform min-max normalization on the training set and the test set respectively. After normalization, set the input sequence with a time window size of 20 steps.

[0008] S3. Combine the Transformer and KAN networks to construct the Transformer-KAN model. The original Transformer model relies on a fully connected layer for output. In the present invention, the KAN network is used to replace the fully connected layer as the output layer. Transformer contains an Encoder-Decoder structure. A linear layer is introduced before the Encoder to map the feature dimensions, and finally the Decoder layer is removed. By directly extracting the features of the last time step, the computational complexity is reduced.

[0009] S4. Define the loss function, optimizer, and set hyperparameters such as the learning rate and batch size

[0010] S5. Input the data into the network, train the model with the training set, and save the trained model

[0011] S6. Load the saved model, make predictions on the test set, and evaluate using various evaluation metrics.

[0012] Compared with traditional time series prediction methods and the original standard Transformer model, Transformer-KAN adopts a learnable activation function, enabling the network to adaptively adjust the non-linear mapping, improving the non-linear modeling ability, and enhancing the expressive ability of stock price prediction. Through the combination of the two, the model can not only capture long-term dependence relationships but also extract short-term trading features. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 . Schematic diagram of the method framework process;

[0014] Figure 2 . Schematic diagram of the deep neural network structure; DETAILED DESCRIPTION OF THE INVENTION

[0015] To more clearly elaborate the purpose, technical solution, and its advantages of the present invention, the present invention will be described in more detail below with reference to the drawings and embodiments. It should be noted that the specific embodiments are only used to explain the present invention and are not intended to limit it.

[0016] S1. In the data preprocessing stage, first, it is necessary to convert the format of the dataset to ensure the correct parsing and standardization of the date field. Specifically, the date field needs to be converted to a unified date and time format for subsequent processing. Secondly, set the date field as the index of the dataset so that the data can be sorted in chronological order to ensure the continuity and consistency of the data. After the data sorting is completed, use isnull().sum() to detect missing values in the dataset. For missing data, an appropriate filling method can be selected according to the specific situation. Here, interpolation filling and mean filling are used. After filling the missing values, the dataset needs to be checked again to ensure its integrity and consistency, laying a solid data foundation for subsequent feature engineering and model training.

[0017] S2. In the further steps of data preprocessing, it is necessary to normalize the complete dataset to improve the training effect of the model and accelerate the convergence speed. First, divide the dataset into a training set (80%) and a test set (20%) in a ratio of 8:2 to ensure that the model can learn the patterns in the data and be evaluated on unseen test data. The division method can use time series splitting, that is, the first 80% of the data is used for training, and the last 20% of the data is used for testing to maintain the temporal dependence of the time series data. Then, perform min-max normalization on the training set and the test set respectively, that is, scale the data to the range of [-1, 1]. The normalization calculation formula (1) is as follows:

[0018]

[0019] where x is the original data, x min and x max are the minimum and maximum values of the data respectively. To prevent data leakage, the maximum and minimum values for normalization need to be calculated based on the training set and applied to the normalization of the test set. After normalization, to meet the input requirements of the Transformer-KAN model, a sliding time window needs to be constructed. Set the time window size to 20, that is, every 20 consecutive time-step data points are used as input features, and the corresponding target variable is the value at the 21st time step. This way can capture the short-term dependencies of the time series and provide sufficient context information for the model to improve the prediction accuracy.

[0020] S3. To make full use of the global attention mechanism of Transformer in time series modeling and the efficient expression ability of KAN in non-linear mapping, we construct a Transformer-KAN combined model to improve the accuracy and generalization ability of stock price market prediction. The formula of the attention mechanism of Transformer is shown in (2):

[0021]

[0022] Among them, Q represents the input vector for which attention is to be calculated currently, generally sourced from the representation of a certain layer of the input sequence. K represents the feature representations of all positions in the input sequence, usually from the same source as Q. V corresponds one-to-one with K and represents the actual content of the input sequence. The principle formula of the KAN network is shown in (3):

[0023]

[0024] Among them, (x1, …, x n ) is the input vector with a dimension of n. represents the input transformation function, represents the activation function, and calculates local feature aggregation, that is, feature combination for each q. represents another non-linear transformation, which is used to further process the aggregated result, and finally the final output is obtained by summation. In the traditional Transformer structure, the output layer usually uses a fully connected layer for feature mapping and prediction. In this study, the KAN network is used to replace the fully connected layer of the Transformer as the output layer, enabling the model to better capture non-linear relationships and improve prediction accuracy. As a neural network based on the Kolmogorov-Arnold representation theorem, KAN can more effectively learn complex patterns in time series data by decomposing complex functions and combining multiple low-dimensional sub-networks. To reduce the computational complexity and optimize the Transformer structure, a linear layer is added before the Encoder, which is used to project the input data into a suitable feature dimension, improve the information expression ability, and match the input requirements of the Transformer. Since the time series prediction task is essentially a one-way causal inference task, the Decoder part in the original Transformer structure is removed, and the Encoder is directly used to extract features, reducing the computational overhead and improving the inference speed. After the Encoder processes the input sequence, the feature of the last time step is directly extracted as the final representation, rather than pooling or weighted averaging the entire sequence. This way can better capture the latest information of the time series and reduce unnecessary calculations.

[0025] The Dropout rate is set to 0.2 to prevent the model from overfitting. The N_layer is set to two, which is used to define the number of layers in the model. The number of heads num_heads in the multi-head attention mechanism is set to 4, and the lr learning rate is set to 0.0005. The number of neurons hidden_space in the hidden layer is set to 16. The number of training epochs is set to 100, and the seed is a random seed used to generate the randomness of the training data. The optimizer used is the ADAM optimizer. These parameters can be passed in through the command line and then parsed and used in the program. This method is very flexible and can adjust the model configuration and training details at runtime without modifying the code. The proportion of training data, the number of layers of the model, or the learning rate of the optimizer can be changed through the command line

[0026] S5. After completing data preprocessing and model construction, the next step is to input the data into the Transformer-KAN network model and use the training set to train the model. Finally, the trained model is saved for subsequent prediction and evaluation. During the training process, we use the training set (Train Set) to optimize the model parameters so that it can accurately fit the data pattern. The input data passes through the Transformer encoder, and KAN is used for non-linear mapping to generate the final prediction result. The model parameters are updated through the gradient descent algorithm to gradually approach the optimal solution. After the model training is completed, to avoid repeated training and improve the subsequent inference efficiency, we save the trained Transformer-KAN model for direct loading and use in the production environment or future experiments

[0027] S6. Load the saved model, make predictions on the test set, and evaluate using various evaluation metrics. The evaluation metrics include Mean Absolute Error (MAE), Mean Square Error (MSE), Mean Absolute Percentage Error (MAPE), and coefficient of determination (R^2). The calculation formulas are as shown in Equations 4-7

[0028] as shown in Equation 4-7:

[0029]

[0030] where N represents the number of samples, that is, the number of data points. y i is the actual value, is the predicted value. represents the absolute error of a single data point. The smaller the MAE, the smaller the model error. MAE can be used to evaluate the accuracy of model prediction and is suitable for situations that are not very sensitive to outliers

[0031]

[0032] Where N represents the number of samples, i.e., the number of data points. y i is the actual value, and is the predicted value.

[0033]

[0034] Where N represents the number of samples, i.e., the number of data points. y i is the actual value, and is the predicted value.

[0035]

[0036] y i is the actual value, is the predicted value, and is the sum of squared residuals, representing the error of the model. 2 is the total sum of squares, representing the degree of fluctuation of the data itself. R 2 reflects the degree to which the model explains the variation of the dependent variable, and its value range is generally between 0 and 1. The closer it is to 1, the stronger the model's predictive ability. R

[0037] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

[0038] Appendix:

[0039] Comparison of experimental results of various methods, where TCN, GRU, and LSTM are deep learning models with relatively high usage frequencies in the field of time series prediction. Three technology stocks are selected as the data set for comparison, and through MAE, MSE, MAPE, R2 From the exponents, it can be seen that the performance indicators of Transformer-KAN are the best.

[0040]

Claims

1. A stock price prediction method based on the Transformer-KAN model, characterized in that: The following steps are involved: S1. Perform data preprocessing. The input features are the stock market opening price, closing price, highest price, lowest price, and number of transactions. The output label is the future stock market closing price. First, read the dataset, then convert the dataset to date format and set the index. Then, sort by date, check and fill in missing values to obtain a complete dataset. S2. Normalize the data set. First, divide the data set into training set and test set in a ratio of 8:

2. Then, normalize the training set and test set to their maximum and minimum values. Finally, set the input sequence with a time window size of 20. S3. Build the Transformer-KAN model. First, delete the Decoder part of the Transformer. Then introduce the KAN network to replace the original fully connected layer of the Transformer. A linear layer is introduced before the Encoder to convert the dimension of the input features. Finally, the features of the last time step are output. S4. Set the loss function, optimizer, learning rate, batch size and other hyperparameters for model training. S5. Input the data into the neural network, train the constructed model with the training set, and save the trained model. S6. Load the saved model, make predictions using the test set, and evaluate using various evaluation metrics.

Citation Information

Cited By

  • Logging data layering method and device based on T-KAN structure and medium

    CN121234072A

  • Logging data stratification method, device and medium based on T-KAN structure

    CN121234072B

  • Machine learning-based multi-output regression incremental learning dynamic adaptation stock price prediction model construction method

    CN121413805A