Stock market emotion prediction method based on multi-modal deep learning model

Through the multimodal deep learning model, the time sequence data of stock market is processed using the Transformer encoder and the long-headed self-attention mechanism, and the cross-modal attention mechanism is combined with the stock and news text data, the gradient problem of the recurrent neural network and the waste of computing resources is solved, and the accuracy and efficiency of stock market sentiment prediction are improved.

CN120277218APending Publication Date: 2025-07-08GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510403754.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has the problem of the explosion or disappearance of the gradient of the recurrent neural network in stock market sentiment prediction, and the computational resources are seriously wasted, and the stock timing data and text data characteristics are insufficiently integrated, which affects the prediction accuracy.

Method used

The multimodal deep learning model is adopted to process stock timing data through the Transformer encoder and the multi-head self-attention mechanism, and combine the cross-modal attention mechanism to integrate stock timing data and news text data, and optimize data processing using position coding and data block division.

Benefits of technology

It improves the accuracy and computing efficiency of stock market sentiment predictions, enhances the ability to converge across modal data, captures long-term dependencies and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277218A_ABST
    Figure CN120277218A_ABST
Patent Text Reader

Abstract

The invention discloses a stock market emotion prediction method based on a multi-modal deep learning model. The method comprises the following steps: collecting stock time sequence data and news text data; the collected data is preprocessed; adding a position code to the preprocessed stock time sequence data of each time step; mapping the stock time sequence data added with the position codes, and converting the stock time sequence data into feature data suitable for model input; inputting the stock time sequence data added with the position codes and mapped into a Transform encoder, and extracting time sequence characteristics of the stock time sequence data; extracting news text data features through a multi-head self-attention mechanism and a feedforward neural network; calculating cross attention scores of the stock time sequence data and the news text data, and determining how to weight and fuse information of different modals; and predicting the emotion of the stock market through a network to obtain a prediction result. According to the invention, prediction accuracy and prediction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and stock market sentiment change prediction, and particularly relates to a stock market sentiment prediction method based on a multimodal deep learning model. Background Art

[0002] The fluctuations in the stock market are not only determined by the fluctuations in the stock prices of companies, but also affected by various factors. The stock market is full of complex non-linearities and uncertainties, and the best stock market prediction methods should cover various factors in the market. In addition to traditional stock data indicators, factors such as the macroeconomic environment, market sentiment, and investors' psychology also need to be considered. Especially in the current economic environment, relying solely on stock price indicators for prediction may not comprehensively reflect the changes in the market. Therefore, in addition to the technical indicators of the stock market, it is necessary to effectively combine macroeconomic data and market sentiment to improve the accuracy of stock market trend prediction.

[0003] In stock market sentiment analysis, the emotional fluctuations of investors have an important impact on the trend of the stock market. Different market sentiments, such as greed, fear, and neutrality, usually directly reflect the market's risk preference and expectations for future trends. By analyzing key technical indicators, the changing trends of market sentiment can be effectively captured.

[0004] In existing time series prediction technologies, many methods still rely on traditional hybrid models such as recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) to process stock time series data. However, due to the strong long-term dependence of sequence data in the stock market, recurrent neural networks are prone to problems such as gradient explosion or disappearance, resulting in difficulties in training the model and affecting the prediction effect. In addition, existing methods using attention mechanisms to process data consume a large amount of computing resources when dealing with time series data with a long time period, resulting in low efficiency. Especially in stock market sentiment change prediction, the time dependence is very strong, and data closer to the target date may have a greater impact on the prediction. Existing methods fail to effectively focus on stock time series data related to the target date, causing waste of computing resources.

[0005] In addition, in text input models, current methods usually have problems with insufficient feature fusion between stock time series data and text data, resulting in the failure of the information of the two to complement each other effectively, thereby affecting the prediction accuracy and the expressive ability of the model. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a stock market sentiment prediction method based on a multimodal deep learning model.

[0007] To achieve the above purpose, the technical solution provided by the present invention is as follows:

[0008] A method for predicting stock market sentiment based on a multimodal deep learning model, including:

[0009] Collect stock time series data and news text data;

[0010] Preprocess the collected data;

[0011] Add positional encoding to the preprocessed stock time series data for each time step to preserve the order information of the time series;

[0012] Map the stock time series data with positional encoding to convert it into feature data suitable for model input;

[0013] Input the stock time series data with positional encoding and mapping into a Transformer encoder to extract the time series features of the stock time series data;

[0014] Extract news text data features through a multi-head self-attention mechanism and a feed-forward neural network;

[0015] Calculate the cross-attention scores of the stock time series data and the news text data to determine how to weight and fuse information from different modalities;

[0016] Predict the stock market sentiment through a network to obtain the prediction result.

[0017] Further, collecting stock time series data and news text data includes:

[0018] Obtain historical stock time series data from financial data sources, including trading volume and turnover rate, volatility index VIX, relative strength index RSI, moving average convergence index MACD, GDP growth rate;

[0019] Collect text data related to stocks from news browsing platforms.

[0020] Further, preprocessing the collected data includes:

[0021] Normalize the stock time series data so that the data has a unified scale and avoid difficulties in model training caused by features with different dimensions;

[0022] First divide the stock time series data into independent channel data to avoid interference of information between different channel data;

[0023] Cut the independent channel data into several data blocks, and each data block contains stock price information within a certain time span;

[0024] Cut a window sequence with a total length of L into N non-overlapping subsequences with a length of p;

[0025]

[0026] Clean the news text data to remove irrelevant characters; use the word2vec model to train and obtain the word vectors of each word in the entire text library, resulting in a word vector matrix with a data volume equal to the size of the entire word library.

[0027] Furthermore, map the stock time series data with positional encoding, and the mapping formula is as follows:

[0028] x d = W p x p + W pos

[0029] where x d is the input of the Transformer encoder, is a matrix with D rows and N columns, where each column represents the feature vector of a time step in the sequence; W p is the weight matrix used to map the stock time series data x p with positional encoding into an embedding space; W pos is the weight matrix of positional encoding used to add positional information to the model.

[0030] Furthermore, extract the time series features of the stock time series data by using the attention mechanism:

[0031] The self-attention mechanism adopts the query-key-value mode to improve the model ability; where Q = [q1,..., q N , K = [k1,..., k N , V = [v1,..., v N are the query vector matrix, key vector matrix, and value vector matrix respectively; the calculation processes of these three matrices are as follows:

[0032]

[0033] W Q is the query weight matrix; W K is the key weight matrix; W V is the value weight matrix; is the transpose of x d ;

[0034] H = [h1,..., h N is the output feature matrix, and the output vector h n is obtained by operating on the query vector matrix, key vector matrix, and value vector matrix:

[0035]

[0036] att is the attention function, qn is the nth query vector, k j is the jth key vector, v j is the jth value vector, s(k j , q n ) is a similarity function used to calculate the similarity between the key vector k j and the query vector q n ;

[0037] The output vector h n obtains the stock time series data feature H after passing through the feed-forward neural network FNN time .

[0038] Furthermore, the news text data features are extracted through the multi-head self-attention mechanism and the feed-forward neural network, including:

[0039] For the news text input T = [t1, t2,..., t m , is the embedding representation of each word, and the corresponding query weight matrix Q T , key weight matrix K T , and value weight matrix V T are calculated. The calculation formulas are as follows:

[0040] Q T = TW Q , K T = TW K , V T = TW V

[0041] The multi-head attention is used to parallelize the attention mechanism, and each head has an independent weight matrix:

[0042] Multi-headAttention(Q, K, V) = Concat(head1, head2,..., head h )W O

[0043] where head i = Attention(Q i , K i , V i ) is the output of the ith attention head, and W O is the linear transformation matrix of the output;

[0044] After the news text extracts features through the multi-head attention mechanism, it is passed to the feed-forward neural network FNN to obtain the news text data feature H text .

[0045] Furthermore, calculating the cross-attention scores between stock time-series data and news text data involves the following process:

[0046] In the cross-attention mechanism, the stock time-series data is used as the query, and the news text data is used as the key and value. The query Q of the stock time-series data time and the key K of the news text data text calculate the attention scores:

[0047]

[0048] d k is the dimension of the key vector;

[0049] Normalize the attention scores through the Softmax operation:

[0050]

[0051] Weight the values V of the news text data using the cross-attention weights text and perform a weighted sum:

[0052] Cross - AttentionOutput = Softmax(Cross - Attention(Q time , K text ))V text .

[0053] Furthermore, predicting the stock market sentiment through a network includes:

[0054] The stock time-series data feature H time and the news text data feature H text are concatenated:

[0055] H concat = concat(H time , H text )

[0056] The concatenated H concat contains the information of the stock time-series data and the news text data, and

[0057] The concatenated feature H concat is input into the fully connected layer FC, and the role of the fully connected layer FC is to map the high-dimensional concatenated features to the space of the target categories; the operation is as follows:

[0058]

[0059] Among them, is the learnable weight matrix, is the bias term, The output result of the network;

[0060] It corresponds to the result of the stock market sentiment classification task; the probability value of each type of sentiment is output through the Softmax layer, and the predicted value is the category corresponding to the maximum probability.

[0061] Compared with the prior art, the principles and advantages of this technical solution are as follows:

[0062] 1. More effective time-dependence processing:

[0063] Existing stock price prediction models usually adopt recurrent neural networks such as RNN or LSTM. These models are prone to problems such as gradient explosion or disappearance when dealing with long-term dependencies, resulting in difficult training and affecting prediction accuracy. However, this technical solution uses a Transformer encoder and the multi-head attention mechanism to process stock time series data, which can better capture long-term dependencies in the time series, avoid the defects of traditional recurrent neural networks, and thus improve the prediction accuracy and model stability.

[0064] 2. Reduce waste of computing resources:

[0065] When existing methods process data with a long time period, they often need to calculate the correlations between all time steps, resulting in waste of computing resources. This technical solution divides the time series data into appropriately sized data blocks (Patches), which can effectively reduce the number of tokens that need to be calculated in the attention mechanism, thereby reducing the computational amount and resource consumption and improving computational efficiency.

[0066] 3. Strong cross-modal data fusion ability:

[0067] When traditional prediction methods fuse time series data and news text data for stocks, they often cannot fully explore the potential correlations between the two, resulting in poor fusion effects. However, this technical solution introduces a cross-modal attention mechanism, enabling the stock time series data features to effectively query key information in the news text data. This mechanism can strengthen the correlation between the stock time series data and the news text data, thereby enhancing the model's comprehensive understanding ability of the two types of data and improving the accuracy of stock market sentiment prediction. Description of the Drawings

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the services required for the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0069] Figure 1Flowchart of the principle of the stock market sentiment prediction method based on the multimodal deep learning model of the present invention (in the figure, text represents news text data; Multihead-Attention represents multi-head attention; FFN represents feed-forward neural network; Cross-Attention represents cross-attention; concat represents concatenation; Transformer Encoder represents encoder; Position Embedding Projection represents position embedding and mapping; univariate input data represents univariate input data; patch represents data block; FC represents fully connected layer). Detailed implementation manners

[0070] The present invention will be further described below in conjunction with specific embodiments:

[0071] The stock market sentiment prediction method based on the multimodal deep learning model described in this embodiment includes the following steps:

[0072] S1. Collect stock time series data and news text data, including:

[0073] Obtain historical stock time series data from financial data sources (such as Yahoo Finance, Google Finance, OECD, etc.), including trading volume and turnover rate, volatility index VIX, relative strength index RSI, moving average convergence index MACD, GDP growth rate;

[0074] Collect text data related to stocks from news browsing platforms (such as Reuters, Eastmoney, etc.).

[0075] S2. Preprocess the collected data, including:

[0076] Standardize the stock time series data so that the data has a unified scale and avoid difficulties in model training caused by features of different dimensions;

[0077] First divide the stock time series data into independent channel data to avoid interference of information between different channel data;

[0078] Cut the independent channel data into several data blocks (Patch), and each data block contains stock price information within a certain time span;

[0079] Cut the window sequence with a total length of L into N non-overlapping subsequences with a length of p;

[0080]

[0081] Clean the news text data to remove irrelevant characters; use the word2vec model to train and obtain the word vectors of each word in the entire text library, resulting in a word vector matrix with a data volume equal to the size of the entire word library.

[0082] S3. Add positional encoding to the preprocessed stock time series data for each time step to preserve the order information of the time series; the positional encoding helps the model understand the time order of the data by assigning a unique encoding to each time point.

[0083] S4. Map the stock time series data with positional encoding to convert it into feature data suitable for model input; the mapping formula is as follows:

[0084] x d =W p x p +W pos

[0085] where x d is the input to the Transformer encoder, is a matrix with D rows and N columns, where each column represents the feature vector of a time step in the sequence; W p is the weight matrix used to map the stock time series data x p to an embedding space; W pos is the weight matrix of the positional encoding used to add the positional information to the model.

[0086] S5. Input the stock time series data with positional encoding and mapping into the Transformer encoder to extract the time series features of the stock time series data. The process includes:

[0087] The self-attention mechanism adopts the query-key-value mode to improve the model's ability; where Q = [q1,..., q N , K = [k1,..., k N , V = [v1,..., v N are the query vector matrix, key vector matrix, and value vector matrix respectively; the calculation processes of these three matrices are as follows:

[0088]

[0089] W Q is the query weight matrix; W K is the key weight matrix; W V is the value weight matrix; is the transpose of x d ;

[0090] H = [h1,..., h Nis the output feature matrix, and the output vector is h n It is obtained through operations using the query vector matrix, key vector matrix, and value vector matrix:

[0091]

[0092] att is the attention function, and q n is the nth query vector, k j is the jth key vector, v j is the jth value vector, and s(k j , q n ) is the similarity function used to calculate the similarity between the key vector k j and the query vector q n ;

[0093] The output vector h n obtains the stock time series data feature H after passing through the fully connected layer FC time .

[0094] S6. Extract the news text data features through the multi-head self-attention mechanism and the feed-forward neural network. The process includes:

[0095] For the news text input T = [t1, t2,..., t m , is the embedding representation of each word, and calculate the corresponding query weight matrix Q T , key weight matrix K T , and value weight matrix V T . The calculation formulas are as follows:

[0096] Q T = TW Q , K T = TW K , V T = TW V

[0097] Use multi-head attention to parallelize the attention mechanism, and each head has an independent weight matrix:

[0098] Multi-headAttention(Q, K, V) = Concat(head1, head2,..., head h )W O

[0099] where head i = Attention(Q i , K i , V i ) is the output of the ith attention head, and W O is the output linear transformation matrix;

[0100] After the news text extracts features through the multi-head attention mechanism, it is passed to the feed-forward neural network FNN to obtain the news text data feature H text 。

[0101] S7. Calculate the cross-attention scores of the stock time-series data and the news text data to determine how to weight and fuse information of different modalities. The process includes:

[0102] In the cross-attention mechanism, the stock time-series data is used as the query, the news text data is used as the key and value, and the query Q of the stock time-series data time and the key K of the news text data text Calculate the attention scores:

[0103]

[0104] d k is the dimension of the key vector;

[0105] Normalize the attention scores through the Softmax operation:

[0106]

[0107] Use the cross-attention weights to perform weighted summation on the value V of the news text data text :

[0108] Cross - AttentionOutput = Softmax(Cross - AttentionQ time ,K text )V text 。

[0109] S8. Predict the stock market sentiment through the network to obtain the prediction result.

[0110] The process of this step includes:

[0111] Concatenate the stock time-series data feature H time and the news text data feature H text :

[0112] H concat = concat(H time ,H text )

[0113] The concatenated H concat contains the information of the stock time-series data and the news text data, and

[0114] Pass the concatenated feature Hconcat It is input into the fully connected layer FC, and the function of the fully connected layer FC is to map the high-dimensional concatenated features to the space of the target category; the operation is as follows:

[0115]

[0116] Among them, is a learnable weight matrix, is a bias term, is the output result of the network;

[0117] corresponds to the result of the stock market sentiment classification task; the probability value of each type of sentiment is output through the Softmax layer, and the predicted value is the category corresponding to the maximum probability.

[0118] In this embodiment, by using the Transformer encoder and utilizing the multi-head attention mechanism to process the stock time series data, it can better capture the long-term dependencies in the time series, avoid the defects of traditional recurrent neural networks, and thus improve the prediction accuracy and the stability of the model. By dividing the time series data into data blocks (Patches) of appropriate sizes, the number of tokens that need to be calculated in the attention mechanism can be effectively reduced, thereby reducing the computational amount and resource consumption and improving the computational efficiency. By introducing the cross-modal attention mechanism, the stock time series data features can effectively query the key information in the news text data. This mechanism can strengthen the correlation between the stock time series data and the news text data, thereby enhancing the model's comprehensive understanding ability of the two types of data and improving the accuracy of stock market sentiment prediction.

[0119] The above-described embodiments are only the preferred embodiments of the present invention and do not limit the scope of implementation of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for predicting stock market sentiment based on a multi-modal deep learning model, characterized in that, Including: Collecting stock time-series data and news text data; Preprocessing the collected data; Adding positional encoding to the preprocessed stock time-series data for each time step to preserve the order information of the time series; Mapping the stock time-series data with positional encoding to convert it into feature data suitable for model input; Inputting the stock time-series data with positional encoding and mapping into a Transformer encoder to extract the time-series features of the stock time-series data; Extracting news text data features through the multi-head self-attention mechanism and the feed-forward neural network; Calculating the cross-attention scores between the stock time-series data and the news text data to determine how to weight and fuse information of different modalities; Predicting the stock market sentiment through a network to obtain the prediction result.

2. The stock market sentiment prediction method based on a multi-modal deep learning model according to claim 1, characterized in that Collecting stock time-series data and news text data, including: Obtaining historical stock time-series data from financial data sources, including trading volume and turnover rate, volatility index VIX, relative strength index PSI, moving average convergence divergence index MACD, GDP growth rate; Collecting text data related to stocks from news browsing platforms.

3. The method for predicting stock market sentiment based on a multi-modal deep learning model according to claim 1, wherein Preprocessing the collected data, including: Normalizing the stock time-series data to make the data have a unified scale and avoid difficulties in model training caused by features with different dimensions; Dividing the stock time-series data into independent channel data first to avoid interference of information between different channel data; Cutting the independent channel data into several data blocks, and each data block contains stock price information within a certain time span; Cutting a window sequence with a total length of L into N non-overlapping subsequences with a length of p; Cleaning the news text data to remove irrelevant characters; training with the word2vec model to obtain the word vectors of each word in the entire text library, and getting a word vector matrix with the data volume of the entire word library size.

4. The method for predicting stock market sentiment based on a multi-modal deep learning model according to claim 1, wherein Mapping the stock time-series data with positional encoding, and the mapping formula is as follows: x d = W p x p + W pos where, x d is the input of the Transformer encoder, is a matrix of D rows and N columns, where each column represents the feature vector of a time step in the sequence; W p is the weight matrix used to map the stock time series data x p added with positional encoding into an embedding space; W pos is the weight matrix of positional encoding used to add positional information into the model.

5. The method for predicting stock market sentiment based on a multi-modal deep learning model according to claim 4, wherein Extracting the time-series features of the stock time-series data by using the attention mechanism: The self-attention mechanism adopts the query-key-value mode to improve the model's ability; among them, Q = [q1,...,q N , K = [k1,...,k N , V = [v1,...,v N are the query vector matrix, the key vector matrix, and the value vector matrix respectively; the calculation processes of these three matrices are as follows: W Q is the query weight matrix; W K is the key weight matrix; W V is the value weight matrix; is the transpose of x d ; H = [h1,..., h N is the output feature matrix, and the output vector h n is obtained by operations using the query vector matrix, the key vector matrix, and the value vector matrix: The attention function is att, and q n is the nth query vector, k j is the jth key vector, v j is the jth value vector, and s(k j , q n ) is the similarity function used to calculate the similarity between the key vector k j and the query vector q n ; Output vector h n After passing through the fully connected layer FC, the stock time series data feature H is obtained time .

6. The method for predicting stock market sentiment based on a multi-modal deep learning model according to claim 5, characterized in that Extracting news text data features through the multi-head self-attention mechanism and the feed-forward neural network, including: For the news text input T = [t1, t2,..., t m , For the embedding representation of each word, calculate the corresponding query weight matrix Q T , key weight matrix K T , and value weight matrix V T , and the calculation formula is as follows: Q T = TW Q , K T = TW K , V T = TW V Using multi-head attention to parallelize the attention mechanism, and each head has an independent weight matrix: Multi-head Attention(Q,K,V)=Concat(head1,head2,...,head h )W O where head i = Attention(Q i , K i , V i ) is the output of the i-th attention head, and W O is the linear transformation matrix of the output; After the news text extracts features through the multi-head attention mechanism, it is passed to the feed-forward neural network FNN to obtain the news text data feature H text 。 7. The method for predicting stock market sentiment based on a multi-modal deep learning model according to claim 6, wherein Calculating the cross-attention scores between the stock time-series data and the news text data, and the process includes: In the cross-attention mechanism, stock time-series data is used as the query, news text data is used as the key and value, and the query Q of the stock time-series data time is calculated with the key K of the news text data text to calculate the attention score: d k is the dimension of the key vector; Normalizing the attention scores through the Softmax operation: Weight the values V of the news text data using cross-attention weights text to perform weighted summation: Cross-Attention Output=Softmax(Cross-Attention(Q time ,K text ))V text 。 8. The method for predicting stock market sentiment based on a multi-modal deep learning model according to claim 7, characterized in that, Predicting the stock market sentiment through a network, including: Concatenate the stock time series data feature H time and the news text data feature H text as follows: H concat = concat(H time , H text ) The spliced H concat contains information on stock time series data and news text data, and Input the concatenated feature H concat into the fully connected layer FC. The role of the fully connected layer FC is to map the high-dimensional concatenated features to the space of the target category. The operation is as follows: Among them, is a learnable weight matrix, is a bias term, is the output result of the network; It is the result corresponding to the stock market sentiment classification task; the probability value of each type of sentiment is output through the Softmax layer, and the predicted value is the category corresponding to the maximum probability.