Q network portfolio method combining long-term and short-term memory and attention mechanism

Through the Q network method combining long and short-term memory and attention mechanism, the prediction problems of nonlinearity and long-term dependence of asset prices in financial investment are solved, and more accurate stock trend prediction and investment strategy are achieved.

CN120013673APending Publication Date: 2025-05-16NAT SUPERCOMPUTING SHENZHEN CENT (SHENZHEN CLOUD COMPUTING CENT)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510079109.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively capture the nonlinear relationship and long-term dependence of asset prices in the field of financial investment, and fails to fully utilize the correlation between individual stocks and market information, resulting in stock trend prediction errors.

Method used

Using a Q network portfolio method combining long and short-term memory and attention mechanism, the Markov chain model and convolutional neural network are constructed, individual stocks, industry and market information are fused into multi-dimensional input vectors, and the time series data is processed using long and short-term memory network and attention network to output Q values ​​to determine the best investment strategy.

Benefits of technology

It realizes a more comprehensive and accurate stock status description, which can better process timing output and information in different time dimensions, and improves the accuracy of stock trend prediction and the effectiveness of investment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013673A_ABST
    Figure CN120013673A_ABST
Patent Text Reader

Abstract

The invention provides a Q network investment portfolio method combining long-term and short-term memory and an attention mechanism, and aims to realize an optimal decision of stock investment by fusing individual stock, industry and market information and utilizing deep Q learning and a long-term and short-term memory attention model. The method comprises the following specific steps: constructing a Markov chain model based on individual stock, industry and market signals, and defining an agent, an environment, a state, a behavior and an award; designing a convolutional neural network to encode the signal features into the input of a long-short-term memory network; constructing a long short-term memory network to process the time sequence data and outputting a hidden state; designing an attention network to carry out weighted summation on the hidden state and outputting a Q value; and optimizing model parameters by adopting a deep Q network training method. According to the method, the stock trend can be predicted more accurately through fusion of multi-dimensional information and an advanced neural network structure. Therefore, a more effective investment strategy can be provided for investors, the market risk is reduced, and the investment income is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of deep neural networks in the field of financial investment, and in particular to a Q network investment portfolio method combining long-term and short-term memory with an attention mechanism. Background Art

[0002] The problem of asset price trend prediction in the financial market has always been a major problem that troubles investors and related researchers. Reasonable prediction can help investors allocate assets reasonably and obtain higher returns. However, asset prices are affected by multiple factors such as company operations, investor mentality and macroeconomic policies, showing characteristics such as volatility, non-stationarity, cyclicality, nonlinearity and long-term dependence.

[0003] Traditional investment solutions are mainly based on time series related methods, such as ARIMA (Autoregressive Moving Average Model) and GARCH (Generalized Conditional Regression Autovariance Model). They can capture the volatility and periodicity of financial time series, but it is difficult to analyze non-stationary series and capture the nonlinear relationship of financial time series. Deep learning-based methods such as RNN (Recurrent Neural Network) can solve this problem well.

[0004] In order to simulate the trading environment to achieve the maximum benefit of the current model, we can learn from the methods of reinforcement learning to build and optimize this process. The goal of reinforcement learning is to achieve the maximum benefit in the environment, which is similar to the goal of trading in the stock market. The combination of reinforcement learning with other algorithms in the stock market has been widely studied.

[0005] Stock price data has the problem of large quantity and large dimensions. Q network learning is more efficient. For example, the combination of Q learning and RNN has achieved good results in traffic. However, RNN has some difficult problems to solve, such as gradient disappearance. This problem can be solved by introducing long short-term memory mechanism, which can replace RNN to process the input of original features and better preserve long-term information. The attention mechanism can give different prediction weights to data in different time dimensions, making the prediction results more accurate.

[0006] A portfolio generation method based on deep reinforcement learning. This method considers market conditions as an independent profit-risk balance module, and at the same time enhances the extraction of cross-asset relationships by learning and using graph structures to characterize the relationships between stocks. Finally, it is optimized through a deep reinforcement learning algorithm to determine whether to go long, short, or not participate in investment operations on a certain stock. This method does not use the Q network for action value function estimation.

[0007] Existing technologies input individual stock information and market information into the network separately, and do not perform long-term and short-term memory and attention encoding processing on individual stock information, which may cause the loss of information related to individual stocks and the market, and the information of a long time may be lost in the encoding network of individual stocks. The lack of attention mechanism will also make it difficult to use information of different time dimensions with reasonable weights. At the same time, industry information is not taken into account in stock prediction, which may cause errors in stock trend prediction.

[0008] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention

[0009] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a Q network investment portfolio method combining long short-term memory with attention mechanism.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] A Q network portfolio method combining long short-term memory and attention mechanism includes the following steps:

[0012] S1: Build a Markov chain model based on individual stocks, industries and market signals, define the agent, environment, state, behavior and reward, integrate individual stocks, industry and market information into a multi-dimensional input vector, the state includes multi-day characteristics, the behavior includes short selling, long selling and no investment, and the reward is calculated based on the stock price increase or decrease and operation;

[0013] S2: Construct a convolutional neural network, input the daily features in chronological order, and encode the signal features as the input of the long short-term memory network through the convolution layer and the fully connected layer;

[0014] S3: Build a long short-term memory network, convert the output of the convolutional neural network into a hidden state, use the forget gate, input gate, unit state and output gate mechanism to update the memory, and output the hidden state as the input of the attention network;

[0015] S4: Construct an attention network, use cosine similarity and normalized exponential function to calculate the attention weight, perform weighted summation on the hidden state, and output the Q value through a neural network linear transformation;

[0016] S5: Use the deep Q network training method to initialize the experience playback and Q network parameters, initialize the state in each round, select actions, execute actions, store experience, sample experience, calculate the target Q value and optimize the Q network parameters in each time step, and update the target Q network parameters regularly.

[0017] Further:

[0018] Step S1 specifically includes: constructing a Markov chain model of market, industry and individual stock signals, defining the agent, environment, state, behavior and reward; the state is composed of a multi-dimensional feature vector that integrates multiple days of individual stock, industry and market information, and the behavior includes short selling, long selling and no investment operations on each stock; the reward is calculated based on the stock's rise and fall and operation behavior, and the return is the equally weighted average of the returns of all investment stocks.

[0019] Step S2 specifically includes: inputting the daily multi-dimensional feature vector into the convolutional neural network in chronological order, encoding the features through the convolutional layer, and further converting the convolutional layer output into a one-dimensional vector through a fully connected layer as the input of the long short-term memory network.

[0020] Step S3 specifically includes: inputting the output of the convolutional neural network into the long short-term memory network, updating the memory state through the forget gate, input gate, unit state and output gate mechanism, and outputting the hidden state of the current time step; the hidden state is used as the input of the attention network and passed to the next time step for calculation.

[0021] Step S4 specifically includes: using cosine similarity and normalized exponential function to calculate the attention weights of hidden states at different time steps, performing weighted summation on the hidden states output by the long short-term memory network to obtain weighted hidden states; linearly transforming the weighted hidden states through a neural network to output the final Q value.

[0022] Step S5 specifically includes: using a deep Q network training method to initialize the experience replay pool and Q network parameters; initializing the state in each round, selecting and executing an action according to the ε-greedy strategy at each time step, obtaining a new state and reward after executing the action, and storing the experience in the replay pool; randomly sampling experience from the replay pool, calculating the target Q value and optimizing the Q network parameters; and regularly updating the target Q network parameters to stabilize the training process.

[0023] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the Q network investment portfolio method combining long short-term memory with an attention mechanism.

[0024] A computer program product includes a computer program, which, when executed by a processor, implements the Q network portfolio method combining long short-term memory with an attention mechanism.

[0025] A computing device for a Q network portfolio, comprising

[0026] Memory for storing computer programs;

[0027] A processor is used to implement the method when executing the computer program.

[0028] A Q network portfolio device combining long short-term memory and attention mechanism, including:

[0029] Markov chain model building module, used to build Markov chain models based on individual stocks, industries and market signals, define agents, environments, states, behaviors and rewards, and integrate individual stocks, industries and market information into multi-dimensional input vectors. The states include multi-day features, and the behaviors include short selling, long selling and non-investment. The rewards are calculated based on the stock price fluctuation and operation.

[0030] Convolutional neural network module, which is used to input daily features in time order and encode signal features as input to the long short-term memory network through convolutional layers and fully connected layers;

[0031] The long short-term memory network module is used to convert the output of the convolutional neural network into a hidden state, update the memory using the forget gate, input gate, unit state and output gate mechanism, and output the hidden state as the input of the attention network;

[0032] The attention network module is used to calculate the attention weight using cosine similarity and normalized exponential function, perform weighted summation on the hidden state, and output the Q value through the linear transformation of the neural network;

[0033] The deep Q network training module is used to initialize experience playback and Q network parameters, initialize the state in each round, select actions, execute actions, store experience, sample experience, calculate the target Q value and optimize the Q network parameters in each time step, and regularly update the target Q network parameters.

[0034] The present invention has the following beneficial effects:

[0035] The present invention proposes a stock investment method based on deep Q learning and long short-term memory attention model. The method innovatively integrates individual stocks, industries and market information, and uses Q network to establish action value function for stock market operation problems, and finally predicts stock trends to determine the best investment plan. In particular, this scheme combines Q network and long short-term memory network attention mechanism for the first time to achieve efficient matching of known signals. The main advantages of the present invention are:

[0036] 1. The input vector of the present invention incorporates multi-dimensional information such as individual stocks, markets, and industries. Compared with existing methods, it can more comprehensively describe the stock status and thus obtain a more effective comprehensive investment strategy.

[0037] 2. The Q network stock investment strategy algorithm based on long short-term memory and attention mechanism proposed in the present invention integrates two network structures that are conducive to processing time series input into the traditional Q network. Compared with existing methods, it can better process time series output and give different weights to data of different time series.

[0038] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 The figure is a flow chart of a Q network investment portfolio method combining long short-term memory with an attention mechanism according to an embodiment of the present invention.

[0040] Figure 2 This is an example diagram of interest rate data of a certain interbank lending market according to an embodiment of the present invention.

[0041] Figure 3 This is a convolutional neural network design diagram of an embodiment of the present invention.

[0042] Figure 4 This is a long short-term memory network design diagram of an embodiment of the present invention.

[0043] Figure 5 This is a design diagram of the attention mechanism and output layer of an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.

[0045] The present invention proposes a Q network investment portfolio method that combines long-term and short-term memory with an attention mechanism. Based on the Q network, the long-term and short-term memory network and the attention model, a set of investment operation control frameworks for stocks in specific industries are designed, including: (i) constructing a Markov chain model based on individual stocks, industries and market signals, and defining elements such as agents, environments, states, behaviors, and rewards. (ii) designing a convolutional neural network to convert signal features into long-term and short-term memory network inputs. (iii) designing a long-term and short-term memory network to convert the output of an artificial neural network into a hidden state. (iv) designing an attention network to finally output a Q value.

[0046] See also Figure 1 The embodiment of the present invention provides a Q network investment portfolio method combining long short-term memory and attention mechanism, comprising the following steps:

[0047] Step S1: Construct a Markov chain model based on individual stocks, industries and market signals, define the agent, environment, state, behavior and reward, integrate individual stocks, industry and market information into a multi-dimensional input vector, the state includes multi-day characteristics, the behavior includes short selling, long selling and no investment, and the reward is calculated based on the stock price increase or decrease and operation.

[0048] In a preferred embodiment, step S1 specifically includes: constructing a Markov chain model based on individual stocks, industries and market signals, defining the agent, environment, state, behavior and reward; the state is composed of a multi-dimensional feature vector that integrates multiple days of individual stocks, industries and market information, and the behavior includes short selling, long selling and no investment operations on each stock; the reward is calculated based on the stock's price fluctuation and operation behavior, and the return is the equally weighted average of the returns of all investment stocks.

[0049] Step S2: Construct a convolutional neural network, input the daily features in chronological order, and encode the signal features as the input of the long short-term memory network through the convolution layer and the fully connected layer.

[0050] In a preferred embodiment, step S2 specifically includes: inputting the daily multi-dimensional feature vector into the convolutional neural network in chronological order, encoding the features through the convolutional layer, and further converting the convolutional layer output into a one-dimensional vector through a fully connected layer as the input of the long short-term memory network.

[0051] Step S3: Build a long short-term memory network, convert the output of the convolutional neural network into a hidden state, use the forget gate, input gate, unit state and output gate mechanism to update the memory, and output the hidden state as the input of the attention network;

[0052] In a preferred embodiment, step S3 specifically includes: inputting the output of the convolutional neural network into the long short-term memory network, updating the memory state through the forget gate, input gate, unit state and output gate mechanism, and outputting the hidden state of the current time step; the hidden state is used as the input of the attention network and passed to the next time step for calculation.

[0053] Step S4: construct an attention network, use cosine similarity and normalized exponential function to calculate the attention weight, perform weighted summation on the hidden state, and output the Q value through a neural network linear transformation;

[0054] In a preferred embodiment, step S4 specifically includes: using cosine similarity and normalized exponential function to calculate the attention weights of hidden states at different time steps, performing weighted summation on the hidden states output by the long short-term memory network to obtain weighted hidden states; linearly transforming the weighted hidden states through a neural network to output the final Q value.

[0055] Step S5: Use the deep Q network training method to initialize the experience playback and Q network parameters, initialize the state in each round, select actions, execute actions, store experience, sample experience, calculate the target Q value and optimize the Q network parameters in each time step, and regularly update the target Q network parameters.

[0056] In a preferred embodiment, step S5 specifically includes: using a deep Q network training method to initialize the experience replay pool and Q network parameters; initializing the state in each round, selecting and executing an action according to the ε-greedy strategy at each time step, obtaining a new state and reward after executing the action, and storing the experience in the replay pool; randomly sampling experience from the replay pool, calculating the target Q value and optimizing the Q network parameters; and regularly updating the target Q network parameters to stabilize the training process.

[0057] In some embodiments, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements the Q network portfolio method combining long short-term memory with an attention mechanism.

[0058] In some embodiments, a computer program product includes a computer program, which, when executed by a processor, implements the Q network portfolio method combining long short-term memory with an attention mechanism.

[0059] In some embodiments, an embodiment of the present invention further provides a computing device for a Q network investment portfolio, comprising: a memory for storing a computer program; and a processor for implementing the method described when executing the computer program.

[0060] In some embodiments, the embodiments of the present invention further provide a Q network investment portfolio device combining long short-term memory with an attention mechanism, including: a Markov chain model construction module, which is used to construct a Markov chain model based on individual stocks, industries and market signals, define intelligent agents, environments, states, behaviors and rewards, and fuse individual stocks, industries and market information into multi-dimensional input vectors. The states include multi-day features, the behaviors include short selling, long selling and no investment, and the rewards are calculated based on stock price fluctuations and operations; a convolutional neural network module, which is used to input daily features in chronological order, and encode signal features as inputs to a long short-term memory network through convolutional layers and fully connected layers; a long The short-term memory network module is used to convert the output of the convolutional neural network into a hidden state, update the memory using the forget gate, input gate, unit state and output gate mechanism, and output the hidden state as the input of the attention network; the attention network module is used to calculate the attention weight using cosine similarity and normalized exponential function, perform weighted summation on the hidden state, and output the Q value through a neural network linear transformation; the deep Q network training module is used to initialize experience playback and Q network parameters, initialize the state in each round, select actions, execute actions, store experience, sample experience, calculate the target Q value and optimize the Q network parameters at each time step, and regularly update the target Q network parameters.

[0061] The specific embodiments of the present invention and algorithm examples thereof are further described below.

[0062] like Figure 1As shown in Figure 1, the Q-network portfolio method combining long short-term memory with attention mechanism includes the following steps:

[0063] (1) Constructing a Markov chain model based on individual stocks, industries and market signals

[0064] The main stock features used are the opening and closing prices of a certain stock on that day:

[0065] Table 1 Closing price and opening price data examples

[0066]

[0067]

[0068] is the opening price of stock s on day t, is the closing price.

[0069] The industry characteristics mainly used in the embodiment of the present invention are the opening price and closing price of the sector index on the day, and the data source is the industry index data of Tonghuashun:

[0070]

[0071] is the opening price of industry i on day t, is the closing price.

[0072] The macroeconomic characteristics mainly used in the embodiment of the present invention are the deposit interest rate of the day. The interbank lending market interest rate of a certain bank is used as an indicator.

[0073] For the input of the embodiment of the present invention, each line is a comprehensive feature of a certain stock, including the opening price, closing price, industry index opening price and closing price of the stock, and the central bank interest rate. The embodiment of the present invention defines the feature of a certain day as

[0074]

[0075] in The opening and closing prices of the stocks with the top m market value in the industry studied in the embodiment of the present invention. Since the embodiment of the present invention needs to study the impact of multi-day data on stock trends, the embodiment of the present invention defines the status as

[0076] s t =[x t ,…,x t-n-1 ]

[0077] The state of day t includes the characteristics of the current day to the previous n-1 days.

[0078] Figure 2An example of a bank's interbank lending market interest rate data is shown.

[0079] For any stock, the embodiment of the present invention defines three operations on a certain day: short selling, long selling, and no investment. Short selling means that the embodiment of the present invention borrows the stock at the opening price at the beginning of the trading day, sells it immediately, and buys an equal amount of stock from the issuer at the closing price at the end of the trading day; long selling means that the embodiment of the present invention buys the stock at the opening price at the beginning of the trading day, and sells the stock at the closing price at the end of the trading day; no investment means that the embodiment of the present invention does not trade the stock on the trading day, so the embodiment of the present invention defines the actions for a certain stock on a certain trading day as Its values ​​are 0, -1, and 1, representing no investment, short selling, and long selling, respectively. The actions on a certain day are the set of operations for all m stocks.

[0080]

[0081] The benefits of the present invention are defined as follows:

[0082]

[0083] in

[0084]

[0085]

[0086] For a stock, the rise and fall on day t+1 can be expressed as

[0087] The day's earnings can be used to At the same time, the embodiment of the present invention will establish an investment portfolio with equal weights for each stock, and the final rate of return obtained is the equal-weighted average rate of return of all stocks invested.

[0088] (2) Design a convolutional neural network to convert signal features into long short-term memory network input

[0089] The daily features x t-n-1 …x t The data are transmitted to the network in the order of tn-1 to t in time sequence. For each time step, x is two-dimensional data, and the embodiment of the present invention further encodes it using a convolutional neural network.

[0090] Figure 3The convolutional neural network design is shown. The convolutional neural network includes a convolutional layer and an output layer. The embodiment of the present invention designs c l×l convolution kernels for the first layer, with a stride of b and a padding of q, to ​​obtain the convolutional layer of the embodiment of the present invention. The dimension of the convolutional layer is After obtaining the convolutional layer, the embodiment of the present invention expands the convolutional layer to obtain a After the fully connected layer is transformed by the one-dimensional artificial neural network again, an output layer of [O, 1] will be obtained. This layer will be used as the input of the subsequent long short-term memory network in the embodiment of the present invention. The embodiment of the present invention names it as y t .

[0091] (3) Design a long short-term memory network to convert the output of the artificial neural network into a hidden state.

[0092] like Figure 4 , the embodiment of the present invention designs a long short-term memory network to convert the output of the embodiment of the present invention into a hidden state, which is then used as the input of the attention model. t is the output of the convolutional neural network at this time step, h t-1 is the output of the LSTM network at the previous time step, C t-1 is the memory of the output of the previous time step. σ is the sigmoid function and tanh is the hyperbolic tangent function.

[0093] f t is the output of the forget gate, and the formula is:

[0094] f t =σ(W f ·[h t-1 ,y t ]+b f )

[0095] i t is the output of the input gate,

[0096] i t =σ(W i ·[h t-1 ,y t ]+b i )

[0097] The unit status of the current input:

[0098]

[0099] C t The updated unit status is:

[0100]

[0101] o t is the output of the output gate:

[0102] o t =σ(W o ·[h t-1 ,y t ]+b o )

[0103] h t is the output of the long short-term memory network at this time step:

[0104] h t =o t *tanh(C t )

[0105] h t , C t It will be passed as input into the calculation of the next time step until the last time step t of the embodiment of the present invention.

[0106] (4) Design the attention network to output the final Q value Figure 5 The attention mechanism and output layer design of an embodiment of the present invention are shown.

[0107] Get h t Afterwards, the embodiment of the present invention uses the attention mechanism to recalculate it. The embodiment of the present invention first uses the cosine similarity and the normalized exponential function to calculate the attention network weight w t

[0108]

[0109] Finally, h' required by the embodiment of the present invention is obtained t

[0110]

[0111] In the embodiment of the present invention, h' t After a layer of neural network linear transformation, the final Q value is output

[0112] Q=W Q h' t +b Q

[0113] (5) Network training method

[0114] The algorithm is trained using the deep Q network training method

[0115] Initialize experience playback D and store each converted experience (s t ,a t ,rt ,s t+1 )

[0116] Initialize Q network (convolution, long short-term memory network), parameter w

[0117] Initialize the target Q network, parameters

[0118] For each round m=1…M:

[0119] Initialization state s1

[0120] For each time step t=1…T:

[0121] For the current state t , select action a with the ε-greedy strategy t

[0122] Execute action a t , get the state s t+1 And reward r t

[0123] Storage(s t ,a t ,r t ,s t+1 ) in D

[0124] Randomly sample N samples from D (s j ,a j ,r j ,s j+1 )

[0125] For each sample, if it ends at step j:

[0126] y j =r j

[0127] If not:

[0128]

[0129] Perform stochastic gradient descent optimization on the following formula

[0130]

[0131] Each C step To update,

[0132] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0133] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0134] An embodiment of the present invention further provides a processor, wherein the processor executes a computer program and at least executes the method described above.

[0135] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0136] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0137] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0138] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0139] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.

[0140] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0141] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0142] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0143] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0144] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A Q network portfolio method combining long short-term memory and attention mechanism, characterized in that: The following steps are involved: S1: Build a Markov chain model based on individual stocks, industries and market signals, define the agent, environment, state, behavior and reward, integrate individual stocks, industry and market information into a multi-dimensional input vector, the state includes multi-day characteristics, the behavior includes short selling, long selling and no investment, and the reward is calculated based on the stock price increase or decrease and operation; S2: Construct a convolutional neural network, input the daily features in chronological order, and encode the signal features as the input of the long short-term memory network through the convolution layer and the fully connected layer; S3: Build a long short-term memory network, convert the output of the convolutional neural network into a hidden state, use the forget gate, input gate, unit state and output gate mechanism to update the memory, and output the hidden state as the input of the attention network; S4: Construct an attention network, use cosine similarity and normalized exponential function to calculate the attention weight, perform weighted summation on the hidden state, and output the Q value through a neural network linear transformation; S5: Use the deep Q network training method to initialize the experience playback and Q network parameters, initialize the state in each round, select actions, execute actions, store experience, sample experience, calculate the target Q value and optimize the Q network parameters in each time step, and update the target Q network parameters regularly.

2. The Q network portfolio method combining long short-term memory and attention mechanism as claimed in claim 1, characterized in that: Step S1 specifically includes: constructing a Markov chain model based on individual stocks, industries and market signals, defining the agent, environment, state, behavior and reward; the state is composed of a multi-dimensional feature vector that integrates multiple days of individual stocks, industries and market information, and the behavior includes short selling, long selling and no investment operations on each stock; the reward is calculated based on the stock's rise and fall and operation behavior, and the return is the equally weighted average of the returns of all investment stocks.

3. The Q network portfolio method combining long short-term memory and attention mechanism as described in claim 1 or 2, characterized in that: Step S2 specifically includes: inputting the daily multi-dimensional feature vector into the convolutional neural network in chronological order, encoding the features through the convolutional layer, and further converting the convolutional layer output into a one-dimensional vector through a fully connected layer as the input of the long short-term memory network.

4. The Q network portfolio method combining long short-term memory and attention mechanism as described in any one of claims 1 to 3, characterized in that: Step S3 specifically includes: inputting the output of the convolutional neural network into the long short-term memory network, updating the memory state through the forget gate, input gate, unit state and output gate mechanism, and outputting the hidden state of the current time step; the hidden state is used as the input of the attention network and passed to the next time step for calculation.

5. The Q network portfolio method combining long short-term memory and attention mechanism as described in any one of claims 1 to 4, characterized in that: Step S4 specifically includes: using cosine similarity and normalized exponential function to calculate the attention weights of hidden states at different time steps, performing weighted summation on the hidden states output by the long short-term memory network to obtain weighted hidden states; linearly transforming the weighted hidden states through a neural network to output the final Q value.

6. The Q network portfolio method combining long short-term memory and attention mechanism as described in any one of claims 1 to 5, characterized in that: Step S5 specifically includes: using a deep Q network training method to initialize the experience replay pool and Q network parameters; initializing the state in each round, selecting and executing an action according to the ε-greedy strategy at each time step, obtaining a new state and reward after executing the action, and storing the experience in the replay pool; randomly sampling experience from the replay pool, calculating the target Q value and optimizing the Q network parameters; and regularly updating the target Q network parameters to stabilize the training process.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the Q network investment portfolio method combining long short-term memory with an attention mechanism is implemented as described in any one of claims 1 to 6.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the Q network investment portfolio method combining long short-term memory with an attention mechanism is implemented as described in any one of claims 1 to 6.

9. A computing device for a Q network investment portfolio, characterized in that: include Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 6 when executing the computer program.

10. A Q network portfolio device combining long short-term memory and attention mechanism, characterized in that: include: The Markov chain model building module is used to build a Markov chain model based on individual stocks, industries and market signals, define the agent, environment, state, behavior and reward, and integrate individual stocks, industries and market information into a multi-dimensional input vector. The state includes multi-day characteristics, and the behavior includes short selling, long selling and no investment. The reward is calculated based on the stock price increase or decrease and operation. Convolutional neural network module, which is used to input daily features in time order and encode signal features as input to the long short-term memory network through convolutional layers and fully connected layers; The long short-term memory network module is used to convert the output of the convolutional neural network into a hidden state, update the memory using the forget gate, input gate, unit state and output gate mechanism, and output the hidden state as the input of the attention network; The attention network module is used to calculate the attention weight using cosine similarity and normalized exponential function, perform weighted summation on the hidden state, and output the Q value through the linear transformation of the neural network; The deep Q network training module is used to initialize experience playback and Q network parameters, initialize the state in each round, select actions, execute actions, store experience, sample experience, calculate the target Q value and optimize the Q network parameters in each time step, and regularly update the target Q network parameters.