Quantitative investment transaction model construction method and device, electronic equipment and storage medium
By constructing a quantitative investment trading model based on multi-agent deep reinforcement learning, the problems of traditional models being unable to take into account the needs of different clients and the low efficiency of asset allocation are solved, achieving dynamic balance in investment decisions and improving the stability of returns.
Patent Information
- Application Number
- CN202511067074.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional quantitative models cannot achieve a dynamic balance between return targets and customer experience, resulting in low investment decision efficiency and return stability. They also fail to meet the differentiated needs of conservative and aggressive clients, leading to inefficient cross-market asset allocation.
A quantitative investment trading model is constructed, which trains agents using a deep reinforcement learning algorithm with multi-agent dual-delay deep deterministic policy gradient. Combined with a temporal convolutional neural network and attention mechanism, it realizes end-to-end mapping from raw data to policy suggestions, supporting multiple types of intelligent decision-making.
It significantly improves decision-making efficiency and return stability in complex market environments. Traders can dynamically adjust strategy weights and optimize investment portfolios through human-machine collaboration mechanisms.
Smart Images

Figure CN120912334A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and more particularly to a construction method and device of a quantitative investment transaction model, an electronic device and a storage medium. BACKGROUND
[0002] Under the current industry background of capital management, bank wealth management business mainly faces two challenges: first, customer risk preference stratification intensifies, conservative investors are significantly more sensitive to product net value drawdown, and require capital safety first, while aggressive customers pursue excess returns, and traditional single strategy cannot meet the differentiated needs of the two customer groups; second, cross-market asset allocation is inefficient, and dynamic rebalancing of equity, fixed income and derivative assets lacks intelligent decision support, manual portfolio adjustment has response lag and strategy convergence risk, resulting in portfolio yield volatility exceeding customer expectations. The current quantitative model mainly uses a static risk budget mechanism, which cannot achieve dynamic balance between yield targets and customer experience, resulting in low decision-making efficiency and yield stability of investment. SUMMARY
[0003] Therefore, the present application provides a construction method and device of a quantitative investment transaction model, an electronic device and a storage medium, which are used to construct a quantitative investment transaction model capable of providing reasonable transaction decisions to operators, so as to improve the decision-making efficiency and yield stability of wealth management transactions.
[0004] In order to achieve the above-mentioned purpose, the present application provides the following scheme:
[0005] A construction method of a quantitative investment transaction model applied to an electronic device, the construction method comprising the steps of:
[0006] modeling the quantitative investment transaction problem to design agent parameters;
[0007] training and updating the agent based on a deep reinforcement learning algorithm of multi-agent double-delay deep deterministic policy gradient to obtain a plurality of agents;
[0008] improving the plurality of agents by using a time sequence convolutional neural network and an attention mechanism to obtain the quantitative investment transaction model.
[0009] Optionally, the agent parameters include part or all of the number of agents, global state space, action space of the agent, state transition probability function, reward function and cumulative reward discount factor.
[0010] Optionally, the agent includes an Actor network and a Critic network, wherein:
[0011] The Actor network is used to output a decision action at the current time according to a state given by a current investment transaction market transaction environment.
[0012] The Critic network is used to take the state at the current time and all the actions given by the Actor networks of the agents as inputs, and output a Q value to evaluate the actions.
[0013] Optionally, the deep reinforcement learning algorithm based on the multi-agent double-delay deep deterministic policy gradient is used to train and update the agents to obtain a plurality of agents, including the steps of:
[0014] Initializing network parameters of all agents;
[0015] Each agent outputs an action according to a state at the current time, and adds Gaussian noise to each action;
[0016] The reward at the current time and the state at the next time are given by the investment transaction market environment;
[0017] The actions, the Gaussian noise, the reward and the state at the next time are saved in an experience replay pool;
[0018] A plurality of sample data are randomly drawn from the experience replay pool for training to obtain the plurality of agents.
[0019] A construction device of a quantitative investment transaction model, applied to an electronic device, the construction device comprising:
[0020] A parameter design module configured to design agent parameters by modeling a quantitative investment transaction problem;
[0021] An agent training module configured to train and update agents based on a deep reinforcement learning algorithm based on a multi-agent double-delay deep deterministic policy gradient to obtain a plurality of agents;
[0022] A construction execution module configured to improve the plurality of agents by using a time sequence convolutional neural network and an attention mechanism to obtain the quantitative investment transaction model.
[0023] Optionally, the agent parameters include some or all of the number of agents, a global state space, an action space of the agents, a state transition probability function, a reward function and a cumulative reward discount factor.
[0024] Optionally, the agent includes an Actor network and a Critic network, wherein:
[0025] The Actor network is used to output a decision action at the current time according to a state given by a current investment transaction market transaction environment.
[0026] The Critic network is configured to take the state at the current time and the actions given by the Actor networks of all the agents as inputs, and output Q values to evaluate the actions.
[0027] Optionally, the model training module is configured to:
[0028] initialize the network parameters of all the agents;
[0029] each of the agents outputs an action according to the state at the current time, and adds Gaussian noise to each of the actions;
[0030] obtain the reward at the current time and the state at the next time from the investment transaction market environment;
[0031] save the actions, the Gaussian noise, the reward and the state at the next time in an experience replay pool;
[0032] randomly draw a plurality of sample data from the experience replay pool for training, to obtain the plurality of agents.
[0033] An electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0034] the memory is configured to store computer programs or instructions;
[0035] the processor is configured to execute the computer programs or instructions, so that the electronic device implements the construction method as described above.
[0036] A computer-readable storage medium applied to an electronic device, the storage medium carrying one or more computer programs, the one or more computer programs being executable by the electronic device, so that the electronic device can implement the construction method as described above.
[0037] From the above technical solution can be seen, the application discloses a kind of quantitative investment transaction model construction method, device, electronic equipment and storage medium, the method and device are applied to electronic equipment, specifically, by modeling quantitative investment transaction problem, design agent parameter;Intelligent agent training and updating are carried out based on the deep reinforcement learning algorithm of multi-agent double-delay deep deterministic policy gradient, obtain multiple intelligent agents;Multiple intelligent agents are improved using time sequence convolutional neural network and attention mechanism, and quantitative investment transaction model is obtained.The scheme breaks the fragmentation mode of "data display-manual research" in traditional scheme, realizes end-to-end mapping from raw data to strategy suggestion by reinforcement learning intelligent agent, supports multiple intelligent decisions in quantitative trading scene.Trader can dynamically adjust strategy weight through man-machine cooperation mechanism, so as to significantly improve the decision efficiency and income stability in complex market environment. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0039] Figure 1 A flow chart of a quantitative investment transaction model construction method according to an embodiment of the present application;
[0040] Figure 2 A schematic diagram of the algorithm architecture of MATD3;
[0041] Figure 3 A schematic diagram of the specific training process of a single intelligent agent according to an embodiment of the present application;
[0042] Figure 4 A schematic diagram of network design of a quantitative investment transaction model according to an embodiment of the present application;
[0043] Figure 5 A block diagram of a quantitative investment transaction model construction device according to an embodiment of the present application;
[0044] Figure 6 A block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0045] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the scope of protection of the present application.
[0046] The technical solutions of the present application relate to the following several technologies:
[0047] CPPI (Constant Proportion Portfolio Insurance) is a dynamic asset allocation strategy designed to protect portfolios from market downturns while retaining some upside potential. Its core idea is to dynamically adjust the allocation ratio of risky assets (such as stocks) and risk-free assets (such as bonds) to ensure that the value of the portfolio does not fall below the preset floor value (Floor), while participating in market gains when the market rises.
[0048] TIPP (Time-Invariant Portfolio Protection) aims to dynamically adjust the floor value to ensure that the portfolio net value is always above the preset safety threshold, while allowing part of the funds to participate in risky assets to obtain market upside gains. Its core idea is to maximize returns while protecting capital, widely used in quantitative trading, pension management and structured financial products.
[0049] TD3 (Twin Delayed Deep Deterministic Policy Gradient) is an improved reinforcement learning algorithm specifically designed to solve control problems in continuous action space. It is an extended version of the DDPG (Deep Deterministic Policy Gradient) algorithm, which improves the stability and performance of the algorithm by introducing three key technologies: double Q-network, delayed policy update, and target policy smoothing, effectively alleviating the overestimation problem commonly encountered in DDPG.
[0050] Figure 1 A flowchart of a construction method of a quantitative investment trading model according to an embodiment of the present application.
[0051] As Figure 1As shown, the construction method provided by the embodiment is applied to an electronic device, and is used for constructing a quantitative investment transaction model. The electronic device can be understood as a computer, a server or a cloud platform with data calculation capability and information processing capability. The construction method of the embodiment includes the following steps:
[0052] S1, parameters of an intelligent agent are designed by modeling a quantitative investment transaction problem.
[0053] The application describes a transaction decision problem in a multi-agent scenario as a stochastic game process, and the core elements are: <N, S, {A i} i=1,…,N}, P, {R i} i=1,…,N , γ>, wherein N represents the number of intelligent agents, S is a global state space, {A i i=1,…,N is an action space of each intelligent agent, P: S x A→ Δ(S) is an environment state transition probability function, {R i i=1,…,N is a reward function of each intelligent agent, and γ is a cumulative reward discount factor. In the model in the present application, N intelligent agents make transactions in the quantitative market through collaborative decision making, and the meanings of N, S, A, P are described as follows:
[0054] The state space of each intelligent agent is s = [p, h, b, s share ], wherein p, h, and b are independent state variables of each intelligent agent, is a price vector of D stocks, is the holding share of each stock, is the remaining cash amount, s share is a public state variable of each intelligent agent, s share = [M t , L t , MEI t ], wherein M t is a market sentiment proxy variable, L t is a market overall liquidity indicator, and MEI t is a macroeconomic indicator.
[0055] Each intelligent agent determines an action a (indicating a holding weight of each stock) according to a strategy π θ , and updates the market state through joint action a = {a 1 ,…, a N}. After the transaction is completed, each intelligent agent obtains a reward r i according to the asset change, and iteratively optimizes the strategy through reinforcement learning to maximize the cumulative discounted revenue:
[0056]
[0057] where θ i is a policy parameter, γ is used to balance immediate reward and long-term planning.
[0058] After the agent gives an action a, it needs to be revised by CPPI or TIPP strategy. Constant Proportion Investment Portfolio Insurance (CPPI) allocates total assets into safe assets (protection layer) and risky assets (buffer layer) by setting an asset protection lower limit F and a risk multiplier k:
[0059] E = k · C = k · (A - F),
[0060] where A is total assets, C is buffer layer funds. Higher k value corresponds to more aggressive investment strategy. Time Invariant Portfolio Protection (TIPP) dynamically adjusts the protection lower limit F t so that it changes with total assets:
[0061] F t = max{φA t ,F t-1}, E t = k(A t -F t )
[0062] where φ is an adjustment coefficient, which ensures that the protection lower limit increases with market rise and remains unchanged when it falls, thereby avoiding the "gap risk" of CPPI.
[0063] At time t, the immediate reward of agent i is:
[0064]
[0065] The meanings of each part are as follows: ΔV t = V t -V t-1 is the net value change term, where represents the total assets at the current time, w t is the holding weight, P t is the stock price vector, b t is the cash balance; λ · Cost t is the transaction cost term, where Cost t = η · |a t | · V t-1 , a t ∈ [-1, 1] D is the weight of each stock buy and sell, and η is the fixed transaction cost; is the risk penalty term, where is the variance based on historical returns, and L is the risk aversion coefficient. δ · Penalty tThe bottom gap penalty term is used to limit the action from exceeding the CPPI / TIPP strategy range:
[0066]
[0067] Where F t is the dynamic bottom value, F t in TIPP t = φ·V prev +(1-φ)·F t-1 , δ is the bottom penalty coefficient, which prevents the net value from falling below the safety threshold.
[0068] S2, the agent is trained and updated based on the deep reinforcement learning algorithm of MATD3.
[0069] That is, the agent is trained and updated according to the deep reinforcement learning algorithm of MATD3, thereby obtaining a plurality of agents. MATD3 (Multi-Agent Double Delay Deep Deterministic Policy Gradient) adopts a centralized training and decentralized execution optimization framework, and its core idea is to improve the stability of the policy by double delay update and Twin network structure, which is suitable for partially observable environment and complex interactive scene, and the MATD3 algorithm architecture is as shown in Figure 2 .
[0070] In the MATD3 algorithm implemented in the present application, each agent has two networks in common: an Actor network μ(s; θ μ ) and a Critic network Q(s, a; θ Q ). The Actor network outputs the decision action a at the current time according to the state s given by the current investment transaction market transaction environment, and the Critic network takes the state s at the current time and the actions a1, a2,..., a N given by the Actor networks of all agents as inputs, and outputs Q values to evaluate the actions. In the algorithm implemented in the present application, each agent has 6 networks in common, and the parameters of the Actor network and the corresponding target Actor network are represented as θ μ , θ μ′ ; the parameters of the two Critic networks and the corresponding target Critic networks are represented as:
[0071] Before the training starts, the network parameters of all neural networks are first initialized. The agent outputs the action a(t) using the Actor network according to the state s(t) at the current time, and adds Gaussian noise to the action to increase the exploration of the action space and improve the stability of the training.
[0072]
[0073] When all agents give actions, the investment transaction market environment gives the reward r(t) of the current time and the state s(t+1) of the next time. The state s(t) of the current time, the action a(t), the reward r(t) and the state s(t+1) of the next time are saved in the experience replay pool Z in the form of a four-tuple. Then, the agent randomly extracts J sample data from the experience replay pool for training. The following takes an agent's training as an example to illustrate the training process:
[0074] Firstly, the Q value of t time is obtained, the extracted sample data is processed, and (s j (t),A j (t)) is obtained as the input of the agent Critic network, A j (t) represents the set of actions of all agents at t time, and two Q values of the current agent action are obtained through the Critc network: Q 1,j (t) and Q 2,j (t).
[0075] Secondly, the target Q value of t+1 time is obtained: the agent obtains the next time state s j (t+1) from the sample data, inputs it into the target Actor network of itself, and outputs the corresponding action a In order to make the algorithm converge more stably, a random noise is added after the output action, that is, The state s j (t+1) and the action vector set A output by all agents are input into two target Critic networks, so as to obtain two target Q values:
[0076]
[0077] Thirdly, the parameters of the Critic network are updated. In order to avoid overestimation, a smaller target Q value is used to calculate the TD target:
[0078]
[0079] Then, the gradient descent algorithm is used to update the parameters of the two Critic networks of the agent, wherein β is the learning rate of the Critic network, and δ 1,j , δ 2,j is the L2 loss:
[0080]
[0081] The update of the Actor network is to maximize the cumulative reward expectation To this end, a gradient ascent strategy is adopted, and to avoid the risk of concentrating on the strategy of multiple agents due to the strategy convergence, the mutual information constraint is added as a regularization term to the gradient update formula of the actor:
[0082]
[0083] where a is the learning rate of the actor network, MI(π i ,π j ) represents the mutual information of the strategy π i and the strategy π j , which is calculated by the KL divergence approximation.
[0084] The actor network and the three target networks are updated every M rounds, and the parameters of the target network are updated by using the "soft update" method, that is, the parameters of the network are periodically copied to the target network. The specific training process of a single agent is shown in Figure 3 .
[0085] S3, using a time series convolutional neural network and attention mechanism to improve multiple agents.
[0086] Through the improvement of multiple agents by using a time series convolutional neural network TCN and an attention mechanism (Attention), the quantitative investment transaction model is obtained.
[0087] The core of TCN is to capture the time series dependence through causal convolution layers, and the specific calculation process is as follows:
[0088]
[0089] And through the residual connection to build the network:
[0090]
[0091] wherein, is the input sequence, T is the time step, D is the feature dimension, is the learnable convolution kernel, k is the number of convolution kernels, represents the causal convolution operation.
[0092] The specific calculation process of the soft attention mechanism is as follows:
[0093] s(X i ,q)=v T tanh(WX i +U q )
[0094]
[0095] wherein, W, U and v are network weights; the attention variable z e [1, N] represents the index position of selecting information, z = i represents the i-th input information, and finally the weighted average of the input data is calculated according to the attention distribution to obtain the output. The specific network design of the quantitative investment transaction model is shown in Figure 4
[0096] It can be seen from the above technical solution that the embodiment provides a construction method of a quantitative investment transaction model. The method is applied to an electronic device, specifically, an agent parameter is designed by modeling a quantitative investment transaction problem; the agent is trained and updated based on a multi-agent double-delay deep deterministic policy gradient deep reinforcement learning algorithm to obtain a plurality of agents; and a time sequence convolutional neural network and an attention mechanism are used to improve the plurality of agents to obtain a quantitative investment transaction model. The scheme breaks through the fragmented mode of "data display-human research and judgment" in the traditional scheme, realizes end-to-end mapping from raw data to strategy suggestions through a reinforcement learning agent, and supports multi-class intelligent decision-making in a quantitative trading scenario. Traders can dynamically adjust the strategy weight through a man-machine cooperation mechanism, thereby significantly improving the decision-making efficiency and revenue stability in a complex market environment.
[0097] The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] Although the operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, and that certain of the operations can be performed in parallel or in different order than the order shown.
[0099] It should be understood that each of the steps in the method embodiments of the present disclosure can be performed in different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0100] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0101] Figure 5 A block diagram of a quantitative investment transaction model construction device according to an embodiment of the present application.
[0102] As shown in Figure 5 , the construction device provided by the embodiment is applied to an electronic device, and is used for constructing a quantitative investment transaction model. The electronic device can be understood as a computer, a server or a cloud platform having data calculation capability and information processing capability. The construction device of the embodiment includes a parameter design module 10, an agent training module 20 and a construction execution module 30.
[0103] The parameter design module is used for designing agent parameters by modeling the quantitative investment transaction problem.
[0104] In the present application, the transaction decision problem in the multi-agent scenario is described as a stochastic game process, and the core elements are: <N, S, {A i} i=1,…,N}, P, {R i} i=1,…,N , γ>, wherein N represents the number of agents, S is a global state space, {A i} i=1,…,N is an action space of each agent, P: SxA-> Delta (S) is an environment state transition probability function, {R i} i=1,…,N is a reward function of each agent, and γ is a cumulative reward discount factor. In the model in the present application, N agents make transactions in the quantitative market through collaborative decision making, and the meanings of N, S, A, P are described as follows:
[0105] The state space of each agent is s = [p, h, b, s share ], wherein p, h and b are independent state variables of each agent, is a price vector of D stocks, for the holding share of each stock, for the remaining cash amount, s share for the public state variable of each agent, s share = [M t , L t , MEI t ], where M t is the market sentiment proxy variable, L t is the market overall liquidity indicator, and MEI t is the macroeconomic indicator.
[0106] Each agent determines the action a (representing the holding weight of each stock) according to the strategy π θ and interacts to update the market state through the joint action a = {a 1 , …, a N}. After the transaction is completed, each agent obtains the reward r i according to the asset changes and iteratively optimizes the strategy through reinforcement learning to maximize the cumulative discounted return:
[0107]
[0108] where θ i is the strategy parameter, and γ is used to balance the immediate return and long-term planning.
[0109] After the agent gives the action a, it needs to be corrected through the CPPI strategy or the TIPP strategy. The constant proportion investment portfolio insurance (CPPI) allocates the total assets into safe assets (protection layer) and risky assets (buffer layer) by setting the asset protection lower limit F and the risk multiplier k:
[0110] E = k · C = k · (A - F),
[0111] where A is the total asset, and C is the buffer layer fund. A higher k value corresponds to a more aggressive investment strategy. The time-invariant portfolio protection (TIPP) dynamically adjusts the protection lower limit F t so that it changes with the total asset:
[0112] F t = max{φA t , F t-1}, E t = k (A t - F t )
[0113] where φ is the adjustment coefficient, which ensures that the protection lower limit increases with the market rise and remains unchanged when it falls, thereby avoiding the "gap risk" of CPPI.
[0114] The immediate reward of agent i at time t is:
[0115]
[0116] Each part means as follows: Delta V t = V t -V t-1 is a net value change term, wherein represents the total assets at the current time, w t is a holding weight, P t is a stock price vector, b t is a cash balance; lambda*Cost t is a transaction cost term, wherein Cost t = eta*absolute value of a t *V t-1 , a t is in [-1, 1] D is a weight of buying and selling each stock, and eta is a fixed transaction cost; is a risk penalty term, wherein is the variance based on historical yield rate, and L is a risk aversion coefficient. Delta*Penalty t is a bottom gap penalty term for limiting the action from exceeding the range specified by the CPPI / TIPP strategy:
[0117]
[0118] Wherein, F t is a dynamic bottom value, F t in TIPP t = phi*V prev +(1-phi)*F t-1 , and delta is a bottom penalty coefficient, preventing the net value from falling below the safety threshold.
[0119] The agent training module is used for agent training and updating based on the deep reinforcement learning algorithm of MATD3.
[0120] That is, the agent is trained and updated according to the deep reinforcement learning algorithm of MATD3, thereby obtaining a plurality of agents. MATD3 (Multi-Agent Double Delay Deep Deterministic Policy Gradient) adopts an optimization framework of centralized training and decentralized execution, and the core idea is to improve the stability of the policy through double delay update and Twin network structure, which is suitable for partially observable environment and complex interactive scene, and the MATD3 algorithm architecture is as shown in Figure 2 .
[0121] In the MATD3 algorithm implemented in the present application, each agent has two networks in common: an Actor network mu(s; theta μ ) and a Critic network Q(s, a; theta Q). Actor network outputs the decision action a at the current time according to the state s given by the current investment transaction market environment, and the Critic network outputs the Q value of the state s at the current time and the action a1, a2,..., a N , given by the Actor network of all agents. The Q value evaluates the action. In the algorithm implemented in the present application, each agent has 6 networks in common, and the parameters of the Actor network and the corresponding target Actor network are represented as θ μ ,θ μ′ ; the parameters of the two Critic networks and the corresponding target Critic network are represented as:
[0122] Before the training starts, the network parameters of all neural networks are first initialized. The agent outputs the action a(t) according to the state s(t) at the current time using the Actor network, and adds Gaussian noise to the action to increase the exploration of the action space and improve the stability of the training.
[0123]
[0124] When all agents give the action, the reward r(t) at the current time and the state s(t+1) at the next time are given by the investment transaction market environment. The state s(t) at the current time, the action a(t), the reward r(t) and the state s(t+1) at the next time are saved in the experience replay pool Z in the form of a four-tuple. Then, the agent randomly extracts J sample data from the experience replay pool for training. The following takes an agent's training as an example to illustrate the training process:
[0125] Firstly, the Q value at time t is obtained, the extracted sample data is processed, and (s j (t), A j (t)) is obtained as the input of the Critic network of the agent, A j (t) represents the set of actions of all agents at time t, and two Q values of the current agent action are obtained through the Critic network: Q 1,j (t) and Q 2,j (t).
[0126] Secondly, the target Q value at time t+1 is obtained: the agent obtains the state s j (t+1) at the next time from the sample data, inputs it into the target Actor network of itself, and outputs the corresponding action as In order to make the algorithm converge more stably, a random noise is added after the output action, that is, The state s j (t+1) and the action vector set output by all agents As the input of two target Critic networks, two target Q values are obtained:
[0127]
[0128] Again, the parameters of the Critic network are updated, and a smaller target Q value is used to calculate the TD target to avoid overestimation:
[0129]
[0130] Then, the parameters of the two Critic networks of the agent are updated using the gradient descent algorithm, where β is the learning rate of the Critic network, and δ 1,j , δ 2,j is the L2 loss:
[0131]
[0132] The update of the Actor network aims to maximize the expected cumulative reward , and adopts the gradient ascent strategy. To avoid the risk of risk concentration caused by policy convergence among multiple agents, mutual information constraints are added as regular terms to the gradient update formula of the Actor:
[0133]
[0134] where a is the learning rate of the Actor network, and MI(π i , π j ) represents the mutual information between policy π i and policy π j , which is calculated by KL divergence approximation.
[0135] The Actor network and the three target networks are updated every M rounds, and the parameters of the target network are updated using the "soft update" method, that is, the parameters of the network are periodically copied to the target network. The specific training process of a single agent is shown in Figure 3 .
[0136] An execution module is constructed to improve multiple agents using a temporal convolutional neural network and an attention mechanism.
[0137] The quantitative investment trading model is obtained by improving multiple agents using a temporal convolutional neural network (TCN) and an attention mechanism (Attention).
[0138] The core of TCN is to capture temporal dependencies through causal convolution layers, and the specific calculation process is as follows:
[0139]
[0140] And the residual connection is used to build the network:
[0141]
[0142] wherein, is the input sequence, T is the time step, D is the feature dimension, is the learnable convolution kernel, k is the number of convolution kernels, represents the causal convolution operation.
[0143] The specific calculation process of the soft attention mechanism is as follows:
[0144] s(X i , q) = v T tanh(WX i +Uq)
[0145]
[0146] wherein, W, U and v are network weights; the attention variable z e [1, N] represents the index position of the selected information, z = i represents the i-th input information, and finally the weighted average of the input data is calculated according to the attention distribution to obtain the output. The network design of the specific quantitative investment transaction model is shown in Figure 4 .
[0147] As can be seen from the above technical solution, the embodiment provides a construction device of a quantitative investment transaction model, which is applied to an electronic device, specifically, by modeling the quantitative investment transaction problem, an agent parameter is designed; the agent is trained and updated based on a deep reinforcement learning algorithm of a multi-agent double-delay deep deterministic policy gradient to obtain a plurality of agents; a time sequence convolutional neural network and an attention mechanism are used to improve the plurality of agents to obtain the quantitative investment transaction model. The present scheme breaks through the fragmented mode of "data display-human research and judgment" in the traditional scheme, realizes the end-to-end mapping from the original data to the policy suggestion through the reinforcement learning agent, and supports the multi-class intelligent decision-making in the quantitative transaction scene. The trader can dynamically adjust the strategy weight through the man-machine cooperation mechanism, so as to significantly improve the decision-making efficiency and the income stability in the complex market environment.
[0148] The units described in the embodiments of the present disclosure can be implemented in the form of software or in the form of hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases, for example, the first acquisition unit can also be described as "a unit for acquiring at least two internet protocol addresses".
[0149] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, an example type of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0150] Figure 6 A block diagram of an electronic device according to an embodiment of the present application.
[0151] Reference will now be made to Figure 6 which shows a structural diagram suitable for use in implementing an electronic device in an embodiment of the present disclosure. The terminal device in an embodiment of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, as well as a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device is merely an example, and should not impose any limitation on the functions and use range of an embodiment of the present disclosure.
[0152] The electronic device can include a processing means (e.g., a central processing unit, a graphic processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded into a random access memory (RAM) 603 from an input device 606. In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing means, the ROM, and the RAM are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0153] In general, the following devices can be connected to the I / O interface: input devices including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the electronic device is shown as having various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0154] The present application also provides a computer-readable storage medium embodiment.
[0155] The computer readable storage medium described above is applied to an electronic device and carries one or more computer programs, when the one or more computer programs are executed by the electronic device, the one or more computer programs make the electronic device design an agent parameter by modeling a quantitative investment transaction problem; perform agent training and updating based on a multi-agent double-delay deep deterministic policy gradient deep reinforcement learning algorithm to obtain a plurality of agents; and improve the plurality of agents by using a time sequence convolutional neural network and an attention mechanism to obtain a quantitative investment transaction model. The scheme breaks the fragmented mode of "data display-human research and judgment" in the traditional scheme, realizes end-to-end mapping from raw data to strategy suggestions through a reinforcement learning agent, and supports multi-class intelligent decision-making in a quantitative transaction scenario. Traders can dynamically adjust the strategy weight through a man-machine cooperation mechanism, thereby significantly improving decision-making efficiency and revenue stability in a complex market environment.
[0156] It should be noted that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0157] In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0158] Various embodiments of the present specification are described in progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between various embodiments can be referred to each other.
[0159] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make further changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to cover all the preferred embodiments and all the changes and modifications falling within the scope of the embodiments of the present application.
[0160] Finally, it should also be noted that, in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or terminal device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or terminal device including the element.
[0161] The above describes the technical solutions provided by the present application in detail, and the principles and implementation manners of the present application are described by applying specific examples; the above embodiment description is only for helping to understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range will have changes, and the above description of the present application should not be understood as a limitation of the present application.
Claims
1. A method for constructing a quantitative investment trading model, applied to electronic devices, characterized in that, The construction method comprises the steps of: designing agent parameters by modeling a quantitative investment transaction problem; training and updating agents based on a deep reinforcement learning algorithm of multi-agent double-delay deep deterministic policy gradient to obtain a plurality of agents; improving the plurality of agents by using a time sequence convolutional neural network and an attention mechanism to obtain the quantitative investment transaction model.
2. The construction method of claim 1, wherein, The agent parameters include some or all of the number of agents, a global state space, an action space of the agents, a state transition probability function, a reward function, and an accumulated reward discount factor.
3. The construction method of claim 1, wherein, The agent comprises an Actor network and a Critic network, wherein: The Actor network is used to output a decision action at the current time according to a state given by a current investment transaction market trading environment; The Critic network is used to take the state at the current time and actions given by the Actor networks of all the agents as inputs, and output Q values to evaluate the actions.
4. The construction method of claim 1, wherein, The deep reinforcement learning algorithm based on the deep deterministic policy gradient of multi-agent double-delay is used to train and update the agents to obtain a plurality of agents, comprising the steps of: initializing network parameters of all the agents; each agent outputs an action according to a state at the current time, and adds Gaussian noise to each action; a reward at the current time and a state at the next time are given by an investment transaction market environment; the action, the Gaussian noise, the reward, and the state at the next time are saved in an experience replay pool; a plurality of sample data are randomly drawn from the experience replay pool for training to obtain the plurality of agents. 5.A device for constructing a quantitative investment transaction model, applied to an electronic device, and having the following characteristics. The construction device comprises: a parameter design module configured to design agent parameters by modeling a quantitative investment transaction problem; an agent training module configured to train and update agents based on a deep reinforcement learning algorithm of multi-agent double-delay deep deterministic policy gradient to obtain a plurality of agents; a construction execution module configured to improve the plurality of agents by using a time sequence convolutional neural network and an attention mechanism to obtain the quantitative investment transaction model.
6. The build apparatus of claim 5, wherein, The agent parameters include some or all of the number of agents, a global state space, an action space of the agents, a state transition probability function, a reward function, and an accumulated reward discount factor.
7. The build apparatus of claim 5, wherein, The agent comprises an Actor network and a Critic network, wherein: The Actor network is used to output a decision action at the current time according to a state given by a current investment transaction market trading environment; The Critic network is used to take the state at the current time and actions given by the Actor networks of all the agents as inputs, and output Q values to evaluate the actions.
8. The build apparatus of claim 5, wherein, The model training module is configured to: initialize network parameters of all the agents; each agent outputs an action according to a state at the current time, and adds Gaussian noise to each action; a reward at the current time and a state at the next time are given by an investment transaction market environment; the action, the Gaussian noise, the reward, and the state at the next time are saved in an experience replay pool; The plurality of agents are trained by randomly sampling a plurality of template data from the experience replay pool.
9. An electronic device, comprising: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is configured to store computer programs or instructions; The processor is configured to execute the computer programs or instructions to enable the electronic device to implement the construction method according to any one of claims 1-4. 10.A storage medium readable by a computer, applied to an electronic device, and having stored thereon a plurality of instructions which, when executed by the electronic device, cause the electronic device to perform the method of any one of claims 1 to 9. The storage medium carries one or more computer programs, which can be executed by the electronic device, so that the electronic device can implement the construction method according to any one of claims 1-4.