Asset distribution system and method applying reinforcement learning model reflecting product trend prediction
The asset allocation system employs an actor-critic reinforcement learning model to minimize prediction errors and align investment strategies with market trends, enhancing accuracy and stability in stock product returns.
Patent Information
- Application Number
- PCT/KR2025/012773
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-28
- Filing Date
- 2025-08-22
- Publication Date
- 2026-03-05
AI Technical Summary
Existing AI models for stock product return prediction suffer from low accuracy, leading to increased investment losses due to market volatility and noise in predicted values, necessitating a method to minimize errors and reflect appropriate market trends.
An asset allocation system using an actor-critic reinforcement learning model that combines actor and critic neural networks to adjust the reflection ratio of historical and critic rewards based on the investment period, minimizing noise and aligning with market conditions.
The system effectively reduces prediction errors and aligns investment strategies with market trends, achieving a balance between return and stability by adjusting the ratio of historical and critic rewards according to the investment period.
Smart Images

Figure KR2025012773_05032026_PF_FP_ABST
Abstract
Description
An asset allocation system and method that applies a reinforcement learning model that reflects product trend prediction.
[0001] The present invention relates to an asset allocation system and method using a reinforcement learning model, and more particularly, to an asset allocation system and method using a reinforcement learning model learned through an actor-critic algorithm that reflects stock product trend prediction.
[0002] The actor-critic algorithm is one of the algorithms used in reinforcement learning, a methodology for learning how an agent selects the optimal action in a given environment through learning.
[0003] The actor-critic algorithm is structured to learn both policy and value functions simultaneously, and the actor neural network learns the policy function that determines the agent's behavior.
[0004] The policy function determines which action an agent should take in a given state. It outputs a probability distribution over possible actions in a given state and selects an action based on this distribution.
[0005] Critic neural networks learn a value function at a given state.
[0006] A value function is a function that estimates the expected value or average reward in a given state.
[0007] Critic neural networks can evaluate the value of an agent's chosen action and use this to update the value of the action.
[0008] The actor-critic algorithm combines these two elements to perform learning and can be used for efficient reinforcement learning.
[0009] Meanwhile, in the case of AI models that predict the returns of existing indexes and stock products, the investment return results of the AI model tend to depend on the predicted values of the returns of the index and stock products, and if the accuracy of the predicted values of the returns of the index and stock products is low, there is a problem that the amount of investment loss increases.
[0010] Therefore, in order to respond to market volatility that is difficult to predict, it is necessary to utilize an actor-critic algorithm that reflects the prediction of stock product trends according to market conditions, thereby minimizing noise resulting from errors in the predicted returns of stock products and exploring ways to reflect appropriate market trends according to the product operation period.
[0011] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide an asset allocation system and method that can minimize noise from errors in the predicted value of the return on stock products and reflect appropriate market trends according to the product operation period by applying an appropriate ratio of stock product trend predictions according to market conditions to the critique in order to respond to market volatility that is difficult to predict.
[0012] In order to achieve the above object, according to one embodiment of the present invention, an asset allocation method includes a step of training a reinforcement learning model through an actor-critic algorithm, wherein the actor-critic algorithm is comprised of an actor, which is a neural network that calculates an investment strategy result for asset allocation based on previously stored financial market-related data, and a critic, which is a neural network that evaluates the appropriateness of the investment strategy result calculated through the actor; and a step of the asset allocation system applying financial market-related data to the learned reinforcement learning model to establish an investment strategy for asset allocation.
[0013] And the critic neural network can predict the average return and volatility of each product and the average return and volatility of each asset group in the current market situation by applying the average return and variance of the product's operating period as a supervised learning technique.
[0014] Additionally, actor neural networks can use data on average returns and Sharpe ratios as rewards to produce asset allocation results.
[0015] And the actor neural network sets the data on average returns and Sharpe Ratio as historical rewards, and sets the average returns and volatility prediction results for each product produced through critics as critic rewards, so that the historical rewards and critic rewards are reflected together as rewards, but in order to reduce the impact of prediction errors, the reflection ratio between the historical rewards and critic rewards can be adjusted according to the set operating period of the investment strategy to be established.
[0016] In addition, the actor neural network can adjust the reflection ratio of historical rewards higher the shorter the set operating period of the investment power to be established, and can adjust the reflection ratio of critic rewards higher the longer the set operating period.
[0017] And the critic neural network is prepared in multiples, and each critic neural network can individually perform supervised learning.
[0018] In addition, each critic neural network produces a prediction result value for an item assigned to each critic neural network, and the assigned item may be any one of an average return prediction item by product, a volatility prediction item by product, an average return prediction item by asset class, and a volatility prediction item by asset class.
[0019] And each critic neural network can produce a predicted result value for the average return and volatility for each product, for which a different operating period is set for each critic neural network, and can produce a representative value of the predicted result value produced by each critic neural network based on the set operating period of the investment power to be established so that it can be set as the critic reward of the actor neural network.
[0020] Meanwhile, according to another embodiment of the present invention, an asset allocation system includes a learning unit that trains a reinforcement learning model through an actor-critic algorithm, which is composed of an actor, which is a neural network that calculates an investment strategy result for asset allocation based on previously stored financial market-related data, and a critic, which is a neural network that evaluates the appropriateness of the investment strategy result calculated through the actor; and an investment strategy generation unit that applies financial market-related data to the learned reinforcement learning model to establish an investment strategy for asset allocation.
[0021] As described above, according to the embodiments of the present invention, noise resulting from errors in the prediction value of the return of a stock product can be minimized, and appropriate market trends according to the product operation period can be reflected, thereby establishing an investment strategy for asset allocation that achieves a balance between return and stability.
[0022] Figure 1 is a drawing provided to explain the configuration of an asset allocation system according to one embodiment of the present invention;
[0023] FIG. 2 is a drawing provided for a more detailed configuration description of a processor according to one embodiment of the present invention;
[0024] Figure 3 is a flowchart provided to explain an asset allocation method according to one embodiment of the present invention;
[0025] Figure 4 is a flowchart providing a more detailed description of the learning process of a reinforcement learning model according to one embodiment of the present invention; and
[0026] FIG. 5 is a diagram provided to explain an actor-critic algorithm according to one embodiment of the present invention.
[0027] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0028] FIG. 1 is a diagram provided to explain the configuration of an asset allocation system according to one embodiment of the present invention.
[0029] The asset allocation system according to this embodiment applies stock product trend predictions based on market conditions to the critique at an appropriate ratio to respond to market volatility that is difficult to predict, thereby minimizing noise resulting from errors in the predicted value of stock product returns and establishing an investment strategy for asset allocation that reflects appropriate market trends according to the product operation period.
[0030] To this end, the asset allocation system may include a communication unit (100), a processor (200), and a storage unit (300).
[0031] The communication unit (100) is equipped with a communication module connected to a network, and can collect or transmit to the outside data (e.g., data related to financial markets) necessary for the processor (200) to operate.
[0032] Here, data related to the financial market may include stock indices, information on each stock product, information on asset classes, bond interest rate information, volatility information, raw material information, correlation coefficients, etc.
[0033] The storage unit (300) is provided to store programs and data required for the processor (200) to operate.
[0034] For example, the storage unit (300) can store financial market-related data and learned reinforcement learning models obtained through the communication unit (100), etc.
[0035] The processor (200) trains a reinforcement learning model through an actor-critic algorithm, which is composed of an actor, which is a neural network that produces investment strategy results for asset allocation based on previously stored financial market-related data, and a critic, which is a neural network that evaluates the appropriateness of the investment strategy results produced through the actor, and can establish an investment strategy for asset allocation by applying previously stored or newly acquired financial market-related data to the trained reinforcement learning model.
[0036] FIG. 2 is a drawing provided for a more detailed configuration description of a processor (200) according to one embodiment of the present invention.
[0037] Referring to FIG. 2, the processor (200) may include a learning unit (210) that learns a reinforcement learning model through an actor-critic algorithm, and an investment strategy generation unit (220) that applies financial market-related data to the learned reinforcement learning model to establish an investment strategy for asset allocation.
[0038] The learning unit (210) can train a reinforcement learning model by utilizing an actor neural network and one or more critic neural networks.
[0039] Here, the actor neural network can use data on average returns and Sharpe ratios as rewards to produce investment strategy results for asset allocation.
[0040] The Critic Neural Network can predict the average return and volatility of each product and the average return and volatility of each asset class in the current market situation by applying the average return and variance of the product's operating period as a supervised learning technique.
[0041] That is, the learning unit (210) can train a reinforcement learning model to produce investment strategy results for asset allocation by reflecting historical rewards and critic rewards together as rewards of the actor neural network.
[0042] For example, an actor neural network can set data on average returns and Sharpe ratios as historical rewards, and set the average returns and volatility prediction results for each product produced through critics as critic rewards, thereby reflecting both historical rewards and critic rewards as rewards. In this case, to reduce the impact of prediction errors, the reflection ratio between historical rewards and critic rewards can be adjusted according to the set operating period of the investment strategy to be established. Here, volatility prediction can produce a predicted value based on the standard deviation of returns.
[0043] And for example, the actor neural network can adjust the reflection ratio of historical rewards higher the shorter the set operating period of the investment power to be established, and can adjust the reflection ratio of critic rewards higher the longer the set operating period.
[0044] This can improve market volatility, which is difficult to predict, and minimize noise from predicted value errors by adjusting the reflection ratio of historical rewards, which reflect the current market situation, to a higher level when the set operating period is relatively short, and adjusting the reflection ratio of critical rewards, which reflect the predicted value for the future market, to a higher level when the set operating period is relatively long.
[0045] In addition, when the learning unit (210) trains a reinforcement learning model by reflecting both historical rewards and critic rewards as rewards of the actor neural network, it can calculate asset allocation results by applying only historical rewards using the actor neural network, evaluate the appropriateness of the investment strategy results calculated using the critic neural network, and then adjust the reflection ratio between the historical rewards and critic rewards based on the appropriateness evaluation results.
[0046] And the learning unit (210) adjusts the reflection ratio between the historical reward and the critic reward, and then uses the actor neural network to apply the historical reward and the critic reward according to the adjusted reflection ratio to produce an asset allocation result, and can re-evaluate the appropriateness of the investment strategy result produced using the critic neural network.
[0047] Figure 3 is a flowchart illustrating an asset allocation method according to one embodiment of the present invention. The asset allocation method according to this embodiment can be implemented by the asset allocation system described above with reference to Figures 1 and 2.
[0048] The asset allocation system trains a reinforcement learning model through an actor-critic algorithm, which is composed of an actor, which is a neural network that produces investment strategy results for asset allocation based on previously stored financial market-related data, and a critic, which is a neural network that evaluates the appropriateness of the investment strategy results produced by the actor (S310), and applies previously stored financial market-related data or newly acquired financial market-related data to the trained reinforcement learning model (S320), thereby establishing an investment strategy for asset allocation (S330).
[0049] FIG. 4 is a flowchart provided for a more detailed description of the learning process of a reinforcement learning model according to one embodiment of the present invention, and FIG. 5 is a diagram provided for a description of an actor-critic algorithm according to one embodiment of the present invention.
[0050] Referring to FIG. 4, when the asset allocation system trains a reinforcement learning model by reflecting both historical rewards and critic rewards as rewards of the actor neural network, the asset allocation result is calculated by applying only the historical rewards using the actor neural network (S410), the appropriateness of the investment strategy result calculated using the critic neural network is evaluated (S420), and the reflection ratio between the historical rewards and critic rewards can be adjusted based on the appropriateness evaluation result (S430).
[0051] And, when the reflection ratio between the historical reward and the critic reward is adjusted, the asset allocation system uses the actor neural network to apply the historical reward and the critic reward according to the adjusted reflection ratio to produce the asset allocation result (S440), and can re-evaluate the appropriateness of the investment strategy result produced using the critic neural network (S450).
[0052] At this time, the asset allocation system can repeatedly execute steps S430 to S450 several times by adjusting the reflection ratio between the historical reward and the critic reward differently, and determine the reflection ratio between the historical reward and the critic reward with the highest appropriateness evaluation result among these as the final reflection ratio, thereby completing the training of the reinforcement learning model.
[0053] Additionally, the asset allocation system can implement multiple critic neural networks that constitute the actor-critic algorithm, and perform supervised learning individually for each critic neural network.
[0054] At this time, each critic neural network can produce predicted results for the items assigned to it. These assigned items may include items predicting average returns by product, items predicting volatility by product, items predicting average returns by asset class, and items predicting volatility by asset class.
[0055] In the case of FIG. 5, a critic neural network that produces a prediction result value for an average return prediction item by product and a critic neural network that produces a prediction result value for an item predicting volatility (standard deviation of return) by product are exemplified, but in this asset allocation system, as described above, critic neural networks that produce prediction result values for an average return prediction item by asset group and a volatility prediction item by asset group can be additionally created.
[0056] Additionally, each critic neural network can produce predicted results for the average return and volatility for each product, with different operating periods set for each critic neural network.
[0057] In this case, each critic neural network can produce a representative value of the predicted result value produced by each critic neural network based on the set operating period of the investment power to be established, so that it can be set as the critic neural network reward of the actor neural network.
[0058] For example, when each critic neural network produces a representative value of the predicted result value for the average return, the shorter the set operating period is, the smaller the weight applied to the predicted result value is given, and the longer the set operating period is, the larger the weight applied to the predicted result value is given, so that the predicted result value for the average return with a relatively short set operating period is reflected more than the predicted result value for the average return with a relatively long set operating period, and the reflection ratio can be adjusted.
[0059] At this time, the reflection ratio of each critic neural network can be adjusted so that the difference between the weights applied to the predicted results according to the operating period increases as the fluctuation range (standard deviation value) compared to the average return is calculated to be large.
[0060] Meanwhile, it goes without saying that the technical idea of the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.
[0061] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A step of training a reinforcement learning model through an actor-critic algorithm, in which the asset allocation system is composed of an actor, which is a neural network that produces investment strategy results for asset allocation based on previously stored financial market-related data, and a critic, which is a neural network that evaluates the appropriateness of the investment strategy results produced through the actor; and An asset allocation method comprising a step of establishing an investment strategy for asset allocation by applying financial market-related data to a learned reinforcement learning model.
2. In claim 1, The critic neural network is An asset allocation method characterized by applying the average return and variance of a product's operating period using a supervised learning technique to predict the average return and volatility of each product and the average return and volatility of each asset group in the current market situation.
3. In claim 2, Actor neural networks, An asset allocation method characterized by calculating asset allocation results using data on average returns and the Sharpe ratio as rewards.
4. In claim 3, Actor neural networks, Data on average returns and Sharpe Ratio are set as historical rewards, and the average returns and volatility prediction results for each product produced through critics are set as critic rewards, so that historical rewards and critic rewards are reflected together as rewards. An asset allocation method characterized by adjusting the reflection ratio between historical rewards and critic rewards according to the set operating period of the investment strategy to be established, in order to reduce the impact of prediction errors.
5. In claim 4, Actor neural networks, An asset allocation method characterized by adjusting the reflection ratio of historical rewards higher the shorter the set operating period of the investment strategy to be established, and adjusting the reflection ratio of critical rewards higher the longer the set operating period.
6. In claim 2, The critic neural network is It is prepared in revenge, Each critic neural network, An asset allocation method characterized by individually performing supervised learning.
7. In claim 6, Each critic neural network, Produces a prediction result for each assigned item for each critic neural network, The assigned items are: An asset allocation method characterized by one of the following items: average return prediction items by product, volatility prediction items by product, average return prediction items by asset class, and volatility prediction items by asset class.
8. In claim 6, Each critic neural network, The predicted results for the average return and volatility for each product, with different operating periods set for each critic neural network, are calculated. An asset allocation method characterized by calculating a representative value of the predicted result value produced by each critic neural network based on the set operating period of the investment power to be established so that it can be set as the critic reward of the actor neural network.
9. A learning unit that trains a reinforcement learning model through an actor-critic algorithm, which is composed of an actor, which is a neural network that produces investment strategy results for asset allocation based on previously stored financial market-related data, and a critic, which is a neural network that evaluates the appropriateness of the investment strategy results produced through the actor; and An asset allocation system including an investment strategy generation unit that applies financial market-related data to a learned reinforcement learning model to establish investment strategies for asset allocation.
Citation Information
Patent Citations
Performance directing system
KR1020230017906A
Apparatus and method for generating asset portfolio based on reinforcement learning
KR102170574B1
Method and device for providing information on products that can be provided with a short delivery period
KR102398602B1
Work Permission System Based on Inference of Hazardous Gas and Oxygen Concentration at Work Site
KR102940479B1
KR20210015580A