Dynamic hedging strategy optimization method based on reinforcement learning
By integrating the hidden Markov model and dynamic risk penalty mechanism, the hedging strategy of the reinforcement learning agent is optimized, the problems of multi-scale market perception and risk control are solved, and the adaptability and interpretability of the strategy are achieved.
Patent Information
- Application Number
- CN202510836952.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-22
- Publication Date
- 2025-09-19
AI Technical Summary
Existing reinforcement learning hedging models lack multi-scale market perception capabilities, have fixed risk sensitivity, and lack strategy interpretability, making it difficult to achieve adaptive risk control in complex financial markets.
By integrating the macro-scenario belief state generated by the hidden Markov model with real-time micro-market data, an enhanced state representation is constructed, combined with a dynamic risk penalty mechanism, and the hedging strategy of the reinforcement learning agent is optimized.
It improves the adaptability of strategies in different market environments, realizes scenario-adaptive adjustment of risk sensitivity, provides explainable decision-making logic, and overcomes the limitations of traditional models.
Smart Images

Figure CN120672478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial derivatives risk management, and in particular to a dynamic hedging strategy optimization method based on reinforcement learning. Background Art
[0002] Hedging with financial derivatives is a core risk management tool, offsetting potential losses by establishing opposing positions and ensuring the stability of investment portfolio value. With increasing market volatility and trading frequency, traditional static hedging strategies are struggling to cope with complex and volatile market conditions. Optimizing intelligent, dynamic hedging strategies has become a hot topic in industry research. Especially in modern financial markets, where high-frequency trading and algorithmic trading dominate, adaptive adjustment of hedging strategies and precise risk control have become key factors in enhancing the competitiveness of financial institutions.
[0003] Existing technologies primarily employ three approaches to dynamic hedging: rule-based approaches execute hedging operations through preset trigger conditions and adjustment rules; statistical model-based approaches utilize time series models such as GARCH to predict volatility and adjust the hedge ratio accordingly; and machine learning-based approaches employ deep reinforcement learning to directly learn optimal hedging strategies from historical market data. Reinforcement learning methods have garnered widespread attention due to their ability to directly optimize long-term cumulative returns. Typical implementations include using deep Q-networks (DQNs) or policy gradient methods to model the hedging decision-making process, taking market observations as state input and outputting hedge position adjustments.
[0004] While existing technologies have achieved some success under specific market conditions, they still have several shortcomings. First, traditional reinforcement learning hedging models rely solely on micro-market data to construct state representations and lack the ability to perceive macro-market scenarios. This results in limited strategy adaptability during market regime transitions (e.g., from low volatility to high volatility). This issue stems from the single timescale limitations of state design, which fail to capture both short-term price fluctuations and long-term market patterns. Second, existing risk control mechanisms often employ fixed-weight penalties, ignoring the varying risk sensitivities across different market scenarios. This can lead to insufficient risk constraints during periods of high volatility and excessive conservatism during periods of low volatility, failing to achieve a dynamic risk-return balance. Finally, mainstream reinforcement learning hedging models generally exhibit "black box" characteristics, making their decision-making logic difficult to interpret and lacking explicit connection to market context. This severely hinders their application in real-world financial scenarios and regulatory compliance. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a dynamic hedging strategy optimization method based on reinforcement learning, which solves the problems of the existing hedging model's lack of multi-scale market perception ability, fixed risk sensitivity and insufficient strategy interpretability.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a dynamic hedging strategy optimization method based on reinforcement learning, comprising the following steps: S1. Obtain the market observation data for the current time step and the hedging action performed by the reinforcement learning agent in the previous time step; S2. Inputting the market observation data and the hedging action of the previous time step into the market scenario inference model, and outputting the probability distribution of the current market being in multiple preset implicit scenarios as the macro-scenario belief state; S3. Combining the macro-scenario belief state with the micro-market state at the current time step into an enhanced state vector, inputting the vector into the reinforcement learning agent to generate a hedging action at the current time step; S4. Based on the execution result of the hedging action, the macro-scenario belief state and the preset risk penalty mechanism, calculate the reward signal and update the parameters of the reinforcement learning agent.
[0007] Preferably, step S1 includes the following steps: The market observation data includes the underlying asset price, realized volatility and trading volume; The hedging impulse at the previous time step is used as the target hedging position ratio output by the reinforcement learning agent at the previous time step; This ratio is enforced by: If it is a live transaction, the target position ratio is converted into actual trading instructions through the order management system and sent to the exchange; If it is simulated trading, the target position ratio is recorded in the virtual trading environment and the simulated portfolio status is updated.
[0008] Preferably, the market observation data is obtained by: Underlying asset price: Obtain the latest transaction price of the underlying asset published by the exchange in real time through the financial data interface; Realized Volatility: Calculates the standard deviation of the logarithmic return based on the underlying asset's historical price series within a preset time window. Trading Volume: Obtain the total trading volume of the underlying asset in the latest time step through the market data subscription service.
[0009] Preferably, the calculation time window for the realized volatility is 30 minutes, and the calculation formula for the logarithmic rate of return is: ; Where, For the moment The logarithmic rate of return; For the moment The price of the underlying asset; For the moment The price of the underlying asset.
[0010] Preferably, the step 2 comprises the following steps: The market scenario inference model is a hidden Markov model, and its observation value is an augmented observation tuple consisting of the market observation data and the hedging action of the previous time step; The macro-scenario belief state is calculated by a forward algorithm, specifically: ; Where, For the moment The market is in The posterior probability of each implicit scenario; For the moment The hidden state variables of is defined in the hidden Markov model. a hidden market scenario; is the augmented observation tuple; For the moment market observation data.
[0011] Preferably, the hidden Markov model is pre-trained by the following steps: Collect historical market data and corresponding historical hedging actions to construct an augmented observation sequence; The Baum-Welch algorithm is used to optimize the model parameters to maximize the likelihood probability of the augmented observation sequence.
[0012] By integrating the macro-level belief states generated by a hidden Markov model with real-time micro-level market data, we construct an enhanced state representation, enabling the reinforcement learning agent to simultaneously capture both short-term market fluctuations and long-term scenario evolution patterns. This design improves the strategy's adaptability across diverse market environments.
[0013] Preferably, step S3 includes the following steps: The micro-market conditions include the price of the underlying asset, the remaining maturity of the derivative, the current hedge position and market volatility; The enhanced state vector is generated by vector concatenation, specifically: ; Where, For the moment The enhanced state vector of is the micro-market state vector, which includes the underlying asset price, remaining time to maturity, current position ratio and market volatility; is the macro-scenario belief state vector.
[0014] Preferably, the actor network of the reinforcement learning agent adopts a gated structure, wherein: The macroscopic scenario belief state vector generates a gating signal through a fully connected layer; The gating signal is used to modulate the characteristic expression of the micro-market state in the strategy network.
[0015] Preferably, step S4 includes the following steps: The calculation formula of the reward signal is: ; Where, For the moment reward signal; is the change in portfolio value; For transaction costs; For the past The profit and loss variance of the step; Indicates the past Rolling variance calculation for time steps; A risk penalty weight based on the macro-scenario belief state.
[0016] Preferably, the risk penalty weight is calculated in the following manner: Preset base weights for each implicit scenario ; Multiply the macro-scenario belief state vector by the basic weight vector to obtain the dynamic weight: ; Where, is the dynamic risk penalty weight; For the The default base risk weights associated with each implicit scenario; For the moment The market is in The probability of a hidden scenario; is the total number of hidden scenarios preset in the hidden Markov model.
[0017] The present invention provides a dynamic hedging strategy optimization method based on reinforcement learning. It has the following beneficial effects: 1. This invention fuses macro-scenario belief states generated by a hidden Markov model with real-time micro-market data to construct an enhanced state representation, enabling a reinforcement learning agent to simultaneously capture both short-term market fluctuations and long-term scenario evolution patterns. This design overcomes the limitations of traditional hedging strategies that rely on a single timescale for analysis, improving the strategy's adaptability across diverse market environments and demonstrating greater robustness during market transitions.
[0018] 2. Dynamically calculate risk penalty weights based on macro-scenario belief states, combining the probability distribution of potential market scenarios with pre-set risk preferences to achieve scenario-adaptive adjustment of risk sensitivity. This mechanism enables the agent to automatically strengthen risk constraints during periods of high volatility and moderately relax its return pursuit during periods of stability, effectively balancing the return-risk trade-off and avoiding the strategic rigidity caused by traditional fixed-weight penalties.
[0019] 3. By explicitly modeling implicit market scenarios through hidden Markov models, the macro-belief state vector is given clear physical meaning (e.g., "high volatility risk-averse period" and "low volatility stable period"). The action generation logic of the agent in the decision-making process is strongly linked to the market scenario probabilities, providing a traceable semantic explanation for strategic behavior and overcoming the compliance barriers faced by traditional black-box reinforcement learning models in practical financial applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a schematic diagram of the market scenario inference model structure of the present invention; Figure 3 This is a schematic diagram of the gating structure of the reinforcement learning agent of the present invention; Figure 4 Schematic diagram of the reward signal calculation mechanism of the dynamic risk penalty of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the specification of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] Please see the attached Figure 1 -Attached Figure 4 , an embodiment of the present invention provides a dynamic hedging strategy optimization method based on reinforcement learning, comprising the following steps: S1. Obtain the market observation data for the current time step and the hedging action performed by the reinforcement learning agent in the previous time step; In this embodiment, step S1 provides basic input for subsequent market scenario inference and strategy optimization through real-time data collection and action execution mechanism. The specific implementation is as follows: Market observation data includes underlying asset prices, realized volatility, and trading volume. Underlying asset prices are obtained in real time through a financial data interface, specifically a standardized API provided by the exchange (such as the FIX protocol or RESTful interface), capturing the latest transaction prices at a preset sampling frequency (e.g., every minute). Realized volatility is calculated based on the historical price series within a sliding time window by taking the natural logarithm of the underlying asset price return at each time step within the window and then calculating the standard deviation of these returns.
[0023] The formula for calculating logarithmic rate of return is: ; Where, For the moment The logarithmic rate of return; For the moment The price of the underlying asset; For the moment The price of the underlying asset.
[0024] The length of the sliding time window can be dynamically adjusted based on the derivative's expiration time, preferably 30 minutes, to balance short-term volatility sensitivity and noise filtering requirements. Trading volume data is obtained in real time through a market data subscription service. Specifically, the total trading volume of the underlying asset in the current time step is calculated cumulatively from the transaction records published by the exchange.
[0025] The target hedging position ratio output by the reinforcement learning agent is the impulse at the previous time step. Its value range is normalized to [-1, 1], where -1 represents a full negative hedge and 1 represents a full positive hedge. The execution method of this ratio is divided into two categories according to the trading environment: In live trading scenarios, the Order Management System (OMS) converts target position ratios into actual trading instructions. Specifically, the required position size is calculated based on the current portfolio value and the underlying asset price, and a limit or market order is generated and sent to the exchange for execution. Transaction costs are calculated based on the actual transaction price and quantity, including commissions, slippage, and market impact costs.
[0026] In simulated trading scenarios, hedge position status is updated within a virtual trading environment. This virtual environment maintains the simulated portfolio's holdings, cash, and historical transaction records, calculates theoretical transaction prices (such as the mid-price or weighted average price) based on target position ratios, and records theoretical transaction costs for reward calculations.
[0027] To improve model training stability, the raw market observation data and hedging actions were standardized. The underlying asset price was converted to a logarithmic return series, the volatility was Z-score normalized, and the trading volume was divided by the historical average trading volume to eliminate dimensionality effects. Before being input into the market scenario inference model, the hedging action ratio was linearly scaled to the range [0, 1] to prevent negative values from interfering with probability calculations.
[0028] The data collection and processing module uses an asynchronous pipeline architecture, distributing and buffering real-time data streams through message queues (such as Kafka), ensuring low latency and high throughput in high-frequency scenarios. Both raw and processed data are persistently stored in a time-series database (such as InfluxDB) for subsequent module queries and offline analysis.
[0029] S2. Inputting the market observation data and the hedging action of the previous time step into the market scenario inference model, and outputting the probability distribution of the current market being in multiple preset implicit scenarios as the macro-scenario belief state; In this embodiment, step S2 uses an agent-aware hidden Markov model (AA-HMM) to dynamically infer the market macro-scenario, fusing market observation data with historical actions into augmented observations to generate a macro-belief probability distribution representing the underlying market state. The specific implementation is as follows: The market scenario inference model uses the Hidden Markov Model (HMM) framework. Its core innovation is to introduce agent actions as an observation component. The model definition includes the following elements: Hidden state set: preset Implicit market scenarios (e.g., “high volatility bull market,” “low volatility volatile market,” “liquidity crisis,” etc.), each scenario Characterizes a specific pattern of market dynamics.
[0030] Augmented Observation Design: Observations are expanded into tuples , where The current market observation data (underlying asset price, volatility, and trading volume) obtained in step S1; is the hedging action performed in the previous time step (normalized position ratio). This design enables the model to perceive the joint dynamics of market state and agent behavior; For the moment The augmented observation tuple of .
[0031] By introducing hedging actions as an observation dimension, the model can capture the dynamic interaction between market status and agent behavior, thereby more accurately identifying potential scenarios.
[0032] The hidden Markov model is defined implicit market scenarios (such as "high volatility risk aversion period" and "low volatility stable period"), whose parameters include the state transition matrix , initial state distribution and augmented observation emission probability . Number of implicit scenarios It can be determined through cluster analysis of historical data, and 3 to 5 categories are preferred to balance model complexity and generalization ability.
[0033] The training of the hidden Markov model is completed based on the historical augmented observation sequence, and specifically includes the following steps: Data preparation: Collect the underlying asset price, volatility, and trading volume data over a historical period, as well as the corresponding historical hedging action sequences (which can be generated by a benchmark strategy or provided by actual trading records).
[0034] Sequence construction: Align historical market data and actions by time steps to construct augmented observation sequences , where The length of historical data.
[0035] Parameter optimization: Use the Baum-Welch algorithm for unsupervised training and iteratively update model parameters ; To maximize the likelihood probability of the observation sequence; Where, is the complete parameter set of the hidden Markov model; is the state transition probability matrix; To augment the observed emission probability distribution; is the initial state probability distribution vector.
[0036] After training, the model is able to map augmented observations to the implicit context space, providing prior knowledge for online inference.
[0037] In the real-time hedging process, the hidden Markov model receives the augmented observation provided in step S1 , recursively calculate the macroscopic scenario belief state at the current moment through the forward algorithm: ; Where, For the moment The market is in The posterior probability of each implicit scenario; For the moment The hidden state variables of is defined in the hidden Markov model. a hidden market scenario; is the augmented observation tuple; For the moment market observation data.
[0038] Forward Probability The calculation formula is: ; Where, For the moment The forward probability of For the Generate augmented observations under implicit scenarios The emission probability density of For the hidden state from The scene shifts to The probability of a scenario; For the moment The forward probability of is the total number of hidden scenarios preset in the hidden Markov model; initial conditions = .
[0039] The macro belief state vector is obtained by normalization: ; Where, For the moment The macro-scenario belief state vector of ; For the moment The forward probability of The sum of the forward probabilities of all implicit scenarios is used for normalization calculation; Index variables for summation; Step S2 couples market data with agent behavior, enabling the macro belief state to dynamically reflect the impact of agent actions on the market, providing adaptive situational awareness capabilities for subsequent hierarchical decision-making.
[0040] S3. Combining the macro-scenario belief state with the micro-market state at the current time step into an enhanced state vector, inputting the vector into the reinforcement learning agent to generate a hedging action at the current time step; In this embodiment, step S3 constructs a multi-dimensional enhanced state representation by integrating the macro-market scenario inference results with the micro-real-time transaction status, and generates dynamic and adaptive hedging actions based on the hierarchical gating strategy network. The specific implementation is as follows: Micro-market state feature extraction: The micro-market state is composed of four key features: Underlying asset price: The latest transaction price of the underlying asset obtained in real time through the financial data interface, reflecting the real-time supply and demand relationship in the market.
[0041] Derivative Remaining Time to Expiration: Calculates the difference between the derivative contract's expiration date and the current time, and maps it to the interval [0, 1] using linear normalization, where 1 represents the contract's initial time and 0 represents the expiration time. This feature is used to characterize the impact of time decay on hedging strategies.
[0042] Current Hedge Position: records the number of hedge positions actually executed in the previous time step, normalized by the proportion of the total portfolio value, and the value range is [-1, 1], where -1 indicates a full negative hedge and 1 indicates a full positive hedge.
[0043] Market volatility: Based on the realized volatility calculated in step S1, it is normalized using the Z-score to eliminate dimensional differences and reflects the short-term market risk level.
[0044] The above features are preprocessed in real time through the feature engineering module to ensure the stability of data distribution and avoid model training deviation due to dimensional differences.
[0045] The microscopic state vector is expressed as: ; Where, For the moment The micro-market state vector of ; For the moment The price of the underlying asset; The remaining time to maturity of the derivative; is the hedge position ratio at the previous time step; is the realized volatility; The macroscopic scenario belief state vector output from step S2 and the micro-market state vector Perform cross-scale fusion to generate an enhanced state vector: ; Where, For the moment The enhanced state vector of is the micro-market state vector, which includes the underlying asset price, remaining time to maturity, current position ratio and market volatility; is the macro-scenario belief state vector.
[0046] By integrating macro and micro features, the intelligent agent can simultaneously perceive the long-term market scenario evolution laws and short-term trading dynamics, providing information support for multi-time scale decision-making.
[0047] The actor network of the reinforcement learning agent adopts a gated structure, which is specifically implemented as follows: Macro feature extraction: The macro belief state vector Input the fully connected layer to generate the gating signal , the formula is: ; Where, is the gate signal vector; is the weight matrix of the macro feature extraction layer; is the macro-scenario belief state vector output by step S2; is the bias term of the macro feature extraction layer; is the rectified linear unit activation function; Microscopic feature extraction: The microscopic state vector Input another fully connected layer to obtain basic features : ; Where, is the microscopic eigenvector; is the weight matrix of the micro-feature extraction layer; is the micro-market state vector; is the bias term of the micro-feature extraction layer; Gated modulation: using a gate signal Microscopic features To perform dynamic modulation: ; Where, is the modulated eigenvector; represents element-wise multiplication; is the Sigmoid function; is the gate signal vector.
[0048] Action generation: Modulated features Input and output layer, generating hedging actions : ; Where, For the generated hedging action; is the weight matrix of the output layer; is the modulated eigenvector; is the bias of the output layer; is the hyperbolic tangent function.
[0049] S4. Calculate a reward signal based on the execution result of the hedging action, the macro-scenario belief state, and a preset risk penalty mechanism, and update the parameters of the reinforcement learning agent; In this embodiment, step S4 drives the reinforcement learning agent to optimize the hedging strategy through the design of multi-dimensional reward signals and dynamic risk penalty mechanism. The specific implementation is as follows: The reward signal It consists of three parts: portfolio value change, transaction costs, and scenario-aware risk penalty. Its calculation formula is: ; Where, For the moment reward signal; is the change in portfolio value; For transaction costs; For the past The profit and loss variance of the step; Indicates the past Rolling variance calculation for time steps; A risk penalty weight based on the macro-scenario belief state.
[0050] The risk penalty weight Determine this by following these steps: Basic weight preset: for each implicit scenario Assigning base risk weights ,For example: High volatility scenario 1=0.5 (high risk penalty); Low volatility scenario 2 = 0.2 (low risk penalty).
[0051] Dynamic weight synthesis: the macro belief state Multiply by the base weight vector to get the dynamic weight: ; Where, is the dynamic risk penalty weight; For the The default base risk weights associated with each implicit scenario; For the moment The market is in The probability of a hidden scenario; is the total number of hidden scenarios preset in the hidden Markov model.
[0052] Through step S5, the present invention can automatically increase the risk penalty intensity in macro-risk scenarios (such as high volatility periods), forcing the agent to adopt a conservative strategy; and reduce the penalty in stable scenarios, allowing the strategy to be moderately aggressive.
[0053] The agent uses the proximal policy optimization (PPO) algorithm to update the network parameters. The specific process is as follows: Experience Collection: Storing Trajectory Data To the experience replay pool; Odds Estimate: Calculate the generalized odds estimate (GAE) using the formula: ; Where, For the moment The generalized advantage estimate of , which quantifies how good the current action is relative to the average policy; is a discount factor used to adjust the weight of future rewards. A smaller value indicates more emphasis on recent gains. is the GAE smoothing coefficient, which is used to balance bias and variance. A larger value indicates greater reliance on multi-step returns. For the moment The timing difference error of The total duration of the round.
[0054] Policy optimization: Update the actor network parameters by maximizing the clipping objective function , the formula is: ; Where, is the objective function of the clipping strategy, which is used to constrain the strategy update range; is the current policy network parameter Select Action probability; The old policy network parameters Next select action probability; is the clipping threshold; is the clipping function; Expressing the moment Expected from empirical data.
[0055] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A dynamic hedging strategy optimization method based on reinforcement learning, characterized by: The following steps are involved: S1. Obtain the market observation data for the current time step and the hedging action performed by the reinforcement learning agent in the previous time step; S2. Inputting the market observation data and the hedging action of the previous time step into the market scenario inference model, and outputting the probability distribution of the current market being in multiple preset implicit scenarios as the macro-scenario belief state; S3. Combining the macro-scenario belief state with the micro-market state at the current time step into an enhanced state vector, inputting the vector into the reinforcement learning agent to generate a hedging action at the current time step; S4. Based on the execution result of the hedging action, the macro-scenario belief state and the preset risk penalty mechanism, calculate the reward signal and update the parameters of the reinforcement learning agent.
2. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 1, characterized in that: The step S1 comprises the following steps: The market observation data includes the underlying asset price, realized volatility and trading volume; The hedging impulse at the previous time step is used as the target hedging position ratio output by the reinforcement learning agent at the previous time step; This ratio is enforced by: If it is a live transaction, the target position ratio is converted into actual trading instructions through the order management system and sent to the exchange; If it is simulated trading, the target position ratio is recorded in the virtual trading environment and the simulated portfolio status is updated.
3. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 2 is characterized in that: The market observation data is obtained in the following ways: Underlying asset price: Obtain the latest transaction price of the underlying asset published by the exchange in real time through the financial data interface; Realized Volatility: Calculates the standard deviation of the logarithmic return based on the underlying asset's historical price series within a preset time window. Trading Volume: Obtain the total trading volume of the underlying asset in the latest time step through the market data subscription service.
4. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 3 is characterized in that: The calculation time window for the realized volatility is 30 minutes, and the calculation formula for the logarithmic return is: ; Where, For the moment The logarithmic rate of return; For the moment The price of the underlying asset; For the moment The price of the underlying asset.
5. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 4 is characterized in that: The step 2 comprises the following steps: The market scenario inference model is a hidden Markov model, and its observation value is an augmented observation tuple consisting of the market observation data and the hedging action of the previous time step; The macro-scenario belief state is calculated by a forward algorithm, specifically: ; Where, For the moment The market is in The posterior probability of each implicit scenario; For the moment The hidden state variables of is defined in the hidden Markov model. a hidden market scenario; is the augmented observation tuple; For the moment market observation data.
6. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 5 is characterized in that: The Hidden Markov Model is pre-trained by the following steps: Collect historical market data and corresponding historical hedging actions to construct an augmented observation sequence; The Baum-Welch algorithm is used to optimize the model parameters to maximize the likelihood probability of the augmented observation sequence.
7. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 6 is characterized in that: The step S3 comprises the following steps: The micro-market conditions include the price of the underlying asset, the remaining maturity of the derivative, the current hedge position and market volatility; The enhanced state vector is generated by vector concatenation, specifically: ; Where, For the moment The enhanced state vector of is the micro-market state vector, which includes the underlying asset price, remaining time to maturity, current position ratio and market volatility; is the macro-scenario belief state vector.
8. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 7 is characterized in that: The actor network of the reinforcement learning agent adopts a gated structure, where: The macroscopic scenario belief state vector generates a gating signal through a fully connected layer; The gating signal is used to modulate the characteristic expression of the micro-market state in the strategy network.
9. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 8, characterized in that: The step S4 comprises the following steps: The calculation formula of the reward signal is: ; Where, For the moment reward signal; is the change in portfolio value; For transaction costs; For the past The profit and loss variance of the step; Indicates the past Rolling variance calculation for time steps; A risk penalty weight based on the macro-scenario belief state.
10. The dynamic hedging strategy optimization method based on reinforcement learning according to claim 9, characterized in that: The risk penalty weight is calculated as follows: Preset base weights for each implicit scenario ; Multiply the macro-scenario belief state vector by the basic weight vector to obtain the dynamic weight: ; Where, is the dynamic risk penalty weight; For the The default base risk weights associated with each implicit scenario; For the moment The market is in the probability of a hidden scenario; is the total number of hidden scenarios preset in the hidden Markov model.