A strategy evaluation system and risk control method based on value distribution environment model
Through a strategy evaluation system based on the value distribution environment model, the problems of low efficiency of risk strategy evaluation and insufficient stability of results in the existing technology are solved, and a fast, stable and reliable strategy evaluation effect is achieved.
Patent Information
- Application Number
- CN202411933299.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing risk strategy assessment methods have problems such as low efficiency in strategy assessment, insufficient results stability and reliability.
A strategy evaluation system based on value distribution environment model is proposed, including screening offline data modules, reward value distribution model building modules based on value distribution, state transfer model building modules, state sequence generation modules and strategy evaluation modules. The model is trained through deep neural networks and gradient backpropagation algorithms, and state sequences are generated and their benefits are evaluated.
It realizes the rapid, stable and reliable evaluation of the effect of intelligent risk control strategies, and improves the accuracy and stability of strategy evaluation.
Smart Images

Figure CN119377624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a strategy evaluation system and a risk management method based on a value distribution environment model. Background Art
[0002] Risks are unavoidable in any field that requires strategic decision-making, but effective risk management can reduce the impact of risks. In recent years, with the rapid development of big data and artificial intelligence technologies, intelligent risk management and control technologies have developed rapidly, and a variety of risk management and control technologies based on rules, machine learning methods, deep learning methods, and other different methods have emerged. How to accurately evaluate the actual effects that various risk management and control methods can achieve in actual scenarios has become an important issue that needs to be solved in the current risk strategy selection.
[0003] At present, the methods of risk strategy evaluation mainly include policy verification methods based on real environments, policy evaluation methods based on virtual environments, policy evaluation methods based on historical data, and policy evaluation methods based on importance sampling. However, the policy verification method based on real environments needs to be carried out in the actual environment, the verification process is time-consuming, and there is a potential risk of damage or loss; the policy evaluation method based on virtual environments is based on virtual environments and cannot truly reflect the status and situation in the actual environment, resulting in low reliability of policy evaluation results; the policy evaluation method based on historical data mainly relies on past environmental data and cannot consider the impact or intervention of its own state changes or behaviors on the environment, resulting in insufficient stability of policy evaluation. Due to data sparsity, the policy evaluation method based on importance sampling has too large a variance in policy evaluation, and the evaluation results are not stable and reliable enough.
[0004] In order to quickly, stably and reliably evaluate the effect of intelligent risk management and control strategies in scenarios and promote the rapid application and deployment of intelligent risk management and control strategies, a strategy evaluation system and risk management method based on value distribution environment model are proposed. Summary of the invention
[0005] The embodiments of the present invention propose a strategy evaluation system and a risk management method based on a value distribution environment model, so as to at least solve the problems of low strategy evaluation efficiency, insufficient result stability and reliability in current strategy evaluation methods.
[0006] According to one embodiment of the present invention, a policy evaluation system based on a value distribution environment model is proposed, comprising:
[0007] Offline data screening module: Screening offline data according to the returns of historical strategies and / or the specificity of historical strategies and generating offline data sets according to the quadruple data format;
[0008] Reward value distribution model construction module based on value distribution: establish a loss function based on value distribution learning and a four-tuple offline data set, and build a reward value distribution model based on value distribution according to the loss function;
[0009] State transfer model building module: uses deep neural network and gradient back propagation algorithm to train the state transfer model based on the four-tuple offline data set;
[0010] State sequence generation module: generates state sequence according to reward value distribution model and state transition model;
[0011] Strategy evaluation module: Evaluate the benefits of the state sequence according to the reward value and / or sequence length and / or model error, and obtain the strategy evaluation result according to the benefits of the state sequence.
[0012] In an exemplary embodiment, the filtering of offline data according to the returns of historical strategies and / or the specificity of historical strategies includes:
[0013] Calculate the return evaluation value of the historical strategy based on the positive return of the historical strategy and / or the return growth rate of the historical strategy;
[0014] Calculating a specificity evaluation value of the historical strategy according to a repetition rate of the historical strategy and / or a similarity of the historical strategy;
[0015] Calculating the historical strategy weight value according to the historical strategy return evaluation value and / or the historical strategy specificity evaluation value;
[0016] Filter offline data based on the sorting of historical policy weight values and the preset policy weight threshold.
[0017] In an exemplary embodiment, the four-tuple data format is: t 、Behavior a t , the next moment state s t+1 , Risk R t >.
[0018] In an exemplary embodiment, the establishing a loss function based on value distribution learning and a four-tuple offline data set includes:
[0019] Define a reward value distribution model; the input of the reward value distribution model is the state s at each moment t and behavior a t , the output is the reward R t Distribution of
[0020] Calculate the cumulative distribution function of the reward function and use it to obtain the quantile probability distribution function;
[0021] The loss function is established based on the difference between the model prediction value and the true value.
[0022] In an exemplary embodiment, constructing a reward value distribution model based on value distribution according to a loss function includes:
[0023] Model the reward value distribution model as a neural network based on the loss function;
[0024] The parameters are updated through the gradient descent method of the loss function to obtain the estimated values of different quantiles of the reward distribution function, thereby obtaining a reward value distribution model based on the value distribution.
[0025] In an exemplary embodiment, the training of the state transition model according to the four-tuple offline data set includes:
[0026] Define a state transition model; the input of the state transition model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 ;
[0027] Construct an optimization model based on the maximum likelihood estimation method;
[0028] Deep neural network and gradient back propagation algorithm are used to optimize the model parameters to obtain the state transfer model.
[0029] In an exemplary embodiment, generating a state sequence according to a reward value distribution model and a state transition model includes:
[0030] Step 1: According to the state s t Randomly sample actions a from the policy function t ~ π(s t );
[0031] Step 2: Set the state s t and behavior a t Substitute into the reward value distribution model to obtain the estimated values of different quantiles of the reward function;
[0032] Step 3: Set the state s t and behavior a t Substitute into the state transition model to get the next moment state s t+1 ;
[0033] Step 4, arranging the four-tuple data in chronological order to form a sequence;
[0034] Step 5: Repeat steps 1 to 4 to generate a state sequence of length k.
[0035] In an exemplary embodiment, the step of evaluating the benefit of a state sequence according to a reward value and / or a sequence length and / or a model error comprises:
[0036] The reward weight value is calculated based on the positive correlation between the estimated values of different quantiles of the reward function and the benefits;
[0037] The sequence length weight value is calculated based on the positive correlation between sequence length and return;
[0038] Calculate the model error weight value according to the error of the reward value distribution model and the state transition model;
[0039] Calculate the benefit evaluation value of the state sequence according to the reward weight value and / or the sequence length weight value and / or the model error weight value;
[0040] Evaluate the payoff of a state sequence based on the payoff evaluation value.
[0041] In an exemplary embodiment, obtaining a strategy evaluation result according to the benefits of the state sequence includes:
[0042] Calculate the return evaluation value of the state sequence at different quantiles of the strategy;
[0043] The strategy evaluation results are obtained by comparing the return evaluation values of the state sequences at different quantiles of the strategy.
[0044] According to another embodiment of the present invention, a risk management method based on a value distribution environment model is provided, characterized in that it includes:
[0045] Determine the risk management strategies to be selected;
[0046] Input the historical data of the candidate risk control strategy into the strategy evaluation system based on the value distribution environment model to obtain the strategy evaluation result;
[0047] Select and execute the optimal risk management strategy based on the strategy evaluation results.
[0048] The advantages of a strategy evaluation system and risk management method based on a value distribution environment model of the present invention are:
[0049] (1) The weight value of each historical strategy is calculated based on the positive returns of the historical strategy and / or the return growth rate of the historical strategy and the repetition rate of the historical strategy and / or the similarity of the historical strategy, and the offline data set is screened accordingly. Compared with traditional risk management and strategy evaluation technical solutions, data with high strategy coverage and high historical returns can be selected without interacting with the actual environment, which facilitates the subsequent rapid and accurate construction of environmental models and strategy evaluation.
[0050] (2) Based on value distribution learning and the four-tuple offline data set, a loss function is established to construct a reward value distribution model. Compared with the traditional environmental model technology solution, it can accurately characterize the reward function valuation under different quantiles, making the evaluation results more robust and able to adapt to different strategy performances in various situations.
[0051] (3) The state transition model is introduced and the behavior of the strategy executor is considered in the model input. Compared with traditional risk management and strategy evaluation technical solutions, it can fully consider the interaction between the strategy executor's behavior and the environment, better simulate the actual environment operation, and improve the accuracy and stability of strategy evaluation.
[0052] (4) The benefits of the state sequence are evaluated based on the reward value and / or sequence length and / or model error to obtain the strategy evaluation result. Compared with the traditional strategy evaluation technical solution, it can comprehensively consider the reward, step length and model error to obtain a more stable and reliable evaluation result, thus avoiding significant differences between the strategy evaluation and the actual results. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a structural schematic diagram of a strategy evaluation system based on a value distribution environment model according to an embodiment of the present invention;
[0054] Figure 2 is a flow chart of a strategy evaluation system according to an embodiment of the present invention for filtering offline data according to the benefits of historical strategies and / or the specificity of historical strategies;
[0055] Figure 3 is a flow chart of establishing a loss function based on value distribution learning and a four-tuple offline data set in a strategy evaluation system according to an embodiment of the present invention;
[0056] Figure 4 is a flow chart of constructing a reward value distribution model based on value distribution according to a loss function in a strategy evaluation system according to an embodiment of the present invention;
[0057] Figure 5 is a flow chart of a state transition model trained according to a four-tuple offline data set in a strategy evaluation system according to an embodiment of the present invention;
[0058] Figure 6 is a flow chart of a strategy evaluation system according to an embodiment of the present invention generating a state sequence according to a reward value distribution model and a state transition model;
[0059] Figure 7 It is a flow chart of evaluating the benefits of a state sequence according to a reward value and / or a sequence length and / or a model error of a strategy evaluation system according to an embodiment of the present invention;
[0060] Figure 8 It is a flow chart of obtaining a strategy evaluation result according to the benefits of a state sequence by a strategy evaluation system according to an embodiment of the present invention;
[0061] Fig. 9 It is a flow chart of a risk management method based on a value distribution environment model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0063] The strategy evaluation system based on the value distribution environment model of the present invention is applicable to scenarios including but not limited to financial decision-making, autonomous driving obstacle identification and warning, behavioral risk identification, climate monitoring and warning, weather forecasting and other risk strategy evaluations in areas that require observation and decision-making.
[0064] Taking financial decision making as an example, a strategy evaluation system based on a value distribution environment model according to an embodiment of the present invention is shown in the structural diagram as follows: Figure 1 As shown, including:
[0065] Offline data screening module: Screening offline data according to the returns of historical strategies and / or the specificity of historical strategies and generating offline data sets according to the quadruple data format;
[0066] Reward value distribution model construction module based on value distribution: establish a loss function based on value distribution learning and a four-tuple offline data set, and build a reward value distribution model based on value distribution according to the loss function;
[0067] State transfer model building module: uses deep neural network and gradient back propagation algorithm to train the state transfer model based on the four-tuple offline data set;
[0068] State sequence generation module: generates state sequence according to reward value distribution model and state transition model;
[0069] Strategy evaluation module: Evaluate the benefits of the state sequence according to the reward value and / or sequence length and / or model error, and obtain the strategy evaluation result according to the benefits of the state sequence.
[0070] The offline data is screened according to the returns of the historical strategies and / or the specificity of the historical strategies. The flowchart is as follows: Figure 2 As shown, the steps include:
[0071] Calculate the return evaluation value of the historical strategy based on the positive return of the historical strategy and / or the return growth rate of the historical strategy;
[0072] Calculating a specificity evaluation value of the historical strategy according to a repetition rate of the historical strategy and / or a similarity of the historical strategy;
[0073] Calculating the historical strategy weight value according to the historical strategy return evaluation value and / or the historical strategy specificity evaluation value;
[0074] Filter offline data based on the sorting of historical policy weight values and the preset policy weight threshold.
[0075] In this embodiment, historical strategy data is collected in a real market environment, and the profit evaluation value of the historical strategy is calculated according to the positive profit of the historical strategy and / or the profit growth rate of the historical strategy, which is: calculating the profit evaluation value of the historical strategy according to the positive correlation between the positive profit of the historical strategy and the profit evaluation value of the historical strategy, calculating the profit evaluation value of the historical strategy according to the positive correlation between the profit growth rate of the historical strategy and the profit evaluation value of the historical strategy, or calculating the profit evaluation value of the historical strategy according to the positive correlation between the positive profit of the historical strategy and the profit growth rate of the historical strategy and the profit evaluation value of the historical strategy, and any one of the profit evaluation values of the historical strategy is represented by the variable a;
[0076] The specific evaluation value of the historical strategy is calculated according to the repetition rate of the historical strategy and / or the similarity of the historical strategy, which is: calculating the specific evaluation value of the historical strategy according to the negative correlation between the repetition rate of the historical strategy and the specific evaluation value of the historical strategy, calculating the specific evaluation value of the historical strategy according to the negative correlation between the similarity of the historical strategy and the specific evaluation value of the historical strategy, or calculating the specific evaluation value of the historical strategy according to the negative correlation between the repetition rate of the historical strategy and the similarity of the historical strategy and the specific evaluation value of the historical strategy, wherein the specific evaluation value of the historical strategy is represented by a variable b;
[0077] The calculation of the historical strategy weight value based on the historical strategy benefit evaluation value and / or the historical strategy specificity evaluation value is calculated based on the positive correlation between the historical strategy weight value and the historical strategy benefit evaluation value and / or the historical strategy specificity evaluation value, and the historical strategy weight value is represented by the variable v.
[0078] Different implementation methods for calculating the historical strategy weight value are shown by using Examples A1 to A3 as follows:
[0079] Embodiment A1: Calculate the historical strategy weight value according to the benefit evaluation value of the historical strategy.
[0080] Collect the return value of the strategy in a historical period of time in the environment, calculate the positive return of the historical strategy and / or the return growth rate of the historical strategy, and calculate the return evaluation value u of the historical strategy based on the positive correlation between the positive return of the historical strategy and / or the return growth rate of the historical strategy and the return evaluation value of the historical strategy; calculate the weight value v of the historical strategy based on the positive correlation between the return evaluation value u of the historical strategy and the weight value v of the historical strategy. In a preferred implementation, the weight value v of the historical strategy is calculated as e1·u e2+e3, where e1, e2 (e1·e2>0), and e3 are calculation coefficients obtained through prior training. In this embodiment, the return values of the strategies in a historical period of time in the market environment are collected, and the average positive return is calculated to be 0.8 (normalized according to the preset return threshold), and the return evaluation value of the historical strategy is calculated to be u=0.8 based on the positive correlation between the positive return of the historical strategy and the return evaluation value of the historical strategy. The calculation coefficients obtained through prior training are e1=1, e2=1, and e3=0, and the weight value v=e1·u of the historical strategy is calculated. e2 +e3 =1×0.8+0=0.8.
[0081] Embodiment A2: Calculate the historical strategy weight value according to the specific evaluation value of the historical strategy.
[0082] The number of repetitions or repetition ratios of the strategies in the historical environment within a period of time are collected, and the repetition rate and / or similarity of the historical strategies are calculated. The specific evaluation value b of the historical strategies is calculated based on the negative correlation between the repetition rate and / or similarity of the historical strategies and the specific evaluation value of the historical strategies; the weight value v of the historical strategies is calculated based on the positive correlation between the specific evaluation value b of the historical strategies and the weight value v of the historical strategies. In a preferred implementation, the weight value v of the historical strategies is calculated as e4·b e5 +e6, where e4, e5 (e4·e5>0), and e6 are calculation coefficients obtained through prior training. In this embodiment, the number of repetitions of a strategy in a historical period of time in the market environment is 5, and the repetition rate of the historical strategy is calculated to be 0.5 (calculated based on the ratio of the number of repetitions to the preset strategy repetition threshold), and the specificity of the historical strategy is calculated based on the negative correlation between the repetition rate of the historical strategy and the specificity evaluation value b of the historical strategy. The specificity of the historical strategy is calculated to be b=0.45 / 0.5=0.9 (where 0.45 is the calculation coefficient obtained through prior training), and the calculation coefficients obtained through prior training are e4=1, e5=1, and e6=0, and the weight value v=e4·b of the historical strategy is calculated. e5 +e6=1×0.9+0=0.9.
[0083] Embodiment A3: Calculate the historical strategy weight value according to the historical strategy benefit evaluation value and the historical strategy specificity evaluation value.
[0084] Collect the return value of the strategy in a historical period of time in the environment, calculate the positive return of the historical strategy and / or the return growth rate of the historical strategy, and calculate the return evaluation value u of the historical strategy based on the positive correlation between the positive return of the historical strategy and / or the return growth rate of the historical strategy and the return evaluation value of the historical strategy; collect the number of repetitions or repetition ratios of the strategy in a historical period of time in the environment, calculate the repetition rate of the historical strategy and / or the similarity of the historical strategy, and calculate the specificity evaluation value b of the historical strategy based on the negative correlation between the repetition rate of the historical strategy and / or the similarity of the historical strategy and the specificity evaluation value of the historical strategy; calculate the weight value v of the historical strategy based on the positive correlation between the return evaluation value u of the historical strategy and the specificity evaluation value b of the historical strategy and the weight value v of the historical strategy. In a preferred embodiment, the weight value v of the historical strategy is calculated as e7·u e8 +e9·b e10 +e11, where e7, e8, e9, e10, and e11 are calculation coefficients obtained by prior training. In this embodiment, the return value of the strategy in a historical period of time in the market environment is collected, and the average positive return is calculated to be 0.8 (normalized according to the preset return threshold), and the return evaluation value u of the historical strategy is calculated based on the positive correlation between the positive return of the historical strategy and the return evaluation value of the historical strategy. 0.8; the number of repetitions of a certain strategy in a historical period of time in the market environment is 5, and the repetition rate of the historical strategy is calculated to be 0.5 (calculated based on the ratio of the number of repetitions to the preset strategy repetition threshold), and the specificity of the historical strategy is calculated based on the negative correlation between the repetition rate of the historical strategy and the specificity evaluation value b of the historical strategy. b=0.45 / 0.5=0.9 (where 0.45 is the calculation coefficient obtained by prior training), the calculation coefficients obtained by prior training are e7=0.6, e8=1, e9=0.4, e10=1, and e11=0, and the weight value v=e7·u of the historical strategy is calculated. e8 +e9·b e10 +e11=0.6×0.8+0.4×0.9+0=0.84. In another preferred implementation, the weight value v of the historical strategy is calculated as: e13 b e14+e15, where e12 (e12>0), e13 (e13>0), e14 (e14>0), and e15 are calculation coefficients obtained by prior training. In this embodiment, the return value of the strategy in a historical period of time in the market environment is collected, and the average positive return is calculated to be 0.8 (normalized according to the preset return threshold), and the return evaluation value u=0.8 of the historical strategy is calculated based on the positive correlation between the positive return of the historical strategy and the return evaluation value of the historical strategy; the number of repetitions of a certain strategy in a historical period of time in the market environment is 5, and the repetition rate of the historical strategy is calculated to be 0.5 (calculated based on the ratio of the number of repetitions to the preset strategy repetition threshold), and the specificity of the historical strategy is calculated based on the negative correlation between the repetition rate of the historical strategy and the specificity evaluation value b of the historical strategy. b=0.45 / 0.5=0.9 (where 0.45 is the calculation coefficient obtained by prior training), the calculation coefficients e12=1.2, e13=1, e14=1, and e15=0 obtained by prior training, and the weight value v=e12·u of the historical strategy is calculated e13 b e14 +e15=1.2×0.8×0.9+0=0.864.
[0085] A strategy weight threshold V is set in advance according to the market environment and historical return level, and the historical strategy weight values calculated by the method described in any of embodiments A1 to A3 are sorted, and the data corresponding to the historical strategies whose historical strategy weight values are greater than the strategy weight threshold V are screened out as offline data.
[0086] The four-tuple data format is: <state s t 、Behavior a t , the next moment state s t+1 , Risk R t >.
[0087] In this embodiment, the offline data set includes a total of four-tuple data at multiple time points (set as T time points in this embodiment), and the definition and specific content of each data element are as follows:
[0088] Time t: according to the market environment requirements of the data, set every day, every hour, every minute, every second, etc. as a time t;
[0089] Status t : The collected environmental status at each time t. The status may include the on-site data and off-site data of the day. The on-site data includes but is not limited to the opening price, closing price, highest price, lowest price, fundamental information, and financial report information. The off-site data includes but is not limited to weather, policy information, bank interest rates, gold prices, oil prices, etc.
[0090] Behavior t: The relevant behavior or action taken at each time t in the environment. The behavior can include a combination of the stock traded, the transaction price, and the transaction volume.
[0091] Risk R t :Risk is defined as the value at risk, which means the maximum expected loss within a given confidence level and a certain time limit.
[0092] The loss function is established based on value distribution learning and four-tuple offline data set. The flowchart is as follows Figure 3 As shown, the steps include:
[0093] Define a reward value distribution model; the input of the reward value distribution model is the state s at each moment t and behavior a t , the output is the reward R t Distribution of
[0094] Calculate the cumulative distribution function of the reward function and use it to obtain the quantile probability distribution function;
[0095] The loss function is established based on the difference between the model prediction value and the true value.
[0096] In this embodiment, the reward value distribution model is trained based on the collected four-tuple offline data set. The input of the reward value distribution model is the state s at each moment t and behavior a t , the output is the reward R t The distribution of , Since the reward distribution is unknown, an approximate value distribution learning is adopted.
[0097] First, the cumulative distribution function of the reward function is . Then the quantile probability distribution function (τth quantile) is as shown in formula (1).
[0098] (1)
[0099] Among them, inf is the infimum function, and the quantile τ∈(0,1).
[0100] The quantile refers to the numerical point that divides the probability distribution range of a random variable into several equal parts, and the commonly used ones are the median (i.e., quantile), quartile, percentile, etc. For example, the quartile represents arranging the values of all possible reward functions from small to large and dividing them into four equal parts, and the probability values of obtaining rewards are less than 25%, 50%, and 75% respectively.
[0101] Secondly, define the loss function as shown in formula (2).
[0102] (2)
[0103] in is an indicative function.
[0104] Then the loss function of the τth quantile can be expressed as the function shown in formula (3).
[0105] (3)
[0106] Where r represents the true reward function value, represents the predicted value of the model, That is, when the model prediction exceeds the true value, the loss is the difference between the predicted value and the true value multiplied by τ; when the predicted value is lower than the true value, the loss is the difference between the predicted value and the true value multiplied by 1-τ.
[0107] The reward value distribution model based on value distribution is constructed according to the loss function. The flow chart is as follows Figure 4 As shown, the steps include:
[0108] Model the reward value distribution model as a neural network based on the loss function;
[0109] The parameters are updated through the gradient descent method of the loss function to obtain the estimated values of different quantiles of the reward distribution function, thereby obtaining a reward value distribution model based on the value distribution.
[0110] In this embodiment, according to the loss function established in the above steps, the reward model can be modeled as a neural network, and the parameters are updated by the gradient descent method of the loss function to obtain estimates of different quantiles of the reward function. Compared with the traditional method of optimizing the model by mean square error, quantile regression can obtain estimates of different quantiles, so it can be more robust to abnormal points.
[0111] The state transfer model is trained according to the four-tuple offline data set, and the flow chart is as follows Figure 5 As shown, the steps include:
[0112] Define a state transition model; the input of the state transition model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 ;
[0113] Construct an optimization model based on the maximum likelihood estimation method;
[0114] Deep neural network and gradient back propagation algorithm are used to optimize the model parameters to obtain the state transfer model.
[0115] In this embodiment, the state transition model T is trained based on the collected four-tuple offline data set. The input is the state s at each moment t and behavior a t , the output is the next moment state s t+1 , Risk R t , according to the maximum likelihood estimation, the model is optimized using the formula shown in formula (4):
[0116] (4)
[0117] In practical applications, a deep neural network can be used as a training state transition model, and the parameters in the neural network can be optimized through gradient back propagation. The likelihood function T(·) uses a Gaussian function, and its logarithm is the mean square error.
[0118] The state sequence is generated according to the reward value distribution model and the state transition model. The flow chart is as follows Figure 6 As shown, including:
[0119] Step 1: According to the state s t Randomly sample actions a from the policy function t ~ π(s t );
[0120] Step 2: Set the state s t and behavior a t Substitute into the reward value distribution model to obtain the estimated values of different quantiles of the reward function;
[0121] Step 3: Set the state s t and behavior a t Substitute into the state transition model to get the next moment state s t+1 ;
[0122] Step 4, arranging the four-tuple data in chronological order to form a sequence;
[0123] Step 5: Repeat steps 1 to 4 to generate a state sequence of length k.
[0124] In this embodiment, according to the randomly sampled initial state s0, behaviors a0~π(s0) are randomly sampled from the strategy function, and the initial state s0 and behavior a0 are substituted into the reward value distribution model (3) to obtain the estimated values r of different quantiles of the reward function: τ , that is, R t; Input the initial state s0 and behavior a0 into the state transition model formula (4) to obtain the next state s1; obtain the first set of four-tuple data <state s0, behavior a0, next state s1, risk R1>, which constitute the first set of data in the sequence; repeat the above steps k times to obtain a state sequence with a length of k starting from the initial state, where k is a variable representing the length of the state sequence.
[0125] The benefit of the state sequence is evaluated according to the reward value and / or sequence length and / or model error. The flowchart is as follows Figure 7 As shown, the steps include:
[0126] The reward weight value is calculated based on the positive correlation between the estimated values of different quantiles of the reward function and the benefits;
[0127] The sequence length weight value is calculated based on the positive correlation between sequence length and return;
[0128] Calculate the model error weight value according to the error of the reward value distribution model and the state transition model;
[0129] Calculate the benefit evaluation value of the state sequence according to the reward weight value and / or the sequence length weight value and / or the model error weight value;
[0130] Evaluate the payoff of a state sequence based on the payoff evaluation value.
[0131] In this embodiment, the greater the estimated values of different quantiles of the reward function (i.e., the reward values of different quantiles), the greater the benefit, that is, the greater the reward weight value. Therefore, the reward weight value is calculated based on the positive correlation between the estimated values of different quantiles of the reward function and the reward weight value, and the reward weight value is represented by the variable x.
[0132] The larger the sequence length, the greater the overall benefit, that is, the larger the sequence length weight value. Therefore, the sequence length weight value is calculated based on the positive correlation between the sequence length and the sequence length weight value, and the sequence length weight value is represented by the variable y.
[0133] The smaller the error between the reward value distribution model and the state transition model, the greater the benefit, that is, the larger the model error weight value. Therefore, the model error weight value is calculated based on the negative correlation between the error between the reward value distribution model and the state transition model and the model error weight value. The model error weight value is represented by the variable z.
[0134] The method of calculating the benefit evaluation value of the state sequence based on the reward weight value and / or the sequence length weight value and / or the model error weight value is to calculate the benefit evaluation value of the state sequence based on the positive correlation between the reward weight value and / or the sequence length weight value and / or the model error weight value and the benefit evaluation value, and the benefit evaluation value of the state sequence is represented by the variable p.
[0135] Different implementation methods of calculating the benefit evaluation value of a state sequence are shown by using Examples B1 to B7 as follows:
[0136] Embodiment B1: The benefit evaluation value of the state sequence is calculated based on the positive correlation between the reward weight value and the benefit evaluation value of the state sequence.
[0137] Specifically, the reward weight value x is calculated based on the positive correlation between the estimated value of different quantiles of the reward function and the reward weight value; the benefit evaluation value p of the state sequence is calculated based on the positive correlation between the reward weight value x and the benefit evaluation value of the state sequence. In a preferred embodiment, the benefit evaluation value p of the state sequence is calculated as w1·x w2 +w3, where w1 (w1>0), w2 (w2>0), and w3 are the calculation coefficients obtained by prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained by prior training is 1), and the calculation coefficients obtained by prior training w1=1, w2=1, and w3=0, then the profit evaluation value of the calculated state sequence p=w1·x w2 +w3=1×2+0=2.
[0138] Embodiment B2: The benefit evaluation value of the state sequence is calculated based on the positive correlation between the sequence length weight value and the benefit evaluation value of the state sequence.
[0139] Specifically, the sequence length weight value y is calculated based on the positive correlation between the sequence length and the sequence length weight value; the state sequence benefit evaluation value p is calculated based on the positive correlation between the sequence length weight value y and the benefit evaluation value of the state sequence. In a preferred embodiment, the benefit evaluation value p of the state sequence is calculated as w4·y w5 +w6, where w4 (w4>0), w5 (w5>0), and w6 are the calculation coefficients obtained by prior training. In this embodiment, the length of a state sequence is 3, and the sequence length weight value y=3 (the calculation coefficient obtained by prior training is 1) is calculated based on its positive correlation with the sequence length weight value. The calculation coefficients obtained by prior training are w4=0.8, w5=1, and w6=0, and the profit evaluation value of the calculated state sequence is p=w4·y w5 +w6=0.8×3+0=2.4.
[0140] Embodiment B3: The benefit evaluation value of the state sequence is calculated based on the positive correlation between the model error weight value and the benefit evaluation value of the state sequence.
[0141] Specifically, the model error weight value z is calculated based on the negative correlation between the error of the reward value distribution model and the state transition model and the model error weight value; the benefit evaluation value p of the state sequence is calculated based on the positive correlation between the model error weight value z and the benefit evaluation value of the state sequence. In a preferred embodiment, the benefit evaluation value p of the state sequence is calculated as w7·z w8 +w9, where w7 (w7>0), w8 (w8>0), and w9 are the calculation coefficients obtained by prior training. In this embodiment, the average error of the reward value distribution model and the state transition model obtained by training is 0.1. According to its negative correlation with the model error weight value, the model error weight value z=0.5 / 0.1=5 (the calculation coefficient obtained by prior training is 1), and the calculation coefficients obtained by prior training are w7=0.4, w8=1, and w9=0, then the profit evaluation value of the calculated state sequence p=w7·z w8 +w9=0.4×5+0=2.
[0142] Embodiment B4: The benefit evaluation value of the state sequence is calculated according to the positive correlation between the reward weight value and the sequence length weight value and the benefit evaluation value of the state sequence.
[0143] Specifically, the reward weight value x is calculated based on the positive correlation between the estimated value of different quantiles of the reward function and the reward weight value; the sequence length weight value y is calculated based on the positive correlation between the sequence length and the sequence length weight value; the profit evaluation value p of the state sequence is calculated based on the positive correlation between the reward weight value x and the sequence length weight value y and the profit evaluation value of the state sequence. In a preferred embodiment, the profit evaluation value p of the state sequence is calculated as w10·x w11 +w12·y w13 +w14, where w10 (w10>0), w11 (w11>0), w12 (w12>0), w13 (w13>0), and w14 are calculation coefficients obtained through prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained through prior training is 1); the length of a certain state sequence is 3, and the sequence length weight value y=3 is calculated based on its positive correlation with the sequence length weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w10=0.6, w11=1, w12=0.3, w13=1, and w14=0, then the profit evaluation value p=w10·x is calculated for the state sequence w11 +w12·y w13 +w14=0.6×2+0.3×3+0=2.1. In another preferred embodiment, the profit evaluation value p of the state sequence is calculated as: w16 ·y w17+w18, where w15 (w15>0), w16 (w16>0), w17 (w17>0), and w18 are calculation coefficients obtained through prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained through prior training is 1); the length of a certain state sequence is 3, and the sequence length weight value y=3 is calculated based on its positive correlation with the sequence length weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w15=0.4, w16=1, w17=1, and w18=0, then the profit evaluation value p=w15·x is calculated for the state sequence w16 ·y w17 +w18=0.4×2×3+0=2.4.
[0144] Embodiment B5: The benefit evaluation value of the state sequence is calculated based on the positive correlation between the reward weight value and the model error weight value and the benefit evaluation value of the state sequence.
[0145] Specifically, the reward weight value x is calculated based on the positive correlation between the estimated value of different quantiles of the reward function and the reward weight value; the model error weight value z is calculated based on the negative correlation between the error of the reward value distribution model and the state transition model and the model error weight value; the benefit evaluation value p of the state sequence is calculated based on the positive correlation between the reward weight value x and the model error weight value z and the benefit evaluation value of the state sequence. In a preferred embodiment, the benefit evaluation value p of the state sequence is calculated as w19·x w20 +w21·z w22 +w23, where w19 (w19>0), w20 (w20>0), w21 (w21>0), w22 (w22>0), and w23 are calculation coefficients obtained through prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained through prior training is 1); the average error of the reward value distribution model and the state transition model obtained through training is 0.1, and the model error weight value z=0.5 / 0.1=5 is calculated based on its negative correlation with the model error weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w19=0.6, w20=1, w21=0.2, w22=1, and w23=0, then the profit evaluation value of the calculated state sequence is p=w19·x w20 +w21·z w22 +w23=0.6×2+0.2×5+0=2.2. In another preferred embodiment, the profit evaluation value p of the state sequence is calculated as w24·x w25 ·z w26+w27, where w24 (w24>0), w25 (w25>0), w26 (w26>0), and w27 are calculation coefficients obtained through prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained through prior training is 1); the average error of the reward value distribution model and the state transition model obtained through training is 0.1, and the model error weight value z=0.5 / 0.1=5 is calculated based on its negative correlation with the model error weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w24=0.2, w25=1, w26=1, and w27=0, then the profit evaluation value of the calculated state sequence is p=w24·x w25 ·z w26 +w27=0.2×2×5+0=2.
[0146] Embodiment B6: The profit evaluation value of the state sequence is calculated based on the positive correlation between the sequence length weight value and the model error weight value and the profit evaluation value of the state sequence.
[0147] Specifically, the sequence length weight value y is calculated based on the positive correlation between the sequence length and the sequence length weight value; the model error weight value z is calculated based on the negative correlation between the error of the reward value distribution model and the state transition model and the model error weight value; the state sequence benefit evaluation value p is calculated based on the positive correlation between the sequence length weight value y and the model error weight value z and the benefit evaluation value of the state sequence. In a preferred embodiment, the benefit evaluation value p of the state sequence is calculated as w28·y w29 +w30·z w31 +w32, where w28 (w28>0), w29 (w29>0), w30 (w30>0), w31 (w31>0), and w32 are calculation coefficients obtained through prior training. In this embodiment, the length of a state sequence is 3, and the sequence length weight value y=3 is calculated based on its positive correlation with the sequence length weight value (the calculation coefficient obtained through prior training is 1); the average error of the reward value distribution model and the state transition model obtained through training is 0.1, and the model error weight value z=0.5 / 0.1=5 is calculated based on its negative correlation with the model error weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w28=0.4, w29=1, w30=0.2, w31=1, and w32=0, then the profit evaluation value p=w28·y is calculated for the state sequence. w29 +w30·z w31 +w32=0.4×3+0.2×5+0=2.2. In another preferred embodiment, the profit evaluation value of the state sequence is calculated as p=w33·y w34 ·z w35+w36, where w33 (w33>0), w34 (w34>0), w35 (w35>0), and w36 are calculation coefficients obtained through prior training. In this embodiment, the length of a state sequence is 3, and the sequence length weight value y=3 is calculated based on its positive correlation with the sequence length weight value (the calculation coefficient obtained through prior training is 1); the average error of the reward value distribution model and the state transition model obtained through training is 0.1, and the model error weight value z=0.5 / 0.1=5 is calculated based on its negative correlation with the model error weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w33=0.15, w34=1, w35=1, and w36=0, then the profit evaluation value p=w33·y is calculated for the state sequence w34 ·z w35 +w36=0.15×3×5+0=2.25.
[0148] Embodiment B7: The profit evaluation value of the state sequence is calculated based on the positive correlation between the reward weight value, the sequence length weight value, the model error weight value and the profit evaluation value of the state sequence.
[0149] Specifically, the reward weight value x is calculated based on the positive correlation between the estimated value of different quantiles of the reward function and the reward weight value; the sequence length weight value y is calculated based on the positive correlation between the sequence length and the sequence length weight value; the model error weight value z is calculated based on the negative correlation between the error of the reward value distribution model and the state transition model and the model error weight value; the profit evaluation value p of the state sequence is calculated based on the positive correlation between the reward weight value x, the sequence length weight value y, the model error weight value z and the profit evaluation value of the state sequence. In a preferred embodiment, the profit evaluation value p of the state sequence is calculated as w37·x w38 +w39·y w40 +w41·z w42+w43, where w37 (w37>0), w38 (w38>0), w39 (w39>0), w40 (w40>0), w41 (w41>0), w42 (w42>0), and w43 are calculation coefficients obtained through prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained by prior training is 1); the length of a certain state sequence is 3, and the sequence length weight value y=3 is calculated based on its positive correlation with the sequence length weight value (the calculation coefficient obtained by prior training is 1); the average error of the reward value distribution model and the state transition model obtained by training is 0.1, and the model error weight value z=0.5 / 0.1=5 is calculated based on its negative correlation with the model error weight value (the calculation coefficient obtained by prior training is 1), and the calculation coefficients obtained by prior training are w37=0.4, w38=1, w39=0.25, w40=1, w41=0.12, w42=1, w43=0, and the profit evaluation value p=w37·x is calculated for the state sequence w38 +w39·y w40 +w41·z w42 +w43=0.4×2+0.25×3+0.12×5+0=2.15. In another preferred embodiment, the profit evaluation value p of the state sequence is calculated as w44·x w45 ·y w46 ·z w47 +w48, where w44 (w44>0), w45 (w45>0), w46 (w46>0), w47 (w47>0), and w48 are calculation coefficients obtained through prior training. In this embodiment, the estimated value of the reward function at a certain quantile is 2, and the reward weight value x=2 is calculated based on its positive correlation with the reward weight value (the calculation coefficient obtained through prior training is 1); the length of a certain state sequence is 3, and the sequence length weight value y=3 is calculated based on its positive correlation with the sequence length weight value (the calculation coefficient obtained through prior training is 1); the average error of the reward value distribution model and the state transition model obtained through training is 0.1, and the model error weight value z=0.5 / 0.1=5 is calculated based on its negative correlation with the model error weight value (the calculation coefficient obtained through prior training is 1), and the calculation coefficients obtained through prior training are w44=0.07, w45=1, w46=1, w47=1, and w48=0, and the profit evaluation value p=w44·x of the calculated state sequence is calculated. w45 ·y w46 ·z w47 +w48=0.07×2×3×5+0=2.1.
[0150] The profit evaluation value of the state sequence is calculated according to the method described in any one of embodiments B1 to B7, and the profit size of the state sequence is evaluated according to the size of the evaluation value.
[0151] The environmental error has an upper bound, for a sequence of length k, when When , the error of the environment model has an upper bound as shown in formula (5):
[0152] (5)
[0153] in, is the total variation distance of the state transition model, ; is the total variation of the value distribution reward function model, that is, T and Z are the state transition and reward function of the real environment respectively.
[0154] In an exemplary embodiment, the strategy evaluation result is obtained according to the benefits of the state sequence, as shown in the flowchart. Figure 8 As shown, the steps include:
[0155] Calculate the return evaluation value of the state sequence at different quantiles of the strategy;
[0156] The strategy evaluation results are obtained by comparing the return evaluation values of the state sequences at different quantiles of the strategy.
[0157] In this embodiment, the profit evaluation value p(τ) of the state sequence at different quantiles of each strategy is calculated according to the method described in any one of Embodiments B1 to B7. i , compare p(τ) i And according to the profit evaluation value p(τ) i Arrange the strategies from largest to smallest, which is the strategy evaluation result.
[0158] A risk management method based on a value distribution environment model according to an embodiment of the present invention is shown in the flowchart as follows: Fig. 9 As shown, the steps include:
[0159] Determine the risk management strategies to be selected;
[0160] Input the historical data of the candidate risk control strategy into the strategy evaluation system based on the value distribution environment model to obtain the strategy evaluation result;
[0161] Select and execute the optimal risk management strategy based on the strategy evaluation results.
[0162] In this embodiment, the profit evaluation value p(τ) i The largest strategy quantile serves as the optimal risk management strategy.
[0163] In another exemplary embodiment, taking the field of autonomous driving as an example, a strategy evaluation system based on a value distribution environment model according to an embodiment of the present invention includes:
[0164] Offline data screening module: Screening offline data according to the returns of historical strategies and / or the specificity of historical strategies and generating offline data sets according to the quadruple data format;
[0165] Reward value distribution model construction module based on value distribution: establish a loss function based on value distribution learning and a four-tuple offline data set, and build a reward value distribution model based on value distribution according to the loss function;
[0166] State transfer model building module: uses deep neural network and gradient back propagation algorithm to train the state transfer model based on the four-tuple offline data set;
[0167] State sequence generation module: generates state sequence according to reward value distribution model and state transition model;
[0168] Strategy evaluation module: Evaluate the benefits of the state sequence according to the reward value and / or sequence length and / or model error, and obtain the strategy evaluation result according to the benefits of the state sequence.
[0169] In this embodiment, offline data is collected in a real driving environment, which includes but is not limited to driving environment, weather environment, traffic operation environment, building indoor environment, etc. The four-tuple offline data set includes four-tuple data at multiple moments (set as T moments in this embodiment), and the four-tuple data at each moment t is: < state s t , Driving behavior t , the next moment state s t+1 , Risk R t >. The definition and specific content of each data element are as follows:
[0170] Time t: according to the data environment requirements, set every day, every half day, every hour, every minute, or every second as a time t;
[0171] Status t : At each time t, the collected driving-related environmental conditions include road surface information, vehicle status information, driver status information, traffic status information, etc.
[0172] Behavior t : Related behaviors or actions performed at each time t in the environment, which may include steering wheel / joystick operation, gear lever operation, light change, horn sounding, rearview mirror movement, seat belt tightening / releasing, seat adjustment, airbag deployment, etc.
[0173] Risk R t:Risk is defined as the value at risk, which means the maximum expected loss at a given confidence level and a certain driving time.
[0174] The benefits of the state sequence are evaluated according to the reward value and / or sequence length and / or model error, and the strategy evaluation results of different driving behaviors under different environmental conditions are obtained according to the benefits of the state sequence, so as to facilitate the automatic driving system or the driver to select the optimal driving strategy, reduce risks and improve safety.
[0175] In another exemplary embodiment, taking the field of climate disaster early warning as an example, a strategy evaluation system based on a value distribution environment model according to an embodiment of the present invention includes:
[0176] Offline data screening module: Screening offline data according to the returns of historical strategies and / or the specificity of historical strategies and generating offline data sets according to the quadruple data format;
[0177] Reward value distribution model construction module based on value distribution: establish a loss function based on value distribution learning and a four-tuple offline data set, and build a reward value distribution model based on value distribution according to the loss function;
[0178] State transfer model building module: uses deep neural network and gradient back propagation algorithm to train the state transfer model based on the four-tuple offline data set;
[0179] State sequence generation module: generates state sequence according to reward value distribution model and state transition model;
[0180] Strategy evaluation module: Evaluate the benefits of the state sequence according to the reward value and / or sequence length and / or model error, and obtain the strategy evaluation result according to the benefits of the state sequence.
[0181] In this embodiment, offline data is collected in a real climate environment, which includes but is not limited to topographic environment, monsoon environment, air pressure environment, temperature and humidity environment, etc. The four-tuple offline data set includes four-tuple data at multiple moments (set as T moments in this embodiment), and the four-tuple data at each moment t is: < state s t 、Behavior a t , the next moment state s t+1 , Risk R t >. The definition and specific content of each data element are as follows:
[0182] Time t: according to the data environment requirements, set every day, every half day, every hour, every minute, or every second as a time t;
[0183] Status t: At each time t, the collected environmental conditions include temperature, humidity, air pressure, wind speed, cloud thickness, precipitation probability, thunderstorm index, etc.
[0184] Behavior t : Related behaviors or actions performed at each moment t in the environment, including climate broadcasts, disaster level determination, climate disaster warnings, emergency notifications, etc.
[0185] Risk R t :Risk is defined as the value at risk, which means the maximum expected loss within a given confidence level and a certain time limit.
[0186] The benefits of the state sequence are evaluated according to the reward value and / or sequence length and / or model error, and the strategy evaluation results of different warning behaviors under different climate conditions are obtained according to the benefits of the state sequence, which is convenient for the disaster warning system or warning personnel to select the optimal climate disaster warning strategy and improve the accuracy of the warning.
[0187] Of course, those skilled in the art should realize that the above embodiments are only used to illustrate the present invention, and are not intended to limit the present invention. As long as they are within the scope of the present invention, any changes or modifications to the above embodiments will fall within the protection scope of the present invention.
Claims
1. A strategy evaluation system based on a value distribution environment model, characterized in that: include: Offline data screening module: Screen offline data according to the historical strategy returns and / or the specificity of the historical strategy and generate an offline data set according to the four-tuple data format; the four-tuple data format is: <state s t 、Behavior a t , the next moment state s t+1 , Risk R t > In the field of autonomous driving, the time t is set every day, every half day, every hour, every minute, or every second as a time t according to the environment requirements of the data; the state s t represents the environmental state related to driving collected at each time t, including road surface information, vehicle status information, driver status information, and traffic status information; the behavior a t represents the relevant behaviors or actions performed at each time t in the environment, including steering wheel operation, joystick operation, gear lever operation, light change, horn sounding, rearview mirror movement, seat belt tightening, seat belt release, seat adjustment, and airbag deployment; the risk R t is the value at risk; Or in the field of climate disaster warning, the time t is set every day, every half day, every hour, every minute, or every second as a time t according to the environment requirements of the data; the state s t Indicates the environmental conditions collected at each time t, including temperature, humidity, air pressure, wind speed, cloud thickness, precipitation probability, and thunderstorm index; the behavior a t It represents the relevant behaviors or actions taken at each time t in the environment, including climate broadcast, disaster level determination, climate disaster warning, and emergency notification; the risk R t is the value at risk; Reward value distribution model construction module based on value distribution: establish a loss function based on value distribution learning and a four-tuple offline data set, and build a reward value distribution model based on value distribution according to the loss function; The loss function is established based on the value distribution learning and the four-tuple offline data set, including: defining a reward value distribution model; the input of the reward value distribution model is the state s at each moment t and behavior a t , the output is the reward R t distribution; calculate the cumulative distribution function of the reward function and use it to obtain the quantile probability distribution function; establish a loss function based on the difference between the model prediction value and the true value; The method of constructing a reward value distribution model based on value distribution according to the loss function includes: modeling the reward value distribution model as a neural network according to the loss function; updating parameters by using the gradient descent method of the loss function to obtain estimated values of different quantiles of the reward function, thereby obtaining a reward value distribution model based on value distribution; State transfer model building module: uses deep neural network and gradient back propagation algorithm to train the state transfer model based on the four-tuple offline data set; State sequence generation module: generates state sequence according to reward value distribution model and state transition model; Strategy evaluation module: evaluates the benefits of the state sequence according to the reward value, sequence length, and model error, and obtains the strategy evaluation result based on the benefits of the state sequence; The method of evaluating the benefits of a state sequence according to a reward value, a sequence length and a model error includes: calculating a reward weight value according to a positive correlation between an estimated value of different quantiles of a reward function and benefits; calculating a sequence length weight value according to a positive correlation between sequence length and benefits; calculating a model error weight value according to an average error between a reward value distribution model and a state transition model; calculating a benefit evaluation value of the state sequence according to the reward weight value, a sequence length weight value and a model error weight value; and evaluating the benefits of the state sequence according to the benefit evaluation value.
2. The strategy evaluation system based on the value distribution environment model according to claim 1 is characterized in that: The filtering of offline data according to the returns of historical strategies and / or the specificity of historical strategies includes: Calculate the return evaluation value of the historical strategy based on the positive return of the historical strategy and / or the return growth rate of the historical strategy; Calculating a specificity evaluation value of the historical strategy according to a repetition rate of the historical strategy and / or a similarity of the historical strategy; Calculating the historical strategy weight value according to the historical strategy return evaluation value and / or the historical strategy specificity evaluation value; Filter offline data based on the sorting of historical policy weight values and the preset policy weight threshold.
3. The strategy evaluation system based on the value distribution environment model according to claim 1, characterized in that: The training of the state transfer model according to the four-tuple offline data set includes: Define a state transition model; the input of the state transition model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 ; Construct an optimization model based on the maximum likelihood estimation method; Deep neural network and gradient back propagation algorithm are used to optimize the model parameters to obtain the state transfer model.
4. The policy evaluation system based on the value distribution environment model according to claim 3 is characterized in that: The generating of the state sequence according to the reward value distribution model and the state transition model includes: Step 1: According to the state s t Randomly sample actions a from the policy function t ~π(s t ); Step 2: Set the state s t and behavior a t Substitute into the reward value distribution model to obtain the estimated values of different quantiles of the reward function; Step 3: Set the state s t and behavior a t Substitute into the state transition model to get the next moment state s t+1 ; Step 4, arranging the four-tuple data in chronological order to form a sequence; Step 5: Repeat steps 1 to 4 to generate a state sequence of length k.
5. The policy evaluation system based on the value distribution environment model according to claim 1, characterized in that: The strategy evaluation result is obtained according to the benefits of the state sequence, including: Calculate the return evaluation value of the state sequence at different quantiles of the strategy; The strategy evaluation results are obtained by comparing the return evaluation values of the state sequences at different quantiles of the strategy.
Citation Information
Patent Citations
Off-line reinforcement learning method based on action space constraint
CN115965050A
Risk management method and system based on offline reinforcement learning, and readable storage medium
CN117593131A