Risk management system based on offline reinforcement learning

Through a risk management system based on offline reinforcement learning, the risk management strategy is optimized using offline data sets and supervised learning models, and the problems of high trial and error cost and low interaction efficiency in the existing technology are solved, and more efficient risk management is achieved.

CN119941408AActive Publication Date: 2025-05-06XQUANT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411693768.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-25
Publication Date
2025-05-06
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

The existing risk management methods based on reinforcement learning are the problems of high trial and error costs of strategy, low interaction efficiency with the real environment, and low strategy optimization efficiency.

Method used

A risk management system based on offline reinforcement learning is proposed to optimize risk management strategies through offline data set generation, supervised learning model training, risk reconstruction and minimum risk strategy generation.

Benefits of technology

Without directly interacting with the real environment, reduce the cost of trial and error of strategy, improve the efficiency of strategy optimization, and enhance the dynamicity and accuracy of risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941408A_ABST
    Figure CN119941408A_ABST
Patent Text Reader

Abstract

The invention discloses a risk management system based on off-line reinforcement learning. The system comprises an off-line data set generation module which generates an off-line data set from off-line data according to a screening strategy and a tetrad data format; the sequence data generation module is used for training a supervised learning model according to the offline data set and generating sequence data according to the supervised learning model; the risk reconstruction module is used for calculating a risk adjustment value according to the change value of the sequence data and the training data and reconstructing a risk function according to the risk adjustment value; the training data comprises any one item or a combination of multiple items of training times, training time and training completion degree; and the minimum risk strategy generation module is used for calculating a minimum risk value according to the reconstructed risk function and inputting the minimum risk value into the supervised learning model to obtain a minimum risk strategy function. According to the method, the problems of high strategy trial and error cost, low interaction efficiency with a real environment and low strategy optimization efficiency are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a risk management system based on offline reinforcement learning. Background Art

[0002] Intelligent risk management is an important means of strengthening risk management and control by applying intelligent technology. Traditional intelligent risk management is to optimize the modeling of value at risk and conditional value at risk by establishing machine learning models, so as to support risk management at different confidence levels. Intelligent risk management is to use artificial intelligence technology to process and analyze massive amounts of data and complex business environments, and to conduct comprehensive assessment and management of different types of risks through data analysis and model building. For example, machine learning and deep learning algorithms can perform risk prediction and decision support by learning and pattern recognition of historical data.

[0003] At present, there are four main intelligent risk management methods: risk control based on volatility, risk control based on correlation, risk assessment model based on machine learning, and risk management based on reinforcement learning. However, the existing methods of risk control based on volatility, risk control based on correlation, and risk assessment model based on machine learning only consider risks in a shorter time frame, and do not consider risks in a longer time frame enough. Risk management based on reinforcement learning requires continuous trial and error with the environment to optimize the strategy. If the strategy is tried and errored directly in the real environment, random exploration may not only bring losses, but also the interaction efficiency with the real environment is very low; if a simulation environment is constructed according to physical theorems, the simulation environment is often not accurate enough, making it difficult to obtain the optimal strategy.

[0004] In order to reduce the strategy trial and error cost of reinforcement learning intelligent risk management and improve the efficiency of interaction with the real environment and strategy optimization, a risk management system based on offline reinforcement learning is proposed. Summary of the invention

[0005] The embodiment of the present invention proposes a risk management system based on offline reinforcement learning to at least solve the problems of high cost of strategy trial and error, low efficiency of interaction with the real environment and low efficiency of strategy optimization.

[0006] According to one embodiment of the present invention, a risk management system based on offline reinforcement learning is proposed, comprising:

[0007] Offline data set generation module: generates offline data sets based on the screening strategy and quadruple data format;

[0008] Sequence data generation module: trains the supervised learning model based on the offline data set and generates sequence data based on the supervised learning model;

[0009] Risk reconstruction module: calculates the risk adjustment value according to the change value of the sequence data and the training data and reconstructs the risk function accordingly; the training data includes any one or a combination of the number of training times, training time, and training completion degree;

[0010] Minimum risk strategy generation module: Calculate the minimum risk value based on the reconstructed risk function, input the minimum risk value into the supervised learning model, and obtain the minimum risk strategy function.

[0011] In an exemplary embodiment, the screening strategy includes any one or a combination of strategy optimization principle and strategy coverage principle; the strategy optimization principle is to select various data corresponding to the optimal strategy under different states, and the strategy coverage principle is to select various data of all strategies under the same or similar states.

[0012] In an exemplary embodiment, the four-tuple data format is: t 、Behavior a t , the next moment state s t+1 , Risk R t >.

[0013] In an exemplary embodiment, the step of training a supervised learning model based on an offline data set includes:

[0014] Randomly sample an initial state s from the offline dataset 0 ;

[0015] Randomly initialize the parameters of the policy function π(·), where the input of the policy function is the state s t , the output is behavior a t ;

[0016] The input of the supervised learning model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 , Risk R t ;

[0017] Supervised learning models are computed using deep neural networks and maximum likelihood estimation algorithms.

[0018] In an exemplary embodiment, generating sequence data according to a supervised learning model includes:

[0019] Step 1: According to the state s t Randomly sample actions a from the policy function t ~ π(s t );

[0020] Step 2: Set the state s t With behavior a tInput into the supervised learning model to calculate the next state s t+1 and risk R t ;

[0021] Step 3: Arrange the quadruple data in chronological order < state s t 、Behavior a t , the next moment state s t+1 , Risk R t >, forming a sequence;

[0022] Step 4: Determine whether the sequence reaches the preset length value. If so, end and obtain the sequence data of the preset length; otherwise, return to step 1.

[0023] In an exemplary embodiment, the step of calculating the risk adjustment value based on the change value of the sequence data and the training data includes:

[0024] The prediction certainty weight value is calculated according to the variance of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the standard deviation of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance and standard deviation of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance and average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the standard deviation and average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance, standard deviation and average change value of the sequence data output by the supervised learning model;

[0025] Calculate the convergence weight value according to the reinforcement learning training time, or calculate the convergence weight value according to the number of reinforcement learning trainings, or calculate the convergence weight value according to the reinforcement learning training completion, or calculate the convergence weight value according to the reinforcement learning training time and the number of trainings, or calculate the convergence weight value according to the reinforcement learning training time and the training completion, or calculate the convergence weight value according to the reinforcement learning training number of trainings and the training completion, or calculate the convergence weight value according to the reinforcement learning training time, the number of trainings and the training completion;

[0026] The risk adjustment value is calculated based on the negative correlation between the predicted certainty weight value and the risk, or the risk adjustment value is calculated based on the positive correlation between the convergence weight value and the risk, or the risk adjustment value is calculated based on the negative correlation between the predicted certainty weight value and the risk and the positive correlation between the convergence weight value and the risk.

[0027] In an exemplary embodiment, the risk function is calculated based on the positive correlation between the risk value and the risk adjustment value in the sequence data.

[0028] In an exemplary embodiment, the step of calculating the minimized risk value according to the reconstructed risk function includes:

[0029] According to the status t Calculate discount factors;

[0030] Calculate the minimized risk value based on the discount factor and SAC algorithm.

[0031] In an exemplary embodiment, the minimum risk strategy function is a combination of different states and different behaviors corresponding to the minimized risk value obtained after the minimized risk value is input into the supervised learning model.

[0032] The risk management system based on offline reinforcement learning of the present invention has the following advantages:

[0033] (1) Collect and filter data in a real environment based on the strategy optimization principle or strategy coverage principle, generate an offline data set, and train a supervised learning model based on the offline data set. Compared with traditional risk management technical solutions, this method can directly optimize efficient and feasible intelligent risk management strategies without interacting with the real environment.

[0034] (2) The prediction certainty weight value is calculated according to the variance and / or standard deviation and / or mean change value output by the supervised learning model, the convergence weight value is calculated according to the reinforcement learning training time and / or training times and / or training completion, the risk adjustment value is calculated according to the negative correlation between the prediction certainty weight value and the risk and the positive correlation between the convergence weight value and the risk, and the risk function is reconstructed according to the positive correlation between the risk value and the risk adjustment value. Compared with the traditional fixed risk function technical solution, this can improve the convergence of the function under dynamic uncertainty and effectively avoid the errors caused by the environmental model.

[0035] (3) The minimized risk value is calculated based on the reconstructed risk function, and the strategy function under different behaviors and different state combinations is trained based on the minimized risk value. Compared with traditional risk management technical solutions, by interacting with the environmental model and considering the uncertainty of time series changes, more accurate decisions can be made based on changes in the environment and the system itself. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of the structure of a risk management system based on offline reinforcement learning according to an embodiment of the present invention;

[0037] Figure 2 is a flow chart of training a supervised learning model according to an offline data set in a sequence data generation module of a risk management system according to an embodiment of the present invention;

[0038] Figure 3is a flow chart of generating sequence data according to a supervised learning model in a sequence data generation module of a risk management system according to an embodiment of the present invention;

[0039] Figure 4 is a flow chart of calculating a risk adjustment value according to a change value of sequence data and training data in a risk reconstruction module of a risk management system according to an embodiment of the present invention;

[0040] Figure 5 It is a flow chart of calculating the minimized risk value according to the reconstructed risk function in the minimum risk strategy generation module of the risk management system of an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0042] The offline reinforcement learning-based risk management system of the present invention is applicable to scenarios including but not limited to financial decision-making, autonomous driving obstacle identification and warning, behavioral risk identification, climate monitoring and warning, weather forecasting and other fields that require observation and decision-making.

[0043] Taking financial decision making as an example, a risk management system based on offline reinforcement learning in an embodiment of the present invention has a structural diagram as shown in FIG. Figure 1 As shown, including:

[0044] Offline data set generation module: generates offline data sets based on the screening strategy and quadruple data format;

[0045] Sequence data generation module: trains the supervised learning model based on the offline data set and generates sequence data based on the supervised learning model;

[0046] Risk reconstruction module: calculates the risk adjustment value according to the change value of the sequence data and the training data and reconstructs the risk function accordingly; the training data includes any one or a combination of the number of training times, training time, and training completion degree;

[0047] Minimum risk strategy generation module: Calculate the minimum risk value based on the reconstructed risk function, input the minimum risk value into the supervised learning model, and obtain the minimum risk strategy function.

[0048] In this embodiment, the risk management system based on offline reinforcement learning according to the embodiment of the present invention obtains the minimum risk strategy function, and executes market decisions according to the minimum risk strategy function, which can implement more accurate market strategies according to the market environment and its own changes.

[0049] The screening strategy includes any one or a combination of the strategy optimization principle and the strategy coverage principle; the strategy optimization principle is to select the data corresponding to the optimal strategy under different states, and the strategy coverage principle is to select the data of all strategies under the same or similar states. In this embodiment, offline data is collected in a real market environment, and the data corresponding to the optimal strategy under different states are selected, or the data of all strategies under the same or similar states are selected.

[0050] The four-tuple data format is: t 、Behavior a t , the next moment state s t+1 , Risk R t >.

[0051] In this embodiment, the offline data set includes a total of four-tuple data at multiple time points (set as T time points in this embodiment), and the definition and specific content of each data element are as follows:

[0052] Time t: according to the market environment requirements of the data, set every day, every half day, every hour, every minute, every second, etc. as a time t;

[0053] Status t : The collected environmental status at each time t. The status may include the on-site data and off-site data of the day. The on-site data includes but is not limited to the opening price, closing price, highest price, lowest price, fundamental information, and financial report information. The off-site data includes but is not limited to weather, policy information, bank interest rates, gold prices, oil prices, etc.

[0054] Behavior t : The relevant behavior or action taken at each time t in the environment. The behavior can include a combination of the stock traded, the transaction price, and the transaction volume.

[0055] Risk R t :Risk is defined as the Value at Risk (VaR), which means the maximum expected loss within a given confidence level and a certain time limit, as shown in formula (1):

[0056] (1)

[0057] Where α is the quantile point, X=(s t ,a t ) is the current state-behavior combination. The risk value means the maximum expected loss at a given confidence level α and a certain market period.

[0058] The supervised learning model is trained according to the offline data set, and the flow chart is as follows Figure 2 As shown, the steps include:

[0059] Randomly sample an initial state s from the offline dataset 0 ;

[0060] Randomly initialize the parameters of the policy function π(·), where the input of the policy function is the state s t , the output is behavior a t ;

[0061] The input of the supervised learning model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 , Risk R t ;

[0062] Supervised learning models are computed using deep neural networks and maximum likelihood estimation algorithms.

[0063] In this embodiment, an initial state s is randomly sampled from the offline data set. 0 ; At the same time, the parameters of the policy function π(·) are randomly initialized. Here, the policy function is modeled by a neural network, and the parameters represent all the parameters of the neural network. The input of the policy function is the state s t , the output is a continuous vector a t .

[0064] According to the collected offline data set, the supervised learning model T is trained. The input of the supervised learning model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 , Risk R t , according to the maximum likelihood estimation, the model is optimized using the formula shown in formula (2):

[0065] (2)

[0066] In practical applications, a deep neural network can be used as a supervised learning model, and the parameters in the neural network can be optimized through gradient back propagation. The likelihood function T(·) uses a Gaussian function, and its logarithm is the mean square error.

[0067] The sequence data is generated according to the supervised learning model, and the flow chart is as follows Figure 3 As shown, including:

[0068] Step 1: According to the state s t Randomly sample actions a from the policy function t ~ π(s t );

[0069] Step 2: Set the state s t With behavior a t Input into the supervised learning model to calculate the next state s t+1 and risk R t ;

[0070] Step 3: Arrange the quadruple data in chronological order < state s t 、Behavior a t , the next moment state s t+1 , Risk R t >, forming a sequence;

[0071] Step 4: Determine whether the sequence reaches the preset length value. If so, end and obtain the sequence data of the preset length; otherwise, return to step 1.

[0072] In this embodiment, according to the initial state s of random sampling 0 , randomly sample actions a from the policy function 0 ~ π(s 0 ), the initial state s 0 With behavior a 0 Input into the supervised learning model (2) to obtain the state s at the next moment 1 and VaR 1 , get the first set of four-tuple data < state s 0 、Behavior a 0 , the next moment state s 1 , Risk R 1 >, forming the first set of data in the sequence; then the state s at the next moment 1 and actions a randomly sampled from the policy function 1 ~ π(s 1 ) is input into the supervised learning model (2) to obtain the state s at the next moment 2 and VaR 2 , get the second set of four-tuple data < state s 1 、Behavior a 1 , the next moment state s 2 , Risk R 2 >, forming the second set of data in the sequence. The preset length value of the sequence is set according to the cumulative error requirement. Here, the preset sequence length value k=5, that is, repeating the above process to train 5 sets of data to obtain a sequence of length 5 starting from the initial state.

[0073] The risk adjustment value is calculated based on the change value of the sequence data and the training data. The flow chart is as follows Figure 4 As shown, including:

[0074] Calculate the prediction certainty weight value according to the variance and / or standard deviation and / or average change value of the sequence data output by the supervised learning model;

[0075] Calculate the convergence weight value according to the reinforcement learning training time and / or the number of trainings and / or the degree of training completion;

[0076] The risk adjustment value is calculated based on the negative correlation between the prediction certainty weight value and the risk and / or the positive correlation between the convergence weight value and the risk.

[0077] In this embodiment, the prediction certainty of the supervised learning model is the prediction certainty weight value, and the prediction certainty weight value is calculated according to the variance and / or standard deviation and / or average change value of the sequence data output by the supervised learning model, which is: the prediction certainty weight value is calculated according to the variance of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the standard deviation of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance and standard deviation of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance and average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the standard deviation and average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance, standard deviation and average change value of the sequence data output by the supervised learning model, and the prediction certainty weight value is represented by the variable p.

[0078] Training convergence is the convergence weight value. The convergence weight value is calculated according to the reinforcement learning training time and / or the number of trainings and / or the training completion degree, which is: calculating the convergence weight value according to the positive correlation between the reinforcement learning training time and the convergence weight value, calculating the convergence weight value according to the positive correlation between the reinforcement learning training number of trainings and the convergence weight value, calculating the convergence weight value according to the positive correlation between the reinforcement learning training completion degree and the convergence weight value, calculating the convergence weight value according to the positive correlation between the reinforcement learning training time and the number of trainings and the convergence weight value, calculating the convergence weight value according to the positive correlation between the reinforcement learning training time and the number of trainings and the convergence weight value, calculating the convergence weight value according to the positive correlation between the reinforcement learning training time and the number of trainings and the training completion degree and the convergence weight value, and calculating the convergence weight value according to the positive correlation between the reinforcement learning training time and the number of trainings and the training completion degree and the convergence weight value. Any one of the convergence weight values ​​is calculated according to the positive correlation between the reinforcement learning training time and the number of trainings and the training completion degree and the convergence weight value. The convergence weight value is represented by the variable q.

[0079] The negative correlation between the prediction certainty weight value and risk means that the smaller the prediction certainty weight value, the higher the risk assigned, and the corresponding risk adjustment value should be larger. The significance of the prediction certainty weight value is that when the prediction certainty of the model is low, a higher risk is assigned (increasing the risk adjustment value), so that the reinforcement learning algorithm tries not to explore high-risk states.

[0080] The positive correlation between the convergence weight value and the risk is that the larger the convergence weight value, the higher the risk, and the corresponding risk adjustment value should be larger. The significance of the convergence weight value is: at the initial moment (weak convergence), reduce the penalty of environmental uncertainty (reduce the risk adjustment value) and encourage reinforcement learning exploration; when the strategy gradually converges (strong convergence), increase the penalty of uncertainty (increase the risk adjustment value) to promote the model to converge as soon as possible.

[0081] The risk adjustment value calculated according to the negative correlation between the predicted certainty weight value and the risk and / or the positive correlation between the convergence weight value and the risk is any one of the following: calculating the risk adjustment value according to the negative correlation between the predicted certainty weight value and the risk adjustment value, calculating the risk adjustment value according to the positive correlation between the convergence weight value and the risk adjustment value, or calculating the risk adjustment value according to the negative correlation between the predicted certainty weight value and the risk adjustment value and the positive correlation between the convergence weight value and the risk, and the risk adjustment value is represented by the variable e.

[0082] Different implementations of calculating the risk adjustment value are:

[0083] Embodiment A1: Calculate the risk adjustment value based on the negative correlation between the prediction certainty weight value and the risk adjustment value.

[0084] The risk adjustment value is calculated based on the negative correlation between the prediction certainty weight value and the risk adjustment value. Specifically, the prediction certainty weight value p is calculated based on the negative correlation between the variance and / or standard deviation and / or average change value output by the supervised learning model and the prediction certainty weight value; the risk adjustment value e is calculated based on the negative correlation between the prediction certainty weight value p and the risk adjustment value. In one embodiment, the calculated risk adjustment value e=o1·p o2 +o3, where o1, o2 (o1·o2<0), and o3 are calculation coefficients obtained by prior training. In this embodiment, the variance, standard deviation, or average change value of the supervised learning model output is statistically calculated. This embodiment takes the variance as an example, calculates the variance s=1.5 of a supervised learning model (Gaussian model) output, and calculates the prediction certainty weight value p=0.8 (p=g1·s g2+g3, where g1, g2, g3 are the calculation coefficients obtained by prior training and g1·g2<0, here g1=1.2, g2=-1, g3=0), the calculation coefficients obtained by prior training o1=1, o2=1, o3=0, and the calculated risk adjustment value e=o1·p o2 +o3=1×0.8+0=0.8.

[0085] Embodiment A2: Calculate the risk adjustment value based on the positive correlation between the convergence weight value and the risk adjustment value.

[0086] The risk adjustment value is calculated based on the positive correlation between the convergence weight value and the risk adjustment value. Specifically, the convergence weight value q is calculated based on the positive correlation between the reinforcement learning training time and / or the number of trainings and / or the training completion and the convergence weight value; the risk adjustment value e is calculated based on the positive correlation between the convergence weight value q and the risk adjustment value. In one embodiment, the calculated risk adjustment value e=o4·q o5 +o6, where o4, o5 (o4·o5>0), and o6 are calculation coefficients obtained through prior training. In this embodiment, the reinforcement learning training time, number of trainings, and training completion are recorded. This embodiment uses training time as a parameter to measure the convergence of training, and sets the convergence weight value q between 0 and 1. As the training time t increases, the convergence weight value q increases. This relationship is realized using a commonly used activation function, i.e., q=sigmoid(t). When the training time t=2, q=0.88, the calculation coefficients o4=1, o5=1, and o6=0 obtained through prior training, and the calculated risk adjustment value e=o4·q o5 +o6=1×0.88+0=0.88.

[0087] Embodiment A3: Calculate the risk adjustment value based on the negative correlation between the prediction certainty weight value and the risk adjustment value and the positive correlation between the convergence weight value and the risk.

[0088] The risk adjustment value is calculated based on the negative correlation between the predicted certainty weight value and the risk adjustment value and the positive correlation between the convergence weight value and the risk. Specifically, the predicted certainty weight value p is calculated based on the negative correlation between the variance and / or standard deviation and / or average change value output by the supervised learning model and the predicted certainty weight value; the convergence weight value q is calculated based on the positive correlation between the reinforcement learning training time and / or the number of trainings and / or the training completion and the convergence weight value; the risk adjustment value e is calculated based on the negative correlation between the predicted certainty weight value p and the risk adjustment value and the positive correlation between the convergence weight value q and the risk. In one embodiment, the calculated risk adjustment value e=o7·p o8 +o9·q o10+o11, where o7, o8 (o7·o8<0), o9, o10 (o9·o10>0), and o11 are calculation coefficients obtained by prior training. In this embodiment, the variance, standard deviation, or average change value of the output of the supervised learning model is statistically calculated. This embodiment takes the variance as an example, calculates the variance s=1.5 of the output of a supervised learning model (Gaussian model), and calculates the prediction certainty weight value p=0.8 (p=g1·s g2 +g3, where g1, g2, g3 are the calculation coefficients obtained by prior training and g1·g2<0, where g1=1.2, g2=-1, g3=0); record the reinforcement learning training time, number of trainings and training completion. This embodiment uses the training time as a parameter to measure the convergence of the training, and sets the convergence weight value q between 0-1. As the training time t increases, the convergence weight value q increases. This relationship is realized using a commonly used activation function, i.e., q=sigmoid(t). When the training time t=2, q=0.88; the calculation coefficients obtained by prior training o7=0.7, o8=1, o9=0.3, o10=1, o11=0, and the risk adjustment value e=o7·p o8 +o9·q o10 + o11 =0.7×0.8+0.3×0.88+0=0.824. In another embodiment, the risk adjustment value is calculated as e=o12·p o13 ·q o14 +o15, where o12, o13 (o12·o13<0), o14 (o13·o14<0), and o15 are calculation coefficients obtained by prior training. In this embodiment, the variance, standard deviation, or average change value of the supervised learning model output is statistically calculated. This embodiment takes the variance as an example, calculates the variance s=1.5 of a supervised learning model (Gaussian model) output, and calculates the prediction certainty weight value p=0.8 (p=g1·s g2 +g3, where g1, g2, g3 are the calculation coefficients obtained by prior training and g1·g2<0, where g1=1.2, g2=-1, g3=0); record the reinforcement learning training time, number of trainings and training completion. This embodiment uses the training time as a parameter to measure the convergence of the training, and sets the convergence weight value q between 0-1. As the training time t increases, the convergence weight value q increases. This relationship is realized using a commonly used activation function, i.e., q=sigmoid(t). When the training time t=2, q=0.88; the calculation coefficients o11=1.2, o12=1, o13=1, o14=0 obtained by prior training, and the risk adjustment value e=o12·p is calculated. o13 ·q o14 +o15 =1.2×0.8×0.88+0=0.845.

[0089] The risk function is calculated based on the positive correlation between the risk value and the risk adjustment value in the sequence data. In this embodiment, the risk function is calculated based on the positive correlation between the risk value and the risk adjustment value in the sequence data. The risk adjustment value e and the risk value R calculated according to any one of the embodiments A1 to A3 are t The risk function is calculated based on the positive correlation between , such as any one of formula (3) or formula (4):

[0090] (3)

[0091] (4)

[0092] Among them, f 1 、f 2 、f 3 、f 4 、f 5 、f 6 It is the calculation coefficient obtained by prior training.

[0093] The minimized risk value is calculated based on the reconstructed risk function. The flow chart is as follows Figure 5 As shown, including:

[0094] According to the status t Calculate discount factors;

[0095] Calculate the minimized risk value based on the discount factor and SAC algorithm.

[0096] In this embodiment, the discount factor γ is calculated according to the state t (obtained through risk assessment training under different states), it can generally be taken as 0.95. SAC is a classic reinforcement learning algorithm. Soft Actor-Critic maximizes the entropy increase reward by learning a random strategy. This strategy maps the state to the action and a Q function. This function estimates the target value of the current strategy and optimizes it through approximate dynamic programming. In this way, the SAC algorithm can maximize the return after entropy reinforcement. In this process, SAC will take the target as absolute truth to derive a better reinforcement learning algorithm with stable performance and sufficiently high sample efficiency. Therefore, the SAC algorithm calculates and minimizes the function shown in formula (5) to obtain the minimized risk value.

[0097] (5)

[0098] The minimum risk strategy function is a combination of different states and different behaviors corresponding to the minimized risk value obtained after the minimized risk value is input into the supervised learning model. Input into the supervised learning model to minimize the risk value The corresponding states and output behaviors constitute the strategy function π(a t |s t ).

[0099] In another exemplary embodiment, taking the field of autonomous driving as an example, a risk management system based on offline reinforcement learning in an embodiment of the present invention includes:

[0100] Offline data set generation module: generates offline data sets according to the screening strategy and the four-tuple data format; in this embodiment, offline data are collected in a real driving environment, which includes but is not limited to the driving environment, weather environment, traffic operation environment, building indoor environment, etc. Select the data corresponding to the optimal strategy under different states, or select the data of all strategies under the same or similar states.

[0101] The offline data set includes four-tuple data at multiple time points (T time points in this embodiment), and the four-tuple data at each time point t is: t , Driving behavior t , the next moment state s t+1 , Risk R t >. The definition and specific content of each data element are as follows:

[0102] Time t: according to the data environment requirements, set every day, every half day, every hour, every minute, or every second as a time t;

[0103] Status t : At each time t, the collected driving-related environmental status. In this embodiment, the status may include driving road information, vehicle status information, driver status information, traffic status information, etc.

[0104] Behavior t : The relevant behaviors or actions performed at each time t in the environment. In this embodiment, the behaviors may include steering wheel / joystick operation, gear lever operation, light change, horn sounding, rearview mirror movement, seat belt tightening / releasing, seat adjustment, airbag deployment, etc.

[0105] Risk R t :Risk is defined as the value at risk, which means the maximum expected loss at a given confidence level and a certain driving time.

[0106] Sequence data generation module: trains the supervised learning model based on the offline data set and generates sequence data based on the supervised learning model;

[0107] Risk reconstruction module: calculates the risk adjustment value according to the change value of the sequence data and the training data and reconstructs the risk function accordingly; the training data includes any one or a combination of the number of training times, training time, and training completion degree;

[0108] Minimum risk strategy generation module: Calculate the minimum risk value based on the reconstructed risk function, input the minimum risk value into the supervised learning model, and obtain the minimum risk strategy function. In this embodiment, the minimum risk value is calculated based on the reconstructed risk function, and the minimum risk value is input into the supervised learning model. The state corresponding to the minimum risk value and the output behavior constitute the minimum risk strategy function under different states and different behavior combinations. Executing autonomous driving according to the minimum risk strategy function can effectively improve the safety of autonomous driving behavior.

[0109] In another exemplary embodiment, taking the field of climate disaster early warning as an example, a risk management system based on offline reinforcement learning in an embodiment of the present invention includes:

[0110] Offline data set generation module: Generate offline data sets according to the screening strategy and the four-tuple data format. In this embodiment, offline data is collected in a real climate environment, and the real climate environment includes but is not limited to topographic environment, monsoon environment, air pressure environment, temperature and humidity environment, etc. Select the data corresponding to the optimal strategy under different states, or select the data of all strategies under the same or similar states.

[0111] The offline data set includes four-tuple data at multiple time points (T time points in this embodiment), and the four-tuple data at each time point t is: t 、Behavior a t , the next moment state s t+1 , Risk R t >. The definition and specific content of each data element are as follows:

[0112] Time t: according to the data environment requirements, set every day, every half day, every hour, every minute, or every second as a time t;

[0113] Status t : The environmental state collected at each time t. In this embodiment, the state may include temperature, humidity, air pressure, wind speed, cloud thickness, precipitation probability, thunderstorm index, etc.

[0114] Behavior t : Related behaviors or actions performed at each time t in the environment. In this embodiment, the behaviors may include weather broadcasting, disaster level determination, weather disaster warning, emergency notification, etc.

[0115] Risk R t:Risk is defined as the value at risk, which means the maximum expected loss within a given confidence level and a certain time limit.

[0116] Sequence data generation module: trains the supervised learning model based on the offline data set and generates sequence data based on the supervised learning model;

[0117] Risk reconstruction module: calculates the risk adjustment value according to the change value of the sequence data and the training data and reconstructs the risk function accordingly; the training data includes any one or a combination of the number of training times, training time, and training completion degree;

[0118] Minimum risk strategy generation module: Calculate the minimized risk value based on the reconstructed risk function, input the minimized risk value into the supervised learning model, and obtain the minimum risk strategy function. In this embodiment, the minimized risk value is calculated based on the reconstructed risk function, and the minimized risk value is input into the supervised learning model. The state corresponding to the minimized risk value and the output behavior constitute the minimum risk strategy function under different states and different behavior combinations. In this embodiment, executing climate disaster warning according to the minimum risk strategy function can effectively improve the accuracy of climate disaster warning.

[0119] Of course, those skilled in the art should realize that the above embodiments are only used to illustrate the present invention, and are not intended to limit the present invention. As long as they are within the scope of the present invention, any changes or modifications to the above embodiments will fall within the protection scope of the present invention.

Claims

1. A risk management system based on offline reinforcement learning, characterized in that: include: Offline data set generation module: generates offline data sets based on the screening strategy and quadruple data format; Sequence data generation module: trains the supervised learning model based on the offline data set and generates sequence data based on the supervised learning model; Risk reconstruction module: calculates the risk adjustment value according to the change value of the sequence data and the training data and reconstructs the risk function accordingly; the training data includes any one or a combination of the number of training times, training time, and training completion degree; Minimum risk strategy generation module: Calculate the minimum risk value based on the reconstructed risk function, input the minimum risk value into the supervised learning model, and obtain the minimum risk strategy function.

2. The risk management system based on offline reinforcement learning according to claim 1, characterized in that: The screening strategy includes any one or a combination of strategy optimization principle and strategy coverage principle; the strategy optimization principle is to select various data corresponding to the optimal strategy under different states, and the strategy coverage principle is to select various data of all strategies under the same or similar states.

3. The risk management system based on offline reinforcement learning according to claim 1, characterized in that: The four-tuple data format is: <state s t 、Behavior a t , the next moment state s t+1 , Risk R t >.

4. The risk management system based on offline reinforcement learning according to claim 3 is characterized in that: The step of training the supervised learning model according to the offline data set includes: Randomly sample an initial state s0 from the offline dataset; Randomly initialize the parameters of the policy function π(·), where the input of the policy function is the state s t , the output is behavior a t ; The input of the supervised learning model is the state s at each moment t and behavior a t , the output is the next moment state s t+1 , Risk R t ; Supervised learning models are computed using deep neural networks and maximum likelihood estimation algorithms.

5. The risk management system based on offline reinforcement learning according to claim 4, characterized in that: The generating sequence data according to the supervised learning model comprises: Step 1: According to the state s t Randomly sample actions a from the policy function t ~ π(s t ); Step 2: Set the state s t With behavior a t Input into the supervised learning model to calculate the next state s t+1 and risk R t ; Step 3: Arrange the quadruple data in chronological order < state s t 、Behavior a t , the next moment state s t+1 , Risk R t >, forming a sequence; Step 4: Determine whether the sequence reaches the preset length value. If so, end and obtain the sequence data of the preset length; otherwise, return to step 1.

6. The risk management system based on offline reinforcement learning according to claim 5, characterized in that: The step of calculating the risk adjustment value according to the change value of the sequence data and the training data includes: The prediction certainty weight value is calculated according to the variance of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the standard deviation of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance and standard deviation of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance and average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the standard deviation and average change value of the sequence data output by the supervised learning model, or the prediction certainty weight value is calculated according to the variance, standard deviation and average change value of the sequence data output by the supervised learning model; Calculate the convergence weight value according to the reinforcement learning training time, or calculate the convergence weight value according to the number of reinforcement learning trainings, or calculate the convergence weight value according to the reinforcement learning training completion, or calculate the convergence weight value according to the reinforcement learning training time and the number of trainings, or calculate the convergence weight value according to the reinforcement learning training time and the training completion, or calculate the convergence weight value according to the reinforcement learning training number of trainings and the training completion, or calculate the convergence weight value according to the reinforcement learning training time, the number of trainings and the training completion; The risk adjustment value is calculated based on the negative correlation between the predicted certainty weight value and the risk, or the risk adjustment value is calculated based on the positive correlation between the convergence weight value and the risk, or the risk adjustment value is calculated based on the negative correlation between the predicted certainty weight value and the risk and the positive correlation between the convergence weight value and the risk.

7. The risk management system based on offline reinforcement learning according to claim 6, characterized in that: The risk function is calculated based on the positive correlation between the risk value and the risk adjustment value in the sequence data.

8. The risk management system based on offline reinforcement learning according to claim 7, characterized in that: The step of calculating the minimized risk value according to the reconstructed risk function comprises: According to the status t Calculate discount factors; Calculate the minimized risk value based on the discount factor and SAC algorithm.

9. The risk management system based on offline reinforcement learning according to claim 1, characterized in that: The minimum risk strategy function is a combination of different states and different behaviors corresponding to the minimized risk value obtained after the minimized risk value is input into the supervised learning model.

Citation Information

Patent Citations

  • Risk prevention and control decision method, device, system and equipment

    CN111325444A

  • Bank risk pricing optimization method and device based on deep reinforcement learning

    CN112488826A

  • Multi-agent depth deterministic strategy gradient method based on course learning

    CN113449458A

  • Systems and methods for risk-sensitive reinforcement learning

    US20210232970A1

  • Computer-based systems, computing components and computing objects configured to implement dynamic outlier bias reduction in machine learning models

    US20220092346A1