Robust intelligent decision-making method and device and electronic equipment thereof

By using value function distribution networks in reinforcement learning to quantify the return distribution of different risk levels and adjust the action probability distribution of the policy network, the problem of performance degradation and easy to fall into local optimality in the prior art when the environment changes slightly, and higher robustness and decision-making reliability are achieved.

CN120218170APending Publication Date: 2025-06-27TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271217.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing reinforcement learning methods degrade the performance of the policy network when the environment changes slightly, and are prone to local optimization and lack robustness.

Method used

By determining the current state in the state space of the target environment and inputting its input value function distribution network, the inverse cumulative distribution function value is obtained, and the return distribution under different risk levels is quantified. The current dynamic risk level is determined based on the inverse cumulative distribution function value, input it into the policy network, and adjust the action probability distribution to reduce the local optimal risk.

Benefits of technology

It improves the robustness of the policy network, reduces the risk of preference for high-risk or low-risk actions, and avoids the strategic network from falling into the dilemma of local optimality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218170A_ABST
    Figure CN120218170A_ABST
Patent Text Reader

Abstract

The invention provides a robust intelligent decision-making method and device and electronic equipment thereof, and relates to the technical field of artificial intelligence. The method comprises the following steps: determining a current state in a state space of a target environment; the current state is input into a value function distribution network, a plurality of inverse cumulative distribution function values output by the value function distribution network are obtained, and the inverse cumulative distribution function values correspond to specified quantiles; determining a current dynamic risk level of the current state based on the inverse cumulative distribution function value, and inputting the current state and the current dynamic risk level into a strategy network to obtain action probability distribution output by the strategy network; and selecting an action based on the action probability distribution, determining a next state in the state space of the target environment according to the action, and completing intelligent decision when the next state is a termination state, so that the risk that a policy network falls into local optimum can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a robust intelligent decision-making method, apparatus, and electronic device thereof. Background Art

[0002] Reinforcement learning is an important branch in the field of artificial intelligence, and remarkable achievements have been made in various fields of sequential decision-making. The success of reinforcement learning depends on generating policy execution trajectories through trial-and-error learning in a specific environment to train the policy network of the agent. However, reinforcement learning depends on training in a specific environment, and when the environment changes slightly, the performance of the policy network will also drop significantly.

[0003] To make the policy network applicable to real-world scenarios, the policy network needs to be endowed with the ability to cope with such perturbations and maintain performance, and this ability is generally referred to as robustness. There are mainly two categories of methods for improving the robustness of the policy network. One category is to apply perturbations to the state, action, transition matrix of the environment, etc. during the training process, and optimize the cumulative reward of the policy network on the premise of considering bounded perturbations, but this will increase additional computational resources and time overhead. The other category is to implicitly improve the robustness of the policy by optimizing risk metrics such as CVaR and Wang of the cumulative reward. However, the optimization of CVaR α only focuses on the worst α proportion of trajectories, resulting in the policy network obtained by training falling into the dilemma of local optimality. How to solve the problem that the policy network obtained by training falls into local optimality is an important issue that the industry urgently needs to solve at present. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention provides a robust intelligent decision-making method, apparatus, and electronic device thereof.

[0005] The present invention provides a robust intelligent decision-making method, including:

[0006] Determine the current state in the state space of the target environment; Input the current state into the value function distribution network to obtain multiple inverse cumulative distribution function values output by the value function distribution network, and the inverse cumulative distribution function values correspond to specified quantiles; Based on the inverse cumulative distribution function values, determine the current dynamic risk level of the current state, input the current state and the current dynamic risk level into the policy network, and obtain the action probability distribution output by the policy network; Select an action based on the action probability distribution, determine the next state in the state space of the target environment according to the action, and complete the intelligent decision-making when the next state is a termination state.

[0007] A robust intelligent decision-making method provided by the present invention, determining the current dynamic risk level of the current state based on the inverse cumulative distribution function value, includes: Determine the dynamic risk budget of the current state; Based on the dynamic risk budget and the inverse cumulative distribution function value, determine the current dynamic risk level of the current state.

[0008] A robust intelligent decision-making method provided by the present invention, determining the dynamic risk budget of the current state, includes: When the current state is not the initial state, obtain the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state; Based on the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state, determine the dynamic risk budget of the current state.

[0009] A robust intelligent decision-making method provided by the present invention, the value function distribution network is a non-crossing quantile network; Based on the dynamic risk budget and the inverse cumulative distribution function value, determining the current dynamic risk level of the current state, includes: When the dynamic risk budget is between the minimum inverse cumulative distribution function value and the maximum inverse cumulative distribution function value, determine the left inverse cumulative distribution function value and the right inverse cumulative distribution function value of the distribution interval where the dynamic risk budget is located; Based on the left inverse cumulative distribution function value and the corresponding first specified quantile, the right inverse cumulative distribution function value and the corresponding second specified quantile, and the dynamic risk budget, determine the current dynamic risk level of the current state based on the inverse cumulative distribution function value.

[0010] A robust intelligent decision-making method provided by the present invention, before inputting the current state into the value function distribution network, the method further includes: Determine the initial value function distribution network and the initial policy network, repeatedly select action steps based on the action probability distribution in the training environment, and obtain multiple actions from the initial state to the termination state; At least determine the reward corresponding to each action, the inverse cumulative distribution function value, the current state, and the current dynamic risk level, and obtain the policy execution trajectory; Repeat the step of obtaining the policy execution trajectory to obtain a preset number of the policy execution trajectories; Based on the policy execution trajectory, adjust the initial value function distribution network to obtain the value function distribution network.

[0011] A robust intelligent decision-making method provided by the present invention, which adjusts the initial value function distribution network based on the policy execution trajectory to obtain the value function distribution network, includes: Based on the reward and the inverse cumulative distribution function value of each state in the policy execution trajectory, determine the cumulative return distribution of each state in the policy execution trajectory; Input each state in the policy execution trajectory into the initial value function distribution network to obtain multiple inverse cumulative distribution function values of each state output by the value function distribution network; Adjust the initial value function distribution network based on the energy distance between the cumulative return distribution of each state and the corresponding inverse cumulative distribution function value to obtain the value function distribution network.

[0012] A robust intelligent decision-making method provided by the present invention, before inputting the current state and the current dynamic risk level into the policy network, the method further includes: Based on the adjusted initial value function distribution network and the dynamic risk level of each state, determine the cumulative return expectation and conditional value at risk of the corresponding state; Determine the optimization objective of the initial policy network and the value function of each state according to the cumulative return expectation and the conditional value at risk; Determine the advantage function corresponding to the policy execution trajectory based on the value function; Adjust the initial policy network according to the optimization objective and the advantage function to obtain the policy network.

[0013] The present invention also provides a robust intelligent decision-making device, including: A current state determination module, configured to determine the current state in the state space of the target environment; A cumulative distribution determination module, configured to input the current state into the value function distribution network to obtain multiple inverse cumulative distribution function values output by the value function distribution network, and the inverse cumulative distribution function values correspond to specified quantiles; An action probability determination module, configured to determine the current dynamic risk level of the current state based on the inverse cumulative distribution function value, input the current state and the current dynamic risk level into the policy network, and obtain the action probability distribution output by the policy network; An intelligent decision-making module, configured to select an action based on the action probability distribution, determine the next state in the state space of the target environment according to the action, and complete the intelligent decision-making when the next state is a termination state.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the robust intelligent decision-making method as described in any one of the above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the robust intelligent decision-making method as described in any one of the above is implemented.

[0016] The robust intelligent decision-making method, device, and electronic device provided by the present invention determine the current state in the state space of the target environment, input the current state into the value function distribution network, and obtain multiple inverse cumulative distribution function values output by the value function distribution network to quantify the return distribution under different risk levels. Then, based on the inverse cumulative distribution function values, the current dynamic risk level of the current state is determined. The current state and the current dynamic risk level are input into the policy network. By adjusting the current dynamic risk level in real time, the preference risk of the policy network for high-risk or low-risk actions is reduced, and the action probability distribution output by the policy network is obtained. An action is selected based on the action probability distribution, and the next state in the state space of the target environment is determined according to the action. When the next state is a termination state, the intelligent decision-making is completed, thereby reducing the risk that the policy network falls into a local optimum. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is one of the flow diagrams of the robust intelligent decision-making method provided by the present invention.

[0019] Figure 2 is the structural diagram of the value function distribution network of the robust intelligent decision-making method provided by the present invention.

[0020] Figure 3 is the schematic diagram of the quantile estimation method of the robust intelligent decision-making method provided by the present invention.

[0021] Figure 4 is the schematic diagram of the current dynamic risk level estimation method of the current state of the robust intelligent decision-making method provided by the present invention.

[0022] Figure 5 is the schematic diagram of the conditional value at risk corresponding value function calculation of the robust intelligent decision-making method provided by the present invention.

[0023] Figure 6 It is the second schematic flow chart of the robust intelligent decision-making method provided by the present invention.

[0024] Figure 7 It is the schematic structural diagram of the robust intelligent decision-making device provided by the present invention.

[0025] Figure 8 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0027] The following combines Figures 1 - 8 to describe the robust intelligent decision-making method, device and its electronic device of the present invention.

[0028] Figure 1 It is one of the schematic flow charts of the robust intelligent decision-making method provided by the present invention. As Figure 1 shown, the method includes the following steps: Step 101: Determine the current state in the state space of the target environment.

[0029] The target environment refers to the environment in which the robust intelligent decision-making method is actually applied, such as a room, or a section of road, etc. The state space refers to the set that describes all possible states in the target environment. Each state represents a specific situation of the target environment. For example, the state can represent the specific position of a robot in a room, or the specific speed of a vehicle on a road, etc. The current state refers to the specific state in which the target environment is located at the current moment.

[0030] It can be understood that the robust intelligent decision-making method in this embodiment can be applied to an intelligent agent to improve the strategy of the intelligent agent. Exemplarily, factors affecting the state of the target environment can be determined to determine the state space of the target environment. The factors can include the intelligent agent, and the intelligent agent can be monitored to determine the current state in the state space of the target environment.

[0031] Step 102: Input the current state into the value function distribution network, and obtain a plurality of inverse cumulative distribution function values output by the value function distribution network. The inverse cumulative distribution function values correspond to specified quantiles.

[0032] The value function distribution network is a pre-trained neural network with the number of neurons in its input layer being the same as the dimension of the state space, and is used to estimate the value function distribution of a specific state. The input of the value function distribution network is a specific state in the state space of the target environment, and the output is the value of the inverse cumulative distribution function corresponding to a specified quantile. Exemplarily, the value function distribution network can be a feed-forward neural network or a convolutional neural network, etc.

[0033] The specified quantiles refer to multiple values representing probabilities preset from 0 to 1, and are used to represent different risk levels of the concerned cumulative return distribution. Among them, the specific number and specific values of the specified quantiles can be set according to actual needs, and this embodiment does not further limit this.

[0034] The value of the inverse cumulative distribution function refers to the cumulative return in the value function distribution of the corresponding state where the probability does not exceed τ when the value of the specified quantile is τ. For example, when τ = 0.1, the corresponding value of the inverse cumulative distribution function is the cumulative return of the worst 10% in the value function distribution.

[0035] It can be understood that if there are M preset specified quantiles, the value function distribution network can output M values of the inverse cumulative distribution function, and the M values of the inverse cumulative distribution function respectively correspond one-to-one to the M specified quantiles.

[0036] Through the values of the inverse cumulative distribution function output by the value function distribution network, the return distribution under different risk levels can be quantified, so that the policy network can take into account both the expected return and the tail risk, and improve the robustness of the final decision. For example, in the autonomous driving scenario, the value of the inverse cumulative distribution function with a low quantile τ = 0.05 output by the value function distribution network can help the final decision avoid high-risk operations such as sharp turns.

[0037] Step 103: Determine the current dynamic risk level of the current state based on the value of the inverse cumulative distribution function, and input the current state and the current dynamic risk level into the policy network to obtain the action probability distribution output by the policy network.

[0038] The current dynamic risk level is a parameter reflecting the risk sensitivity of the current state, and is used to reduce the preference risk of the policy network for high-risk or low-risk actions. For example, τ = 0.05 indicates that the current policy network needs to focus on the risk of the worst 5% of the cumulative return. The current state and its current dynamic risk level can be combined and called the augmented state.

[0039] The policy network is a pre-trained neural network with the number of neurons in its input layer being the same as the dimension of the state space and the number of neurons in its output layer being the same as the dimension of the action space. The specific dimensions of each intermediate hidden layer can be set according to actual situations, and this embodiment does not further limit this.

[0040] The action probability distribution output by the policy network. Specifically, in a discrete action space, the policy network outputs the logarithmic probabilities of each action; in a continuous action space, the policy network outputs the means of each dimension of the action. Moreover, in a continuous action space, the policy network also maintains a unified variance for each dimension of the action. Exemplarily, the policy network can be a feedforward neural network or a convolutional neural network, etc.

[0041] In this step, through the real-time adjustment of the dynamic risk level, it is convenient to adaptively select actions according to the risk distribution of the current state, so as to reduce the preference risk of the final decision for high-risk or low-risk actions and improve the robustness of the final decision.

[0042] Step 104: Select an action based on the action probability distribution, and determine the next state in the state space of the target environment according to the action. When the next state is a termination state, the intelligent decision-making is completed.

[0043] Exemplarily, in a discrete action space, an action can be selected from sample actions based on the action probability distribution. In a continuous action space, a continuous action can be generated based on the action probability distribution in combination with Gaussian distribution parameters, so as to determine the next state in the state space of the target environment according to the action. The target environment can determine whether the next state is a termination state according to the task completion situation. Specifically, the target environment can determine that the next state is a termination state according to task success or task failure, etc.

[0044] The robust intelligent decision-making method provided by the embodiments of the present invention determines the current state in the state space of the target environment, inputs the current state into the value function distribution network, and obtains multiple inverse cumulative distribution function values output by the value function distribution network to quantify the return distribution under different risk levels. Then, based on the inverse cumulative distribution function values, the current dynamic risk level of the current state is determined, and the current state and the current dynamic risk level are input into the policy network. By real-time adjustment of the current dynamic risk level, the preference risk of the policy network for high-risk or low-risk actions is reduced, and the action probability distribution output by the policy network is obtained. An action is selected based on the action probability distribution, and the next state in the state space of the target environment is determined according to the action. When the next state is a termination state, the intelligent decision-making is completed, thereby reducing the risk that the policy network falls into a local optimum.

[0045] In one embodiment, the value function distribution network is a non-crossing quantile network. As Figure 2 shown, the value function distribution network may include a multilayer perceptron (MLP), an interval output head a position output head and a bias output head ; Among them, the interval output head, the position output head, and the deviation output head can each be a fully connected layer or a stack of multiple fully connected layers.

[0046] The multi-layer perceptron can be used to process the input state to obtain an embedding vector , and the embedding vector can be respectively input into the interval output head , the position output head , and the deviation output head to obtain the value function distribution of this input state . Specifically, the value function distribution of this input state can be calculated through the following formula : : It can be understood that Figure 2 in, FC refers to a fully connected layer, and FC + Softmax refers to a fully connected layer followed by a Softmax activation function.

[0047] Exemplarily, the multi-layer perceptron can input the embedding vector into the position output head and the deviation output head based on the first fully connected layer in cooperation with Relu, and can input the embedding vector into the interval output head based on the second connection layer and Softmax in cooperation with Cumsum. By performing a multiplication operation with , and then performing an addition operation with , the value function distribution of the state is obtained.

[0048] As Figure 2 shown, M specified quantiles can be preset. After inputting the state into the value function distribution network and obtaining the value function distribution output by the value function distribution network, the value function CDF values corresponding to the specified quantiles can be determined . The value function CDF values at the M quantiles can be mixed with Dirac functions to obtain the value function distribution under the specified quantiles, so as to output the inverse cumulative distribution function values of the value function distribution at the specified M quantiles. The M inverse cumulative distribution function values correspond one-to-one with the M specified quantiles.

[0049] Based on any of the above embodiments, determining the current dynamic risk level of the current state based on the inverse cumulative distribution function value includes: Determining the dynamic risk budget of the current state; Based on the dynamic risk budget and the inverse cumulative distribution function value, determining the current dynamic risk level of the current state.

[0050] The dynamic risk budget refers to a risk control parameter that is adjusted in real time as the state changes and is used to quantify the risk threshold that can be tolerated in the current state.

[0051] Compared with a fixed risk threshold, in this embodiment, the time-varying risk threshold in the current state after environmental changes can be determined through the dynamic risk budget, and the current dynamic risk level of the current state can be determined by combining the inverse cumulative distribution function value, which can adaptively adjust the risk preference, increase the diversity of risk preferences, and thus avoid falling into the dilemma of local optimality.

[0052] In one embodiment, when the current state is the initial state, the value at risk of the initial state can be determined, and the value at risk is determined as the dynamic risk budget of the initial state.

[0053] Exemplarily, the of the initial state can be denoted as , then The estimation method of Figure 3 is shown as: Let , then Wherein, is the initial state, is the quantile serial number to be interpolated, is the predefined risk level, is the previous quantile serial number to be interpolated, is the next quantile serial number to be interpolated, is the number of quantiles, is the inverse cumulative distribution function value at the Mth quantile in the initial state.

[0054] For example, is the inverse cumulative distribution function value corresponding to the quantile serial number M in the initial state.

[0055] Based on any of the above embodiments, determining the dynamic risk budget of the current state includes: When the current state is not the initial state, obtaining the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state; Based on the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state, determining the dynamic risk budget of the current state.

[0056] A reward refers to the immediate feedback generated by the target environment after executing an action, which is used to evaluate the quality of the action in a specific state, so as to guide the final decision to complete the task by maximizing the cumulative reward, etc. For example, in autonomous driving, a positive reward is obtained for safe driving, and a negative reward is obtained for a collision, so as to guide the final decision of safe driving by maximizing the cumulative reward, etc.

[0057] In this embodiment, the actual risk consumption of the action is reflected by the reward, and according to the risk threshold that can be tolerated in the current state of the actual risk consumption, the policy network can be prompted to explore actions with different risk levels.

[0058] In one embodiment, the time step can be associated with the state, so as to determine the information of different states based on the time step.

[0059] In one embodiment, a discount factor can be introduced to determine the dynamic risk budget of the current state, as shown in the following formula: Among them, is the dynamic risk budget of the current state, is the dynamic risk budget of the previous state, is the reward obtained after executing the action in the previous state, is the discount factor.

[0060] The discount factor can be set according to actual needs, and this embodiment does not make further limitations on this. In one embodiment, .

[0061] In this embodiment, the discount factor reduces the influence weight of the long-term risk on the dynamic risk budget of the current state, weakens the influence of the long-term risk, prompts the policy network to pay more attention to the recent risk, and avoids overly conservative decisions.

[0062] In one embodiment, the value function distribution network is a non-crossing quantile network. If the dynamic risk budget is not greater than the inverse cumulative distribution function value at the first quantile, then the current dynamic risk level of the current state is determined to be 1; if the dynamic risk budget is not less than the inverse cumulative distribution function value at the Mth quantile, then the current dynamic risk level of the current state is determined to be 0.

[0063] Based on any of the above embodiments, the value function distribution network is a non-crossing quantile network; Determining the current dynamic risk level of the current state based on the dynamic risk budget and the inverse cumulative distribution function value includes: When the dynamic risk budget is between the minimum inverse cumulative distribution function value and the maximum inverse cumulative distribution function value, determine the left inverse cumulative distribution function value and the right inverse cumulative distribution function value of the distribution interval where the dynamic risk budget is located; Based on the left inverse cumulative distribution function value and the corresponding first specified quantile, the right inverse cumulative distribution function value and the corresponding second specified quantile, and the dynamic risk budget, determine the current dynamic risk level of the current state based on the inverse cumulative distribution function value.

[0064] It can be understood that the inverse cumulative distribution function values output by the non-crossing quantile network meet the increasing requirement. The minimum inverse cumulative distribution function value is the inverse cumulative distribution function value at the 1st quantile, and the maximum inverse cumulative distribution function value is the inverse cumulative distribution function value at the Mth quantile.

[0065] The distribution interval is a continuous numerical range formed by two adjacent inverse cumulative distribution function values. The left inverse cumulative distribution function value of the distribution interval is the inverse cumulative distribution function value that is exactly lower than the dynamic risk budget, and the right inverse cumulative distribution function value of the distribution interval is the inverse cumulative distribution function value that is exactly higher than the dynamic risk budget.

[0066] As Figure 4 shown, when the value function distribution network is a non-crossing quantile network, the current dynamic risk level of the current state can be determined based on the following formula : Among them, is the current dynamic risk level, is the left inverse cumulative distribution function value of the distribution interval, is the right inverse cumulative distribution function value of the distribution interval, is the first specified quantile, is the second specified quantile; is the dynamic risk budget at the current moment, is the inverse cumulative distribution function value at the 1st quantile, is the inverse cumulative distribution function value at the Mth quantile.

[0067] Based on any of the above embodiments, before inputting the current state into the value function distribution network, the method further includes: Determine the initial value function distribution network and the initial policy network, and repeatedly select action steps based on the action probability distribution in the training environment to obtain multiple actions from the initial state to the termination state; At least determine the reward, the inverse cumulative distribution function value, the current state, and the current dynamic risk level corresponding to each action to obtain a policy execution trajectory; Repeat the steps of obtaining the policy execution trajectory to obtain a preset number of the policy execution trajectories; Adjust the initial value function distribution network based on the policy execution trajectory to obtain the value function distribution network.

[0068] The training environment refers to the scenario constructed for generating the policy execution trajectory. It can be understood that generally, there are differences between the target environment where the robust intelligent decision-making method is actually applied and the training environment.

[0069] In some embodiments, the dynamic risk budget corresponding to each action is also determined. Exemplarily, starting from the initial state , the initial value function distribution network and the initial policy network are determined. In the training environment, repeat the step of selecting actions based on the action probability distribution to obtain the current dynamic risk level of the initial state , and then the policy network is based on to select actions , and based on the training environment to obtain , update the dynamic risk budget , based on the dynamic risk budget and the initial state to determine the current dynamic risk level according to the corresponding inverse cumulative distribution function value ; record the information of the state . This process can be repeated to obtain the information of the state . Until the termination state is obtained from the training environment, a policy execution trajectory is obtained based on the information of each state. of the information until the termination state is obtained from the training environment, and a policy execution trajectory is obtained based on the information of each state.

[0070] Exemplarily, the policy execution trajectory can be stored in a preset storage space. During the training process, when the remaining space in the preset storage space is not enough to accommodate a new policy execution trajectory, it is determined that the preset number of the policy execution trajectories has been obtained. The capacity of the preset storage space can be set according to requirements, and no further limitation is made in this embodiment. For example, the capacity of the preset storage space can be set according to the difficulty level of the task corresponding to the policy.

[0071] Based on any of the above embodiments, adjusting the initial value function distribution network based on the policy execution trajectory to obtain the value function distribution network includes: Based on the reward and the inverse cumulative distribution function value of each state in the policy execution trajectory, determine the cumulative return distribution of each state in the policy execution trajectory; Input each state in the policy execution trajectory into the initial value function distribution network to obtain multiple inverse cumulative distribution function values of each state output by the value function distribution network; Adjust the initial value function distribution network based on the energy distance between the cumulative return distribution for each state and the corresponding inverse cumulative distribution function value to obtain the value function distribution network.

[0072] The cumulative return distribution determined based on the rewards and inverse cumulative distribution function values for each state in the policy execution trajectory is the cumulative return distribution actually obtained by policy execution. Therefore, adjusting the parameters of the initial value function distribution network based on this cumulative return distribution can optimize the initial value function distribution network and improve the accuracy of the multiple inverse cumulative distribution function values output for each state by the initial value function distribution network.

[0073] Exemplarily, as shown in the following formula, the energy distance between the cumulative return distribution for each state and the corresponding inverse cumulative distribution function value can be minimized to adjust the parameters of the initial value function distribution network.

[0074] Where is the loss function of the value function distribution network with parameter ; is the expectation at different time steps; is the inverse cumulative distribution function value of the previous state at time step t; is the first error term of the difference between and the inverse cumulative distribution function value is the second error term for measuring the difference between the cumulative return distribution of the current state at time step t and the inverse cumulative distribution function value ; is the error term for measuring the difference between the cumulative return distribution of the current state at time step t and the inverse cumulative distribution function value ; is the number of quantiles.

[0075] Compared with avoiding relying only on the expected return, in this embodiment, by using the cumulative return distribution for each state in multiple policy execution trajectories, the cumulative return for each state is quantified, enhancing the comprehensiveness of risk perception. In this embodiment, by minimizing the energy distance between the cumulative return distribution and the inverse cumulative distribution function value output by the initial value function distribution network, the modeling bias of the initial value function distribution network is reduced, increasing the reliability of the final decision.

[0076] Moreover, compared with focusing on the single-point expectation, in this embodiment, by quantifying the return distribution under different risk levels, the utilization efficiency of the policy execution trajectory can be improved.

[0077] Based on any of the above embodiments, before inputting the current state and the current dynamic risk level into the policy network, the method further includes: Determine the cumulative return expectation and conditional value at risk of the corresponding state based on the adjusted initial value function distribution network and the dynamic risk level of each state; Determine the optimization objective of the initial policy network and the value function of each state according to the cumulative return expectation and the conditional value at risk; Determine the advantage function corresponding to the policy execution trajectory based on the value function; Adjust the initial policy network according to the optimization objective and the advantage function to obtain the policy network.

[0078] Exemplarily, the optimization objective of the initial policy network can be determined by linearly combining the cumulative return expectation and the conditional value at risk based on the following formula: Wherein, is the cumulative return expectation of the policy, is the conditional value at risk of the policy, is the weight factor.

[0079] It can be understood that by adjusting the value of , the preference of the policy network for cumulative return and risk can be adjusted. The value of can be adjusted based on the dynamic risk level.

[0080] The value function distribution at a specified quantile can be obtained according to the adjusted initial value function distribution network and , according to and calculate and determine the value function as shown in the following formula : Wherein, is the state value function under the policy , is the CVaR state value function under the policy .

[0081] The value function corresponding to the conditional value at risk (CVaR) can be calculated in combination with Figure 5 .

[0082] The advantage function can be determined by the following formula: Among them, is the TD-error (Temporal Difference Error), is the state at the next moment, is the dynamic risk budget at the next moment, is the TD-error at time t, is a numerical parameter, is the length of the policy execution trajectory.

[0083] Exemplarily, can take 0.97.

[0084] Taking the policy network as as an example, the parameters of the policy network can be updated based on the following formula : Among them, is the collected policy execution trajectory, is 's clipping function, is a numerical parameter, are the parameters of the policy network.

[0085] Exemplarily, can take 0.2.

[0086] In this embodiment, by determining the optimization objective of the initial policy network through the cumulative return expectation and the conditional value at risk, the risk and return can be weighed in real time, which is convenient for dynamically adjusting the policy network to explore actions with different risk levels, thereby reducing the risk of falling into local optima.

[0087] Figure 6 is the second flow diagram of the robust intelligent decision-making method provided by the present invention. As Figure 6 shown, in order to specifically illustrate the function of the robust intelligent decision-making method provided by this embodiment, a specific example is provided below.

[0088] Determine the initial policy network and initialize the value function distribution network; Based on the initial policy network, operate the policy in the training environment to collect the policy execution trajectory; Adjust the initial value function distribution network based on the policy execution trajectory. Specifically, determine each state in the policy execution trajectory and the multiple cumulative rewards of the corresponding states in multiple policy execution trajectories, and obtain the cumulative reward distribution of each state in the policy execution trajectory based on the multiple cumulative rewards; input each state in the policy execution trajectory into the initial value function distribution network to obtain multiple inverse cumulative distribution function values of each state output by the value function distribution network; adjust the initial value function distribution network based on the energy distance between the cumulative reward distribution of each state and the corresponding inverse cumulative distribution function value to obtain the value function distribution network; Adjust the initial policy network based on the policy execution trajectory. Specifically, determine the expected cumulative reward and conditional value at risk of the corresponding state based on the adjusted initial value function distribution network and the dynamic risk level of each state; determine the optimization objective of the initial policy network and the value function of each state according to the expected cumulative reward and conditional value at risk; determine the advantage function of the corresponding policy execution trajectory based on the value function; adjust the initial policy network according to the optimization objective and the advantage function to obtain the policy network; Determine whether the policy network converges. If not, repeat the above training process. If so, end the training; wherein, the convergence of the policy network can be that the fluctuation of the average cumulative reward based on the intelligent decision made by the policy network is less than the preset fluctuation threshold for a continuously specified number of times.

[0089] The robust intelligent decision-making device provided by the present invention will be described below. The robust intelligent decision-making device described below can be correspondingly referred to the robust intelligent decision-making method described above.

[0090] Figure 7 is a schematic structural diagram of the robust intelligent decision-making device provided by the present invention, as Figure 7 shown. The device includes: The current state determination module 701 is used to determine the current state in the state space of the target environment; The cumulative distribution determination module 702 is used to input the current state into the value function distribution network to obtain multiple inverse cumulative distribution function values output by the value function distribution network, and the inverse cumulative distribution function values correspond to specified quantiles; The action probability determination module 703 is used to determine the current dynamic risk level of the current state based on the inverse cumulative distribution function value, input the current state and the current dynamic risk level into the policy network, and obtain the action probability distribution output by the policy network; The intelligent decision-making module 704 is used to select an action based on the action probability distribution, determine the next state in the state space of the target environment according to the action, and complete the intelligent decision-making when the next state is a termination state.

[0091] Based on any of the above embodiments, the action probability determination module 703 includes: A dynamic risk budget determination unit for determining the dynamic risk budget of the current state; A current dynamic risk level determination unit for determining the current dynamic risk level of the current state based on the dynamic risk budget and the inverse cumulative distribution function value.

[0092] Based on any of the above embodiments, the dynamic risk budget determination unit is specifically configured to: When the current state is not the initial state, obtain the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state; Determine the dynamic risk budget of the current state based on the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state.

[0093] Based on any of the above embodiments, the value function distribution network is a non-crossing quantile network; The current dynamic risk level determination unit is specifically configured to: When the dynamic risk budget is between the minimum inverse cumulative distribution function value and the maximum inverse cumulative distribution function value, determine the left inverse cumulative distribution function value and the right inverse cumulative distribution function value of the distribution interval where the dynamic risk budget is located; Determine the current dynamic risk level of the current state based on the left inverse cumulative distribution function value and the corresponding first specified quantile, the right inverse cumulative distribution function value and the corresponding second specified quantile, and the dynamic risk budget based on the inverse cumulative distribution function value.

[0094] Based on any of the above embodiments, the robust intelligent decision-making device further includes a value function distribution network optimization module for: Determine an initial value function distribution network and an initial policy network, repeatedly select action steps based on the action probability distribution in the training environment, and obtain a plurality of actions from the initial state to the termination state; At least determine the reward, the inverse cumulative distribution function value, the current state, and the current dynamic risk level corresponding to each action to obtain a policy execution trajectory; Repeat the step of obtaining the policy execution trajectory to obtain a preset number of the policy execution trajectories; Adjust the initial value function distribution network based on the policy execution trajectory to obtain the value function distribution network.

[0095] Based on any of the above embodiments, the value function distribution network optimization module is specifically configured to: determine the cumulative return distribution of each state in the policy execution trajectory based on the reward and the inverse cumulative distribution function value of each state in the policy execution trajectory; Input each state in the policy execution trajectory into the initial value function distribution network to obtain multiple inverse cumulative distribution function values of each state output by the value function distribution network; Adjust the initial value function distribution network based on the energy distance between the cumulative return distribution of each state and the corresponding inverse cumulative distribution function value to obtain the value function distribution network.

[0096] Based on any of the above embodiments, the robust intelligent decision-making device further includes a policy network training module, which is used for: Based on the adjusted initial value function distribution network and the dynamic risk level of each state, determine the cumulative return expectation and conditional value at risk of the corresponding state; Determine the optimization objective of the initial policy network and the value function of each state according to the cumulative return expectation and the conditional value at risk; Determine the advantage function corresponding to the policy execution trajectory based on the value function; Adjust the initial policy network according to the optimization objective and the advantage function to obtain the policy network.

[0097] Figure 8 Illustrates a schematic structural diagram of an electronic device, as Figure 8 shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the robust intelligent decision-making method, which includes: determining the current state in the state space of the target environment; inputting the current state into the value function distribution network to obtain multiple inverse cumulative distribution function values output by the value function distribution network, and the inverse cumulative distribution function values correspond to specified quantiles; determining the current dynamic risk level of the current state based on the inverse cumulative distribution function values, inputting the current state and the current dynamic risk level into the policy network to obtain the action probability distribution output by the policy network; selecting an action based on the action probability distribution, and determining the next state in the state space of the target environment according to the action. When the next state is a termination state, the intelligent decision-making is completed.

[0098] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0099] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the robust intelligent decision-making method provided by the above-mentioned various methods. The method includes: determining the current state in the state space of the target environment; inputting the current state into the value function distribution network to obtain multiple inverse cumulative distribution function values output by the value function distribution network, and the inverse cumulative distribution function values correspond to specified quantiles; determining the current dynamic risk level of the current state based on the inverse cumulative distribution function values, inputting the current state and the current dynamic risk level into the policy network to obtain the action probability distribution output by the policy network; selecting an action based on the action probability distribution, and determining the next state in the state space of the target environment according to the action. When the next state is a termination state, the intelligent decision-making is completed.

[0100] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the robust intelligent decision-making method provided by the above-mentioned various methods. The method includes: determining the current state in the state space of the target environment; inputting the current state into the value function distribution network to obtain multiple inverse cumulative distribution function values output by the value function distribution network, and the inverse cumulative distribution function values correspond to specified quantiles; determining the current dynamic risk level of the current state based on the inverse cumulative distribution function values, inputting the current state and the current dynamic risk level into the policy network to obtain the action probability distribution output by the policy network; selecting an action based on the action probability distribution, and determining the next state in the state space of the target environment according to the action. When the next state is a termination state, the intelligent decision-making is completed.

[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A robust intelligent decision-making method, characterized in that: include: Determine the current state in the state space of the target environment; Inputting the current state into a value function distribution network to obtain a plurality of inverse cumulative distribution function values ​​output by the value function distribution network, wherein the inverse cumulative distribution function values ​​correspond to specified quantiles; Determine a current dynamic risk level of the current state based on the inverse cumulative distribution function value, input the current state and the current dynamic risk level into a policy network, and obtain an action probability distribution output by the policy network; An action is selected based on the action probability distribution, and a next state in the state space of the target environment is determined according to the action. When the next state is a terminal state, intelligent decision-making is completed.

2. The robust intelligent decision-making method according to claim 1, characterized in that: The determining the current dynamic risk level of the current state based on the inverse cumulative distribution function value comprises: determining a dynamic risk budget for said current state; A current dynamic risk level for the current state is determined based on the dynamic risk budget and the inverse cumulative distribution function value.

3. The robust intelligent decision-making method according to claim 2, characterized in that: Determine a dynamic risk budget for the current state, including: When the current state is not the initial state, obtaining the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state; The dynamic risk budget of the current state is determined based on the dynamic risk budget of the previous state and the reward of the corresponding action of the previous state.

4. The robust intelligent decision-making method according to claim 2, characterized in that: The value function distribution network is a non-crossing quantile network; Determining a current dynamic risk level of the current state based on the dynamic risk budget and the inverse cumulative distribution function value includes: When the dynamic risk budget is between the minimum inverse cumulative distribution function value and the maximum inverse cumulative distribution function value, determining the left inverse cumulative distribution function value and the right inverse cumulative distribution function value of the distribution interval where the dynamic risk budget is located; A current dynamic risk level of the current state is determined based on the left inverse cumulative distribution function value and the corresponding first specified quantile, the right inverse cumulative distribution function value and the corresponding second specified quantile, and the dynamic risk budget based on the inverse cumulative distribution function value.

5. The robust intelligent decision-making method according to claim 1, characterized in that: Before inputting the current state into the value function distribution network, the method further includes: Determine an initial value function distribution network and an initial policy network, and repeatedly select action steps based on the action probability distribution in a training environment to obtain multiple actions from an initial state to a terminal state; At least determining the reward corresponding to the action in each state, the inverse cumulative distribution function value, the current state, and the current dynamic risk level to obtain a strategy execution trajectory; Repeat the step of obtaining the strategy execution track to obtain a preset number of the strategy execution tracks; The initial value function distribution network is adjusted based on the strategy execution trajectory to obtain the value function distribution network.

6. The robust intelligent decision-making method according to claim 5, characterized in that: Adjusting the initial value function distribution network based on the strategy execution trajectory to obtain the value function distribution network includes: Determining a cumulative reward distribution for each state in the policy execution trajectory based on the reward and the inverse cumulative distribution function value for each state in the policy execution trajectory; Input each state in the strategy execution trajectory into the initial value function distribution network to obtain multiple inverse cumulative distribution function values ​​of each state output by the value function distribution network; The value function distribution network is obtained by adjusting the initial value function distribution network based on the energy distance between the cumulative reward distribution of each state and the corresponding inverse cumulative distribution function value.

7. The robust intelligent decision-making method according to claim 6, characterized in that: Before inputting the current state and the current dynamic risk level into the policy network, the method further includes: Determine the cumulative return expectation and conditional risk value of the corresponding state based on the adjusted initial value function distribution network and the dynamic risk level of each state; Determining the optimization target of the initial strategy network and the value function of each state according to the cumulative return expectation and the conditional risk value; Determining an advantage function corresponding to the strategy execution trajectory based on the value function; The initial policy network is adjusted according to the optimization objective and the advantage function to obtain a policy network.

8. A robust intelligent decision-making device, characterized in that: include: A current state determination module, used to determine the current state in the state space of the target environment; A cumulative distribution determination module, configured to input the current state into a value function distribution network to obtain a plurality of inverse cumulative distribution function values ​​output by the value function distribution network, wherein the inverse cumulative distribution function values ​​correspond to specified quantiles; an action probability determination module, configured to determine a current dynamic risk level of the current state based on the inverse cumulative distribution function value, input the current state and the current dynamic risk level into a policy network, and obtain an action probability distribution output by the policy network; The intelligent decision-making module is used to select an action based on the action probability distribution, determine the next state in the state space of the target environment according to the action, and complete the intelligent decision when the next state is a terminal state.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the robust intelligent decision-making method as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robust intelligent decision-making method as claimed in any one of claims 1 to 7 is implemented.