Double-loop application method and system for partially observable Markov decision problems
By adopting a dual-cycle application method in the decision-making and control process of intelligent cars, using the gradient information of the maximum likelihood estimation and hidden state model, the problems of noise and uncertainty in the decision-making process of some observable Markov are solved, and more stable optimal strategy solutions and better driving performance are achieved.
Patent Information
- Application Number
- CN202210897910.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-28
AI Technical Summary
In the decision-making and control process of smart cars, there is system noise and uncertainty in part of the observed Markov decision-making process, which makes existing methods unstable and difficult to achieve effective driving performance when solving optimal strategies.
Using the dual-cycle application method, by obtaining the observation data of the automobile, constructing historical information representation based on the maximum likelihood estimation, and using the gradient information of the hidden state model as an intrinsic reward function, the optimal strategy is encouraged to learn the uncertainty in the environment, thereby solving the optimal strategy for partially observable Markov decision-making.
It improves the robustness of solving the optimal strategy in some observable and observed uncertain scenarios with noise, and meets the corresponding driving performance requirements.
Smart Images

Figure CN115356923B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of learning decision control applications for intelligent vehicles, and in particular to a double-loop application method and system for partially observable Markov decision problems. Background Art
[0002] In common industrial control problems, it is often impossible to obtain the true state of the system. Instead, the observed state with system noise is obtained. In the perception system of intelligent vehicles, there are often a lot of perception noise, perception jitter and other problems, which in turn affect the decision-making and control process of intelligent vehicles. Partially Observable Markov Decision Process (POMDP) is often used to model this type of problem.
[0003] General reinforcement learning methods usually assume that the state of the environment is fully observable at each time step, but this assumption does not hold true in reality. Due to the existence of system noise and other uncertainties, only the observed state can be obtained, while the true state is a hidden state that cannot be observed. At present, there are two common solution methods:
[0004] One is to reconstruct historical information by characterizing the function and reconstruct the hidden state z t Directly carry out the strategy π(z t ) learning. Among them, the choice of representation function is often strongly related to the structure of the input data. For example, convolutional neural networks are suitable for image input, and recurrent neural networks are suitable for serialized data input. However. End-to-end training methods often add too much burden to the representation function. The representation function needs to not only extract effective features from historical information, but also meet the requirements of maximizing the reward function when solving the optimal strategy. If it is not designed properly, it may lead to instability in the learning process. In addition, the representation function commonly used for historical information input is the Recurrent Neural Network (RNN), which often has problems with gradient disappearance and training difficulties. Its update process does not meet the principle of Bayesian posterior update, and can only obtain deterministic hidden state encoding.
[0005] The other is to learn the observed state from the data (o t ) model and the real state (s t ) model, and obtain the estimate of the true state as the belief state b through Bayesian estimation t This type of method uses the belief state b in the belief state space t Approximate true state And drive based on belief state b tHowever, this process has a trade-off between model error and policy optimality. If we rely on the learned belief state b t The model drives the MDP process forward, which may cause state error accumulation, resulting in a gap between the environment updated by the strategy and the real environment, such as Figure 1 As shown. If we reduce the belief state b t If the model is dependent on the original model, it will fall into the same dilemma as the first method. Summary of the invention
[0006] The present application provides a dual-loop application method, system, electronic device and storage medium for a partially observable Markov decision problem, which can solve the optimal strategy and meet the corresponding driving performance in uncertain scenarios with partial observability and noise in the observation, and improve the robustness of the solution method to observation uncertainty.
[0007] The first aspect of the present application provides a dual-loop application method for a partially observable Markov decision problem, comprising the following steps: obtaining observation data of a car; in the partially observable Markov decision process, constructing a historical information representation based on the observation data by the maximum likelihood estimation method, while using the gradient information of the hidden state model as an intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision; and generating the optimal control strategy of the car based on the optimal strategy of the partially observable Markov decision.
[0008] Optionally, in one embodiment of the present application, while constructing a historical information representation based on the observation data by the maximum likelihood estimation method, the gradient information of the hidden state model is used as the intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision, including: based on the observation data, determining the hidden state space, action space, observation space, reward function, state transition model in the hidden state space, state transition model between the hidden state and the observed state, and discount factor; constructing the partially observable Markov decision process according to the hidden state space, the action space, the observation space, the reward function, state transition model in the hidden state space, state transition model between the hidden state and the observed state, and the discount factor.
[0009] Optionally, in one embodiment of the present application, both the observed state transfer model of the outer loop driving the partially observable Markov decision process and the hidden state transfer model of the inner loop driving the partially observable Markov decision process are unknown.
[0010] Optionally, in one embodiment of the present application, after constructing the partially observable Markov decision process, it also includes: in the outer loop, selecting a sequence of observed states as input, and defining the observed state transfer model in the outer loop as a likelihood function of the hidden state space, so as to drive the external Markov process through the likelihood hidden state of historical observations.
[0011] Optionally, in one embodiment of the present application, after constructing the partially observable Markov decision process, it also includes: in the inner loop, establishing a random hidden state transfer model, sampling the probability distribution obeyed by the parameter each time forward inference is performed, and using the sampled parameter value as the parameter value for this forward inference, and using the variational lower bound in the interaction data pool to update the random hidden state transfer model in the inner loop, so as to drive the internal Markov process through the random hidden state transfer model.
[0012] Optionally, in one embodiment of the present application, the optimal strategy for solving the partially observable Markov decision includes: using the information gain in the inner loop as the intrinsic reward function of the intelligent agent, superimposing it with the reward function obtained by interacting with the real environment in the outer loop, and obtaining the reconstructed reward function as the update target of the value network.
[0013] Optionally, in one embodiment of the present application, the optimal strategy based on the partially observable Markov decision generates the optimal control strategy of the vehicle, including: loading a neural network parameter file corresponding to the optimal strategy, and inputting an observation state sequence to map and obtain the optimal control strategy.
[0014] The second aspect of the present application provides a dual-loop application system for a partially observable Markov decision problem, including: an acquisition module for acquiring observation data of a car; a solution module for, in a partially observable Markov decision process, constructing a historical information representation based on the observation data by a maximum likelihood estimation method, while using the gradient information of the hidden state model as an intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision; and an application module for generating an optimal control strategy for the car based on the optimal strategy of the partially observable Markov decision.
[0015] Optionally, in one embodiment of the present application, the solution module is further used to determine, based on the observation data, a hidden state space, an action space, an observation space, a reward function, a state transition model in the hidden state space, a state transition model between the hidden state and the observed state, and a discount factor; and construct the partially observable Markov decision process according to the hidden state space, the action space, the observation space, the reward function, the state transition model in the hidden state space, the state transition model between the hidden state and the observed state, and the discount factor.
[0016] Optionally, in one embodiment of the present application, both the observed state transfer model of the outer loop driving the partially observable Markov decision process and the hidden state transfer model of the inner loop driving the partially observable Markov decision process are unknown.
[0017] Optionally, in one embodiment of the present application, after constructing the partially observable Markov decision process, it also includes: a processing module, which is used to select a sequence of observed states as input in the outer loop, and define the observed state transfer model in the outer loop as a likelihood function of the hidden state space, so as to drive the external Markov process through the likelihood hidden state of historical observations.
[0018] Optionally, in one embodiment of the present application, after constructing the partially observable Markov decision process, it also includes: an update module, which is used to establish a random hidden state transfer model in the inner loop, sample the probability distribution obeyed by the parameter during each forward inference, and use the sampled parameter value as the parameter value for this forward inference, and use the variational lower bound in the interaction data pool to update the random hidden state transfer model in the inner loop to drive the internal Markov process through the random hidden state transfer model.
[0019] Optionally, in one embodiment of the present application, the optimal strategy for solving the partially observable Markov decision includes: using the information gain in the inner loop as the intrinsic reward function of the intelligent agent, superimposing it with the reward function obtained by interacting with the real environment in the outer loop, and obtaining the reconstructed reward function as the update target of the value network.
[0020] Optionally, in one embodiment of the present application, the optimal strategy based on the partially observable Markov decision generates the optimal control strategy of the vehicle, including: loading a neural network parameter file corresponding to the optimal strategy, and inputting an observation state sequence to map and obtain the optimal control strategy.
[0021] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the double-loop application method for the partially observable Markov decision problem as described in the above embodiment.
[0022] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the double-loop application method for the partially observable Markov decision problem as described above.
[0023] The dual-loop application method and system for the partially observable Markov decision problem of the embodiment of the present application include an outer loop and an inner loop. The outer loop aims to construct a historical information representation based on the maximum likelihood estimation method to avoid reducing the accumulation of state errors caused by driving the MDP process using the learned model. The inner loop aims to use the gradient information of the hidden state model as the intrinsic reward function to motivate the optimal strategy to learn the uncertainty in the environment and guide action exploration. As a result, it is possible to solve the optimal strategy in uncertain scenarios with partial observability and noise in the observation, and meet the corresponding driving performance, thereby improving the robustness of the solution method to observation uncertainty.
[0024] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0026] Figure 1 This is a schematic diagram of the model state error accumulation;
[0027] Figure 2 A flowchart of a double-loop application method for a partially observable Markov decision problem provided according to an embodiment of the present application;
[0028] Figure 3 An architectural diagram of a double-loop application method for a partially observable Markov decision problem provided according to an embodiment of the present application;
[0029] Figure 4 A schematic diagram of a double-loop application system structure for a partially observable Markov decision problem provided according to an embodiment of the present application;
[0030] Figure 5 Schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0032] The following describes the double-loop application method, system, electronic device and storage medium of the partially observable Markov decision problem of the embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology center, the present application provides a double-loop application method for partially observable Markov decision problems. In this method, different from the problem setting in the traditional planning field, it is assumed that the observation model and state transition model of the system are unknown, and how to find the optimal strategy under the partially observable problem setting is discussed.
[0033] Specifically, Figure 2 The present invention is a flowchart of a double-loop application method for a partially observable Markov decision problem provided according to an embodiment of the present application.
[0034] like Figure 2 As shown, the double-loop application method for this partially observable Markov decision problem includes the following steps:
[0035] In step S101 , observation data of the car is acquired.
[0036] In the embodiment of the present application, based on the observation data and taking into account the vehicle driving performance requirements, the optimal control strategy under the partially observable problem setting is obtained through offline learning and training using a deep neural network as a carrier, and the optimal control strategy is applied online. It is possible to solve the optimal strategy in uncertain scenarios with partial observability and noise in the observations, and meet the corresponding driving performance requirements, thereby improving the robustness of the solution method to observation uncertainty.
[0037] In step S102, in the partially observable Markov decision process, while constructing a historical information representation based on the observed data by the maximum likelihood estimation method, the gradient information of the hidden state model is used as the intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision.
[0038] In an embodiment of the present application, while constructing a historical information representation based on the observation data by the maximum likelihood estimation method, the gradient information of the hidden state model is used as the intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision, including: determining the hidden state space, action space, observation space, reward function, state transition model in the hidden state space, state transition model between hidden state and observed state, and discount factor based on the observation data; constructing a partially observable Markov decision process according to the hidden state space, action space, observation space, reward function, state transition model in the hidden state space, state transition model between hidden state and observed state, and discount factor.
[0039] Specifically, we construct a partially observable Markov decision process characterized by a seven-tuple, with seven tuples Among them, is the hidden state space, Represents the unobservable true state at time step t (defined as the hidden state). is the action space, represents the action at time step t. is the observation space, Represents the observed state at time step t. is the reward function, which means that in the observed state o t Next, perform action a t Then transfer to the next observation state o t+1 Rewards received. In the hidden state space The state transition model on . Hidden state With observation status The state transition model between the hidden state and the observed state of the agent is z t+1 and t+1 The probability of . γ is the discount factor, γ∈(0,1).
[0040] In an embodiment of the present application, both the observed state transfer model of the outer loop driving the partially observable Markov decision process and the hidden state transfer model of the inner loop driving the partially observable Markov decision process are unknown.
[0041] Furthermore, unlike the definition of partially observable Markov decision process problems in the traditional planning field, the embodiments of the present application use the hidden state transfer model Observational state transfer model Set to unknown. Defined by The driving process is an external loop, which is composed of The driving process is an internal loop.
[0042] In an embodiment of the present application, after constructing a partially observable Markov decision process, it also includes: in an outer loop, selecting a sequence of observed states as input, and defining the observed state transfer model in the outer loop as a likelihood function of the hidden state space, so as to drive the external Markov process through the likelihood hidden state of historical observations.
[0043] Since the hidden state space It cannot be obtained directly. In the embodiment of the present application, a historical observation state sequence of length H with Markov properties is selected. t:t+h As input, abbreviated as o t,h . Define the observation state transition model in the outer loop is the likelihood function of the hidden state space, p(ζ t | t,h ). t is the likelihood hidden state of the historical observations encoded in the outer loop. Therefore, the likelihood hidden state ζ obtained by the historical observations in the outer loop t Drives an external Markov process.
[0044] Assuming that the observed states satisfy the independence assumption and parameterizing the likelihood function with φ, we obtain the following approximate relationship:
[0045] logp(ζ t |o t,h )≈log∏ i∈[0,h] q φ (o t+i ) (1)
[0046] Based on the Markov random field, the following relationship is obtained:
[0047] log∏ i∈[0,h] q φ (o t+i ) (2)
[0048] =logΦ(o t ,o t )+Σ i∈[1,H] log∫q φ (o t )logΨ(o t ,o t+i )do t +c t
[0049] Among them, Φ(·) and Ψ(·) are potential energy functions defined in Markov random fields, respectively expressed as f 1 (·) and f 2 (·)express. is a permutation invariant operator. By iterating k times in the hidden state space, we can obtain the likelihood function of the hidden state space satisfying the following relationship:
[0050]
[0051] In practice, f 1 (·) and f 2 (·) By approximating the neural network, we can get the likelihood hidden state ζ of the historical observations t distribution.
[0052] Furthermore, it is assumed that in each step of the interaction between the agent and the environment, the agent tries to obtain useful information from the environment. t , get the hidden state z corresponding to the new observation state t+1 After that, the information gain obtained by the agent in the hidden state space is defined as:
[0053]
[0054] Where ω is a random variable representing the parameters of the hidden state model. t+1 According to the hidden state transfer model The hidden state at time t+1 is obtained, ζ t It is the observation state transition model Hidden state information encoding obtained from historical information posteriors. and They represent information entropy and information gain respectively. Information entropy A measure of the complexity of a random variable (i.e., the size of its uncertainty). The more unstable a system is, or the higher the uncertainty of an event, the higher its entropy. Information Gain To measure the t+1 After that, the degree to which information complexity (uncertainty) is reduced. Therefore, Represents the degree to which information complexity (uncertainty) is reduced during an interaction.
[0055] The agent chooses actions that maximize the reduction of information complexity (uncertainty) during each interaction, that is, it encourages the exploration of uncertainty:
[0056]
[0057] Taking information gain as the intrinsic reward function of the agent and superimposing it with the reward function obtained by interacting with the real environment in the outer loop, the reconstructed reward function can be obtained as follows:
[0058]
[0059] rout = r(ζ t ,at)
[0060]
[0061] Where η is the weight coefficient. out is the outer loop reward, r in It is an inner cycle reward.
[0062] In an embodiment of the present application, after constructing a partially observable Markov decision process, it also includes: in an inner loop, establishing a random hidden state transfer model, sampling the probability distribution obeyed by the parameter each time forward inference is performed, and using the sampled parameter value as the parameter value of this forward inference, and using the variational lower bound in the interactive data pool to update the random hidden state transfer model in the inner loop, so as to drive the internal Markov process through the random hidden state transfer model,
[0063] Furthermore, in the inner loop, a random hidden state transfer model is established Defined as the likelihood hidden state ζ t With action a t The parameter θ of the hidden state transfer model is a random variable that follows a Gaussian probability distribution. Each time forward inference is performed, First, the probability distribution of the parameter is sampled, and the sampled parameter value ω is used as the parameter value of this forward inference. is a random model that can satisfy the many-to-many mapping relationship between observations and hidden states, written as:
[0064]
[0065] Among them, θ:={μ,σ} and k are the parameters of the Gaussian model, μ and σ are the expectation and variance respectively, and k represents the number of Gaussian distributions, which together characterize Parameter distribution. Furthermore, to ensure that the variance is positive, σ=log(1+exp(ρ)). Combined with the reparameterization technique, the parameter ω of a single sampling satisfies:
[0066]
[0067] At this time, the random hidden state transfer model The parameter update of can be achieved by maximizing the variational lower bound (ELBO) Get, that is:
[0068]
[0069]
[0070]
[0071] Among them, q(ω|θ) is used as a parameterized distribution to approximate the unknown actual posterior distribution p r (ω) is the prior probability distribution. Representative data. Specifically, formula (9) can be calculated as follows:
[0072]
[0073]
[0074] Among them, in the embodiment of the present application, the expectation of the likelihood function is estimated by using Monte Carlo sampling. M is the number of updated data pairs. Furthermore, the random hidden state transfer model in the inner loop Perform gradient descent update, then The gradient of satisfies the following relationship:
[0075]
[0076]
[0077] The gradient of the model parameters μ and ρ can be written as:
[0078]
[0079]
[0080] Therefore, the update of the parameters of the random hidden state transfer model in the inner loop satisfies the following relationship:
[0081]
[0082]
[0083] Among them, α μ With α ρ are the update step sizes of μ and ρ respectively.
[0084] In summary, the variational lower bound in the interactive data pool is used to model the random hidden state transfer in the inner loop. Therefore, the inner loop transfers the model through random hidden state Drives the internal Markov process.
[0085] In an embodiment of the present application, the optimal strategy for solving partially observable Markov decisions includes: using the information gain in the inner loop as the intrinsic reward function of the intelligent agent, superimposing it with the reward function obtained by interacting with the real environment in the outer loop, and obtaining the reconstructed reward function as the update target of the value network.
[0086] Furthermore, the intrinsic reward of each state transition in the outer loop is calculated. Then the intrinsic reward defined in equation (4) can be written as:
[0087]
[0088] Among them, θ t+1 With θ t Get the new hidden state z t+1 Parameters of the front and back, hidden state models.
[0089] Therefore, the reconstructed reward function is:
[0090] r′(ζ t ,a t ,ζ t+1 )=r(ζ t ,a t )+ηD KL [q(ω;θ t+1 )||q(ω;θ t )] (15)
[0091] In order to avoid single-step updates to the hidden state model, the Newton method is used in the embodiment of the present application to approximate the intrinsic reward of single-step state transfer, and θ is defined. t+1 Satisfies the following relationship:
[0092] θ t+1 =θ t +λΔθ (16)
[0093] The embodiment of the present application considers that the update change of the hidden state model caused by single-step data is small, so it is assumed that Δθ→0, and then equation (14) can be written as:
[0094]
[0095] Among them, the second-order gradient The following numerical calculation relationship is satisfied:
[0096]
[0097]
[0098] First-order gradient can be obtained through the automatic derivation mechanism. Therefore, r incontains the gradient information in the hidden state transfer model, and the single-step reward function can be written as:
[0099]
[0100] In the outer loop t Establish the state value function and execute the strategy evaluation process to obtain Calculate the one-step TD-error as the objective function for updating the value network and the observed state transition model in the outer loop:
[0101]
[0102] Among them, r′(ζ t ,a t ,ζ t+1 ) contains the gradient information of the hidden state model. It is used to guide action exploration when the hidden state model is inaccurate in the early stage of training, and to evaluate the observation state with uncertainty when the hidden state model is more accurate in the later stage of training to avoid overestimation. Derivative the objective function and update the policy function parameters:
[0103]
[0104] in, For parameters The update step size of .
[0105] In the outer loop t Establish a strategy function and execute the strategy improvement process:
[0106]
[0107] ψ′←ψ-β ψ dψ
[0108] Among them, β ψ is the update step size of the parameter ψ. The optimal strategy for partially observable Markov decision problems is obtained by iterating the strategy evaluation and strategy improvement process.
[0109] In step S103, an optimal control strategy for the vehicle is generated based on the optimal strategy of the partially observable Markov decision.
[0110] In an embodiment of the present application, an optimal control strategy for a car is generated based on an optimal strategy of partially observable Markov decision making, including: loading a neural network parameter file corresponding to the optimal strategy, inputting an observation state sequence, and mapping to obtain the optimal control strategy.
[0111] like Figure 3As shown, the method of the embodiment of the present application includes an outer loop and an inner loop, and the gradient information of the random hidden state in the inner loop is used to guide the updating process of the strategy in the outer loop. The agent selects the action that maximizes the reduction of information complexity (uncertainty) in each interaction process, that is, encourages the exploration of uncertainty. Further, the information gain in the inner loop is used as the intrinsic reward function of the agent, and the reward function obtained by interacting with the real environment in the outer loop is superimposed to obtain the reconstructed reward function as the update target of the value network. During offline learning and training, a batch fully connected neural network is used for training, and the stochastic gradient descent method is used to autonomously update the neural network parameters based on the reconstructed reward function information with the goal of minimizing the cost function. Through continuous learning and training, until the change in the objective function converges to the preset error upper limit ε, an approximate optimal strategy is obtained, and the approximate optimal strategy parameters are saved as a local file. When applied online, by loading the neural network parameter file for application, the observation state sequence is input, and the control action can be mapped to achieve a double-loop efficient solution and real-time application of partially observable Markov decision problems.
[0112] According to the dual-loop application method for partially observable Markov decision problems proposed in the embodiment of the present application, it includes an outer loop and an inner loop. The outer loop aims to construct a historical information representation based on the maximum likelihood estimation method to avoid reducing the accumulation of state errors caused by driving the MDP process using the learned model. The inner loop aims to use the gradient information of the hidden state model as the intrinsic reward function to motivate the optimal strategy to learn the uncertainty in the environment and guide action exploration. As a result, it is possible to solve the optimal strategy in uncertain scenarios with partial observability and noise in the observation, and meet the corresponding driving performance, thereby improving the robustness of the solution method to observation uncertainty.
[0113] Next, a double-loop application system for partially observable Markov decision problems proposed in an embodiment of the present application is described with reference to the accompanying drawings.
[0114] Figure 4 A schematic diagram of the structure of a double-loop application system for a partially observable Markov decision problem provided according to an embodiment of the present application.
[0115] like Figure 4 As shown, the double-loop application system 10 for the partially observable Markov decision problem includes: an acquisition module 100 , a solution module 200 and a solution module 300 .
[0116] The acquisition module 100 is used to acquire the observation data of the car. The solution module 200 is used to construct the historical information representation based on the observation data by the maximum likelihood estimation method in the partially observable Markov decision process, and use the gradient information of the hidden state model as the intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision. And the application module 300 is used to generate the optimal control strategy of the car based on the optimal strategy of the partially observable Markov decision.
[0117] Optionally, in one embodiment of the present application, the solution module is further used to determine the hidden state space, action space, observation space, reward function, state transition model in the hidden state space, state transition model between hidden state and observed state, and discount factor based on the observation data; and construct a partially observable Markov decision process according to the hidden state space, action space, observation space, reward function, state transition model in the hidden state space, state transition model between hidden state and observed state, and discount factor.
[0118] Optionally, in one embodiment of the present application, both the observed state transfer model of the outer loop driving the partially observable Markov decision process and the hidden state transfer model of the inner loop driving the partially observable Markov decision process are unknown.
[0119] Optionally, in one embodiment of the present application, after constructing a partially observable Markov decision process, it also includes: a processing module for selecting a sequence of observed states as input in an outer loop, and defining the observed state transfer model in the outer loop as a likelihood function of the hidden state space, so as to drive the external Markov process through the likelihood hidden state of historical observations.
[0120] Optionally, in one embodiment of the present application, after constructing a partially observable Markov decision process, it also includes: an update module, which is used to establish a random hidden state transfer model in an inner loop, sample the probability distribution obeyed by the parameter during each forward inference, and use the sampled parameter value as the parameter value for this forward inference, and use the variational lower bound in the interactive data pool to update the random hidden state transfer model in the inner loop to drive the internal Markov process through the random hidden state transfer model.
[0121] Optionally, in one embodiment of the present application, the optimal strategy for solving the partially observable Markov decision includes: using the information gain in the inner loop as the intrinsic reward function of the intelligent agent, superimposing it with the reward function obtained by interacting with the real environment in the outer loop, and obtaining the reconstructed reward function as the update target of the value network.
[0122] Optionally, in one embodiment of the present application, an optimal control strategy for a car is generated based on an optimal strategy of a partially observable Markov decision, including: loading a neural network parameter file corresponding to the optimal strategy, and inputting an observation state sequence to map and obtain the optimal control strategy.
[0123] It should be noted that the above explanation of the embodiment of the double-loop application method for partially observable Markov decision problems is also applicable to the double-loop application system for partially observable Markov decision problems of this embodiment, which will not be repeated here.
[0124] According to the dual-loop application system for partially observable Markov decision problems proposed in the embodiment of the present application, it includes an outer loop and an inner loop. The outer loop aims to construct a historical information representation based on the maximum likelihood estimation method to avoid reducing the accumulation of state errors caused by driving the MDP process using the learned model. The inner loop aims to use the gradient information of the hidden state model as an intrinsic reward function to motivate the optimal strategy to learn the uncertainty in the environment and guide action exploration. As a result, it is possible to solve the optimal strategy and meet the corresponding driving performance in uncertain scenarios with partial observability and noise in the observation, thereby improving the robustness of the solution method to observation uncertainty.
[0125] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0126] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .
[0127] When the processor 502 executes the program, the double-loop application method for the partially observable Markov decision problem provided in the above embodiment is implemented.
[0128] Furthermore, the electronic device further comprises:
[0129] The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0130] The memory 501 is used to store computer programs that can be executed on the processor 502 .
[0131] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0132] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0133] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0134] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0135] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned double-loop application method for the partially observable Markov decision problem.
[0136] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0137] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0138] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0139] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0140] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0141] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0142] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0143] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A double-loop application method for partially observable Markov decision problems, characterized in that: The following steps are involved: Get the observation data of the car; In the partially observable Markov decision process, based on the observed data, the historical information representation is constructed by the maximum likelihood estimation method, and the gradient information of the hidden state model is used as the intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment, and solve the optimal strategy of the partially observable Markov decision, which specifically includes: Based on the observation data, determine a hidden state space, an action space, an observation space, a reward function, a state transition model on the hidden state space, a state transition model between the hidden state space and the observation space, and a discount factor; Constructing the partially observable Markov decision process according to the hidden state space, the action space, the observation space, the reward function, the state transition model on the hidden state space, the state transition model between the hidden state space and the observation space, and the discount factor; The observed state transfer model driving the outer loop of the partially observable Markov decision process and the hidden state transfer model driving the inner loop of the partially observable Markov decision process are both unknown, wherein, After constructing the partially observable Markov decision process, it also includes: In the outer loop, a sequence of observed states is selected as input, and the observed state transition model in the outer loop is defined as the likelihood function of the hidden state space to drive the external Markov process through the likelihood hidden state of historical observations; In the inner loop, a random hidden state transfer model is established. During each forward inference, the probability distribution obeyed by the parameter is sampled, and the sampled parameter value is used as the parameter value of this forward inference. The random hidden state transfer model in the inner loop is updated using the variational lower bound in the interactive data pool, so as to drive the internal Markov process through the random hidden state transfer model; and Generating the optimal control strategy of the vehicle based on the optimal strategy of the partially observable Markov decision specifically includes: loading the neural network parameter file corresponding to the optimal strategy, inputting the observation state sequence, and mapping to obtain the optimal control strategy.
2. The method according to claim 1, characterized in that The optimal strategy for solving the partially observable Markov decision includes: The information gain in the inner loop is used as the intrinsic reward function of the intelligent agent, and is superimposed with the reward function obtained by interacting with the real environment in the outer loop. The reconstructed reward function is used as the update target of the value network.
3. A double-loop application system for partially observable Markov decision problems, characterized in that: include: An acquisition module is used to obtain observation data of the car; A solution module is used to construct a historical information representation based on the observed data by the maximum likelihood estimation method in a partially observable Markov decision process, and at the same time, use the gradient information of the hidden state model as the intrinsic reward function to encourage the optimal strategy to learn the uncertainty in the environment and solve the optimal strategy of the partially observable Markov decision process, wherein, The solution module is further used to: Based on the observation data, determine a hidden state space, an action space, an observation space, a reward function, a state transition model on the hidden state space, a state transition model between the hidden state space and the observation space, and a discount factor; Constructing the partially observable Markov decision process according to the hidden state space, the action space, the observation space, the reward function, the state transition model on the hidden state space, the state transition model between the hidden state space and the observation space, and the discount factor; The observed state transfer model driving the outer loop of the partially observable Markov decision process and the hidden state transfer model driving the inner loop of the partially observable Markov decision process are both unknown; After constructing the partially observable Markov decision process, it also includes: A processing module is used to select a sequence of observation states as input in an outer loop, and define the observation state transfer model in the outer loop as a likelihood function of a hidden state space, so as to drive an external Markov process through the likelihood hidden state of historical observations; An update module is used to establish a random hidden state transfer model in the inner loop, sample the probability distribution obeyed by the parameter each time forward inference is performed, and use the sampled parameter value as the parameter value of this forward inference, and use the variational lower bound in the interactive data pool to update the random hidden state transfer model in the inner loop, so as to drive the internal Markov process through the random hidden state transfer model; and An application module is used to generate an optimal control strategy for the vehicle based on the optimal strategy of the partially observable Markov decision, specifically comprising: The neural network parameter file corresponding to the optimal strategy is loaded, and the observation state sequence is input to map and obtain the optimal control strategy.
4. The system according to claim 3, characterized in that The optimal strategy for solving the partially observable Markov decision includes: The information gain in the inner loop is used as the intrinsic reward function of the intelligent agent, and is superimposed with the reward function obtained by interacting with the real environment in the outer loop. The reconstructed reward function is used as the update target of the value network.
5. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the double-loop application method for the partially observable Markov decision problem as described in any one of claims 1 to 2.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the double-loop application method for the partially observable Markov decision problem as described in any one of claims 1-2.
Citation Information
Patent Citations
Method for predicting indoor movement trajectory data based on HMM model
CN108882172A
Anthropomorphic automatic driving car-following model based on deep reinforcement learning
CN109733415A