A hierarchical risk control deduction device system for unknown open environment

By combining incentive-based exploration and experience-based strategies in unknown open environments, and utilizing methods such as Markov decision chains and Bellman dynamic programming, hierarchical adaptive behaviors are constructed. This solves the problem of rapidly inferring the deductive characteristics of hierarchical behaviors in unknown environments, thereby improving the adaptability and task completion efficiency of intelligent robots.

CN115841155BActive Publication Date: 2026-03-27FUDAN UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately infer hierarchical behavioral derivation characteristics in unknown open environments, especially when dealing with sparse space decision-making problems, and lack adaptability and lifelong learning capabilities.

Method used

This paper adopts the idea of ​​combining encouragement-based exploration with empirical strategies to approximate the stochastic conditional probability distribution in open systems. By iteratively calculating the confidence level of the dominant strategy in the intermediate incremental buffer, it constructs hierarchical adaptive behavior under unknown risk environments, including the steps of observation layer, analysis layer, judgment layer, confidence layer and iterative reinforcement layer. It uses Markov decision chain, Gibbs sampling and Bellman dynamic programming methods for risk assessment and strategy updating.

Benefits of technology

It enables faster and more accurate inference of hierarchical behavioral derivation characteristics in unknown open environments, improves the adaptability and lifelong learning ability of intelligent robots in complex problems, finds a balance between exploration and utilization, and enhances environmental adaptability and task completion efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115841155B_ABST
    Figure CN115841155B_ABST
Patent Text Reader

Abstract

The present application relates to a hierarchical risk control deduction device system for an unknown open environment, which approximates the risk random probability distribution in the open system by fine-tuning the inference conditional random field transition state, and maximally evaluates the adaptive confidence level of the dominant advantage strategy. The present application clarifies the internal relationship between the target condition policy and the prediction process, iteratively calculates the incremental buffer and modifies its response to the open environment, and enables it to handle randomness throughout the autonomous learning process. In the open environment, this hierarchical structure is easier to implement because the efficiency of the scheme is higher and the amount of calculation consumed is less. The present application realizes the effective abstraction of the potential risk of the environment, proves the actual potential of risk estimation and reasoning in the robot task, and further improves the exploration efficiency and improves the effectiveness and interpretability of the hierarchical architecture.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a hierarchical risk control deduction device system for an unknown open environment, belonging to the technical field of artificial intelligence robots. BACKGROUND

[0002] In the field of artificial intelligence, autonomous learning systems are usually designed from three aspects of environment cognition, behavior strategy, interaction and reasoning, and their adaptability to unknown environmental changes is enhanced through continuous learning. The ability of humans and many intelligent organisms to solve complex problems is usually manifested as a process of learning from hierarchical cognitive mechanisms, meaning that hierarchical adaptive responses are crucial to the development of autonomous learning and reasoning in biology and cognition.

[0003] In the real world, building an effective autonomous learning not only needs feedback of environmental rewards, but also considers the uncertainty of the environment. Establishing effective identification and cognitive reasoning of unknown environments is a necessary step for autonomous learning. The advantage of open systems is that intelligent robots can judge whether there are similar risk features in the current scene, learn from the obtained experience sequence, and then generalize to multi-modal scenes, thereby guiding the specific reinforcement of certain functions, such as robot exploration tasks, Go games, and morphological evolution. Hierarchical prediction planning can handle the balance between sparse rewards and sufficient risk cognition based on continuous space deduction behaviors with intrinsic motivation, which helps to find environment-adaptive behaviors that can improve overall benefits when dealing with large-scale state-action sparse space decision problems. However, this new perspective is currently only applicable to tasks that are the same or similar to previous tasks or tasks in simple generative domains. SUMMARY

[0004] The present application is made to solve the above problems, and aims to provide a hierarchical risk control deduction device system for an unknown open environment, which can more accurately, quickly and widely infer hierarchical behavior deduction characteristics under different environmental risks. To this end, the present application provides the following technical solutions:

[0005] The present application provides a hierarchical risk control deduction device system for an unknown open environment, which combines the ideas of encouraging exploration and experience strategy to approximate the random conditional probability distribution in the open system, iteratively calculates the confidence level of the dominant advantage strategy in the intermediate incremental buffer, and constructs hierarchical adaptive behavior under unknown risk environment. It has the following characteristics, including the following steps:

[0006] Step S1, observation layer: import real-time sampling observation sequence of environment information;

[0007] Step S2, analysis layer: construct historical experience sequence of action observation;

[0008] Step S3, judgment layer: risk event trigger detection and failure judgment;

[0009] Step S4, confidence layer: generate inference model and update confidence interval;

[0010] Step S5, iterative reinforcement layer: cross-validation of real and simulation in complex multi-modal system, constantly backtrack and evaluate confidence interval, and feedback real-time sampling observation sequence to step S1 for repeated iteration.

[0011] In the hierarchical risk control deduction device system for unknown open environment provided by the application, the observation sequence in step S1 can be a decentralized partially observable Markov decision chain G, which includes state s∈S, action a∈A, and observation sampling sequence is the sequence of state i→j transition in O(s,a):S×A→Z at time t, reward function R t ∈R(s,a) and transition condition

[0012] In the hierarchical risk control deduction device system for unknown open environment provided by the application, the observation sequence in step S1 can be a decentralized partially observable Markov decision chain G, which includes state s∈S, action a∈A, and observation sampling sequence

[0013] Step S2-1, in the case of significant non-stationary, due to environmental bias and exponential state space calculation, the cumulative reward is as follows:

[0014]

[0015] Where γ is the step discount factor;

[0016] Step S2-2, according to the observable local historical action-observation experience sequence Encode the action trajectory, the goal is to generate the strategy π(a|s)∝exp{Q(s t ,a t )}, input the action-observation history of the local trajectory, and estimate to produce a joint action;

[0017] Step S2-3, assuming that the traversal joint probability distribution of the recommended action performed in response to the risk in the experience sequence is: So that the random strategy π(x)=(π1,π2,…) satisfies The conditional transition probability, there exists a unique normalized distribution such that the approximation holds, that is:

[0018]

[0019] Step S2-4, for the observable space (S, A), the initial Gibbs sampling sequence is Xn = (x i : i = 1, 2, …, n), at time t, according to the action-observation history generate A t (s, u t | τ t ) as the optimal response action.

[0020] In the hierarchical risk control deduction device system for unknown open environment provided by the application, the step S3 can further include the following sub-steps.

[0021] Step S3-1, in the open system, it is assumed that each A t (s, u t | τ t ) follows a Bernoulli distribution, and the utility value of the joint action space (u t | τ t ) in the state space s ~ z i→j is approximately equal to the sum of the utility values of the action-observation history sequence τ t , and the proposal distribution of the optimal response is as follows:

[0022]

[0023] wherein, is the global reward Q n obtained in a limited N steps, and the confidence degree derived by iteration is irrelevant to the state space s;

[0024] Step S3-2, in order to solve the problems of environmental bias and exponential state space calculation in the non-stationary scene, the idea of combining encouraging exploration with experience strategy is adopted, and a group of weights for exploration interaction is allocated to the sparse environmental reward of each iteration:

[0025]

[0026] wherein, the weight value is determined by real-time opportunistic environmental exploration, and the global reward under the current risk condition is updated as follows:

[0027]

[0028] wherein, by sampling , the strategy can be generated (note: the confidence degree: ), since is always true, so the loss expectation of unknown environmental changes estimated by sampling samples is proportional to the gradient of the real unknown risk;

[0029] Step S3-3, when and only when the advantage policy of the historical action observation sequence is no longer applicable, the exploration degree of the agent to the external environment should be increased.

[0030] In the hierarchical risk control deduction device system for unknown open environment provided by the application, the step S4 can further include the following sub-steps:

[0031] Step S4-1, based on sufficient observation of the local historical action sequence , the guide policy is updated from the part of the trajectory with better expected return, so as to improve the utilization of existing experience, and the high-level advantage policy is effectively used for the joint policy update of the next iteration;

[0032] Step S4-2, since it is difficult to approach the ideal global Pareto optimal value on the sparse action space, the introduction of the maximum average difference can improve the utilization of data, and the real-time policy is evaluated and improved through self-imitation learning, so as to ensure that the algorithm amplifies the most certain advantage experience in the self-training process, and improves the collective knowledge of the unmarked target domain.

[0033]

[0034] Among them, The calculation is simplified by the Gaussian kernel without odd sub-terms, φ(·) is a classifiable target function, and K(·, ·) can be simply estimated by kernel density estimation, And Sampling from individual p and A respectively helps to reduce the high dimension;

[0035] Step S4-3, by defining the Beta distribution confidence of the cumulative reward as follows, the expression can match the k-level best response In step S4-1, and ensure that the expected return of each action is not affected by random noise, so as to construct the non-prior perfect Bayesian condition:

[0036]

[0037] Among them, b is an incremental buffer for evaluating the confidence level of the joint strategy, and is subject to binomial Beta b (win,lose). This means By selecting the binomial distribution between the joint advantage policy And the random strategy, the credibility of the reward obtained by the current joint strategy is calculated. Among them, the random strategy Refers to the case that the environment reward cannot be effectively obtained under the condition of configuring the prior parameter.

[0038] In the hierarchical risk control deduction device system for unknown open environment provided by the application, the step S5 can further comprise the following sub-steps:

[0039] In step S5-1, the agent can use the Bellman dynamic programming equation to further promote the thorough exploration of the unknown open scene, deduce the real probability distribution conforming to the current state and make the best response based on the real probability distribution;

[0040] In step S5-2, the action with higher environmental reward is given priority under the same confidence condition, and the Boltzmann distribution with a non-zero likelihood is allocated to update the joint advantage policy;

[0041] In step S5-3, the joint advantage policy is updated iteratively by the probability distribution of the step cumulative utility each time, and finally converges to an equilibrium state through continuous interaction with the environment.

[0042] Effects of the application

[0043] According to the hierarchical risk control deduction device system for unknown open environment, analogy reasoning learning is performed based on the historical sequence of action observation, and the hierarchical deduction device helps the intelligent robot to interact with the environment in the experience-cognition combination when predicting the risk response strategy, so as to find a proper balance between exploration and utilization of the two learning elements. In addition, the experimental results show that the application allows the intelligent robot to evaluate the possible conditional probability transition state and the best response confidence level by finding the similarity distance between different domain migration tasks, so as to continuously integrate the hierarchical cognitive mechanism and adapt to the environment and improve the skills of solving complex problems, so the application has adaptive and lifelong learning intelligent behavior mode. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a flowchart of the hierarchical risk control deduction device system in the embodiment of the application;

[0045] Figure 2 is a flowchart of the intelligent robot assembly kit as a model training and risk deduction device in the embodiment of the application;

[0046] Figure 3 is a flowchart of step S2 of the hierarchical risk control deduction device system in the embodiment of the application;

[0047] Figure 4 is a cognitive effect diagram of the current environment potential risk and available reward in the embodiment of the application;

[0048] Figure 5is a flow chart of step S3 of the hierarchical risk control deduction device system in the embodiment of the present application;

[0049] Figure 6 is a flow chart of step S4 of the hierarchical risk control deduction device system in the embodiment of the present application;

[0050] Figure 7 is a flow chart of step S5 of the hierarchical risk control deduction device system in the embodiment of the present application; and

[0051] Figure 8 is an effect schematic diagram of the best response (see arrow) of the hierarchical risk control device system in the embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the technical means, creative features, purposes and effects of the present application easy to understand, the following embodiments will be specifically described in combination with the drawings for the hierarchical risk control deduction device system for unknown open environment.

[0053] <EMBODIMENT>

[0054] Figure 1 is a flow chart of the hierarchical risk control deduction device system in the embodiment of the present application.

[0055] Figure 2 is a flow chart of the intelligent robot assembly kit as a model training and risk deduction device in the embodiment of the present application;

[0056] As Figure 2 shown, an embodiment application scene of the hierarchical risk control deduction device system for unknown open environment selects the intelligent robot assembly kit under the Jetson Nano system platform as the model training and risk deduction device. In the embodiment, the intelligent robot has a stable low-delay system, is provided with four sensor switching modules, can obtain dynamic images and sensor data information in the environment, is configured with a McLaren universal wheel at the bottom to realize autonomous travel control, is equipped with a lifting flexible telescopic mechanical claw at the top, can realize intelligent obstacle avoidance and environment perception, and demonstrates the generation of hierarchical deduction behavior in an unknown risk scene.

[0057] In the embodiment, the idea of combining encouraging exploration with experience strategy is adopted to approximate the random condition probability distribution in the open system, and the confidence level of the dominant advantage strategy in the intermediate incremental buffer is iteratively calculated to construct the hierarchical adaptive behavior in the unknown risk environment, which specifically includes the following steps:

[0058] Step S1, observation layer: import real-time sampling observation sequence of environment information.

[0059] Import a decentralized partially observable Markov decision process G as an observation sequence, which includes state s∈S, action a∈A, and observation sampling sequence is the sequence of state i→j transitions in O(s,a):S×A→Z at time t, and reward function R t ∈R(s,a) and transition condition P(s,a).

[0060] Figure 3 is the flow chart of step S2 of the hierarchical risk control deduction device system in the embodiment of the application.

[0061] As shown in the figure, step S2, analysis layer: build a history experience sequence of action observation. Figure 3

[0062] Step S2-1, in the case of significant non-stationarity, due to environmental bias and exponential state space calculation, the cumulative reward is as follows:

[0063]

[0064] Where γ is the step discount factor.

[0065] Step S2-2, encode the action trajectory according to the observable local history action-observation experience sequence. The goal is to generate a strategy π(a|s)∝exp{Q(s t ,a t )} which takes the local trajectory's action-observation history as input and estimates to produce a joint action.

[0066] Step S2-3, assume that the traversed joint probability distribution of the recommended action performed in response to the risk in the experience sequence is: So that the random strategy π(x)=(π1,π2,…) satisfies the conditional transition probability, then there exists a unique normalized distribution such that the approximation holds, that is:

[0067]

[0068] Figure 4 is the cognitive effect diagram of the potential risk and available reward of the current environment in the embodiment of the application, wherein the intelligent robot is initially located at A, needs to detect the possible routes of the potential risks O1 and O2 (such as Figure 2 circles) to go out in real time, and complete the complex task of picking up the water bottle at B and placing it in the grid at C.

[0069] Step S2-4, for the observable space (S, A), the initial Gibbs sampling sequence is X n =(x i ​: i = 1, 2, …, n), at time t, according to the action-observation history generated A t (s, u t | τ t ) as the best response action performed by the intelligent robot as shown in Figure 2 .

[0070] Figure 5 is a flow chart of step S3 of the hierarchical risk control deduction device system in the embodiments of the present application.

[0071] As shown in Figure 5 , step S3, judge layer: risk event trigger detection and failure judgment.

[0072] Step S3-1, in an open system, it is assumed that each A t (s, u t | τ t ) follows a Bernoulli distribution. The utility value of the joint action space (u t | τ t ) in the state space s ~ z i→j can be approximated as the sum of the utility values of the action-observation history sequence τ t . Then the proposal distribution of the best response is as follows:

[0073]

[0074] wherein, is the global reward Q n obtained within a limited N steps, and the confidence degree derived by iteration is irrelevant to the state space s.

[0075] Step S3-2, in order to solve the problems of environmental bias and exponential state space calculation in non-stationary scenarios, the idea of combining encouraging exploration with experience strategy is adopted, and a group of weights for exploration interaction is assigned to the sparse environmental reward of each iteration:

[0076]

[0077] The value is determined by real-time opportunistic environmental exploration. Then the global reward under the current risk condition is updated accordingly as follows:

[0078]

[0079] wherein, by sampling , the strategy (remark: confidence degree: b ∈ B) can be generated.

[0080] Since The constant is established, so that the loss expectation of the unknown environmental change estimated by sampling the sample is proportional to the gradient of the true unknown risk.

[0081] Step S3-3, when and only when the advantage policy of the historical action observation sequence is no longer applicable, the exploration degree of the intelligent robot to the external environment should be increased.

[0082] Figure 6 The flowchart of step S4 of the hierarchical risk control deduction device system in the embodiment of the application.

[0083] As shown in Figure 6 Step S4, confidence layer: generate inference model and update confidence interval;

[0084] Step S4-1, based on sufficient observation of the local historical action sequence , the update is carried out from the part of the trajectory with better than expected returns by guiding the policy, so as to improve the utilization of existing experience, so as to effectively utilize the high-level advantage policy for the joint policy update of the next iteration.

[0085] Step S4-2, since it is difficult to approach the ideal global Pareto optimal value on the sparse action space, the introduction of the maximum average difference can improve the utilization of data, and the real-time policy is evaluated and improved through self-imitation learning, so as to ensure that the algorithm amplifies the most certain advantage experience in the self-training process, so as to improve the collective knowledge of the unmarked target domain.

[0086]

[0087] Among them, The calculation is simplified by a Gaussian kernel without odd-numbered sub-terms, φ(·) is a classifiable target function, and K(·,·) can be simply estimated by kernel density estimation. And Sampling from individuals p and A respectively helps to reduce the high dimension.

[0088] Step S4-3, by defining the Beta distribution confidence of the cumulative reward as follows, the expression can be matched with the k-level best response A * in step S4-1, and it is ensured that the expected return of each action is not affected by random noise, so as to build a non-prior perfect Bayesian condition.

[0089]

[0090] Among them, b is an incremental buffer for evaluating the confidence level of the joint strategy, and is subject to binomial Beta b (win,lose). This means that B selects the joint advantage policy The confidence of the reward obtained by the current joint strategy is calculated between the binomial distribution of the random strategy. refers to the case that the environment reward cannot be effectively obtained under the condition of configuring the prior parameters.

[0091] Figure 7 is the flowchart of step S5 of the hierarchical risk control deduction device system in the embodiment of the present application;

[0092] Figure 8 is the effect diagram of the optimal response (see arrow) of the hierarchical risk control device system of the preferred embodiment of the present application.

[0093] As shown in Figure 7 , step S5, an iterative reinforcement layer based on the iteration reinforcement layer as shown in Figure 3 is constructed: real and simulation cross-validation is performed in a complex multi-modal system, and the confidence interval is constantly traced back and evaluated, and the real-time observation sequence is fed back to step S1 for repeated iteration.

[0094] Step S5-1, the intelligent robot can use the Bellman dynamic programming equation to further promote the thorough exploration of the unknown open scene, such as detecting the potential risk change as shown in Figure 8 the upper left corner O1, deducing the real probability distribution conforming to the current state (i.e. the probability of O1 starting to move outward and the possible travel route), and formulating the hierarchical optimal response as shown in Figure 8 .

[0095] Step S5-2, under the same confidence condition, the action with higher environmental reward is given priority, and a non-zero Boltzmann distribution is allocated to update its joint advantage strategy Specifically, in the present embodiment, the water bottle at B needs to be picked up by the intelligent robot and placed at C, so the L1 arrow represents the optimal solution under the current O1 risk; at the same time, considering the possible travel time and movement route of O1, the buffer in the BUFFER block provides a safer suboptimal strategy for the intelligent robot, that is, the intelligent robot picks up the water bottle at B by advancing, then enters the buffer according to the direction of the L2 arrow, and further observes the current position of the risk O1 and predicts its next travel direction, and then uses the universal go-kart to formulate a relocation route from the buffer to the end point C, selects the shortest path (such as L2 arrow) to reach C and puts the water bottle into the grid, and completes the task.

[0096] Step S5-3, the joint advantage strategy at each time is the probability distribution of the step cumulative utility Iterative update is performed, and finally, through continuous interaction with the environment, the intelligent robot converges to an equilibrium state, successfully completes the complex task of picking and placing by identifying potential risks in the environment.

[0097] Effects of the embodiments

[0098] According to the hierarchical risk control deduction device system for an unknown open environment, the cognitive mechanism integrating combined abstraction and predictive processing performs analogical reasoning learning based on the historical sequence of action observation, and the hierarchical deduction device helps the intelligent robot to interact with the environment in the experience-cognition combination when predicting the risk response strategy, so as to find a proper balance between exploration and utilization of the two learning elements. In addition, the embodiment results show that the intelligent robot can evaluate the possible conditional probability transition state and the optimal response confidence level by finding the similarity distance between different domain migration tasks, so as to continuously integrate the hierarchical cognitive mechanism and adapt to the environment and improve the skills of solving complex problems, so the intelligent behavior mode of the application has adaptability and lifelong learning.

[0099] In summary, the hierarchical risk control deduction device system for an unknown open environment in the embodiment is easier to realize in an open environment because the efficiency of the scheme is higher and the amount of calculation consumed is less. The application realizes effective abstraction of potential risks in the environment, proves the actual potential of risk estimation and reasoning in robot tasks, and further improves the exploration efficiency and improves the effectiveness and interpretability of the hierarchical architecture.

[0100] The above implementation is a preferred case of the application and does not limit the protection scope of the application.

Claims

1. A hierarchical risk control deduction device system for unknown open environment, which combines the idea of encouraging exploration and experience strategy to approximate the random condition probability distribution in the open system, iteratively calculates the confidence level of the dominant advantage strategy in the intermediate incremental buffer to construct the hierarchical adaptive behavior in the unknown risk environment, characterized in that, Includes the following steps: ​ Step S1, Observation Layer: Import real-time sampling observation sequences of environmental information; Step S2, Analysis Layer: Constructing a historical experience sequence of action observations; Step S3, Judgment Layer: Risk Event Trigger Detection and Failure Judgment; Step S4, Confidence Layer: Generate the inference model and update the confidence interval; Step S5, Iterative Reinforcement Layer: Performing physical and simulation interactions in complex multimodal systems. Cross-validation is performed, continuously backtracking and evaluating the confidence interval, and the real-time sampling observation sequence is fed back into step S1 for repeated iterations. Step S3 specifically includes the following sub-steps: Step S3-1, in open systems, assume each follows Bernoulli distribution, joint action space The utility value under the state space The utility value of the observed history sequence The proposal distribution of the best response is as follows: = , wherein, is finite Global reward is obtained within steps While the confidence derived from the iterative inference is independent of the state space ​ Step S3-2: To address the issues of environmental bias and exponential state space computation in non-stationary scenarios, a combination of incentive-based exploration and empirical strategies is adopted. In each iteration, the sparse environmental reward is assigned a set of weights for exploring interactions. , The weight value is determined by real-time opportunistic environment exploration, and the global reward under the current risk conditions is updated accordingly as follows: , wherein the sampling A strategy can be generated Confidence: , Because Constantly true, so that the loss of the unknown environment changes estimated by sampling the sample is directly proportional to the gradient of the true unknown risk; Step S3-3: If and only if the advantageous strategy based on the historical action observation sequence is no longer applicable, the agent should increase its exploration of the external environment. Specifically, step S5 includes the following sub-steps: In step S5-1, the agent can use Bellman dynamic programming equations to further promote the thorough exploration of the unknown open scenario, deduce the true probability distribution that conforms to the current state, and formulate the best response accordingly. Step S5-2: Under equal confidence conditions, prioritize actions with higher environmental rewards and update their joint advantage strategies by assigning a Boltzmann distribution with non-zero likelihood. In step S5-3, each joint advantage strategy is iteratively updated by accumulating the probability distribution of the utility by step size, and finally converges to an equilibrium state through continuous interaction with the environment.

2. The hierarchical risk control analysis device system for unknown open environments according to claim 1, characterized in that: wherein, The observation sequence in the step S1 is a decentralized part of an observable Markov decision chain wherein the state is , the action is , the observation sampling sequence is a sequence of states transitions according to the reward function and the transition condition .

3. The hierarchical risk control analysis device system for unknown open environments according to claim 1, characterized in that: wherein Step S2 specifically includes the following sub-steps: Step S2-1: Under significantly non-stationary conditions, due to environmental bias and exponential state-space calculation, the cumulative reward is as follows: , wherein, is a step discount factor; Step S2-2, encode the action-observation history sequence to generate a policy taking the action-observation history of the local trajectory as input and estimating a joint action; Step S2-3, assuming the traversed joint probability distribution of the recommended actions to be performed in response to risks in the empirical sequence is: such that the random policy satisfies the condition transition probabilities, then there exists a unique normalizing distribution such that the approximation holds, i.e.: , Step S2-4, for the observable space , the initial Gibbs sampling sequence is , at time , the action-observation history is generated , as the best response action.

4. The hierarchical risk control analysis device system for unknown open environments according to claim 1, characterized in that: wherein Step S4 specifically includes the following sub-steps: Step S4-1, based on sufficient observation of the local history action sequence , the utilization of existing experience is improved to effectively utilize the high-level advantage policy for the joint policy update of the next iteration by updating the guide policy from the part of the trajectory with better-than-expected returns. Step S4-2: Since it is difficult to approach the ideal global Pareto optimal value in a sparse action space, the introduction of the maximum mean difference can improve the utilization of data. Real-time policy evaluation and improvement are carried out through self-imitation learning to ensure that the algorithm amplifies the experience of the most certain advantage during self-training, so as to improve the collective knowledge of the unlabeled target domain. , where, Simplifying the computation by using a Gaussian kernel without odd sub-items, is a classifiable objective function, while can be simply estimated by kernel density estimation, respectively from individual and A sampling, which helps to reduce the high dimensionality; Step S4-3, the cumulative reward is defined as distribution confidence, the expression can be matched with the level best response each action is not affected by random noise, thus building a perfect Bayesian condition without priori: , wherein, As an incremental buffer to assess the confidence level of the joint policy, the binomial This means By selecting the joint dominant policy The credibility of the reward obtained by the current joint policy is calculated by the binomial distribution between the joint policy and the random policy, wherein the random policy It refers to the case that the environment reward cannot be effectively obtained under the condition of configuring the prior parameter.

Citation Information

Patent Citations

  • Incentive control for multi-agent systems

    US20210319362A1