Self-adaptive decision-making method based on posterior and diversity collaborative task sampling

By combining task risk prediction model and diversity regularization screening criteria, the problems of low efficiency and insufficient robustness of adaptive decision-making in the existing technology are solved, efficient adaptive decision-making in complex scenarios such as robot control are achieved, and the adaptability and robustness of the system are improved.

CN120258079APending Publication Date: 2025-07-04TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510289738.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing adaptive decision-making methods have problems with inefficient efficiency and insufficient robustness in fields such as robotic control, especially when performing strategy evaluation in large-scale environments and are difficult to adapt quickly in unseen but similar scenarios.

Method used

Adaptive decision-making method based on posterior and diversity collaborative task sampling is adopted, combined with task risk prediction model and diversity regularization screening criterion, task risk assessment is optimized through Bayesian criterion and rheological inference, and strategies are dynamically adjusted to improve the adaptability and robustness of the system.

Benefits of technology

It improves the system performance and adaptability in complex scenarios such as robot control, can efficiently make adaptive decisions in diversified task scenarios, reduces additional environmental interaction costs, and improves the robustness and generalization capabilities of the strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258079A_ABST
    Figure CN120258079A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive decision-making method based on posterior and diversity collaborative task sampling, and the method comprises the steps: carrying out the sampling from robot control task distribution in each decision-making model training, obtaining candidate training tasks, carrying out the sampling of each screened training task through employing a current sampling strategy, and generating training data, the parameters of the decision model are updated by using the training data; inputting the training data into the task risk prediction model so as to calculate an approximate evidence lower bound ELBO loss based on an encoder-decoder architecture, and updating parameters of the task risk prediction model according to an ELBO loss function obtained through calculation; and inputting test data into the updated decision model, and optimizing decision information based on a task risk assessment result fed back by the updated task risk prediction model so as to output a final task decision result of the robot control task. In a complex application scene of robot control, the overall performance and adaptability of the system can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and particularly to an adaptive decision-making method based on posterior and diversity collaborative task sampling. Background Art

[0002] Deep reinforcement learning (RL) has made remarkable progress in solving complex sequential decision-making problems in the past few years. However, one existing challenge is how to effectively transfer the reinforcement learning policy to unseen but similar scenarios without having to learn from scratch, i.e., adaptive decision-making. Adaptive decision-making has important applications in fields including robot control, autonomous driving, etc.

[0003] A common strategy in adaptive decision-making is to randomize the environment. For example, a distribution is imposed on the Markov decision process (MDP) to perform policy search in a zero-shot or few-shot manner. This has promoted the rise of the research paradigms of domain randomization (DR) and meta-reinforcement learning (Meta-RL), which train adaptive policies through periodic learning of tasks. At the same time, there is an increasing concern about the adaptation robustness to the worst-case scenario, because most real-world decision-making scenarios are inherently risk-sensitive, and adaptation failure may lead to catastrophic consequences, such as robot damage or accidents in autonomous driving.

[0004] The prospect of active inference in adaptive robust decision-making: When the risk aversion principle is integrated into domain randomization (DR) and meta-reinforcement learning (Meta-RL) to enhance adaptation robustness, preferentially selecting challenging tasks during the optimization process usually requires a large number of iterations in a large-scale environment, resulting in intensive and expensive policy evaluations. To overcome this efficiency bottleneck, a method has been proposed previously to construct a risk prediction model for actively inferring the difficulty of the Markov decision process (MDP) and selecting the worst-case subset in policy search. This method is identified as the robust active task sampling (RATS) paradigm.

[0005] A typical work of this paradigm is model predictive task sampling (MPTS), which adopts a task scenario learning training paradigm. In each iteration, first MDPs are randomly selected from the task space, then the risk prediction model is used to predict the difficulty of each MDP, and a task acquisition function is constructed based on the upper confidence bound UCB rule to screen out tasks from MDPs, and these A task training strategy is proposed to better learn the adaptive strategy without increasing the additional cost of environmental interaction. This method demonstrates the great potential of RATS in efficient and robust adaptive decision-making, especially when it is difficult to conduct exhaustive policy evaluation in a large task space.

[0006] Although RATS shows great potential in decision-making, several problems can still be found in its latest state-of-the-art method, MPTS: (i) Currently, no general tool suitable for theoretical analysis has been developed, such as the concept of robustness in optimization, which is an indispensable consideration in risk-averse scenarios. (ii) This method requires specific pseudo-batch sizes / candidate task batch sizes and other configurations. This makes it crucial to properly mix random sampling and predictive sampling during the task subset selection process, otherwise, the generalization ability and robustness will be reduced. (iii) In the research of RATS, the discussion on sampling principles is still insufficient. The currently adopted upper confidence bound (UCB) principle needs to balance the worst-case scenario and uncertainty, and its hyperparameters are difficult to adjust in practice. These defects make it difficult to apply the existing work to practical adaptive decision-making scenarios represented by robot control. Summary of the Invention

[0007] The present invention aims to solve at least one of the technical problems in the related art to some extent.

[0008] The present invention proposes an adaptive decision-making method based on posterior and diversity collaborative task sampling, which combines posterior and diversity collaborative task sampling, a dynamic adjustment mechanism, and risk assessment optimization, enabling the system to perform excellently in diverse and uncertain task scenarios and possessing higher adaptability and robustness.

[0009] Another object of the present invention is to propose an adaptive decision-making device based on posterior and diversity collaborative task sampling.

[0010] To achieve the above object, on the one hand, the present invention proposes an adaptive decision-making method based on posterior and diversity collaborative task sampling, including:

[0011] Obtain the robot control task distribution in the adaptive decision-making system;

[0012] In each training of the decision model, sample candidate training tasks from the robot control task distribution, for each selected training task, sample training data using the current sampling strategy, and update the parameters of the decision model using the training data; wherein, the training data includes task features and behavioral patterns under the current sampling strategy.

[0013] Using Bayesian criteria and variational inference to obtain an approximate evidence lower bound, and inputting training data into a task risk prediction model to calculate the approximate evidence lower bound ELBO loss based on an encoder-decoder architecture, and updating the parameters of the task risk prediction model according to the calculated ELBO loss function;

[0014] Inputting test data into the updated decision model, and optimizing decision information based on the task risk assessment result fed back by the updated task risk prediction model to output the final task decision result of the robot control task.

[0015] The adaptive decision-making method based on posterior and diversity collaborative task sampling according to the embodiment of the present invention may further have the following additional technical features:

[0016] In an embodiment of the present invention, the method further includes modeling the robot control task process as a Markov decision process; wherein, the state space is feasible machine learning parameters; the action space includes a training task subset; the transition dynamics are updated policy parameters transferred based on the previous moment's policy parameters and training task batches; the reward function quantifies the improvement of adaptability robustness after state transition.

[0017] In an embodiment of the present invention, the method further includes a task acquisition function enhanced based on diversity regularization, and according to selecting the worst task, adaptation risk, and increasing the task candidate set Evaluating the task batch includes:

[0018] Constructing a diversity maximization problem:

[0019]

[0020] where τ i represents the task identifier of task i, represents the training task batch, represents the candidate task batch, represents the task subset, represents the candidate task set, represents the task subset the acquisition function value of, measures the diversity of the subset by calculating through the pairwise distance d of the identifiers: γ represents the diversity trade-off coefficient.

[0021] In an embodiment of the present invention, the posterior sampling strategy is adopted as the acquisition principle:

[0022] z t ~q φ (z t |Ht )

[0023]

[0024] Among them, H t is the historical information at time t; φ is the encoder model in the risk prediction model, which maps the historical information H t to a latent variable z t , q represents the corresponding distribution; l represents the true task risk, represents the predicted task risk; represents the task identifier of task i in the candidate task set; ψ is the decoder model in the risk prediction model, which maps the latent variable t and the task identifier to the task risk prediction distribution of task i, and p represents the corresponding distribution; represents the optimal training task batch obtained at time t + 1.

[0025] In one embodiment of the present invention, the approximate evidence lower bound ELBO is obtained by using Bayes' criterion and variational inference as the training objective:

[0026]

[0027] Among them is the fixed conditional prior from the last update, β is the penalty weight, D KL represents the Kullback-Leibler divergence.

[0028] It can be understood that the present invention improves the system performance and adaptability in complex scenarios such as robot control through efficient training task screening. The core lies in combining the task risk prediction model and unique screening criteria to achieve precise selection of training tasks. Based on this, the specific steps are as follows:

[0029] First, determine the robot control task distribution in the adaptive decision-making system, providing a basis for subsequent training task sampling.

[0030] In each training of the decision-making model, candidate training tasks are sampled from the task distribution. The task risk prediction model is used to evaluate the risks of the candidate tasks and predict the risk levels of each task. According to the output of the task risk prediction model and the screening criteria designed in the present invention (such as selecting the worst tasks, adapting to risks, and increasing the task candidate set, etc.), the most suitable training task batch is screened out from the candidate tasks. This process ensures that the training tasks are not only challenging but also cover a wider task space, thereby improving the adaptability and robustness of the system.

[0031] For each selected training task, use the current sampling strategy to generate training data, including task features and behavior patterns.

[0032] Use the generated training data to update the parameters of the decision-making model and optimize the model performance.

[0033] Input the training data into the task risk prediction model. Based on the encoder-decoder architecture, use the Bayesian criterion and variational inference to calculate the approximate evidence lower bound (ELBO) loss.

[0034] Update the parameters of the task risk prediction model according to the ELBO loss function to further improve its prediction accuracy.

[0035] In the test phase, input the test data into the updated decision-making model. Based on the task risk assessment results feedback by the task risk prediction model, optimize the decision-making information and output the final task decision result.

[0036] Among them, the task risk prediction model, which is the core tool for training task screening, provides a quantitative basis for the screening process by predicting the risk level of tasks. It can identify the most challenging tasks, thus guiding the system to focus on these tasks during the training process and improving the adaptability and robustness of the system.

[0037] Among them, the present invention designs diverse screening criteria, including selecting the worst tasks, adapting to risks, and increasing the task candidate set, etc. These criteria not only consider the difficulty of tasks but also take into account the diversity and coverage of tasks, ensuring that the system can perform well in a wide range of scenarios.

[0038] By combining the task risk prediction model and the screening criteria, the present invention can efficiently screen out the most suitable training tasks, thereby achieving efficient adaptive decision-making in complex application scenarios.

[0039] To achieve the above object, on the other hand, the present invention proposes an adaptive decision-making device based on posterior and diversity collaborative task sampling, including:

[0040] A task distribution determination module for obtaining the robot control task distribution in the adaptive decision-making system;

[0041] A decision model training module for sampling candidate training tasks from the robot control task distribution in each training of the decision model, sampling and generating training data for each selected training task using the current sampling strategy, and updating the parameters of the decision model using the training data; wherein, the training data includes task features and behavior patterns under the current sampling strategy;

[0042] A risk prediction model update module, which is used to obtain an approximate evidence lower bound by using Bayesian criteria and variational inference, and input training data into the task risk prediction model to calculate the approximate evidence lower bound ELBO loss based on the encoder-decoder architecture, and update the parameters of the task risk prediction model according to the calculated ELBO loss function;

[0043] A task decision output module, which is used to input test data into the updated decision model, and optimize decision information based on the task risk assessment result feedback by the updated task risk prediction model, so as to output the final task decision result of the robot control task.

[0044] The adaptive decision-making method and device based on posterior and diversity collaborative task sampling in the embodiments of the present invention improve the task coverage and diversity, and ensure that the model can cope with diverse actual task requirements. Especially in complex application scenarios of robot control, it performs particularly well and can effectively improve the overall performance and adaptability of the system.

[0045] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0046] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0047] Figure 1 is a flowchart of an adaptive decision-making method based on posterior and diversity collaborative task sampling according to an embodiment of the present invention;

[0048] Figure 2 is a structural diagram of an adaptive decision-making device based on posterior and diversity collaborative task sampling according to an embodiment of the present invention. Detailed Embodiments

[0049] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] An adaptive decision-making method and device based on posterior and diversity collaborative task sampling according to an embodiment of the present invention will be described below with reference to the accompanying drawings.

[0052] Figure 1 is a flowchart of an adaptive decision-making method based on posterior and diversity collaborative task sampling according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0053] S1. Obtain the robot control task distribution in the adaptive decision-making system;

[0054] S2. In each training of the decision-making model, sample candidate training tasks from the robot control task distribution, generate training data by sampling for each selected training task using the current sampling strategy, and update the parameters of the decision-making model using the training data; wherein, the training data includes task features and behavior patterns under the current sampling strategy;

[0055] S3. Use Bayesian criteria and variational inference to obtain an approximate evidence lower bound, and input the training data into the task risk prediction model to calculate the approximate evidence lower bound ELBO loss based on the encoder-decoder architecture, and update the parameters of the task risk prediction model according to the calculated ELBO loss function;

[0056] S4. Input the test data into the updated decision-making model, and optimize the decision-making information based on the task risk assessment result feedback by the updated task risk prediction model to output the final task decision result of the robot control task.

[0057] It can be understood that in terms of theoretical understanding, the present invention first abstracts the general task episodic learning process into a Markov decision process (MDP), constructs an infinite multi-armed bandit (i-MAB), and demonstrates that MPTS is a particular solution of the i-MAB.

[0058] In terms of implementation, for the adaptive decision-making scenario of robot control, the present invention aims to explore a better task subset using a risk prediction model: more meeting the worst-case requirements and wider task coverage. Specifically, the present invention enhances the sampling function through diversity regularization, solves the concentration problem under a larger scale, thereby achieving stronger exploration of the task space. To simplify amortized evaluation and utilize stochastic optimism, the present invention adopts a posterior sampling strategy to search for the optimal task subset, and this design also simplifies the hyperparameter design. Based on these modifications, the present invention proposes a posterior and diversity collaborative task sampling method (PDTS) as an active task sampling (RATS) method for robust and efficient adaptive decision-making.

[0059] In one embodiment of the present invention, theoretical modeling based on infinite multi-armed bandits (i-MABs): The present invention discovers that the task sampling process in task-robust scenario learning can be modeled as a Markov decision process. The state space is the feasible machine learning parameters; the action space contains a series of subsets of training tasks; the transition dynamics are the updated policy parameters transferred based on the previous moment's policy parameters and training task batches; the reward function quantifies the improvement in adaptability robustness after state transition. Through this theoretical tool, the present invention can analyze that the RATS method approximates to solving i-MAB. Specifically,

[0060] State space: The state space is defined as the feasible machine learning parameters. These parameters include the state and configuration of the current model, reflecting the current situation of the system. In an adaptive decision-making system represented by robot control, these parameters help the system evaluate the current task environment and the effectiveness of the policy.

[0061] Action space: The action space contains subsets of training tasks. Each subset represents a specific task or a group of tasks for training and optimizing the decision-making model. In an adaptive decision-making system represented by robot control, these task subsets can be different operation instructions or path planning tasks, aiming to improve the performance of the robot in various scenarios.

[0062] Transition dynamics: The transition dynamics describe the process of transferring to the updated policy parameters based on the previous moment's policy parameters and training task batches. This means that each time, adjustments are made according to the current task batch and policy parameters to generate new policy parameters. In an adaptive decision-making system represented by robot control, this dynamic adjustment mechanism ensures that the system can quickly respond to environmental changes and continuously optimize its behavior strategy.

[0063] Reward function: The reward function quantifies the improvement in adaptability robustness after state transition. In this way, the present invention can evaluate the impact of each state transition on the system performance, thereby guiding the system to develop in a more optimal direction. In an adaptive decision-making system represented by robot control, the reward function can help the system identify which policy adjustments are most helpful in improving the task success rate and the ability to cope with complex environments.

[0064] Through this modeling approach, in an adaptive decision-making system represented by robot control, the task sampling process can be better understood and optimized. For example, in the application scenario of robot control, this method can help the system dynamically adjust its behavior strategy and improve its ability to cope with complex environments. Specifically, when the robot faces different tasks, the system can select the optimal task subset for training according to the current state and task requirements, and then update the policy parameters to ensure that the system can make optimal decisions and maintain high adaptability and robustness in various environments.

[0065] In one embodiment of the present invention, in the robust active task sampling (RATS) of adaptive decision-making, the present invention expects the task subset to be collected to be as difficult as possible and to cover the task space as much as possible, which brings two requirements for the task acquisition function and the task sampling process: i. The worst tasks should be selected, ii. The task candidate set should be increased as much as possible. However, the previous work MPTS only considered the former, and may encounter serious performance breakdowns when increasing This result can be attributed to the concentration of the selected subset within a narrow range.

[0066] The present invention proposes a task acquisition function enhanced by diversity regularization to simultaneously meet the above two requirements. Specifically, instead of simply selecting a single candidate task based on the acquisition function score of each task, the present invention evaluates the task batch based on both i. adaptation risk and ii. task diversity. Formally, the new design of the present invention constructs a diversity maximization problem:

[0067]

[0068] where τ i represents the task identifier of task i, represents the training task batch, represents the candidate task batch, represents the task subset, represents the candidate task set, represents the task subset of the acquisition function value, measures the diversity of the subset calculated by the pairwise distance d of the identifiers, for example γ represents the diversity trade-off coefficient.

[0069] It can be proved that using this diversity-enhanced acquisition function, the worst subset can be explored from among many candidate arms while maintaining a certain degree of coverage of the task space by the subset.

[0070] It is understandable that the upper confidence bound (UCB) acquisition rule used in the previous MPTS method requires multiple random forward passes for each task for evaluation, and the computational cost increases with the increase of . It also requires calibration of balancing the weights of exploration and exploitation in subset search.

[0071] To reduce unnecessary computations and maintain uncertainty optimism, the present invention adopts the posterior sampling strategy as the acquisition principle. Posterior sampling is an extension of Thompson sampling in solving MDPs with infinite actions. The reward or action value is regarded as a random function, and the value of each arm is sampled once in the posterior and used for action selection. Specifically, it is shown by the following formula:

[0072] z t ~q φ (z t |H t )

[0073]

[0074] where H t is the historical information at time t; φ is the encoder model in the risk prediction model, which maps the historical information H t to a latent variable z t , q represents the corresponding distribution; e represents the true task risk, represents the predicted task risk; represents the task identifier of task i in the candidate task set; ψ is the decoder model in the risk prediction model, which maps the latent variable z t and the task identifier to the task risk prediction distribution of task i, and p represents the corresponding distribution; represents the optimal training task batch obtained at time t + 1.

[0075] The specific implementation process of the present invention in combination with the above method is as follows:

[0076] The present invention is based on adaptive decision-making in the robot control scenario, and is mainly applied to scenarios of adaptive decision-making in a zero-shot or few-shot manner, including meta-reinforcement learning, domain randomization, etc. The main modules include a decision-making model and a task risk prediction model.

[0077] During the training process, for each (relaxable) training update of the decision-making model: i. First, sample a candidate task batch using a specific task distribution ii. According to the method of the present invention described above, screen out the training task batch iii. For each task, use policy sampling to obtain training data for training.

[0078] During the training process, for the task risk prediction model, the present invention uses the model architecture and training method of MPTS, that is, adopts an encoder-decoder architecture, and uses Bayesian criteria and variational inference to obtain the approximate evidence lower bound (ELBO) as the training objective:

[0079]

[0080] where is the fixed conditional prior from the last update, β is the penalty weight, D KL represents the Kullback-Leibler divergence.

[0081] During testing, the decision-making module is directly used for decision-making.

[0082] It can be understood that the present invention improves the system performance and adaptability in complex scenarios such as robot control through efficient training task screening. The core lies in combining the task risk prediction model and unique screening criteria to achieve precise selection of training tasks. Based on this, the specific steps are as follows:

[0083] First, determine the distribution of robot control tasks in the adaptive decision-making system, providing a basis for subsequent training task sampling.

[0084] In each training of the decision-making model, candidate training tasks are sampled from the task distribution. The task risk prediction model is used to evaluate the risks of the candidate tasks and predict the risk levels of each task. According to the output of the task risk prediction model and the screening criteria designed by the present invention (such as selecting the worst tasks, adapting to risks, and increasing the task candidate set, etc.), the most suitable batch of training tasks is screened out from the candidate tasks. This process ensures that the training tasks are not only challenging but also cover a wider task space, thereby improving the adaptability and robustness of the system.

[0085] For each screened training task, training data including task features and behavior patterns are generated using the current sampling strategy.

[0086] The parameters of the decision-making model are updated using the generated training data to optimize the model performance.

[0087] The training data is input into the task risk prediction model. Based on the encoder-decoder architecture, the approximate evidence lower bound (ELBO) loss is calculated using Bayesian criteria and variational inference.

[0088] The parameters of the task risk prediction model are updated according to the ELBO loss function to further improve its prediction accuracy.

[0089] In the testing phase, the test data is input into the updated decision-making model. Based on the task risk assessment results feedback by the task risk prediction model, the decision-making information is optimized, and the final task decision-making result is output.

[0090] Among them, the task risk prediction model, which is the core tool for training task screening. By predicting the risk level of tasks, it provides a quantitative basis for the screening process. It can identify the most challenging tasks, thus guiding the system to focus on these tasks during the training process, enhancing the adaptability and robustness of the system.

[0091] Among them, the present invention designs diverse screening criteria, including selecting the worst tasks, adapting to risks, and increasing the task candidate set, etc. These criteria not only consider the difficulty of tasks, but also take into account the diversity and coverage of tasks, ensuring that the system can perform well in a wide range of scenarios.

[0092] By combining the task risk prediction model and the screening criteria, the present invention can efficiently screen out the most suitable training tasks, thus achieving efficient adaptive decision-making in complex application scenarios.

[0093] Specifically: In an adaptive decision-making system represented by robot control, each training update of the decision-making model samples a batch of candidate training tasks from a preset task distribution. The task (robot control task) sampling process in scenario learning is modeled as a Markov decision process (MDP). Specifically, the state space is the feasible machine learning parameters; the action space contains subsets of training tasks; the transition dynamics are the updated policy parameters transferred based on the previous moment's policy parameters and the training task batch; the reward function quantifies the improvement of adaptability robustness after the state transition. These tasks represent different scenarios or task types that the robot may encounter. Then, the present invention screens out the most suitable batch of training tasks, ensuring diversity and representativeness. For each screened training task, samples are taken using the current policy to generate corresponding training data. These training data not only reflect the specific characteristics of the tasks, but also contain the behavior patterns under the current policy. Finally, the generated training data is used to update the parameters of the decision-making model, gradually optimizing the model performance to better adapt to various task requirements, thereby enhancing the robustness of the entire adaptive decision-making system.

[0094] Furthermore, in order to further improve the robustness and adaptability of the adaptive decision-making system represented by robot control, the present invention introduces a task risk prediction model. This model uses Bayesian criteria and variational inference to optimize the evidence lower bound (ELBO). The specific steps are as follows: First, the training data is input into the task risk prediction model, and the ELBO loss is calculated based on the encoder-decoder architecture. The ELBO loss function comprehensively considers the fitting degree of the model to the data and the difference between the posterior distribution and the prior distribution. By optimizing this loss function, the present invention can update the parameters of the task risk prediction model so that it can accurately evaluate the risk levels of different tasks. The results of such risk assessment help the decision-making model to make more robust choices in complex environments, thereby enhancing the performance of the entire adaptive decision-making system. In addition, by modeling the task sampling process as an MDP, the present invention can better understand the task dynamics and their impact on system performance, further improving the adaptability of the system.

[0095] Furthermore, in the test phase, the present invention inputs the actual test data into the updated decision-making model. At this time, the adaptive decision-making system represented by robot control not only relies on its own training results but also refers to the task risk assessment results feedback by the task risk prediction model. Based on this information, the decision-making model can output the final task decision result to ensure optimal choices can be made in various environments. For example, in the application scenario of robot control, this process can help the robot dynamically adjust its action strategy according to the current task risk assessment when facing unknown environments, thereby improving the task success rate and safety. By modeling the task sampling process as an MDP, the system can better understand and respond to task dynamic changes, further enhancing the robustness and adaptability of decision-making. In this way, the entire system can achieve efficient adaptive decision-making in complex and changeable task environments.

[0096] According to the adaptive decision-making method based on posterior and diversity collaborative task sampling of the embodiments of the present invention, the design of active task sampling in adaptive decision-making is guided by the i-MAB theoretical interpretation framework. The designed diversity regularization sampling function solves the concentration problem, allows exploration within a wider range of task sets, and ensures almost worst-case MDP robustness. PDTS adopts posterior sampling-based stochastic optimization, is easy to implement, and can better support adaptive decision-making, especially outstanding in complex application scenarios such as robot control, and can effectively improve the overall performance and adaptability of the system.

[0097] To implement the above embodiments, as Figure 2 shown, the present embodiment also provides an adaptive decision-making device 10 based on posterior and diversity collaborative task sampling, including:

[0098] A task distribution determination module 100, configured to obtain the robot control task distribution in an adaptive decision-making system;

[0099] A decision model training module 200, configured to sample candidate training tasks from the robot control task distribution in each training of the decision model, sample and generate training data for each selected training task using the current sampling strategy, and update the parameters of the decision model using the training data; wherein, the training data includes task features and a behavior pattern under the current sampling strategy;

[0100] A risk prediction model update module 300, configured to obtain an approximate evidence lower bound using Bayesian criteria and variational inference, and input the training data into a task risk prediction model to calculate an approximate evidence lower bound ELBO loss based on an encoder-decoder architecture, and update the parameters of the task risk prediction model according to the calculated ELBO loss function;

[0101] A task decision output module 400, configured to input test data into the updated decision model, and optimize decision information based on the task risk assessment result fed back by the updated task risk prediction model, so as to output the final task decision result of the robot control task.

[0102] Further, it further includes modeling the robot control task process as a Markov decision process; wherein, the state space is feasible machine learning parameters; the action space includes a subset of training tasks; the transition dynamics are updated policy parameters transferred based on the previous moment's policy parameters and a batch of training tasks; the reward function quantifies the improvement of adaptability robustness after state transition.

[0103] Further, it further includes a task acquisition function enhanced based on diversity regularization, and evaluates a batch of tasks according to selecting the worst task, adaptation risk, and increasing the task candidate set Evaluating a batch of tasks includes:

[0104] Constructing a diversity maximization problem:

[0105]

[0106] wherein, τ i represents the task identifier of task i, represents a batch of training tasks, represents a batch of candidate tasks, represents a subset of tasks, represents a set of candidate tasks, represents a subset of tasks the acquisition function value of, measures the diversity of the subset and is calculated by the pairwise distance d of the identifiers: γ represents the diversity trade-off coefficient.

[0107] Furthermore, the posterior sampling strategy is adopted as the acquisition principle:

[0108] z t ~q φ (z t |H t )

[0109]

[0110] where H t is the historical information at time t; φ is the encoder model in the risk prediction model, which maps the historical information H t to a latent variable z t , q represents the corresponding distribution; l represents the true task risk, represents the predicted task risk; represents the task identifier of task i in the candidate task set; ψ is the decoder model in the risk prediction model, which maps the latent variable z t and the task identifier to the task risk prediction distribution of task i, and p represents the corresponding distribution; represents the optimal training task batch obtained at time t+1.

[0111] Furthermore, the Bayesian criterion and variational inference are used to obtain the approximate evidence lower bound ELBO as the training objective:

[0112]

[0113] where is the fixed conditional prior from the last update, β is the penalty weight, D KL represents the Kullback-Leibler divergence.

[0114] According to the adaptive decision-making device based on posterior and diversity collaborative task sampling of the embodiments of the present invention, the design of active task sampling in adaptive decision-making is guided by the i-MAB theoretical interpretation framework. The designed diversity regularization sampling function solves the concentration problem, allows exploration within a wider range of task sets, and ensures almost worst-case MDP robustness. PDTS adopts posterior sampling-based stochastic optimization, is easy to implement, and can better support adaptive decision-making, and performs particularly well in complex application scenarios such as robot control, and can effectively improve the overall performance and adaptability of the system.

[0115] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0116] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

Claims

1. An adaptive decision-making method based on posterior and diversity collaborative task sampling, characterized in that, Including: Obtain the robot control task distribution in the adaptive decision-making system; In each training of the decision-making model, sample candidate training tasks from the robot control task distribution. For each selected training task, sample and generate training data using the current sampling strategy, and update the parameters of the decision-making model using the training data; wherein, the training data includes task features and the behavior pattern under the current sampling strategy; Use the Bayesian criterion and variational inference to obtain an approximate evidence lower bound, and input the training data into the task risk prediction model to calculate the approximate evidence lower bound ELBO loss based on the encoder-decoder architecture, and update the parameters of the task risk prediction model according to the calculated ELBO loss function; Input the test data into the updated decision-making model, and optimize the decision-making information based on the task risk assessment result feedback by the updated task risk prediction model, so as to output the final task decision result of the robot control task.

2. The method according to claim 1, wherein The method further includes modeling the robot control task process as a Markov decision process; wherein, the state space is the feasible machine learning parameters; the action space contains a subset of training tasks; the transition dynamics are the updated policy parameters transferred based on the previous moment's policy parameters and the training task batch; the reward function quantifies the improvement of the adaptive robustness after the state transition.

3. The method according to claim 1, characterized in that, The method further includes a task acquisition function enhanced based on diversity regularization, and selects the worst task, adaptation risk, and enlarges the task candidate set Evaluating the task batch includes: Construct a diversity maximization problem: Among them, τ i represents the task identifier of task i, represents the training task batch, represents the candidate task batch, represents the task subset, represents the candidate task set, represents the task subset 's acquisition function value, measures the diversity of the subset by calculating the pairwise distance d of the identifiers: γ represents the diversity trade-off coefficient.

4. The method according to claim 1, wherein Adopt the posterior sampling strategy as the acquisition principle: z t ~q φ (z t |H t ) Among them, H t is the historical information at time t; φ is the encoder model in the risk prediction model, which maps the historical information H t to a latent variable z t , q represents the corresponding distribution; e represents the true task risk, represents the predicted task risk; represents the task identifier of task i in the candidate task set; ψ is the decoder model in the risk prediction model, which maps the latent variable z t and the task identifier to the task risk prediction distribution of task i, and p represents the corresponding distribution; represents the optimal training task batch obtained at time t+1.

5. The method according to claim 1, characterized in that Use the Bayesian criterion and variational inference to obtain the approximate evidence lower bound ELBO as the training objective: where is the fixed-condition prior from the last update, β is the penalty weight, D KL represents the Kullback-Leibler divergence.

6. An adaptive decision-making device based on posterior and diversity collaborative task sampling, characterized in that, Including: A task distribution determination module, configured to obtain the robot control task distribution in the adaptive decision-making system; A decision-making model training module, configured to sample candidate training tasks from the robot control task distribution in each training of the decision-making model, sample and generate training data using the current sampling strategy for each selected training task, and update the parameters of the decision-making model using the training data; wherein, the training data includes task features and the behavior pattern under the current sampling strategy; A risk prediction model update module, configured to use the Bayesian criterion and variational inference to obtain an approximate evidence lower bound, and input the training data into the task risk prediction model to calculate the approximate evidence lower bound ELBO loss based on the encoder-decoder architecture, and update the parameters of the task risk prediction model according to the calculated ELBO loss function; A task decision output module, configured to input the test data into the updated decision-making model, and optimize the decision-making information based on the task risk assessment result feedback by the updated task risk prediction model, so as to output the final task decision result of the robot control task.

7. The device according to claim 5, characterized in that, It further includes modeling the robot control task process as a Markov decision process; wherein, the state space is the feasible machine learning parameters; the action space contains a subset of training tasks; the transition dynamics are the updated policy parameters transferred based on the previous moment's policy parameters and the training task batch; the reward function quantifies the improvement of the adaptive robustness after the state transition.

8. The device according to claim 5, characterized in that It also includes a task acquisition function enhanced based on diversity regularization, and selects the worst task, adapts to risks, and enlarges the task candidate set Evaluating the task batch, including: Construct a diversity maximization problem: Among them, τ i represents the task identifier of task i, represents the training task batch, represents the candidate task batch, represents the task subset, represents the candidate task set, represents the task subset of the acquisition function value, measures the diversity of the subset by calculating the pairwise distance d of the identifiers: γ represents the diversity trade-off coefficient.

9. The device according to claim 5, characterized in that, Adopt the posterior sampling strategy as the acquisition principle: z t ~q φ (z t |H t ) Among them, H t is the historical information at time t; φ is the encoder model in the risk prediction model, which maps the historical information H t to a latent variable z t , q represents the corresponding distribution; l represents the real task risk, represents the predicted task risk; represents the task identifier of task i in the candidate task set; ψ is the decoder model in the risk prediction model, which maps the latent variable z t and the task identifier to the task risk prediction distribution of task i, p represents the corresponding distribution; represents the optimal training task batch obtained at time t + 1.

10. The device according to claim 5, characterized in that, Use the Bayesian criterion and variational inference to obtain the approximate evidence lower bound ELBO as the training objective: where is the fixed-condition prior from the last update, β is the penalty weight, D KL represents the Kullback-Leibler divergence.