Apparatus for machine learning, computer program, and computer-implemented method

JP2023016762A5Pending Publication Date: 2025-06-06ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022116210
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-22
Filing Date
2022-07-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing machine learning systems struggle to effectively transfer prior knowledge from observed tasks to new tasks, particularly in lifelong learning scenarios, leading to inefficiencies and suboptimal performance.

Method used

A computer-implemented method that utilizes a PAC Bayesian generalization bound as an objective function to update hyperparameter posterior distributions, allowing for the transfer of prior knowledge from observed multi-armed or contextual bandit tasks to new tasks, by maximizing a lower bound on expected rewards and incorporating observable training data.

Benefits of technology

This approach enhances machine learning efficiency by ensuring that prior knowledge is effectively utilized, improving task performance and robustness, and maintaining low expected error rates in new tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an apparatus for machine learning, a computer program, and a computer-implemented method.SOLUTION: A method includes: providing a task including an action space of a multi-armed bandit problem or a contextual bandit problem and a distribution over rewards adjusted for actions (200); providing a hyperparameter prior distribution which is a distribution over the action space (202); and determining a hyperparameter posterior distribution which is a distribution over the action space, related to the hyperparameter prior distribution (224). For the hyperparameter posterior distribution, in using prior distributions sampled from the hyperparameter posterior distribution, a lower limit for the rewards expected for future bandit tasks has a value as large as possible.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Background Art The present invention relates to an apparatus, a computer program, and a computer-implemented method for machine learning.

Background Art

[0002] "A PAC-Bayesian bound for Lifelong Learning" (by Anastasia Pentina, Christoph H. Lampert, arXiv:1311.2838) discloses aspects of a lifelong learning setting for machine learning.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Disclosure of the Invention The computer-implemented method, apparatus, and computer program according to the independent claims provide improved machine learning, particularly based on an objective function for training a lifelong learning system.

Means for Solving the Problems

[0005] A computer-implemented method for machine learning includes providing a task that includes an action space for a multi-armed bandit problem or a contextualized bandit problem and a distribution over rewards adjusted for those actions; providing a hyperparameterized prior distribution, which is the distribution over the action space; and determining a hyperparameterized posterior distribution, which is the distribution over the action space in relation to the hyperparameterized prior distribution, wherein the lower bound on the expected rewards for future bandit tasks, when using a prior distribution sampled from the hyperparameterized posterior distribution, is as large as possible. This means that these tasks are either a multi-armed bandit problem or a contextualized bandit problem. The lower bound is the PAC Bayesian generalization limit, which is the lower bound on the expected rewards for future bandit tasks, when using a prior distribution sampled from the hyperparameterized posterior distribution. The lower bound is used as an objective function for lifelong learning, where the hyperparameterized posterior distribution is updated to approximately maximize the lower bound after observation of the task. The lower bound includes observable quantities and can therefore be calculated using training data from observed tasks. In this way, prior knowledge from the observed set of multi-armed bandit tasks or contextualized bandit tasks is applied to new multi-armed bandit tasks or contextualized bandit tasks.

[0006] Determining the hyperparameter posterior distribution may involve determining the hyperparameter posterior distribution that maximizes the lower bound on the expected reward. The hyperparameter posterior distribution that maximizes the lower bound is expected to assign a high probability to a prior distribution that will consequently result in less error in future tasks.

[0007] Using this method, prior knowledge from an observed set of multi-armed bandit tasks or contextualized bandit tasks can be applied to new multi-armed bandit tasks or contextualized bandit tasks. Many problems can be expressed as either multi-armed bandit problems or contextualized bandit problems. Therefore, prior knowledge can be applied to many types of tasks using this method.

[0008] This delicious, In particular, processing sensor data, especially digital image data or audio data, in relation to a prior distribution sampled from a hyperparameter posterior distribution, for the purpose of classifying sensor data. Detecting the presence of an object in sensor data or performing semantic segmentation on sensor data, or When sampling a prior distribution from a hyperparameter posterior distribution, the robustness of machine learning can be measured, particularly the probability that the expected error for the next task will not exceed a predetermined value, or To detect anomalies in sensor data in relation to a prior distribution sampled from a hyperparameter posterior distribution, or, Learning policies for controlling a physical system and determining control signals for controlling the physical system in relation to prior distributions sampled from hyperparameter posterior distributions. This may include the following. Therefore, this method can be used to apply prior knowledge to new problems, particularly new classification problems, new robustness problems, new anomaly detection problems, and / or new control problems.

[0009] The lower bound is also an indicator of the robustness of machine learning. The hyperparameter posterior distribution also provides information about the expected error for the next task when the prior distribution is sampled from the hyperparameter posterior distribution. The expected error will not exceed a certain value, for example, when the prior distribution is sampled from a particular hyperparameter posterior distribution.

[0010] This method may involve determining the posterior distribution of hyperparameters in multiple iterations and sampling the prior distribution of an iteration from the posterior distribution of hyperparameters in a preceding iteration. Therefore, in machine learning iterations, prior knowledge is repurposed from preceding iterations.

[0011] This method may involve sampling iterative tasks from a distribution across tasks. In particular, machine learning results are improved by processing a large number of diverse tasks.

[0012] This method preferably involves initializing the hyperparameter posterior distribution with the hyperparameter prior distribution. Any available hyperparameter prior distribution is a good starting point for machine learning.

[0013] This method may include initializing the posterior distribution with a prior distribution; determining an action policy from a set of task action policies, including a distribution over actions with probability masses related to the task posterior distribution; randomly sampling or selecting actions from the probability masses; sampling rewards in relation to actions from a distribution over rewards; determining a task dataset containing actions and rewards; and updating the task posterior distribution to include the task dataset. This results in highly efficient machine learning for multi-armed bandit tasks or contextualized bandit tasks, where the task dataset is stored as knowledge of preceding iterations by the generated posterior distribution.

[0014] Providing a task preferably involves providing a task that includes a state space and a distribution over initial states, and this method further involves randomly sampling or selecting initial states from the distribution over initial states, and the distribution over rewards is adjusted to the states of the action and state space. This results in highly efficient machine learning for contextual bandit tasks.

[0015] This method may involve initializing the task dataset with an empty set, and then updating the task posterior distribution over a predetermined number of rounds. Thus, a lower bound is calculated from the training data observed from the tasks.

[0016] Determining the posterior distribution of hyperparameters may involve determining a lower bound in relation to the Kullback-Leible divergence of the posterior and prior distributions of hyperparameters. The desired improved machine learning is achieved when the prior distribution is optimized so that it yields a similar effect on both the observed task and the new task. The Kullback-Leible divergence of the posterior and prior distributions of hyperparameters provides an excellent method for determining a lower bound from training data from the observed task.

[0017] The present invention provides an apparatus for machine learning, characterized by being configured to carry out the steps in the method described in any one of claims 1 to 10.

[0018] The present invention provides a computer program that, when executed on a computer, includes computer-readable instructions for causing a computer to perform the method described in any one of claims 1 to 10.

[0019] Further advantageous embodiments can be derived from the following description and drawings. [Brief explanation of the drawing]

[0020] [Figure 1] This is a schematic diagram showing some of the equipment used for machine learning. [Figure 2] This diagram shows the steps of a machine learning method. [Modes for carrying out the invention]

[0021] FIG. 1 schematically shows a part of an apparatus 100 for machine learning. The apparatus 100 includes at least one processor 102 and at least one storage device 104. The at least one storage device 104 may store a computer program including computer-readable instructions for causing a computer to implement the method described below with reference to FIG. 2 when executed on the computer. The apparatus 100 is particularly configured to perform the steps of the method when at least one processor 102 executes the instructions of the computer program.

[0022] The machine learning method in this example uses a basic learning algorithm Q = Q(D, P) and starts from step 200. The basic learning algorithm returns a posterior distribution Q. Since the basic learning algorithm uses a dataset D and a prior distribution P, the posterior distribution is described herein as Q(D, P) to clarify the relationship to D and P.

[0023] In step 200, the method includes providing a task T i including.

[0024] The task T i is a multi-armed bandit problem including an action space A, i.e., a set of actions a, and a distribution ρ i (r│a) over the reward r adjusted for the action a T i =(A, ρ i ) and may be a multi-armed bandit problem that is a task represented as.

[0025] The task T i is a quadruple including an action space A, i.e., a set of actions a, a state space S, a distribution ρ i (r|s, a) over the reward r adjusted for the action a and the state s, and a distribution over the initial state μ(s) of the context bandit problem T i =(S, A, μ i , ρ i) This can be expressed as a contextualized bandit problem, which is a task.

[0026] The reward r is assumed to be between 0 and 1, and task T i It is assumed that these are iid (independent and identically distributed) samples taken from environment T.

[0027] Then, step 202 is executed.

[0028] In step 202, this method includes providing a hyperparameter prior distribution P.

[0029] The hyperparameter prior distribution P is a distribution over a set of possible prior distributions. Each prior distribution is a distribution over the action space A.

[0030] Then, step 204 is executed.

[0031] In step 204, this method involves initializing the hyperparameter posterior distribution Q with the hyperparameter prior distribution P. The hyperparameter posterior distribution Q is another distribution over the prior distribution.

[0032] Then, step 206 is executed.

[0033] In step 206, this method obtains the distribution T over task T from the iteration i of task T i This includes sampling.

[0034] Task T i In, m i Task T including individual action policies i Set of behavioral policies

number

[0035] Action policy b ijIn the case of the multi-armed bandit problem, the probability mass b ij (a) includes a distribution across behavior a. In this example, behavior policy b ij This is the previously observed training data

number

number

[0036] The elements of the training dataset are, in the multi-armed bandit task,

number

[0037] Action policy b ij In the case of a contextualized bandit problem, the probability mass b ij (a|s ij Includes a distribution over behavior a adjusted for state s, accompanied by behavior policy b. ij This is the previously observed training data

number

[0038] In contextualized bandit tasks, the elements of the training dataset are as follows:

number

[0039] Then, step 208 is executed.

[0040] In step 208, this method involves sampling the prior distribution P of iteration i from the hyperparameter posterior distribution Q of the previous iteration i-1.

[0041] Then, step 210 is executed.

[0042] In step 210, this method includes initializing the posterior distribution Q with the prior distribution P.

[0043] Then, step 212 is executed.

[0044] In step 212, this method uses an empty set to create dataset D i This includes initializing it.

[0045] Then, step 214 is executed.

[0046] In step 214, this method is used for task T i Action policy B ij From the set, the action policy b associated with the task posterior distribution Q ij This includes making a decision.

[0047] Then, step 216 is executed.

[0048] In step 216, this method is action policy b ij Action a from the probability mass ij This includes randomly sampling or selecting.

[0049] Then, step 218 is executed.

[0050] In step 218, this method, in the case of the multi-armed bandit problem, involves the distribution ρ over the reward r. i (r│a) Action a ij In relation to this, reward r ij This includes sampling.

[0051] In step 218, this method, in the case of a contextualized bandit problem, gives the distribution over the initial state μ(s) to the initial state s ij Randomly sampling or selecting and a distribution ρ over the reward r adjusted for the action a and the state s in the state space S i Action a from (r│s,a) ij In relation to the reward r ij This includes sampling.

[0052] Then, step 220 is executed.

[0053] In step 220, this method is action a ij and reward r ij Dataset D containing i This includes making a decision.

[0054] Dataset D i In the case of the multi-armed bandit problem, m i This is a training set containing individual observable training data pairs:

number

[0055] Dataset D i In the case of the contextualized bandit problem, m i This is a training set containing a triple (three-part) of observable training data: D i =((s i1 ,a i1 ,r i1 ),…,(s,a ij-1 ,r ij-1 ))

[0056] Then, step 222 is executed.

[0057] In step 222, this method is used for dataset D i This includes updating the posterior distribution Q to include it.

[0058] For this purpose, a deterministic fundamental learning algorithm is used in dataset D i It takes the prior distribution P as input and the posterior distribution Q = Q(D i It causes P).

[0059] Task T i The goal is to find the posterior distribution Q that maximizes the expected reward. In the case of the multi-armed bandit problem, the expected rewards are as follows:

number

[0060] In the example of the multi-armed bandit problem, dataset D i Using this, the expected reward is estimated as follows:

number

number

[0061] In the example of the contextualized bandit problem, dataset D i Using this, the expected reward is estimated as follows:

number

[0062] In this example, the method involves updating the posterior distribution Q over a predetermined number of rounds m. This means that the method continues through step 214 for m rounds.

[0063] Then, step 224 is executed.

[0064] In step 224, the method includes determining a hyperparameter posterior distribution Q in relation to a hyperparameter prior distribution P, such that the lower bound of the expected reward L for a future bandit task is as large as possible when prior information P sampled from this hyperparameter posterior distribution Q is used.

[0065] In this example, determining the hyperparameter posterior distribution Q involves determining the hyperparameter posterior distribution Q that maximizes the lower bound on the expected reward L.

number

[0066] In this example, the lower bound (L(Q')), which is a function for the multi-armed bandit problem, is one of the following objective functions:

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0067] In the case of a contextualized bandit problem, these objective functions are L i to L i (cb) Replaced by,

number

number

number

[0068] The performance of a hyperparameter posterior distribution Q is, for example, that of a basic learning algorithm Q(D) with prior distributions P~Q. n+1 When using ;P), a new task T n+1 It is measured by marginal transfer reward, which is the expected reward for the action.

[0069] In the case of the multi-armed bandit problem, the peripheral transfer reward is as follows:

number

[0070] In the contextualized bandit problem, the peripheral transfer reward is as follows:

number

[0071] This means that determining the hyperparameter posterior distribution Q involves determining and approximating the expected reward L in relation to the Kullback-Leibra divergence DKL(Q||P) of the hyperparameter posterior distribution Q and the hyperparameter prior distribution P.

[0072] In this example, the hyperparameter posterior distribution Q is determined in n iterations i.

[0073] This means that this method continues to step 206 for n iterations. Then, optionally, step 226 is performed.

[0074] In this example, in step 226, the method includes processing sensor data, in particular digital image data or audio data, in relation to a prior distribution sampled from a hyperparameter posterior distribution Q. The sensor may be a microphone, video, radar, LiDAR, ultrasonic sensor, motion sensor, thermal imaging sensor, or sonar sensor.

[0075] Sensor data is processed, for example, to classify the sensor data, to detect the presence of objects within the sensor data, or to perform semantic segmentation on the sensor data. Sensor data classification is framed, for example, as a contextualized bandit problem. Images may be classified, for example, based on the presence or type of traffic signs, road surface, pedestrians, or vehicles.

[0076] Alternatively or additionally, step 226 may include learning a policy for controlling the physical system and determining control signals for controlling the physical system in relation to a prior distribution sampled from a hyperparameter posterior distribution Q.

[0077] For example, when sampling a prior distribution from a hyperparameter posterior distribution Q, the probability that the expected error for the next task will not exceed a predetermined value may be determined.

[0078] Alternatively or additionally, step 226 may include determining a metric for the robustness of the machine learning model. In this example, the robustness metric is the expected error to the prior distribution sampled from the hyperparameter posterior distribution Q.

[0079] A prior distribution may be used to process the sensor data. This is, for example, when the expected error is less than a threshold; otherwise, a different prior distribution may be sampled to process the sensor data or to determine a control signal.

[0080] Alternatively or additionally, step 226 may include detecting anomalies in the sensor data in relation to a prior distribution sampled from a hyperparameter posterior distribution Q. In this embodiment, the anomaly detection problem is framed as a multi-armed bandit problem.

[0081] In this example, the process then ends.

[0082] The following is an example of the steps performed by a computer in a machine learning method for the multi-armed bandit problem:

number

Claims

1. 1. A computer-implemented method for machine learning, comprising: The method comprises: Providing a task (200) that includes an action space of a multi-armed bandit problem or a contextual bandit problem and a distribution over the rewards adjusted for the actions; providing a hyperparameter prior distribution (202), the hyperparameter prior distribution being a distribution over the action space; determining 224 a hyperparameter posterior distribution, which is a distribution over the action space relative to the hyperparameter prior distribution; Including, For the hyperparameter posterior distribution, a lower bound on expected rewards for future bandit tasks when using a prior distribution sampled from the hyperparameter posterior distribution has the largest possible value. A method comprising:

2. Determining 224 the hyperparameter posterior distribution includes determining the hyperparameter posterior distribution that maximizes the lower bound on the expected reward. The method of claim 1.

3. Processing (226) the sensor data, in particular digital image data or audio data, in relation to a prior distribution sampled from the hyper-parameter posterior distribution, in particular to classify the sensor data; Detecting the presence of objects in the sensor data or performing semantic segmentation on the sensor data; or Determining (226) a measure of the robustness of the machine learning when sampling a prior distribution from the hyperparameter posterior distribution, in particular the probability that the expected error for a next task will not exceed a predetermined value; or Detecting anomalies in the sensor data relative to a prior distribution sampled from the hyper-parameter posterior distribution (226); or learning a policy for controlling a physical system and determining a control signal for controlling the physical system in relation to a prior distribution sampled from the hyper-parameter posterior distribution. Including, The method of claim 1.

4. Determining 224 the hyperparameter posterior distribution over multiple iterations; Sampling the prior distribution for an iteration from the hyper-parameter posterior distribution of a previous iteration (208); Including, The method of claim 1.

5. Sampling the tasks of the iteration from a distribution over tasks (206). The method according to claim 4.

6. initializing the hyper-parameter posterior distribution with the hyper-parameter prior distribution (204); The method of claim 1.

7. initializing (210) a posterior distribution with the prior distribution; determining (214) from a set of behavior policies for the task, a behavior policy that includes a distribution over actions with a probability mass associated with the task posterior distribution; randomly sampling (216) or selecting actions from the probability mass; Sampling 218 a reward from the distribution over rewards associated with the behavior; determining 220 a dataset comprising the actions and the rewards; updating (222) the posterior distribution to include the task data set; Including, The method of claim 1.

8. Providing 200 the task includes providing the task including a state space and a distribution over initial states; The method further includes randomly sampling (218) or selecting an initial state from the distribution over initial states; the distribution over rewards is adjusted for the actions and states of the state space; The method according to claim 7.

9. Initializing the data set with an empty set (212); then updating the posterior distribution in a predetermined number of rounds; Including, The method of claim 1.

10. Determining the hyperparameter posterior distribution (224) includes determining and approximating the expected reward associated with the Kullback-Leibler divergence of the hyperparameter posterior distribution and the hyperparameter prior distribution. The method of claim 1.

11. An apparatus (100) for machine learning, characterized in that it is configured to carry out the steps of the method according to any one of the claims 1 to 10.

12. A computer program comprising computer readable instructions which, when executed on a computer, causes the computer to carry out the method according to any one of claims 1 to 10.