Method and apparatus for learning a strategy, and method and apparatus for operating a strategy

By employing reinforcement learning with Guided Policy Search to adapt parameters of evolutionary algorithms like CMA-ES, the method addresses the challenge of optimizing parameters across diverse problems, enhancing performance and reducing the need for derivative information.

JP7695836B2Active Publication Date: 2025-06-19ROBERT BOSCH GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021120505
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-23
Filing Date
2021-07-21
Publication Date
2025-06-19
Estimated Expiration
2041-07-21

AI Technical Summary

Technical Problem

Existing evolutionary algorithms, such as CMA-ES, face challenges in adaptively optimizing parameters for diverse optimization problems without requiring significant adaptation or derivative information.

Method used

A computer-implemented method using reinforcement learning, specifically Guided Policy Search, to learn a strategy for optimally adapting parameters of evolutionary algorithms like CMA-ES, based on state information and reward signals, allowing for adaptive parameter tuning without explicit derivative information.

Benefits of technology

The method enables efficient adaptation of parameters for evolutionary algorithms, improving their performance across a wide range of optimization problems with minimal interaction and without requiring derivative information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695836000008
    Figure 0007695836000008
  • Figure 0007695836000009
    Figure 0007695836000009
  • Figure 0007695836000010
    Figure 0007695836000010
Patent Text Reader

Abstract

To provide a method (20) for learning a strategy (π) for optimally making at least one of evolution algorithms (σ) adaptive.SOLUTION: The method includes the steps of: initiating a strategy of calculating a parameter display (A) of a parameter (σ) on the basis of state information (S); and learning a strategy (π) by using a reinforcement learning. Which parameter display is optimum for possible state information is learned from a parameter display determined by a CMA-ES algorithm and the strategy depending on the state information (S), and from the interaction between a question instant (14) and a reward signal (R).SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method, as well as a computer program and a machine-readable storage medium, for learning and operating a strategy configured to parameterize an evolutionary algorithm. [Background technology]

[0002] Prior Art Evolutionary strategies (ES) are stochastic and derivative-free methods for numerically solving optimization problems. They belong to the class of evolutionary algorithms and evolutionary computation. Evolutionary algorithms are based on the laws of biological evolution as a whole, namely the repeated interplay of change (by recombination and mutation) and selection. At each generation (so-called iteration), new individuals (so-called candidate solutions, x) are generated by a often stochastic change of the current parent individuals. Some individuals are then selected based on their fitness or their objective function value F(x) to become parents of the next generation. In this way, individuals with progressively better fitness are generated over the course of the generations.

[0003] CMA-ES( C ovariance M Atrix A Adaptation E evolution S The covariance matrix adaptation evolution strategy (Covariance Matrix Adaptation Evolution Strategy) is an evolutionary algorithm for optimizing continuous "black box functions". The algorithm uses a special type of strategy for numerical optimization, where new candidate solutions x follow a multivariate normal distribution and are scaled by R n where evolutionary recombination is achieved by determining an adapted median for the distribution. The pairwise dependencies between variables in the distribution are represented by a covariance matrix. Covariance matrix adaptation (CMA) is a method for updating the covariance matrix of the distribution.

[0004] For the learning of the random sampling distribution, in CMA-ES, only the hierarchy among candidate solutions is utilized, and neither the derivative nor the function value itself is required for the method. Summary of the Invention Problems to be Solved by the Invention

[0005] Advantages of the Invention With respect to such prior art, the present invention has the advantage that CMA-ES can be superior because the multivariate normal distribution is automatically adapted according to the conditions. Furthermore, from experiments, it has been found that the generation characteristics of the approach according to the present invention are high, and the method can be directly applied to a wide range of other unexpected application cases that vary widely, and no adaptation of the strategy for new application cases is required.

[0006] Furthermore, it has been found that the process of the present invention is also applicable to other evolutionary algorithms, such as differential evolution. Means for Solving the Problems

[0007] Disclosure of the Invention In a first aspect, the present invention is a computer-implemented method for learning a strategy for optimally adapting at least one parameter of an evolutionary algorithm, particularly the CMA-ES algorithm or the differential evolution algorithm, the method including initializing a strategy for calculating a parameter representation of the parameter depending on state information regarding a problem instance. Then, the learning of the strategy is performed using reinforcement learning. In this case, from the interaction between the CMA-ES algorithm, the problem instance, the parameter representation of the parameter determined using the strategy depending on the state information, and the reward signal R, it is learned which parameter representation is optimal for any possible state information.

[0008] Preferably, the parameters are the step size of the CMA-ES algorithm, the number of candidate solutions (population-size) (λ), the mutation rate (mutation-rate), and / or the crossover rate (crossover-rate).

[0009] Here, it is proposed that the state information includes at least one value including the current parameter display of the parameter, the cumulative distance, and / or the difference between the function values of the function to be optimized for the problem instance in the current iteration and the previous iteration.

[0010] Preferably, the state information includes the parameter displays over a plurality of previous iterations. Particularly preferably, it is the last 40 iterations.

[0011] Furthermore, it is proposed that the state information also includes context information regarding the problem instance, particularly the function to be optimized. The advantage here is that the strategies can be made different among various problem instances, whereby good results can be achieved for individual problem instances.

[0012] The parameter display is the value assigned to the parameter.

[0013] Furthermore, it is proposed that reinforcement learning using "Guided Policy Search" (GPS) is used, the parameter display of the parameter is determined using configurable heuristics, a plurality of characteristics of the parameter display of the parameter are characterized therefrom, and a teacher for learning the strategy is provided using GPS from the characteristics.

[0014] In a second aspect of the present invention, a method of operating a strategy learned according to the first aspect of the present invention is proposed. For this purpose, state information about the problem instance here is calculated. The learned strategy calculates a parameter representation depending on the state information, and CMA-ES is applied over at least one iteration together with the parameter representation. Preferably, two steps are alternately repeated until an optimum value is obtained.

[0015] Here, it is proposed that the problem instance is a training process for machine learning, and the hyperparameters of the training process are optimized using CMA-ES with a strategy, or the problem instance is vehicle route optimization or task planning in manufacturing / production, and the route or task planning is optimized using CMA-ES with a strategy. The task planning may be, for example, one operation of a manufacturing step.

[0016] Furthermore, it is proposed that a deep neural network is learned using the optimized hyperparameters. Preferably, the neural network calculates an output quantity depending on the sensor quantity detected by a sensor, and is then learned so that this can be used for calculating a control quantity using a control unit.

[0017] The control quantity can be used to control an actuator of a technical system. The technical system may be, for example, at least a partially autonomous machine, at least a partially autonomous vehicle, a robot, a tool, a machine tool, or an aircraft, for example a drone. The input quantity can be calculated, for example, depending on the detected sensor data and can be supplied. The sensor data can be detected by a sensor of the technical system, for example a camera, or can alternatively be received from the outside.

[0018] In other aspects, the present invention relates to a computer program configured to implement the above-described method, and a machine-readable storage medium storing the computer program.

[0019] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0021] Description of Embodiments In the CMA-ES algorithm, two main methods are used to adapt the parameters of the search distribution.

[0022] The first method is a maximum likelihood method based on the idea of increasing the probability of valid candidate solutions and search steps. The median of the distribution is updated so that the probability of the previous valid candidate solution is maximized. The covariance matrix of the distribution is updated (incrementally) so that the probability of the previous valid search step increases. The two updates can be interpreted as natural gradient ascent.

[0023] The second method is characterized by two paths, called search paths or development paths, of the temporal evolution of the mean of the strategy distribution. These paths contain important information about the correlation between successive steps. In particular, when successive steps follow a similar direction, the evolution path becomes longer. The evolution path is used in two forms. One path is used instead of individual valid search steps for the adaptation process of the covariance matrix, enabling as fast a variance increase (Varianzerhoehung) as possible in a favorable direction. The other path is used to perform additional step size control. The step size control aims to orthogonalize successive movements of the mean of the distribution in the prediction. The step size control effectively prevents the prior covariance and enables a fast covariance towards the optimum.

[0024] Mathematically, CMA-ES can be represented as an optimization algorithm for a continuous function f: R n →R by randomly drawing candidate solutions from a multivariate normal distribution N(m,σ 2 c), where m corresponds to the predicted value, σ 2 corresponds to the step width, and c corresponds to the covariance matrix.

[0025] The covariance matrix can be initialized as the identity matrix and m and σ 2 can be initialized by random values, for example, or set by the user. Then, CMA-ES recursively updates the normal distribution N so that the probability of valid candidate solutions increases.

[0026] For each iteration, also referred to as generation G, first, x of λ candidate solutions is drawn, and only the best μ candidate solutions are used as parents for the next generation g + 1. The number λ of candidate solutions x can be initially set to a set value. Then, the predicted value m is adapted as follows: (g+1) where c is a learning rate and can be set equal to 1, for example. (Equation 1):

Number

[0027] Then, as is also known as cumulative step length adaptation (CSA), the step width c is adapted as follows: (Equation 2):

Number

Number

Number

Number

[0028] Note that there are a number of other various means of adapting the step width σ, and it is generally significant that a reciprocal switching occurs among these various means during the execution of CMA - ES. Regarding the adaptation of the covariance matrix c, it is shown in the literature on CMA - ES.

[0029] To enhance the performance of CME - ES, the adaptation of the step width σ (g+1) is based on the state information s by (Equation 2) gis executed depending on this, and for this purpose a policy π(s g ) is used. That is, a policy π that minimizes an (arbitrary) cost function is searched for, where the cost function L characterizes how well CMA-ES is truncated using the policy.

[0030] Mathematically, this can be expressed as (Equation 4): [Number] As shown. Here, it is proposed to determine the policy using reinforcement learning. For this purpose, an action space that is preferably continuous and has possible values of step width σ (g) is defined.

[0031] The state information s g includes each of the following states, namely · The current step width σ (g) · The current path [Number] , and / or · The difference in the values of the continuous function f to be optimized from at least generations (g) and (g - 1) is included.

[0032] Preferably, the state information includes the history of a plurality of previous step widths. Particularly preferably, for this purpose, the last 40 step widths are used. Additionally, the history of the difference in the current function value of the function f to be optimized from generation (g) can also be added to a plurality of values from previous generations. If there are not enough past values for the history, here, for example, zeros can be filled.

[0033] It is proposed to use the cost function L of the negative function value f(x) as the best truncated candidate inside the current iteration. This has the advantage that the so-called always performance is optimized.

[0034] In the following embodiments, the question is raised as to how significantly the sampling efficiency can be improved. The basic idea here is not to learn the strategy from scratch, but rather to have a teacher present. Through simulation, it has been found that about 10,000 CMA-ES runs are required for learning the strategy by reinforcement learning, but when the teacher is involved, this value can be reduced to as few as 1,100 CMA-ES runs. Therefore, only a slight interaction between the CMA-ES algorithm and its environment is required.

[0035] In the first embodiment, the teacher is used in the form of self-adaptive heuristics, where (Equation 2) is used as the heuristic. Note that other equations for determining the step size, such as "two-point step-size adaptation", can also be used as heuristics instead of (Equation 2).

[0036] In a preferred embodiment, the teacher exists in the form of a guiding trajectory, also known by the name "Guided policy search" (GPS). In this case, the strategy is adapted using supervised learning such that its output trajectory is very similar to the guiding trajectory.

[0037] For further details regarding GPS, see Levine, S., Abbeel, P.: Learning neural network policies with guided policy search under unknown dynamics, In: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K. (eds.) Proceedings of the 28th International Conference on Advances in Neural Information Processing Systems (NeurIPS’14), pp.1071-1079 (2014), which is available online at https: / / papers.nips.cc / paper / 5444-learning-neural-network-policies-with-guided-policy-search-under-unknown-dynamics.pdf, or Levine, S., Koltun, V.: Guided policy search, In: Dasgupta, S., McAllester, D. (eds.), Proceedings of the 30th International Conference on Machine Learning (ICML’13), pp.1 (2013), which is available online at http: / / proceedings.mlr.press / v28 / levine13.pdf.

[0038] It is proposed to use the step width σ determined by CMA-ES, which consists of successively consecutive generations g, as a guiding trajectory for GPS. That is, a plurality of guiding trajectories are determined, and then a teacher is created therefrom. For this, refer to the literature on GPS mentioned above. The teacher is a "teacher distribution" (GPS: "trajectory distribution" / "guiding distribution") that is determined so that the reward is maximized and, in particular, the deviation from the strategy π is minimized in GPS. Since GPS updates the teacher over time to improve the reward, there is a possibility that the teacher may deviate from the step width determined by CMA-ES (in particular, according to Equation 2) over time. This is because only the student (strategy π) and the teacher are forced to stay close to each other by GPS. If both the teacher and the student deviate significantly from the step width σ determined by CMA-ES, it may happen that the learned strategy π cannot reproduce the behavior of the case of the step width σ determined by CMA-ES.

[0039] Therefore, the inventors propose to use, in addition to the teacher, other sampled trajectories for the step width σ determined by CMA-ES in order to obtain various educational experiences. That is, alternatively, in order not to deviate too far from the teacher's trajectory, the student can be limited by a rigidity divergence criterion for the teacher, and other sampled characteristics of the step width σ determined by CMA-ES are used as trajectories. In the following, this is also referred to as an additional teacher. The additional teacher can be calculated as described above for the teacher.

[0040] To use an additional teacher with optimal action, it is proposed to extend GPS by introducing a sampling rate. The sampling rate characterizes at what rate GPS uses the teacher's carrier and the additional teacher for strategy learning. The sampling rate can be said to determine how similar the strategy is to the occurrence of CMA-ES. From experiments, it has been found that a sampling rate of 0.3 for the additional teacher yields the best results regardless of the type of function f. In other words, this means that during the training trajectory test, the teacher is used with a probability of 0.7, and the trajectory test of the additional teacher is taken out with a probability of 0.3.

[0041] Note that the teacher or the additional teacher may be other heuristics that determine the step size of CMA-ES.

[0042] Furthermore, to ensure extrapolation during learning, note that the maximum value of one or more initial steps with respect to the median of the starting distribution of the function in the training set is randomly taken from a uniform distribution.

[0043] Preferably, as the strategy π, a neural network consisting of two hidden layers each containing 50 hidden units and a ReLu activation part is used.

[0044] Note that not only the step size but also other parameters, such as the population size of CMA-ES, can be adapted based on the strategy if the strategy is optimized for this purpose. Preferably, the parameters are initially determined using SMAC in advance to obtain a proper initialization.

[0045] SMAC is described in the publication Hutter, F., Hoos, H., Leyton-Brown, K., “Sequential model-based optimization for general algorithm conguration”, In: Coello, C.(ed.), Proceedings of the Fifth International Conference on Learning and Intelligent Optimization (LION’11), Lecture Notes in Computer Science, vol 6683, pp.507 (2011).

[0046] FIG. 1 shows, by way of example, an apparatus for executing the above-described method and an apparatus for operating a strategy π. On the one hand, an agent 10 for reinforcement learning receives state information S and a reward signal R and outputs an action A. The action A corresponds in this case to the step width σ. During the learning of the agent 10, GPS is executed to learn the strategy π depending on the state information S and the reward signal R. During the operation of the agent 10, the strategy is applied, i.e., the action A is determined depending on the state information S.

[0047] The agent 10 receives the state information S and the reward signal R from the environment 11. The environment 11 here includes the following elements, namely, a reward signal generator 13, an internal state calculation circuit 12, and a CMA-ES applied to the problem instance 14.

[0048] The reward signal generator 13 calculates the reward signal R as described above, preferably from the action A, depending on the state information S. The state calculation circuit 12 calculates the above-described state and in particular the teacher and the additional teacher. By the action A, the CMA-ES of the problem instance 14 is configured and its regular steps, namely, "forming the next generation of candidates", "evaluating the fitness", "adapting the median m according to Equation 1", and "adapting the covariance matrix c" are executed.

[0049] Further, FIG. 1 shows a calculation unit 15 configured to execute calculations for the device, and a memory 16.

[0050] FIG. 2 schematically shows a flowchart 20 of an embodiment of the method according to the present invention.

[0051] The method starts at step S21. Here, CMA-ES is applied to a given problem instance 14, and a step width σ is calculated and characterized over all generations. For this purpose, the state calculation circuit 12 can be used.

[0052] Then, it proceeds to step S22, where the strategy π is initialized. The strategy may be, for example, a neural network that calculates the step width σ depending on the state information and randomly initializes its parameters, particularly the weights, as already described.

[0053] Then, it proceeds to step S23. In this step, reinforcement learning is applied to optimize the strategy π. This can be done by the agent 10 exploring the problem instance 14 and learning how to select the action A depending on the state information S in order to obtain the most robust reward signal R or the sum of the reward signals R as much as possible. The state information S can be obtained, for example, by the state calculation circuit 12.

[0054] After step S23 ends, it proceeds to step S24. After the strategy π is learned, it can be used to solve a new problem instance 14.

[0055] In one embodiment of the present invention, the problem instance 14 may be the optimization of hyperparameters. In this case, by the CMA-ES and the strategy π, the parameter representation of the hyperparameters can be optimized. The optimization of hyperparameters can be used, for example, for the hyperparameters of a learning algorithm for machine learning. Here, the hyperparameters may be, for example, the learning rate or the weight of the regularization term of the cost function used by the learning algorithm. For example, when the learning rate is optimized, additional teachers for GPS can be used, for example, as heuristics. That is, the cosine-annealing method, the exponential decaying learning rate, or the step-wise decaying learning rate can be used.

[0056] Preferably, the learning algorithm of machine learning is used for the learning of a machine learning system, particularly for the learning of a neural network, and is used, for example, for computer vision based on a computer. That is, it is used, for example, for object recognition or object localization or semantic segmentation.

[0057] In other embodiments, the problem instance 14 may be Vehicle Routing. In this case, by the CMA-ES and the strategy π, a route for the vehicle is determined.

[0058] Other problem instances may be, for example, a time planning problem and a completion problem. That is, for example, it can be determined which manufacturing machine is responsible for and executes which manufacturing task / production task.

[0059] The method of using the strategy for improving the evolutionary algorithm disclosed above is applicable to other evolutionary algorithms. For example, instead of CMA-ES, a differential evolution (DE) algorithm can also be used. In this case, the differential weight of differential evolution (DE) can be adapted by the strategy. When GPS is used for learning the strategy, the following heuristics, namely, DE-APC, ADP, SinDE, DE-random, or SaMDE, are used as the teacher / additional teacher.

[0060] The above-described machine learning system has learned using the method according to the present invention and can be used as follows.

[0061] Preferably, at regular time intervals, the surroundings are detected using sensors, particularly imaging sensors such as video sensors, and sensors that can be provided by a plurality of sensors, for example, using a stereo camera. Other imaging sensors, such as radar, ultrasonic, or LIDAR, are also conceivable. A thermographic camera is also conceivable. The sensor signal S of the sensor (or, in the case of a plurality of sensors, the respective sensor signals S) is transmitted to the control system. Here, the control system receives a series of sensor signals S. The control system calculates a drive signal A from the sensor signal S and transmits this to the actuator.

[0062] The control system receives a series of sensor signals S of the sensor with an optimal receiving unit, and this optimal receiving unit converts the sensor signal S into a series of input images x (alternatively, it is also possible to directly receive each sensor signal S as the input image x). The input image x can be, for example, a cutout of the sensor signal S or a further processed sensor signal S. The input image x includes individual frames of a video display. In other words, the input image x is calculated depending on the sensor signal S. A series of input images x are supplied to a machine learning system, which is an artificial neural network in this embodiment.

[0063] The artificial neural network calculates the output force y from the input image x. The output force y may particularly include the classification and / or semantic segmentation of the input image x. The output force y is supplied to an optimal deformation unit, from which a drive signal A is calculated and supplied to the actuator to drive and control the actuator 10 correspondingly. The output force y includes information about the object detected by the sensor. The actuator receives the drive signal A and is driven correspondingly to execute the corresponding action.

[0064] FIG. 3 shows how a control system 40 that controls at least a partially autonomous robot, here a at least a partially autonomous vehicle 100, can be used.

[0065] The sensor 30 may be, for example, preferably a video sensor arranged inside the vehicle 100.

[0066] The artificial neural network 60 is configured to safely identify an object from the input image x.

[0067] The actuator 10, preferably arranged inside the vehicle 100, may be, for example, the brake, drive mechanism or steering of the vehicle 100. The drive signal A can be calculated in this case such that one or more actuators 10 are driven and controlled, in particular when a specific class of object is, for example, a pedestrian, so as to avoid a collision between the vehicle 100 and the object safely identified by, for example, the artificial neural network 60.

[0068] Alternatively, the at least partially autonomous robot may be another mobile robot (not shown), for example, a robot propelled by flight, navigation, diving, or land movement. The mobile robot may be, for example, at least a partially autonomous lawn mower or at least a partially autonomous cleaning robot. Also, in this case, the drive signal A can be calculated so that the drive mechanism and / or steering unit of the mobile robot are driven and controlled to avoid a collision between the at least partially autonomous robot and an object identified, for example, by the artificial neural network 60.

[0069] In other preferred embodiments, the control system 40 includes one or more processors 45 and at least one machine-readable storage medium storing instructions for causing the control system 40 to implement the method according to the present invention when executed on the processor 45.

[0070] In an alternative embodiment, a display unit 10a is also provided instead of or in addition to the actuator 10.

[0071] Alternatively or additionally, the display unit 10a can be driven by the drive signal A, for example, the calculated safety area is displayed. Also, for example, in the vehicle 100, since non-autonomous steering can also be performed, when it is determined that there is a risk of collision between the vehicle 100 and an object safely identified, the display unit 10a is driven by the drive signal A to send out an optical warning signal or an acoustic warning signal.

[0072] FIG. 4 shows an example in which a control system 40 for driving and controlling the manufacturing machine 11 is used by driving and controlling the actuator 10 that drives and controls the manufacturing machine 11 of the manufacturing system 200. The manufacturing machine 11 may be, for example, a punching machine, a saw, a boring machine, and / or a cutting machine.

[0073] The sensor 30 may be, for example, an optical sensor for detecting the characteristics of the production products 12a, 12b. The production products 12a, 12b may be movable. The actuator 10 that controls the manufacturing machine 11 can be driven and controlled depending on the assignment of the detected production products 12a, 12b so that the manufacturing machine 11 correspondingly executes the next processing step for the correct one of the production products 12a, 12b. Also, by identifying the correct characteristics of the production products 12a, 12b (i.e., without incorrect assignment), the manufacturing machine 11 can be adapted according to the same manufacturing steps for the processing of the next production product.

[0074] FIG. 5 shows an embodiment in which a control system 40 that controls the access system 300 is used. The access system 300 may include physical access monitoring, for example, a door 401. A video sensor 30 is provided for detecting personnel. The object identification system 60 can interpret the detected image. When a plurality of personnel are detected simultaneously, the mutual assignment of the personnel (i.e., objects), for example, by analyzing their movements, can calculate the ID of the personnel, for example, with particularly high reliability. The actuator 10 may be a lock that releases or prohibits access monitoring depending on the drive signal A, for example, a lock that opens or prohibits opening of the door 401. For this purpose, the drive signal A can be selected depending on the interpretation of the object identification system 60, for example, depending on the calculated ID of the personnel. Instead of physical access monitoring, logical access monitoring can also be performed.

[0075] FIG. 6 shows an example in which a control system 40 that controls the monitoring system 400 is used. The difference from the example shown in FIG. 5 is that in this example, a display unit 10a is provided instead of the actuator 10, and this is driven and controlled by the control system 40. For example, the artificial neural network 60 calculates the ID of the target recorded by the video sensor 30 with high reliability, and based on this, for example, a suspicious person can be estimated. In this case, the drive signal A is selected so that such a target is highlighted in color by the display unit 10a.

[0076] FIG. 7 shows an example in which a control system 40 for controlling the personal assistant 250 is used. The sensor 30 is preferably an optical sensor that receives an image of the gesture of the user 249.

[0077] Depending on the signal of the sensor 30, the control system 40 calculates the drive signal A of the personal assistant 250, for example, by performing gesture recognition in a neural network. In this case, the calculated drive signal A is transmitted to the personal assistant 250, and thereby, the corresponding drive control is performed. The calculated drive signal A is particularly selected to correspond to the desired drive control assumed by the user 249. The assumed desired drive control can be calculated depending on the gesture identified by the artificial neural network 60. The control system 40 can be selected to transmit the drive signal A to the personal assistant 250 depending on the assumed desired drive control, and / or can be selected to transmit the drive signal A to the personal assistant according to the assumed desired drive control.

[0078] The corresponding drive control may include, for example, the personal assistant 250 calling information from the database and reconfiguring it in an adjustable manner for the user 249.

[0079] Instead of the personal assistant 250, for corresponding drive control, home appliances (not shown), in particular, a washing machine, a range, an oven, a microwave oven, or a dishwasher may be provided.

[0080] FIG. 8 shows an embodiment in which a control system 40 for controlling a medical imaging system 500, for example, an MRT device, an X-ray device, or an ultrasonic device is used. The sensor 30 may be, for example, a sensor provided by an imaging sensor, and the display unit 10a is driven and controlled by the control system 40. For example, the neural network 60 can determine whether there is a point requiring attention in the area recorded by the imaging sensor, and in this case, the drive signal A is selected such that the area is highlighted in color by the display unit 10a.

[0081] The concept of "computer" includes any device for processing configurable computing protocols. Here, the computing protocol may exist in the form of software, or in the form of hardware, or in a mixed form of software and hardware.

Claims

A computer-implemented method (20) for learning a strategy (π) that optimally adapts at least one step size (σ) of a CMA-ES algorithm, comprising: Initializing the strategy (π) that calculates a parameter representation (A) of the step size (σ) depending on state information (S) regarding a problem instance (14); Learning the strategy (π) using reinforcement learning; wherein from an interaction between the CMA-ES algorithm, the parameter representation (A) determined using the strategy (π) depending on the state information (S), the problem instance (14), and a reward signal (R), it is learned which parameter representation is optimal for any possible state information. A method according to claim 1, Claim 2 wherein the state information (S) includes at least one value including the current parameter representation (A) of the step size (σ), the cumulative distance (p σ ), and / or the difference between function values of a function (f) to be optimized of the problem instance (14) in the current iteration and the previous iteration. A method according to claim 1. Claim 3 wherein reinforcement learning using "Guided Policy Search" (GPS) is used, the parameter representation (A) of the step size (σ) of the problem instance (14) is determined using configurable heuristics, from which a plurality of characteristics of the parameter representation (A) of the step size (σ) are characterized, and a teacher for learning a strategy is provided using GPS from the characteristics. A method according to claim 1 or 2. Claim 4 wherein a sampling rate is supplemented to GPS, the sampling rate characterizing with what probability a strategy is learned by the teacher or an additional teacher, the additional teacher including other characteristics of values of the step size (σ). A method according to claim 3.

5. The sampling rate is 0.3, The method according to claim 4.

6. A method for operating a strategy learned by the method according to any one of claims 1 to 5, The state information of the problem instance is determined, The learned strategy calculates a parameter display depending on the state information, CMA-ES is applied together with the parameter display for at least one iteration, Method.

7. The problem instance is a training process for machine learning, and the hyperparameters of the training process are optimized using CMA-ES using a strategy, or, The problem instance is vehicle route optimization or task planning in manufacturing / production, and the route or task planning is optimized using CMA-ES using a strategy, The method according to claim 6.

8. A computer program configured to implement the method according to any one of claims 1 to 7.

9. A machine-readable storage medium storing the computer program according to claim 8.

10. An apparatus comprising the machine-readable storage medium according to claim 9.

Citation Information

Patent Citations

  • Advertisement putting method, device and system, computing equipment and storage medium

    CN111222902A

  • Optimization method and optimization program

    JP2005202960A

  • Learning control system and learning control method

    JP2019046422A

  • Robot movement apparatus and related methods

    WO2020062002A1