META-SCHOLAR EVOLUTION STRATEGY BLACKBOX OPTIMIZATION CLASSIFIERS

DE102021204943B4Active Publication Date: 2025-08-14ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102021204943
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-29
Filing Date
2021-05-17
Publication Date
2025-08-14
Estimated Expiration
2041-05-17

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computational method (400) for training a meta-learned evolutionary strategy black-box optimization classifier to learn an actuator control command for actuating a computer-controlled machine, the method comprising: Receiving (402) one or more training functions F and one or more initial metalearning parameters Θ of the metalearning evolutionary strategy blackbox optimization classifier; Sampling (404) a generation of λ samples z1,...,z λ a sampled objective function f ∈ F of the one or more training functions F and an initial mean m (0) of the sampled objective function f ∈ F, where the generation of λ samples z1,...,z λ with a multivariate normal distribution N(m (0) = 0, e (0) C (0) = 1) with initial mean m (0) = 0 and an initial covariance matrix C (0) = 1 in samples x1,...,xλ is shifted and scaled using the equation: xi = m ( t ) + σ ( t ) C ( t ) 1 2 zi , where σ (t) a step size of a number T steps in t = 1,...,T; for the number of T steps in t = 1,...,T, calculating (406) a set of T means m (1) ,...,m (T) by running the meta-scholarly evolutionary strategy black-box optimization classifier on the sampled objective function f ∈ F using the initial mean m (0) ; Calculate (408) a loss function L(f(m 0 ), ..., f(m T )) of the set of T means m (1) ,...,m (T) ; and Updating (410) the one or more initial metalearning parameters Θ of the meta-learned evolutionary strategy black-box optimization classifier in response to a characteristic of the loss function L to obtain an updated meta-learned evolutionary strategy black-box optimization classifier that includes a weighted combination of the generation of λ samples with larger weights for samples with smaller values ​​for the objective function f ∈ F; Sending input signals received from a sensor into the updated meta-scholarly evolutionary strategy black-box optimization classifier to obtain output signals designed to characterize a classification of the input signals; and Sending an actuator control command to an actuator (14) of the computer-controlled machine in response to the output signals; and Actuating the computer-controlled machine in response to the actuator control command.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to computational methods and computer systems for training and deploying meta-learning evolutionary strategy black-box optimization classifiers (e.g., machine learning or ML algorithms). BACKGROUND

[0002] A black-box function is a function whose analytic form is unknown. Black-box functions may be unknown or too complex to be directly modeled. Optimization models have been developed to optimize black-box functions. An existing family of black-box optimizers is called evolutionary strategies. An evolutionary strategy is an optimization technique based on evolutionary concepts. Evolutionary strategies are used with nonlinear or nonconvex continuous optimization problems.

[0003] A well-known evolution strategy is xNES (exponential natural evolution strategy). xNES involves an exponential parameterization of a search distribution to guarantee invariance. xNES is designed to compute a natural gradient without the need for an explicit Fisher information matrix. Another evolution strategy is called CMA-ES (covariance matrix adaptation evolution strategy). CMA-ES uses a maximum likelihood principle that increases the probability of successful candidate solutions and search steps. CMA-ES also records two different paths of the time evolution of the strategy's distribution mean, which are also referred to as search or evolution paths. Unlike other optimization methods, CMA-ES requires fewer assumptions regarding the nature of the black-box function.CMA-ES does not require derivatives or the function values ​​themselves, but instead ranks candidate solutions to find the best one.

[0004] Patrick, M., Craig, AP, Cunniffe, NJ, Parry, M., and Gilligan, CA, Testing stochastic software using pseudo-oracles, in Proceedings of the 25 th International Symposium on Software Testing and Analysis (ISSTA 2016), Association for Computing Machinery, New York, NY, USA, 235-246 (2016) disclose a search-based technique for testing implementations of stochastic models by maximizing the differences between the implementation and a pseudo-oracle.

[0005] The Wikipedia article on the Covariance Matrix Fitting Evolution Strategy (CMA-ES) from April 13, 2020, available online at https: / / en.wikipedia.org / w / index.php?title=CMA-ES&oldid=950756813 [accessed March 14, 2025], discloses the principles for fitting parameters of a search distribution using a CMA-ES algorithm.

[0006] Slowik, A., and Kwasnicka, H, Evolutionary algorithms and their applications to engineering problems, Neural Computing and Applications, 32, 12363-12379 (2020) reveal a family of evolutionary algorithms, their main properties and applications in practice. SUMMARY

[0007] The invention is defined by the independent claims. The dependent claims define advantageous embodiments. According to one embodiment, a computational method for training a meta-scholarly evolutionary strategy black-box optimization classifier is disclosed. The method comprises receiving one or more training functions and one or more initial meta-learning parameters of the meta-scholarly evolutionary strategy black-box optimization classifier. The method further comprises sampling a sampled objective function from the one or more training functions and an initial mean of the sampled objective function. The method further comprises calculating a set of T number of means by running the meta-scholarly evolutionary strategy black-box optimization classifier on the sampled objective function using the initial mean for a number T of steps in t = 1,...,T.The method further comprises calculating a loss function from the set of T-number of means. The method further comprises updating the one or more initial metalearning parameters of the metalearning, evolutionary strategy black-box optimization classifier in response to a characteristic of the loss function.

[0008] In another embodiment, a computational method for learning an actuator control command from a meta-learned evolutionary strategy black-box optimization classifier having one or more parameters. The computational method comprises sampling and transforming a generation of samples λ into a generation of transformed samples λ. The computational method further comprises ranking the generation of transformed samples λ in response to one or more function evaluations of the generation of transformed samples λ. The method further comprises updating the one or more parameters of the meta-learned evolutionary strategy black-box optimization classifier in response to the ranked generation of transformed samples λ and one or more learned parameters.The one or more learned parameters can be trained on a set of objective functions that are similar to (e.g., share a functional characteristic with) the learned evolutionary strategy black-box optimization classifier. The method further includes sending input signals obtained from a sensor to the updated meta-learned evolutionary strategy black-box optimization classifier to obtain output signals configured to characterize a classification of the input signals. The method further includes sending an actuator control command to an actuator of a computer-controlled machine in response to the output signals.

[0009] In yet another embodiment, a computational method for training and using a meta-learned evolutionary strategy black-box optimization classifier is disclosed. The computational method includes receiving one or more learned parameters trained on a set of objective functions that are similar to (e.g., share a functional characteristic with) the learned evolutionary strategy black-box optimization classifier. The computational method further includes updating the meta-learned evolutionary strategy black-box optimization classifier with the one or more learned parameters in response to a generation of samples λ. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 shows a schematic representation of an interaction between a computer-controlled machine and a control system according to one embodiment. Fig. Figure 2 shows a schematic representation of the control system of Fig. 1, which is designed to control a vehicle which may be a partially autonomous vehicle or a partially autonomous robot. Fig. Figure 3 shows a schematic representation of the control system of Fig. 1, which is designed to control a manufacturing machine, such as a punching machine, a cutter or a pistol drill, of a manufacturing system, e.g. as part of a production line. Fig. Figure 4 shows a schematic representation of the control system of Fig. 1, which is designed to control a power tool, such as a drill or a drive with an at least partially autonomous mode. Fig. Figure 5 shows a schematic representation of the control system of Fig. 1, which is designed to control an automated personal assistant. Fig. Figure 6 shows a schematic representation of the control system of Fig. 1, which is designed to control a surveillance system, such as an access control system or a surveillance system. Fig. Figure 7 shows a schematic representation of the control system of Fig. 1, which is designed to control an imaging system, for example an MRI device, an X-ray imaging device or ultrasound device. Fig. 8 shows a schematic representation of a training system for training a classifier according to one or more embodiments. Fig. 9 shows a flowchart of a computational method for training a classifier (e.g., a black box algorithm) according to one or more embodiments. Fig. 10 shows a flowchart of a computational method for using a classifier (e.g., a black box algorithm) using a meta-learning evolution strategy according to one embodiment. DETAILED DESCRIPTION

[0010] Embodiments of the present disclosure are described herein. It should be understood, however, that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of specific components. Therefore, specific structural and functional details disclosed herein are not to be considered limiting, but merely as a representative basis for teaching those skilled in the art to variously employ the embodiments. As will be appreciated by those of ordinary skill in the art, various features illustrated and described with reference to any of the figures may be combined with features illustrated in one or more other figures to produce embodiments not expressly illustrated or described.The combinations of features illustrated provide representative embodiments for typical applications. However, various combinations and modifications of the features consistent with the teachings of the present disclosure may be desired for specific applications or implementations.

[0011] Fig. 1 shows a schematic representation of an interaction between a computer-controlled machine 10 and a control system 12. The computer-controlled machine 10 includes an actuator 14 and a sensor 16. The actuator 14 may include one or more actuators, and the sensor 16 may include one or more sensors. The sensor 16 is configured to detect a state of the computer-controlled machine 10. The sensor 16 may be configured to encode the detected state into sensor signals 18 and send the sensor signals 18 to the control system 12. Non-limiting examples of the sensor 16 would be video, radar, LiDAR, ultrasonic, and motion sensors. In one embodiment, the sensor 16 is an optical sensor configured to detect optical images of an environment proximate the computer-controlled machine 10.

[0012] The control system 12 is configured to receive sensor signals 18 from the computer-controlled machine 10. As explained below, the control system 12 may be further configured to learn actuator control commands 20 depending on the sensor signals and send the actuator control commands 20 to the actuator 14 of the computer-controlled machine 10.

[0013] As in Fig. 1, the control system 12 includes a receiving unit 22. The receiving unit 22 may be configured to receive sensor signals 18 from the sensor 30 and to transform the sensor signals 18 into input signals x. In an alternative embodiment, the sensor signals 18 are received directly as input signals x without the receiving unit 22. Each input signal x may be a part of each sensor signal 18. The receiving unit 22 may be configured to process each sensor signal 18 to produce each input signal x. The input signal x may include data corresponding to an image recorded by the sensor 16.

[0014] The control system 12 includes the classifier 24. The classifier 24 may be configured to learn actuator control commands 20 from input signals x using a machine learning (ML) algorithm, such as a neural network or RNN (recurrent neural network). The control system 12 may be configured to train the ML algorithm. The ML algorithm may be a meta-learning evolution strategy for black-box optimization, as disclosed in one or more present embodiments.

[0015] The classifier 24 is configured to be parameterized by one or more parameters. The parameters may be stored in and provided by non-volatile memory 26. The classifier 24 is configured to determine output signals y from input signals x. Each output signal y includes information that assigns one or more labels to each input signal x. The classifier 24 may send the output signals y to the conversion unit 28. The conversion unit 28 is configured to convert the output signals y into actuator control commands 20. The control system 12 is configured to send the actuator control commands 20 to the actuator 14, which is configured to actuate the computer-controlled machine 10 in response to the actuator control commands 20. In another embodiment, the actuator 14 is configured to actuate the computer-controlled machine 10 directly based on the output signals y.

[0016] Upon receipt of the actuator control commands 20 by the actuator 14, the actuator 14 is configured to perform an action corresponding to the respective actuator control command 20. The actuator 14 may include control logic configured to transform the actuator control commands 20 into a second actuator control command used to control the actuator 14. In one or more embodiments, the actuator control commands 20 may be used to control a display instead of or in addition to an actuator.

[0017] In another embodiment, the control system 12 includes the sensor 16 instead of or in addition to the computer-controlled machine 10, which includes the sensor 16. The control system 12 may also include the actuator 14 instead of or in addition to the computer-controlled machine 10, which includes the actuator 14.

[0018] As in Fig. 1, the control system 12 further includes a processor 30 and memory 32. The processor 30 may include one or more processors. The memory 32 may include one or more storage devices. The classifier 24 (e.g., ML algorithms) of one or more embodiments may be implemented by the control system 12, which includes non-volatile storage 26, the processor 30, and the memory 32.

[0019] Non-volatile storage 26 may include one or more persistent data storage devices, such as a hard drive, an optical drive, a tape drive, a non-volatile semiconductor device, cloud storage, or any other device capable of persistently storing information. Processor 300 may include one or more devices selected from among high-performance computing (HPC) systems, including high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units, field-programmable gate arrays, programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other devices that manipulate signals (analog or digital) based on computer-executable instructions residing in memory 32.The memory 32 may comprise a single memory device or a number of memory devices, including, but not limited to, RAM (random access memory), volatile memory, non-volatile memory, SRAM (static random access memory), DRAM (dynamic random access memory), flash memory, cache memory, or any other device capable of storing information.

[0020] The processor 30 may be configured to read into the memory 32 and execute computer-executable instructions residing in the non-volatile storage 26 and implementing one or more machine learning algorithms and / or methodologies of one or more embodiments. The non-volatile storage 26 may include one or more operating systems and applications. The non-volatile storage 26 may store compiled and / or interpreted computer programs created using a variety of programming languages ​​and / or technologies, including, without limitation and either alone or in combination, Java, C, C++, C#, Objective C, Fortran, Pascal, Java Script, Python, Perl, and PL / SQL.

[0021] When executed by processor 30, the computer-executable instructions of non-volatile storage 26 may cause control system 12 to implement one or more of the ML algorithms and / or methodologies disclosed herein. Non-volatile storage 26 may also include ML data (including data parameters) that support the functions, features, and processes of one or more embodiments described herein.

[0022] The program code implementing the algorithms and / or methodologies described herein may be distributed individually or collectively as a program product in a variety of different forms. The program code may be distributed using a computer-readable storage medium having computer-readable program instructions embodied thereon for causing a processor to perform aspects of one or more embodiments. Computer-readable storage media, which are non-transitory in nature, may include volatile and non-volatile, and removable and non-removable tangible media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.Computer-readable storage media may further include: RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other semiconductor storage technology, portable CD-ROM (Compact Disc Read-Only Memory) or other optical storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be read by a computer. Computer-readable program instructions may be downloaded into a computer, other type of programmable data processing apparatus, or other device from a computer-readable storage medium, or to an external computer or storage device over a network.

[0023] Computer-readable program instructions stored on a computer-readable medium can be used to direct a computer, other types of programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored on the computer-readable medium produce an article of manufacture including instructions that implement the functions, steps, and / or operations specified in the flowcharts or diagrams. In certain alternative embodiments, consistent with one or more embodiments, the functions, steps, and / or operations specified in the flowcharts and diagrams can be reordered, processed serially, and / or processed concurrently.Additionally, consistent with one or more embodiments, any of the flowcharts and / or illustrations may include more or fewer nodes or blocks than those illustrated.

[0024] The processes, methods or algorithms may be implemented in whole or in part using suitable hardware components such as ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), automatons, controllers or other hardware components or devices or a combination of hardware, software and firmware components.

[0025] Fig. 2 shows a schematic representation of the control system 12, which is designed to control a vehicle 50, which may be an at least partially autonomous vehicle or an at least partially autonomous robot. As in Fig. 2, the vehicle 50 includes an actuator 14 and a sensor 16. The sensor 16 may include one or more video sensors, radar sensors, ultrasonic sensors, LidDAR sensors, and / or position sensors (e.g., GPS). One or more of the one or more specific sensors may be integrated into the vehicle 50. Alternatively, or in addition to one or more specific sensors identified above, the sensor 16 may include a software module configured, when executed, to determine a state of the actuator 14. A non-limiting example of a software module would be a weather information software module configured to determine a current or future state of the weather in the vicinity of the vehicle 50 or at another location.

[0026] The classifier 24 of the control system 12 of the vehicle 50 can be configured to detect objects in the environment of the vehicle 50 based on the input signals x. In such an embodiment, the output signal y can include information characterizing the environment of objects in relation to the vehicle 50. The actuator control command 20 can be determined based on this information. The actuator control command 20 can be used to avoid collisions with the detected objects.

[0027] In embodiments where the vehicle 50 is an at least partially autonomous vehicle, the actuator 14 may be implemented in a brake, a drive system, an engine, a powertrain, or a steering system of the vehicle 50. Actuator control commands 20 may be determined such that the actuator 14 is controlled such that the vehicle 50 avoids collisions with detected objects. Detected objects may also be classified depending on what the classifier 24 considers them most likely to be, such as pedestrians or trees. The actuator control commands 20 may be determined depending on the classification.

[0028] In other embodiments where the vehicle 50 is an at least partially autonomous robot, the vehicle 50 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, or walking. The mobile robot may be an at least partially autonomous lawnmower or an at least partially autonomous cleaning robot. In such embodiments, the actuator control command 20 may be determined such that a drive unit, a steering unit, and / or a braking unit of the mobile robot can be controlled such that the mobile robot can avoid collisions with identified objects.

[0029] In another embodiment, the vehicle 50 is an at least partially autonomous robot in the form of a gardening robot. In such an environment, the vehicle 50 may use an optical sensor, such as sensor 16, to determine a condition of plants in an environment near the vehicle 50. The actuator 14 may be a nozzle configured to spray chemicals. Depending on the identified species and / or an identified condition of the plants, the actuator control command 20 may be determined to cause the actuator 14 to spray the plants with an appropriate amount of suitable chemicals.

[0030] The vehicle 50 may be an at least partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances would be a washing machine, a stove, an oven, a microwave, or a dishwasher. In such a vehicle 50, the sensor 16 may be an optical sensor configured to detect a condition of an object to be processed by the household appliance. For example, if the household appliance is a washing machine, the sensor 16 may detect a condition of the laundry in the washing machine. The actuator control command 20 may be determined based on the detected condition of the laundry.

[0031] Fig. Figure 3 shows a schematic representation of the control system 12 configured to control a manufacturing machine 100, such as a punching machine, a cutter, or a pistol drill, of a manufacturing system 102, e.g., as part of a production line. The control system 12 may be configured to control the actuator 14 configured to control the production machine 100.

[0032] The sensor 16 of the production machine 100 can be an optical sensor configured to detect one or more properties of a manufactured product 104. The classifier 24 can be configured to determine a state of the manufactured product 104 from one or more of the detected properties. The actuator 14 can be configured to control the production machine 100 for a subsequent manufacturing step of the manufactured product 104 depending on the determined state of the manufactured product 104. The actuator 14 can be configured to control functions of the production machine 100 on the subsequent manufactured product 106 of the production machine 100 depending on the determined state of the manufactured product 104.

[0033] Fig. Figure 4 shows a schematic representation of the control system 12 configured to control a power tool 150, such as a drill or a drive with an at least partially autonomous mode. The control system 12 may be configured to control an actuator 14 configured to control the power tool 150.

[0034] The sensor 16 of the power tool 150 may be an optical sensor configured to detect one or more characteristics of the work surface 152 and / or the fastener 154 driven into the work surface 152. The classifier 24 may be configured to determine, from one or more of the detected characteristics, a condition of the work surface 152 and / or the fastener 154 relative to the work surface 152. The condition may be that the fastener 154 is flush with the work surface 152. Alternatively, the condition may be the hardness of the work surface 154. The actuator 14 may be configured to control the power tool 150 such that the drive function of the power tool 150 is adjusted depending on the determined condition of the fastener 154 relative to the work surface 152 or one or more detected characteristics of the work surface 154.For example, the actuator 14 may terminate the drive function when the state of the fastener 154 is flush relative to the work surface 152. As another non-limiting example, the actuator 14 may apply more or less torque depending on the hardness of the work surface 152.

[0035] Fig. Figure 5 shows a schematic representation of the control system 12 configured to control an automated personal assistant 200. The control system 12 may be configured to control the actuator 14 configured to control the automated personal assistant 200. The automated personal assistant 200 may be configured to control a household appliance, such as a washing machine, a stove, an oven, a microwave, or a dishwasher.

[0036] Sensor 16 may be an optical sensor and / or an audio sensor. The optical sensor may be configured to receive video images of gestures 204 from user 202. The audio sensor may be configured to receive a voice command from user 202.

[0037] The control system 12 of the automated personal assistant 200 may be configured to determine actuator control commands 20 configured for the control system 12. The control system 12 may be configured to determine actuator control commands 20 according to sensor signals 18 from the sensor 16. The automated personal assistant 200 is configured to send sensor signals 18 to the control system 12. The classifier 24 of the control system 12 may be configured to execute a gesture recognition algorithm to identify the gesture 204 performed by the user 202 to determine the actuator control commands 20 and send the actuator control commands 20 to the actuator 14. The classifier 24 may be configured to retrieve information from the non-volatile storage in response to the gesture 204 and output the retrieved information in a form suitable for receipt by the user 202.

[0038] Fig. Figure 6 shows a schematic representation of control system 12 configured to control a surveillance system 250. Surveillance system 250 may be configured to physically control access through door 252. Sensor 16 may be configured to detect a scene relevant to determining whether to grant access. Sensor 16 may be an optical sensor configured to generate and transmit image and / or video data. Such data may be used by control system 12 to detect a person's face.

[0039] The classifier 24 of the control system 12 of the surveillance system 250 may be configured to interpret the image and / or video data by comparing identities of known individuals stored in the non-volatile storage 26 to thereby determine an individual's identity. The classifier 12 may be configured to generate an actuator control command 20 in response to the interpretation of the image and / or video data. The control system 12 is configured to send the actuator control command 20 to the actuator 12. In this embodiment, the actuator 12 may be configured to lock or unlock the door 252 in response to the actuator control command 20. In other embodiments, other non-physical, logical access control is also possible.

[0040] The monitoring system 250 may also be an observation system. In such an embodiment, the sensor 16 may be an optical sensor configured to detect a scene being observed, and the control system 12 is configured to control the display 254. The classifier 24 is configured to determine a classification of a scene, e.g., whether the scene detected by the sensor 16 is suspicious. The control system 12 is configured to send an actuator control command 20 to the display 254 in response to the classification. The display 254 may be configured to adjust the displayed content in response to the actuator control command 20. For example, the display 254 may highlight an object deemed suspicious by the classifier 24.

[0041] Fig. 7 shows a schematic representation of the control system 12 configured to control an imaging system 300, for example, an MRI device, an x-ray imaging device, or ultrasound device. The sensor 16 may be, for example, an imaging sensor. The classifier 24 may be configured to determine a classification of all or a portion of the acquired image. The classifier 24 may be configured to determine or select an actuator control command 20 in response to the classification. For example, the classifier 24 may interpret a region of an acquired image as potentially abnormal. In this case, the actuator control command 20 may be determined or selected to cause the display 302 to display the imaging and highlight the potentially abnormal region.

[0042] Evolutionary strategies have been used as black-box function optimizers. Evolutionary strategies use iterative black-box optimization methods. An evolutionary strategy attempts to find a global minimizer of an unknown (e.g., black-box) loss function f in a feasible space X. Equation (1) shows this calculation expressed algebraically. argminx∈X f(x)

[0043] Examples of existing evolution strategies include xNES (exponential natural evolution strategy) and CMA-ES (covariance matrix adaptation evolution strategy). Evolution strategies are designed to perform optimization on multiple objective functions without tuning hyperparameters. Existing evolution strategies use predefined heuristics and hand-tuned hyperparameters to perform their parameter updates. Standard evolution strategies do not utilize any knowledge of the structure of a function being optimized. Therefore, optimization performance cannot be improved with this knowledge. Accordingly, computational optimization methods are needed that incorporate prior knowledge of an optimization problem by first training on a set of similar objective functions.

[0044] In one or more embodiments, computational methods and computer systems are presented that exploit knowledge of the structure of a function to improve optimization performance. In one or more embodiments, the computational methods and computer systems meta-learn a black-box optimization algorithm using ML, e.g., deep learning. An RNN (recurrent neural network) can be trained to perform the black-box optimization. In one embodiment, the computational methods and computer systems use deep meta-learning to improve the performance of an optimization algorithm on a specific class of functions.Such a class may refer broadly to the properties or characteristics of the functions in the class, such as the amount of noise in the evaluation of the function or the degree of a set of polynomials, or may be very specific, such as functions resulting from numerical simulations of a particular experiment. The computational methods and computer systems may include meta-learning parameters θ designed to define how the evolutionary strategy updates its parameters. For example, static hyperparameters (e.g., learning rates, step sizes, momentum-controlling coefficients, and / or coefficients controlling the weight given to each term in an update of their parameters) in the evolutionary strategy algorithm can be translated into learnable parameters. As another example, heuristic updates to m and σ can be achieved using ML, e.g.,a neural network, where in this case θ represents the weights and / or biases of the neural network.

[0045] In one or more embodiments, since the operations performed by an evolutionary strategy are differentiable, the evolutionary strategy algorithm of the computational methods and computer systems is implemented in an auto-differentiation framework, where the algorithm computes a value and automatically constructs a procedure for computing derivatives of the value. The metalearning parameters θ can then be trained using a supervised gradient descent learning method.

[0046] In an evolutionary strategy of one or more embodiments, the objective function f: R n→ R is stored in non-volatile memory 354. Memory 358 includes instructions that, when executed by processor 356, execute an evolution strategy designed to optimize f by iteratively updating one or more of the following parameters: m ∈ R n , σ ∈ R, C ∈ R n×n a multivariable Gaussian bell curve N(m,σC). m is a mean. σ is a step size. C is a covariance matrix. The iterative steps may be defined as t = 1, ..., T. During each step t, the following steps may be executed iteratively by the processor 356. The first step may have two iterative components. The two components are executed iteratively for a number of samples, defined as i = 1, ..., λ. The first component is sampling a number λ of samples z i~ N(0,1), which is a multivariable normal distribution with zero mean and unit covariance matrix, and is otherwise called generation. A vector distributed according to N(0,1) has independent (0,1)-normally distributed components. The second component is scaling and shifting the λ number of samples z. i ~ N(0,1) to x i Samples using the following equation: xi=m+σC12zi

[0047] In a second iterative step, the samples x1,...,x λ classified according to their function evaluations, so that the classified samples f(x1) ≤ f(x2) ≤...≤ f(x λ ). In a third iterative step, the evolution parameters m ∈ R n , σ ∈ R, C ∈ R n×n according to x1,...,x λand one or more parameters θ, which is determined using a training method of one or more embodiments. As part of the updating step, the mean m of the Gaussian bell can be updated by replacing it with a weighted combination of samples in the current generation, giving more weight to samples with better (i.e., smaller) function values.

[0048] The optimization algorithm for the function f may be trained on a training set of functions, as described herein with reference to one or more embodiments. The optimization algorithm for the function f may be repeated (e.g., the three iterative steps) until a termination criterion is reached. For example, the process may terminate when the best function value does not change after a certain number of iterations, such as 100 or 1000 iterations, or any number of iterations in between.

[0049] Fig. Figure 8 shows a schematic representation of a training system 350 for training the classifier 24. In one or more embodiments, computational methods and computer systems of one or more embodiments are presented for training the classifier 24, which may be used in one or more of the Fig. 2-7. For example, the classifier 24 can be implemented with a robot to learn to perform a specific task or to learn the best parameters for a laser or other production tool.

[0050] The classifier 24 may be a meta-evolution strategy algorithm. A set of training functions (F) and initial meta-learning parameters θ of an evolution strategy may be stored in the non-volatile memory 354. The training system 350 is configured to execute a training procedure for finding optimized meta-learning parameters θ. The memory 358 includes instructions that, when executed by the processor 356, execute the training procedure. The training system 350 is configured to execute the training procedure in a series of steps. The first step may be sampling a function f ∈ F and an initial mean m {(0)}~ U[-1,1] n The second step can be calculating the mean values ​​m (1) ,...,m (T) by executing one of the optimization algorithms of one or more embodiments on the objective function f with the initial mean value m (0) for a number T of steps in t = 1, ..., T. The third step can be calculating a loss L from the evaluated means f(M 0 ),...,f(m T ) to obtain a scalar loss. The loss L can be expressed in the form of a loss function L(f(m 0 ),...,f(m T )). The fourth step can be updating the parameters θ using the gradients of the loss function ∇ θ L. Backpropagation can be performed to train the parameters θ with respect to the loss.

[0051] Fig. Figure 9 shows a flowchart 400 of a computational method for training classifier 24 (e.g., a black-box algorithm) using a meta-learning evolution strategy according to one embodiment. The computational method may be performed using training system 350. The computational method for training classifier 24 may be identified by a training procedure.

[0052] In step 402, inputs for the training procedure are received. In one embodiment, the inputs include a set of training functions F and initial meta-learning parameters θ of the evolution strategy. In step 404, one or more of the set of training functions F are used.

[0053] In step 404, a function f ∈ F and an initial mean m {(0)} ~ U[-1,1] nsampled. Function sampling and initial mean sampling are used in step 406 as set forth below.

[0054] In step 406, mean values ​​m (1) ,...,m (T) by running a meta-scholarly evolutionary strategy algorithm (e.g. the one in Fig. 10 identified algorithm) on the objective function f with the initial mean value m {(0)} calculated. As in Fig. As shown in Figure 9, step 406 is repeated for a number of T steps. The calculated mean values ​​after a number of T steps are used in step 408 as set forth below.

[0055] In step 408, a loss function is calculated from the means calculated in step 406. Step 408 may calculate a loss L from the evaluated means f(M 0 ), ... , f(m T ) to obtain a scalar loss. The loss L can be expressed as a loss function L(f(M 0),...,f(m T )). This loss function calculation may include other parameters.

[0056] In step 410, the meta-learned parameters θ of the evolution strategy are calculated using the gradients of the loss function ∇ θ L is updated. Although step 410 uses gradients of the loss function, in other embodiments, this step may be performed by gradient descent or other deep learning optimizers, such as Adam.

[0057] Steps 404, 406, 408, and 410 are repeated until the meta-scholarly optimization algorithm converges, as set forth in step 412. The last updated values ​​of the meta-scholarly parameters θ from step 414 are used in the meta-scholarly optimization algorithm, as set forth below.

[0058] As in Fig. 9, one function is used to train θ at a time. In one or more other embodiments, the training procedure may be performed for a minibatch of functions, with a minibatch size of up to or greater than 128 functions. In other embodiments, the number of functions in the minibatch may be between 20 and 30. In still other embodiments, the number of functions in the minibatch may exceed 128. In alternative embodiments, samples from a distribution, e.g., samples from a Gaussian process, may be used instead of from a prescribed set of training functions.

[0059] Fig. 10 shows a flowchart 450 of a computational method for using a classifier 24 (e.g., a black-box algorithm) using a meta-learning evolution strategy according to one embodiment. The computational method may be executed using the control system 12.

[0060] In step 452, for example i ~ N(0,1) for i = 1,...,λ. In step 454, z i using equation (2) by scaling and shifting z i transformed. As in Fig. As shown in Figure 10, steps 452 and 454 are repeated for i = 1,...,λ. The transformed samples x i are used by step 456.

[0061] In step 456, the x i -Samples are classified according to their function evaluations, such that the classified samples f(x1) ≤ f(x2) ≤...≤ f(x λ ). The classified samples x iare used by step 458.

[0062] In step 458, the ranked samples x i and the last updated values ​​of the meta-learned parameters θ from step 414 of Fig. 9 is used to update the evolution strategy parameters m, σ and C of the black box optimization.

[0063] Although, as in Fig. 10, the metaevolution strategy algorithm is run for a fixed number of T steps, it can also be run until a stopping criterion is reached. Fig. Figure 10 also shows a fixed number of λ samples per generation, but this number can be varied, or samples from previous generations can be used.

[0064] Although exemplary embodiments are described above, these embodiments are not intended to describe all possible forms encompassed by the claims. The words used in the specification are words of description, not limitation, and it is understood that various changes may be made without departing from the spirit and scope of the disclosure. As described above, the features of various embodiments may be combined to form further embodiments of the invention that may not be expressly described or illustrated.Although various embodiments may have been described as providing advantages or being preferred over other embodiments or prior art implementations with respect to one or more desired characteristics, those of ordinary skill in the art will appreciate that one or more features or characteristics may be compromised to achieve desired overall system attributes depending on the specific application and implementation. These attributes may include, but are not limited to, cost, resilience, durability, life cycle cost, marketability, appearance, packaging, size, maintainability, weight, manufacturability, ease of assembly, etc.To the extent that any embodiments are described as less desirable than other embodiments or prior art implementations with respect to one or more features, these embodiments are not outside the scope of the disclosure and may be desirable for particular applications.

Claims

[1] A computational method (400) for training a meta-learned evolutionary strategy black-box optimization classifier to learn an actuator control command for actuating a computer-controlled machine, the method comprising: Receiving (402) one or more training functions F and one or more initial metalearning parameters Θ of the metalearning evolutionary strategy blackbox optimization classifier; Sampling (404) a generation of λ samples z1,...,z λ a sampled objective function f ∈ F of the one or more training functions F and an initial mean m (0) of the sampled objective function f ∈ F, where the generation of λ samples z1,...,z λ with a multivariate normal distribution N(m (0) = 0, e (0) C (0) = 1) with initial mean m (0) = 0 and an initial covariance matrix C (0) = 1 in samples x1,...,xλ is shifted and scaled using the equation: xi=m(t)+σ(t)C(t)12zi, where σ (t) a step size of a number T steps in t = 1,...,T; for the number of T steps in t = 1,...,T, calculating (406) a set of T means m (1) ,...,m (T) by running the meta-scholarly evolutionary strategy black-box optimization classifier on the sampled objective function f ∈ F using the initial mean m (0) ; Calculate (408) a loss function L(f(m 0 ), ..., f(m T )) of the set of T means m (1) ,...,m (T) ; and Updating (410) the one or more initial metalearning parameters Θ of the meta-learned evolutionary strategy black-box optimization classifier in response to a characteristic of the loss function L to obtain an updated meta-learned evolutionary strategy black-box optimization classifier that includes a weighted combination of the generation of λ samples with larger weights for samples with smaller values ​​for the objective function f ∈ F; Sending input signals received from a sensor into the updated meta-scholarly evolutionary strategy black-box optimization classifier to obtain output signals designed to characterize a classification of the input signals; and Sending an actuator control command to an actuator (14) of the computer-controlled machine in response to the output signals; and Actuating the computer-controlled machine in response to the actuator control command. [2] A computational method according to claim 1, wherein the characteristic of the loss function L gradient ∇ θ L is the loss function L. [3] A computational method according to claim 1, wherein the characteristic of the loss function L is gradient descent of the loss function L. [4] The computational method of claim 1, wherein the updating step (410) is performed by a deep learning optimizer. [5] The computational method of claim 1, wherein the sampling step (404), the first calculation step (406), the second calculation step (408) and the updating step (410) are executed interactively in a loop until a stop condition is met. [6] The computational method of claim 5, wherein the stopping condition is convergence of the meta-learned evolutionary strategy black-box optimization classifier. [7] A computational method for learning an actuator control command from a meta-learned evolutionary strategy black-box optimization classifier for actuating a computer-controlled machine, the method comprising: Sampling (404) and transforming a generation of λ samples z1,...,z λ into a generation of λ transformed samples, where the generation of λ samples z1,...,z λ with a multivariate normal distribution N(m (0) = 0, σ (0) C (0) = 1) with initial mean m (0) = 0 and an initial covariance matrix C (0) = 1 in samples x1,...,x λ is shifted and scaled using the equation: xi=m(t)+σ(t)C(t)12zi, x i where σ (t) a step size of a number T steps in t = 1, ..., T; ranking the generation of λ-transformed samples in response to one or more function evaluations of the generation of λ-transformed samples to obtain a ranked generation of λ-transformed samples; Updating (410) one or more parameters of the meta-learned evolutionary strategy black-box optimization classifier in response to the ranked generation of λ-transformed samples and one or more learned parameters trained on a set of objective functions that share a function characteristic with the learned evolutionary strategy black-box optimization classifier to obtain an updated meta-learned evolutionary strategy black-box optimization classifier that includes a weighted combination of the ranked generation of λ-transformed samples with larger weights for the transformed samples of the ranked generation of λ-transformed samples with smaller values ​​for the one or more function evaluations; Sending input signals received from a sensor into the updated meta-scholarly evolutionary strategy black-box optimization classifier to obtain output signals designed to characterize a classification of the input signals; and Sending an actuator control command to an actuator of the computer-controlled machine in response to the output signals; and Actuating the computer-controlled machine in response to the actuator control command. [8] A computational method according to claim 7, wherein the number of parameters m (t) , σ (t) , and C (t) includes. [9] A computational method according to claim 8, wherein a neural network is used to update m (t) and σ (t) is used. [10] A computational method according to claim 9, wherein the number of learned parameters represents weights and / or biases of the neural network. [11] A computational method according to claim 7, wherein the generation of λ transformed samples from a multivariate normal distribution N(m (t) , σ (t) C (t) ) is produced. [12] A computational method according to claim 7, wherein the one or more function evaluations are one or more objective function evaluations. [13] A computational method according to claim 7, wherein the ranked generation of λ transformed samples satisfies f(x1) ≤ f(x2) ≤ ...≤ f(x1). [14] A computational method for training and using a meta-learned evolutionary strategy black-box optimization classifier to learn an actuator control command for actuating a computer-controlled machine, the method comprising: Receiving (402) one or more learned parameters trained on a set F of objective functions f ∈ F that share a function characteristic with the learned evolutionary strategy black-box optimization classifier and sampled (404) from a generation of λ samples z1,...,z λ , where the generation of λ samples z1,...,z λ with a multivariate normal distribution N(m (0) = 0,σ (0) C (0) = 1) with initial mean m (0) = 0 and an initial covariance matrix C (0) = 1 in samples x1,...,x λ is shifted and scaled using the equation: xi=m(t)+σ(t)C(t)12zi, where σ (t) denotes a step size of a number of T steps in t = 1,...,T; and Updating (410) the meta-learned evolutionary strategy black-box optimization classifier with the one or more learned parameters in response to a generation of λ samples to obtain an updated meta-learned evolutionary strategy black-box optimization classifier that includes a weighted combination of the generation of λ samples with larger weights for samples with smaller values ​​for the one or more function evaluations; Sending input signals received from a sensor into the updated meta-scholarly evolutionary strategy black-box optimization classifier to obtain output signals designed to characterize a classification of the input signals; and Sending an actuator control command to an actuator of a computer-controlled machine in response to the output signals; and Actuating the computer-controlled machine in response to the actuator control command. [15] The computational method of claim 14, wherein the number of learned parameters comprises a learning rate. [16] The computational method of claim 14, wherein the one or more learned parameters comprise a neural network. [17] A computational method according to claim 14, wherein the generation of λ samples from a multivariate normal distribution N(m (t) , σ (t) C (t) ) is produced. [18] The computational method of claim 14, wherein the updating step (410) comprises updating the meta-learned evolutionary strategy black-box optimization classifier with the one or more learned parameters and one or more static parameters in response to the generation λ of samples λ.