Meta-learning evolutionary strategy black-box optimization classifier

By training the meta-learning evolution strategy black box optimization classifier on a set of similar objective functions, the machine learning meta-learning method is used to optimize the parameters of the evolution strategy, and the problem of assumptions of multiple and derivative information needs for black box functions in the existing technology is solved, and efficient black box function optimization is achieved.

CN113743442BActive Publication Date: 2025-05-13ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110590519.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-29
Filing Date
2021-05-28
Publication Date
2025-05-13
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

Existing evolutionary strategies require more assumptions and derivative information when optimizing black box functions, and it is difficult to effectively utilize the knowledge of function structure to improve optimization performance.

Method used

The black box optimization classifier is adopted for the meta-learning evolution strategy, and by training on a similar set of objective functions, the machine learning meta-learning method is used to optimize the parameters of the evolution strategy, reducing the assumptions of the black box function, and optimizing without the need for derivatives.

Benefits of technology

It realizes black box function optimization without the need for explicit Fisher information matrix and derivative, improves optimization performance, and can effectively utilize functional structure knowledge to improve optimization effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113743442B_ABST
    Figure CN113743442B_ABST
Patent Text Reader

Abstract

A computational method for training a meta-learning evolution strategy black-box optimization classifier. The method includes: receiving one or more training functions and one or more initial meta-learning parameters of the meta-learning evolution strategy black-box optimization classifier. The method further includes: sampling sampled target functions from the one or more training functions and initial means of the sampled functions. The method also includes: for t =1,…, T In T The number of steps is calculated by running the meta-learning evolution strategy black-box optimization classifier on the sampled objective function using the initial mean T The method also includes: T The method further comprises: updating one or more initial meta-learning parameters of the meta-learning evolution strategy black-box optimization classifier in response to the characteristics of the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computational methods and computer systems for training and providing meta-learning evolution strategy black-box optimization classifiers (eg, machine learning (ML) algorithms). Background Art

[0002] A black box function is a function for which no analytical form is known. Black box functions may be unknown or too complex to be modeled directly. Optimization models have been developed to optimize black box functions. One existing family of black box optimizers is called evolution strategies. Evolution strategies are an optimization technique based on the concept of evolution. Evolution strategies are used in nonlinear or non-convex continuous optimization problems.

[0003] One known evolution strategy is the exponential natural evolution strategy (xNES). xNES includes an exponential parameterization of the search distribution to guarantee invariance. xNES is configured to compute the natural gradient without the need for an explicit Fisher information matrix. Another evolution strategy is called the covariance matrix adaptive evolution strategy (CMA-ES). CMA-ES uses the maximum likelihood principle, which increases the probability of successful candidate solutions and search steps. CMA-ES also records two different paths of the time evolution of the mean of the distribution of the strategy, which are otherwise referred to as search or evolution paths. Compared to other optimization methods, CMA-ES requires fewer assumptions about the properties of the black box function. CMA-ES does not require derivatives or function values ​​themselves, but rather ranks candidate solutions to find the best solution. Summary of the invention

[0004] According to one embodiment, a computational method for training a meta-learning evolution strategy black-box optimization classifier is disclosed. The method includes: receiving one or more training functions and one or more initial meta-learning parameters of the meta-learning evolution strategy black-box optimization classifier. The method further includes: sampling the sampled target function from the one or more training functions and the initial mean of the sampled function. The method also includes: for t =1,…, T In T The number of steps is calculated by running the meta-learning evolution strategy black-box optimization classifier on the sampled objective function using the initial mean T The method also includes: T The method further comprises: updating one or more initial meta-learning parameters of the meta-learning evolution strategy black-box optimization classifier in response to the characteristics of the loss function.

[0005] In another embodiment, a computational method for learning actuator control commands from a meta-learning evolution strategy black-box optimization classifier having one or more parameters includes: λ The generation of samples is sampled and transformed into transformed samples λ The calculation method further comprises: in response to the transformed sample λ One or more function evaluations of the generation of transformed samples λ The method further comprises: responding to the transformed samples λ The method further comprises: sending an input signal obtained from a sensor to the updated meta-learning evolution strategy black box optimization classifier to obtain an output signal configured to characterize the classification of the input signal. The method further comprises: transmitting an actuator control command to an actuator of a computer-controlled machine in response to the output signal.

[0006] In yet another embodiment, a computational method for training and using a meta-learned evolution strategy black-box optimization classifier is disclosed. The computational method includes: receiving one or more learning parameters, the one or more learning parameters being trained on a set of objective functions similar to the learned evolution strategy black-box optimization classifier (e.g., sharing functional characteristics with the learned evolution strategy black-box optimization classifier). The computational method also includes: responding to a sample λ The generation of a meta-learning evolution strategy black-box optimization classifier is updated using one or more learned parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 A schematic diagram of the interaction between a computer-controlled machine and a control system is depicted according to one embodiment.

[0008] Figure 2 Depicted Figure 1 Schematic diagram of a control system configured to control a vehicle, which may be a partially autonomous vehicle or a partially autonomous robot.

[0009] Figure 3 Depicted Figure 1 Schematic diagram of a control system configured to control a manufacturing machine (such as a punch tool, a tool, or a gun drill) of a manufacturing system (such as a portion of a production line).

[0010] Figure 4 Depicted Figure 1 Schematic diagram of a control system configured to control a power tool (such as a drill or driver) having an at least partially autonomous mode.

[0011] Figure 5 Depicted Figure 1 Schematic diagram of a control system configured to control an automated personal assistant.

[0012] Figure 6 Depicted Figure 1 Schematic diagram of a control system configured to control a monitoring system (such as a controlled access system or a surveillance system).

[0013] Figure 7 Depicted Figure 1 Schematic diagram of a control system configured to control an imaging system (such as an MRI device, an x-ray imaging device, or an ultrasound device).

[0014] Figure 8 Depicted is a schematic diagram of a training system for training a classifier in accordance with one or more embodiments.

[0015] Fig. 9 Depicted is a flow diagram of a computational method for training a classifier (eg, a black-box algorithm) in accordance with one or more embodiments.

[0016] Fig.10 Depicted is a flow diagram of a computational method for using a classifier (eg, a black-box algorithm) by using a meta-learning evolution strategy, according to one embodiment. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure are described herein. However, it is to be understood that the disclosed embodiments are merely examples, and other embodiments may take various alternative forms. The drawings are not necessarily drawn to scale; some features may be exaggerated or minimized to show the details of a particular component. Therefore, the specific structural and functional details disclosed herein should not be interpreted as restrictive, but merely as a representative basis for teaching those skilled in the art to adopt the embodiments in various ways. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any one of the drawings may be combined with the features illustrated in one or more other drawings to produce embodiments that are not explicitly illustrated or described. The combination of the illustrated features provides a representative embodiment of a typical application. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of the present disclosure may be desired.

[0018] Figure 1A schematic diagram of the interaction between a computer-controlled machine 10 and a control system 12 is depicted. The computer-controlled machine 10 includes an actuator 14 and a sensor 16. The actuator 14 may include one or more actuators, and the sensor 16 may include one or more sensors. The sensor 16 is configured to sense a condition of the computer-controlled machine 10. The sensor 16 may be configured to encode the sensed condition into a sensor signal 18 and transmit the sensor signal 18 to the control system 12. Non-limiting examples of the sensor 16 include video, radar, lidar, ultrasound, and motion sensors. In one embodiment, the sensor 16 is an optical sensor that is configured to sense an optical image of the environment near the computer-controlled machine 10.

[0019] The control system 12 is configured to receive sensor signals 18 from the computer controlled machine 10. As explained below, the control system 12 may be further configured to learn actuator control commands 20 dependent on the sensor signals and transmit the actuator control commands 20 to the actuators 14 of the computer controlled machine 10.

[0020] like Figure 1 As shown in , the control system 12 includes a receiving unit 22. The receiving unit 22 can be configured to receive the sensor signal 18 from the sensor 30 and transform the sensor signal 18 into an input signal x. In an alternative embodiment, the sensor signal 18 is directly received as the input signal x without the receiving unit 22. Each input signal x can be a portion of each sensor signal 18. The receiving unit 22 can be configured to process each sensor signal 18 to generate each input signal x. The input signal x can include data corresponding to an image recorded by the sensor 16.

[0021] The control system 12 includes a classifier 24. The classifier 24 can be configured to learn the actuator control command 20 from the input signal x using a machine learning (ML) algorithm, such as a neural network or a recurrent neural network (RNN). The control system 12 can be configured to train the ML algorithm. As disclosed in one or more embodiments herein, the ML algorithm can be a meta-learning evolution strategy for black-box optimization.

[0022] The classifier 24 is configured to be parameterized by one or more parameters. These parameters may be stored in and provided by the non-volatile storage device 26. The classifier 24 is configured to determine the output signal y from the input signal x. Each output signal y includes information that assigns one or more labels to each input signal x. The classifier 24 may transmit the output signal y to a conversion unit 28. The conversion unit 28 is configured to convert the output signal y into an actuator control command 20. The control system 12 is configured to transmit the actuator control command 20 to the actuator 14, and the actuator 14 is configured to actuate the computer-controlled machine 10 in response to the actuator control command 20. In another embodiment, the actuator 14 is configured to actuate the computer-controlled machine 10 directly based on the output signal y.

[0023] When the actuator 14 receives the actuator control command 20, the actuator 14 is configured to perform an action corresponding to the associated actuator control command 20. The actuator 14 may include control logic configured to transform the actuator control command 20 into a second actuator control command for controlling the actuator 14. In one or more embodiments, the actuator control command 20 may be used to control a display instead of or in addition to the actuator.

[0024] In another embodiment, the control system 12 includes the sensor 16 instead of or in addition to the computer-controlled machine 10 including the sensor 16. The control system 12 may also include the actuator 14 instead of or in addition to the computer-controlled machine 10 including the actuator 14.

[0025] like Figure 1 As shown in , the control system 12 also includes a processor 30 and a memory 32. The processor 30 may include one or more processors. The memory 32 may include one or more memory devices. The classifier 24 (e.g., ML algorithm) of one or more embodiments may be implemented by the control system 12, which includes a non-volatile storage device 26, a processor 30, and a memory 32.

[0026] The non-volatile storage 26 may include one or more persistent data storage devices, such as a hard drive, an optical drive, a tape drive, a non-volatile solid-state device, a cloud storage device, or any other device capable of persistently storing information. The processor 30 may include one or more devices selected from a high-performance computing (HPC) system, including a high-performance core, a microprocessor, a microcontroller, a digital signal processor, a microcomputer, a central processing unit, a field programmable gate array, a programmable logic device, a state machine, a logic circuit, an analog circuit, a digital circuit, or any other device that manipulates signals (analog or digital signals) based on computer-executable instructions residing in the memory 32. The memory 32 may include a single memory device or multiple memory devices, including but not limited to random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information.

[0027] The processor 30 may be configured to read into the memory 32 and execute computer executable instructions that reside in the non-volatile storage 26 and embody one or more ML algorithms and / or methods of one or more embodiments. The non-volatile storage 26 may include one or more operating systems and applications. The non-volatile storage 26 may store content compiled and / or interpreted from a computer program created using various programming languages ​​and / or techniques, including but not limited to, and either alone or in combination, Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL.

[0028] When executed by the processor 30, the computer executable instructions of the non-volatile storage device 26 may cause the control system 12 to implement one or more ML algorithms and / or methods disclosed herein. The non-volatile storage device 26 may also include ML data (including data parameters) that support the functions, features, and processes of one or more embodiments described herein.

[0029] The program code embodying the algorithms and / or methods described herein can be distributed individually or collectively as a program product in a variety of different forms. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon, which is used to cause a processor to perform aspects of one or more embodiments. Inherently non-transitory computer-readable storage media may include volatile and non-volatile and removable and non-removable tangible media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium may further include: RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technology, portable compact disk read-only memory (CD-ROM) or other optical storage device, cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or any other medium that can be used to store desired information and can be read by a computer. The computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing device, or another device, or downloaded to an external computer or external storage device via a network.

[0030] The computer-readable program instructions stored in the computer-readable medium can be used to instruct a computer, other types of programmable data processing devices, or other devices to operate in a particular manner so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions, actions, and / or operations specified in the flowchart or diagram. In certain alternative embodiments, the functions, actions, and / or operations specified in the flowchart and diagram can be reordered, processed serially, and / or processed simultaneously in accordance with one or more embodiments. In addition, any of the flowchart and / or diagram may include more or fewer nodes or blocks than those illustrated in accordance with one or more embodiments.

[0031] These processes, methods or algorithms may be embodied in whole or in part using suitable hardware components, such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), state machines, controllers or other hardware components or devices, or a combination of hardware, software and firmware components.

[0032] Figure 2 1 depicts a schematic diagram of a control system 12 configured to control a vehicle 50, which may be an at least partially autonomous vehicle or an at least partially autonomous robot. Figure 2As shown in , the vehicle 50 includes an actuator 14 and a sensor 16. The sensor 16 may include one or more video sensors, radar sensors, ultrasonic sensors, lidar sensors, and / or position sensors (e.g., GPS). One or more of the one or more specific sensors may be integrated into the vehicle 50. Instead of or in addition to the one or more specific sensors identified above, the sensor 16 may include a software module configured to determine the state of the actuator 14 when executed. A non-limiting example of a software module includes a weather information software module that is configured to determine current or future weather conditions near the vehicle 50 or other location.

[0033] The classifier 24 of the control system 12 of the vehicle 50 may be configured to detect an object in the vicinity of the vehicle 50 depending on the input signal x. In such an embodiment, the output signal y may include information characterizing the vicinity of the object of the vehicle 50. The actuator control command 20 may be determined based on this information. The actuator control command 20 may be used to avoid a collision with the detected object.

[0034] In embodiments where the vehicle 50 is at least partially autonomous, the actuators 14 may be embodied in the brakes, propulsion system, engine, transmission, or steering of the vehicle 50. The actuator control commands 20 may be determined such that the actuators 14 are controlled such that the vehicle 50 avoids a collision with a detected object. The detected objects may also be classified according to what the classifier 24 considers them to be most likely to be, such as a pedestrian or a tree. The actuator control commands 20 may be determined depending on the classification.

[0035] In other embodiments in which the vehicle 50 is an at least partially autonomous robot, the vehicle 50 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and walking. The mobile robot may be an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. In such embodiments, the actuator control commands 20 may be determined such that a propulsion unit, a steering unit, and / or a braking unit of the mobile robot may be controlled such that the mobile robot may avoid a collision with the identified object.

[0036] In another embodiment, the vehicle 50 is an at least partially autonomous robot in the form of a gardening robot. In such an embodiment, the vehicle 50 may use an optical sensor as the sensor 16 to determine the state of plants in the environment near the vehicle 50. The actuator 14 may be a nozzle configured to spray a chemical. Depending on the identified species and / or the identified state of the plant, the actuator control command 20 may be determined to cause the actuator 14 to spray the plant with an appropriate amount of an appropriate chemical.

[0037] The vehicle 50 may be an at least partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances include a washing machine, a stove, an oven, a microwave, or a dishwasher. In such a vehicle 50, the sensor 16 may be an optical sensor configured to detect the state of an object to be processed by the household appliance. For example, in the case where the household appliance is a washing machine, the sensor 16 may detect the state of the laundry inside the washing machine. The actuator control command 20 may be determined based on the detected state of the laundry.

[0038] Figure 3 A schematic diagram of a control system 12 is depicted, the control system 12 being configured to control a manufacturing machine 100 (such as a punch tool, a tool, or a gun drill) of a manufacturing system 102 (such as a portion of a production line). The control system 12 may be configured to control an actuator 14 configured to control the manufacturing machine 100.

[0039] The sensor 16 of the manufacturing machine 100 may be an optical sensor configured to capture one or more attributes of the manufactured product 104. The classifier 24 may be configured to determine a state of the manufactured product 104 from the one or more captured attributes. The actuator 14 may be configured to control the manufacturing machine 100 depending on the determined state of the manufactured product 104 for a subsequent manufacturing step of the manufactured product 104. The actuator 14 may be configured to control a function of the manufacturing machine 100 on a subsequent manufactured product 106 of the manufacturing machine 100 depending on the determined state of the manufactured product 104.

[0040] Figure 4 A schematic diagram of a control system 12 configured to control a power tool 150 (such as a drill or driver) having at least a partially autonomous mode is depicted. The control system 12 may be configured to control an actuator 14 configured to control the power tool 150 .

[0041] The sensor 16 of the power tool 150 may be an optical sensor configured to capture one or more properties of the working surface 152 and / or the fastener 154 driven into the working surface 152. The classifier 24 may be configured to determine the state of the working surface 152 and / or the state of the fastener 154 relative to the working surface 152 from the one or more captured properties. The state may be that the fastener 154 is flush with the working surface 152. Alternatively, the state may be the hardness of the working surface 154. The actuator 14 may be configured to control the power tool 150 so that the driving function of the power tool 150 is adjusted depending on the determined state of the fastener 154 relative to the working surface 152, or one or more captured properties of the working surface 154. For example, if the state of the fastener 154 is flush with the working surface 152, the actuator 14 may suspend the driving function. As another non-limiting example, the actuator 14 may apply additional or less torque depending on the hardness of the working surface 152.

[0042] Figure 5 A schematic diagram of a control system 12 is depicted, which is configured to control an automated personal assistant 200. The control system 12 may be configured to control an actuator 14, which is configured to control the automated personal assistant 200. The automated personal assistant 200 may be configured to control a household appliance, such as a washing machine, a stove, an oven, a microwave, or a dishwasher.

[0043] Sensor 16 may be an optical sensor and / or an audio sensor. The optical sensor may be configured to receive a video image of gesture 204 of user 202. The audio sensor may be configured to receive a voice command of user 202.

[0044] The control system 12 of the automated personal assistant 200 may be configured to determine an actuator control command 20, which is configured to the control system 12. The control system 12 may be configured to determine the actuator control command 20 based on the sensor signal 18 of the sensor 16. The automated personal assistant 200 is configured to transmit the sensor signal 18 to the control system 12. The classifier 24 of the control system 12 may be configured to: execute a gesture recognition algorithm to identify a gesture 204 made by the user 202, determine the actuator control command 20, and transmit the actuator control command 20 to the actuator 14. The classifier 24 may be configured to retrieve information from the non-volatile storage device in response to the gesture 204, and output the retrieved information in a form suitable for receipt by the user 202.

[0045] Figure 6A schematic diagram of a control system 12 is depicted, the control system 12 being configured to control a monitoring system 250. The monitoring system 250 may be configured to physically control entry through a door 252. The sensor 16 may be configured to detect a scene relevant to a decision as to whether entry is granted. The sensor 16 may be an optical sensor configured to generate and transmit image and / or video data. The control system 12 may use such data to detect the face of a person.

[0046] The classifier 24 of the control system 12 of the monitoring system 250 can be configured to interpret the image and / or video data by matching the identities of known persons stored in the non-volatile storage device 26, thereby determining the identity of the person. The classifier 12 can be configured to generate an actuator control command 20 in response to the interpretation of the image and / or video data. The control system 12 is configured to transmit the actuator control command 20 to the actuator 12. In this embodiment, the actuator 12 can be configured to lock or unlock the door 252 in response to the actuator control command 20. In other embodiments, non-physical logical access control is also possible.

[0047] Monitoring system 250 may also be a surveillance system. In such an embodiment, sensor 16 may be an optical sensor configured to detect a scene under surveillance, and control system 12 is configured to control display 254. Classifier 24 is configured to determine a classification of a scene, for example, whether a scene detected by sensor 16 is suspicious. Control system 12 is configured to transmit actuator control commands 20 to display 254 in response to the classification. Display 254 may be configured to adjust the displayed content in response to actuator control commands 20. For example, display 254 may highlight an object that is considered suspicious by classifier 24.

[0048] Figure 7 A schematic diagram of a control system 12 is depicted, the control system 12 being configured to control an imaging system 300 (e.g., an MRI device, an x-ray imaging device, or an ultrasound device). The sensor 16 may be, for example, an imaging sensor. The classifier 24 may be configured to determine a classification of all or part of a sensed image. The classifier 24 may be configured to determine or select an actuator control command 20 in response to the classification. For example, the classifier 24 may interpret an area of ​​the sensed image as being potentially abnormal. In this case, the actuator control command 20 may be determined or selected to cause the display 302 to display the imaging and highlight the potentially abnormal area.

[0049] Evolution strategies have been used as black-box function optimizers. Evolution strategies use iterative black-box optimization methods. Evolution strategies try to find the optimal solution in the feasible space. X Find the unknown (e.g., black-box) loss function in fThe global minimizer of . Equation (1) shows this calculation in algebraic terms.

[0050] (1)

[0051] Examples of existing evolution strategies include exponential natural evolution strategies (xNES) and covariance matrix adaptive evolution strategies (CMA-ES). Evolution strategies are configured to optimize a wide variety of objective functions without adjusting hyperparameters. Existing evolution strategies use predefined heuristics and manually tuned hyperparameters to perform updates to their parameters. Standard evolution strategies do not use any knowledge about the structure of the function being optimized. Therefore, using this knowledge may not improve optimization performance. Therefore, a computational optimization method is needed that incorporates prior knowledge of the optimization problem by first training on a set of similar objective functions.

[0052] In one or more embodiments, computational methods and computer systems are presented that exploit knowledge about function structure to improve optimization performance. In one or more embodiments, the computational methods and computer systems use ML to metalearn black-box optimization algorithms (e.g., deep learning). Recursive neural networks (RNNs) can be trained to perform black-box optimization. In one embodiment, the computational methods and computer systems use deep meta-learning to improve the performance of optimization algorithms on specific classes of functions. Such a class can broadly refer to properties or characteristics of functions in the class, such as the amount of noise in the evaluation of a function or the degree of a set of polynomials, or it can be very specific, such as functions generated from numerical simulations of specific experiments. The computational methods and computer systems can include meta-learning parameters θ , the meta-learning parameters θ is configured to define how the evolution strategy updates its parameters. For example, static hyperparameters in the evolution strategy algorithm (e.g., learning rate, step size, coefficients controlling momentum, and / or coefficients controlling the weight given to each term in the update of its parameters) can be converted into learnable parameters. As another example, m and σ The heuristic update of can be performed using ML (e.g., neural networks), in which case, θ Represents the weights and / or biases of a neural network.

[0053] In one or more embodiments, because the operations performed by the evolution strategy are differentiable, the evolution strategy algorithm of the computational method and computer system is implemented into an automatic differentiation framework, where the algorithm computes a value and automatically constructs a process for computing the derivative of the value. Then, the meta-learning parameters can be trained in a supervised learning manner using gradient descent. θ .

[0054] In an evolution strategy of one or more embodiments, the objective function f : R n → R is stored in non-volatile storage 354. Memory 358 includes instructions that, when executed by processor 356, execute an evolution strategy configured to optimize by iteratively updating one or more of the following parameters: f : Multivariate Gaussian N ( m , σC )of m ∈ R n , σ ∈ R , C ∈ R n×n . m is the mean. σ is the step length. C is the covariance matrix. The iteration step can be defined as t =1,…, T During each step t, the processor 356 may iteratively perform the following steps. The first step may have two iterative sub-parts. For the sub-parts defined as i =1,…, λ The first sub-part is to iterate over multiple samples of λ The number of samples to be sampled z i ∼ N (0, I ),Should z i ∼ N (0, I ) is a multivariate normal distribution with zero mean and unit covariance matrix, and is otherwise referred to as a generation. N (0, I ) has independent (0,1) normally distributed components. The second subsection is to use the following equation to convert λ Number of samples z i ∼ N (0, I ) scale and shift x i sample:

[0055] (2)

[0056] In the second iteration, according to the sample x 1,…, x λ The function evaluation sorts them so that the sorted samples satisfy f ( x 1) ≤ f ( x 2) ≤ ⋯ ≤ f ( x λ ). In the third iteration step, according to x 1,…, x λ and one or more parameters θ To update the evolution parameters m ∈ R n , σ ∈ R , C ∈ R n×n ,parameter θ is determined using the training method of one or more embodiments. As part of the update step, the mean of the Gaussian can be replaced by a weighted combination of the samples in the current generation m To update the mean m , where more weight is placed on samples with better (i.e., smaller) function values.

[0057] For functions f The optimization algorithm of can be trained on a training set of functions, as described herein with respect to one or more embodiments. f The optimization algorithm (e.g., three iteration steps) is continued until a termination criterion is reached. For example, if the best function value does not change after a certain number of iterations (such as 100 or 1000 iterations, or any number of iterations in between), the process can end.

[0058] Figure 8 A schematic diagram of a training system 350 for training a classifier 24 is depicted. In one or more embodiments, one or more embodiments of a computing method and a computer system are presented for training a classifier that can be Figure 2-7 . For example, the classifier 24 may be implemented using a robot to learn to perform a task or to learn optimal parameters for a laser or other manufacturing tool.

[0059] The classifier 24 may be a meta-evolution strategy algorithm. A set of training functions for the evolution strategy ( F ) and initial meta-learning parameters θThe training system 350 is configured to perform a training process to find the optimal meta-learning parameters. θ The memory 358 includes instructions that, when executed by the processor 356, perform a training process. The training system 350 is configured to perform the training process in a series of steps. The first step may be to perform a training on the function f ∈ F and the initial mean m {(0)} ∼ U [-1,1] n The second step can be: t =1,…, T In T number of steps, by having an initial mean m (0) The objective function f One of the optimization algorithms of one or more embodiments is executed to calculate the mean m (1) ,…, m (T) The third step can be to find the mean of the values f ( m 0 ),…, f ( m T ) to calculate the loss L to obtain the scalar loss. L The loss function can be used L ( f ( m 0 ),…, f ( m T )). The fourth step may include using the gradient of the loss function ∇ θ L To update the parameters θ . Backpropagation can be performed to train the parameters with respect to this loss θ .

[0060] Fig. 9 A flowchart 400 is depicted of a computational method (eg, a black-box algorithm) for training a classifier 24 using a meta-learning evolution strategy according to one embodiment. The computational method may be performed using a training system 350. The computational method for training a classifier 24 may be demonstrated by a training process.

[0061] In step 402, input is received for the training method. In one embodiment, the input includes a set of training functions of the evolution strategy F and initial meta-learning parameters θ In step 404, one or more of the set of training functions F are used.

[0062] In step 404, the function f ∈ F and the initial mean m {(0)} ∼ U [-1,1] n Sampling is performed. In step 406 as described below, function sampling and initial mean sampling are used.

[0063] In step 406, by m {(0)} The objective function f Running a meta-learning evolution strategy algorithm on Fig.10 The algorithm identified in m (1) ,…, m (T) .like Fig. 9 As shown in T Step 406 is repeated by using the number of T The mean calculated after the step .

[0064] In step 408, a loss function is calculated from the mean values ​​calculated in step 406. Step 408 may be a process of calculating the mean value of the loss function. f ( m 0 ),…, f ( m T ) to calculate the loss L to obtain the scalar loss. L The loss function can be used L ( f ( m 0 ),…, f ( m T )) in the form of. Other parameters may be included in the loss function calculation.

[0065] In step 410, the gradient of the loss function ∇ θ L To update the meta-learning parameters of the evolution strategy θ Although step 410 uses the gradient of the loss function, in other embodiments, this step may be performed by gradient descent or other deep learning optimizers such as Adam.

[0066] Steps 404, 406, 408, and 410 are repeated until the meta-learning optimization algorithm converges, as described in step 412. The last updated meta-learning parameters of step 414 are used in the meta-learning optimization algorithm as described below. θ value.

[0067] like Fig. 9 As shown in , we train one function at a time θ In one or more other embodiments, the training process may be run for minibatches of functions, where the minibatch size is up to or exceeds 128 functions. In other embodiments, the number of functions in a minibatch may be between 20 and 30. In still other embodiments, the number of functions in a minibatch may exceed 128. In alternative embodiments, samples from a distribution (e.g., samples from a Gaussian process) may be utilized instead of samples from a prescribed set of training functions.

[0068] Fig.10 A flowchart 450 is depicted of a computational method for using a classifier 24 (eg, a black-box algorithm) with a meta-learning evolution strategy, according to one embodiment. The computational method may be performed using the control system 12 .

[0069] In step 452, for i =1,…, λ ,right z i ∼ N (0, I ) is sampled. In step 454, equation (2) is used to pass z i Zoom and shift to z i Transform. Fig.10 As shown in i =1,…, λ , repeat steps 452 and 454. Step 456 uses the transformed sample x i .

[0070] In step 456, according to x i The function evaluation of the sample is x i The samples are sorted so that the sorted samples satisfy f ( x 1) ≤ f ( x 2) ≤ ⋯ ≤ f ( x λ ). Sorted samples xi Used by step 458.

[0071] In step 458, the sorted samples are used x i ,as well as Fig. 9 The last updated meta-learning parameters of step 414 are θ Value to update the evolution strategy parameters of black box optimization m , σ and C .

[0072] Although Fig.10 As shown in , the meta-evolutionary strategy algorithm runs a fixed number of T steps, but it may also run until a stopping criterion is reached. Fig.10 It also presents the fixed number of each generation λ samples, but this number can be varied, or samples from previous generations can be utilized.

[0073] Although exemplary embodiments are described above, these embodiments are not intended to describe all possible forms covered by the claims. The words used in the specification are descriptive words, not restrictive words, and it is understood that various changes can be made without departing from the spirit and scope of the present disclosure. As previously described, the features of various embodiments can be combined to form other embodiments of the present invention that may not be explicitly described or illustrated. Although various embodiments may have been described as providing advantages or being preferred over other embodiments or prior art implementations with respect to one or more desired characteristics, it is recognized by those of ordinary skill in the art that one or more features or characteristics can be compromised to achieve the desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, applicability, weight, manufacturability, ease of assembly, etc. Thus, to the extent that any embodiment is described as less desirable than other embodiments or prior art implementations with respect to one or more characteristics, these embodiments are not outside the scope of the present disclosure and may be desirable for a particular application.

Claims

1. A computational method for training a meta-learning evolution strategy black-box optimization classifier, the method comprising: Receiving one or more training functions and one or more initial meta-learning parameters of a meta-learning evolution strategy black-box optimization classifier; Sampling the sampled objective function from the one or more training functions and the initial mean of the sampled objective function to obtain a generation of sample λ; For T number of steps in t=1, ..., T, computing a set of T number of means by running a meta-learning evolution strategy black-box optimization classifier on the sampled objective function using the initial means; Compute a loss function from the set of means of said T number; In response to the characteristics of the loss function, updating one or more initial meta-learning parameters of the meta-learning evolution strategy black-box optimization classifier to obtain an updated meta-learning evolution strategy black-box optimization classifier, wherein, as part of the updating, the mean is updated by replacing the mean with a weighted combination of samples in the current generation, wherein greater weights are placed on samples with better function values; Sending an input signal obtained from a sensor to an updated meta-learning evolution strategy black-box optimization classifier to obtain an output signal configured to characterize a classification of the input signal; and In response to the output signal, an actuator control command is transmitted to an actuator of the computer controlled machine.

2. The calculation method according to claim 1, wherein: The characteristic of the loss function is the gradient of the loss function.

3. The calculation method according to claim 1, wherein: The property of the loss function is the gradient descent of the loss function.

4. The calculation method according to claim 1, wherein: The update step is performed by a deep learning optimizer.

5. The calculation method according to claim 1, wherein: The sampling step, the first calculation step, the second calculation step and the updating step are performed interactively in a loop until a stop condition is met.

6. The calculation method according to claim 5, wherein: The stopping condition is the convergence of the meta-learning evolution strategy black-box optimization classifier.

7. A computational method for learning actuator control commands from a meta-learning evolutionary strategy black-box optimization classifier, the method comprising: Sampling the sampled objective function and the initial mean of the sampled objective function from one or more training functions of the meta-learning evolution strategy black-box optimization classifier to obtain a generation of sampled λ, and transforming it into a transformed generation of sampled λ; In response to evaluating the one or more functions of the generations of transformed samples λ, sorting the generations of transformed samples λ to obtain sorted generations of transformed samples λ; updating one or more parameters of the meta-learned evolution strategy black-box optimization classifier in response to the ordered generations of transformed samples λ and one or more learning parameters trained on a set of objective functions that share functional properties with the learned evolution strategy black-box optimization classifier to obtain an updated meta-learned evolution strategy black-box optimization classifier, wherein, as part of the updating, the mean of the objective function is updated by replacing the mean with a weighted combination of samples in the current generation, wherein greater weights are placed on samples with better function values; sending an input signal obtained from a sensor to an updated meta-learning evolution strategy black-box optimization classifier to obtain an output signal configured to characterize a classification of the input signal; and In response to the output signal, an actuator control command is transmitted to an actuator of the computer controlled machine.

8. The calculation method according to claim 7, wherein: The multiple parameters include m, σ and C, where m is the mean, σ is the step size and C is the covariance matrix.

9. The calculation method according to claim 8, wherein: Use a neural network to update m and σ.

10. The calculation method according to claim 9, wherein: The multiple learned parameters represent one or both of the weights and biases of the neural network.

11. The calculation method according to claim 7, wherein: The generation of transformed samples λ is generated from a multivariate Gaussian N(m,σC), where m is the mean, σ is the step size and C is the covariance matrix.

12. The calculation method according to claim 7, wherein: The one or more function evaluations are one or more objective function evaluations.

13. The calculation method according to claim 7, wherein: The sorted generations of transformed samples λ satisfy f(x1)≤f(x2)≤…≤f(x λ ).

14. A computational method for training and using a meta-learning evolutionary strategy black-box optimization classifier, the method comprising: receiving one or more learning parameters trained on a set of objective functions that share functional properties with a learned evolution strategy black-box optimization classifier; Sampling the sampled objective functions and the initial means of the sampled objective functions from one or more training functions of the meta-learning evolution strategy black-box optimization classifier to obtain a generation of sample λ; In response to a generation of samples λ, updating a meta-learning evolution strategy black-box optimization classifier using the one or more learning parameters, wherein, as part of the updating, the mean is updated by replacing the mean of the objective function with a weighted combination of samples in the current generation, wherein greater weights are placed on samples with better function values; Sending an input signal obtained from a sensor to an updated meta-learning evolution strategy black-box optimization classifier to obtain an output signal configured to characterize a classification of the input signal; and In response to the output signal, an actuator control command is transmitted to an actuator of the computer controlled machine.

15. The calculation method according to claim 14, wherein: The multiple learning parameters include a learning rate.

16. The calculation method according to claim 14, wherein: The one or more learning parameters comprise a neural network.

17. The calculation method according to claim 14, wherein: The generation of samples λ is generated from a multivariate Gaussian N(m,σC), where m is the mean, σ is the step size and C is the covariance matrix.

18. The calculation method according to claim 14, wherein: The updating step includes updating the meta-learning evolution strategy black-box optimization classifier using one or more learning parameters and one or more static parameters in response to the generation of samples λ.

Citation Information

Patent Citations

  • Generating robust automatic learning systems and testing trained automatic learning systems

    CN110554602A

  • Black-box optimization using neural networks

    CN110832509A