Fine-Grained Random Neural Architecture Search

Optimize the neural network architecture through the random neural architecture search system to generate high-performance and small-size neural networks, solving the problem of waste of model size and computing resources in the existing technology, and is suitable for hardware platforms with limited computing resources.

CN115066689BActive Publication Date: 2025-07-08GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180012814.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-07
Filing Date
2021-02-08
Publication Date
2025-07-08
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently optimize neural network architectures to meet specific task requirements, resulting in waste of model size and computing resources.

Method used

Through the random neural architecture search system, the sparsity in candidate neural network architecture is encouraged by using trainable random masks, and the architecture and parameters of the neural network are optimized in combination with the training data set to generate high-performance, small-size neural networks.

Benefits of technology

The generated neural networks exhibit performance competing with state-of-the-art models on specific tasks, while reducing model size and computing resource requirements, suitable for deployment on hardware platforms with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115066689B_ABST
    Figure CN115066689B_ABST
Patent Text Reader

Abstract

Methods, systems, and devices for determining a neural network architecture, including a computer program encoded on a computer storage medium. One of the methods includes receiving training data; receiving architecture data; allocating to each of a plurality of network operators an exploitation variable indicative of the likelihood that the network operator is utilized in the neural network; generating an optimized neural network for performing a neural network task, including repeatedly performing the following operations: sampling a selected set of network operators; and training a neural network having an architecture defined by the selected set of network operators, wherein the training includes: computing an objective function that evaluates (i) a measure of the computational cost of the neural network and (ii) a measure of the performance of the neural network for the neural network task associated with the training data; and adjusting the respective current values of the exploitation variables and the respective current values of the neural network parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority to U.S. Provisional Application No. 62 / 971,866, filed on February 7, 2020. The disclosure of the prior application is hereby incorporated by reference in its entirety as part of the disclosure of this application. Technical Field

[0003] This specification relates to determining an architecture for a neural network. Background Art

[0004] A neural network is a machine learning model that employs one or more layers of non - linear units to predict an output for received input. In addition to an output layer, some neural networks also include one or more hidden layers. The output of each hidden layer is used as input to the next layer in the network (i.e., the next hidden layer or the output layer). Each layer of the network generates an output from the received input according to the current values of a corresponding set of parameters. Summary of the Invention

[0005] This specification describes a neural network architecture optimization system that is implemented as a computer program on one or more computers in one or more locations, and the system determines an optimal network architecture for a neural network configured to perform a specific machine learning task. Depending on the task, the neural network can be configured to receive any type of digital data input and generate any type of score, classification, or regression output based on the input.

[0006] Generally, one innovative aspect of the subject matter described in this specification can be embodied in a method that includes: receiving training data for training a neural network to perform a neural network task, the training data including a plurality of training examples and corresponding target outputs for each training example; receiving architecture data that defines a plurality of network operators; assigning to each of the plurality of network operators an exploitation variable that indicates the likelihood of the network operator being utilized in the neural network; generating an optimized neural network for performing the neural network task, including repeatedly performing the following operations: sampling a selected set of network operators from the plurality of network operators and according to the respective current values of the exploitation variables; and training a neural network having an architecture defined by the selected set of network operators on the training data to perform the neural network task, where the training includes: computing an objective function that evaluates (i) a measure of the computational cost of the neural network and (ii) a measure of the performance of the neural network for the neural network task associated with the training data; and adjusting the respective current values of the exploitation variables and the respective current values of the neural network parameters based on the determined gradient of the objective function.

[0007] The architecture data can be initialized from one or more predetermined neural network architectures. The method may further include: removing redundant network operators from a plurality of network operators. The plurality of operators may include neural network layers. The neural network layer may include at least one of a convolutional layer, a fully connected layer, a normalization layer, or an activation layer. The plurality of operators may further include filters in a convolutional layer, or neurons in a fully connected layer. A measure of the computational cost of a neural network may include at least one of size, floating point operations per second (FLOPS), or latency. Generating an optimized neural network for performing a neural network task may further include: inserting a zero masking layer after each operator that is not one of a selected set of network operators. Generating an optimized neural network for performing a neural network task may further include: using a concat aggregator to combine corresponding outputs of a selected set of network operations. Each exploitation variable may be defined by one or more distribution parameters. The method may further include calculating a determined gradient of an objective function with respect to one or more distribution parameters. Adjusting the respective values of the exploitation variables may include: backpropagating the determined gradient of the objective function into one or more distribution parameters through the exploitation variables. One or more distribution parameters may define a binary concrete distribution.

[0008] Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method. A system of one or more computers can be configured to perform particular operations or actions by software, firmware, hardware, or any combination thereof installed on the system, which in operation can cause the system to perform these actions. One or more computer programs can be configured to perform particular operations or actions by including instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.

[0009] Particular embodiments capable of implementing the subject matter described in this specification achieve one or more of the following advantages.

[0010] The techniques allow a neural network architecture optimization system to effectively and automatically determine a neural network architecture from a search space, which will result in a small-sized (i.e., parameter-efficient) but high-performance neural network for a particular task. Specifically, during the search process, the system utilizes trainable random masks to encourage sparsity in candidate neural network architectures, thereby reducing the runtime latency, memory footprint, or both of the neural network with the resulting architecture. Specialized hardware can be optimized to efficiently and with minimal runtime latency perform sparse operations, e.g., by not computing multiplications involving 0. Sparse matrices can be efficiently stored in memory, i.e., by not explicitly storing values that are 0, reducing the memory footprint of the resulting architecture.

[0011] The technology also allows the system to jointly determine the training parameter values of a neural network with a selected neural network architecture by training the neural network on a training data set associated with a specific task. More importantly, the specific task can be any neural network task, and the search space can be initialized from any existing neural network architecture. Thus, the system can automatically generate the resulting trained neural networks, which can compete with or outperform state-of-the-art models to perform a wide range of tasks, while having a relatively small model size and thus being suitable for deployment on a hardware platform with limited computational resources, such as mobile devices and embedded systems.

[0012] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 An example neural architecture search system is shown.

[0014] Figure 2 is a flowchart of an example process for searching for an architecture of a neural network.

[0015] Figure 3 is a flowchart of an example process for training a neural network having an architecture defined by a selected set of network operators.

[0016] Figures 4A to 4B An example illustration of a search space is shown.

[0017] Figure 5 An example illustration of inserting a mask corresponding to the use of a variable into a network operator is shown.

[0018] Like reference numerals and names in the different figures represent like elements. DETAILED DESCRIPTION

[0019] This specification describes a system implemented as a computer program on one or more computers in one or more locations that determines an architecture of a task neural network configured to perform a specific neural network task.

[0020] A neural network can be trained to perform any type of machine learning task, i.e., can be configured to receive any type of digital data input and generate any type of score, classification, or regression output based on the input.

[0021] In some cases, the neural network is a neural network configured to perform an image processing task (i.e., receive an input image and process the input image to generate a network output of the input image). For example, the task can be image classification, and the output generated by the neural network for a given image can be a score for each of a set of object classes, where each score represents an estimated likelihood that the image contains an object belonging to that class. As another example, the task can be image embedding generation, and the output generated by the neural network can be a digital embedding of the input image. As yet another example, the task can be object detection, and the output generated by the neural network can identify the locations in the input image where objects of a particular type are depicted. As yet another example, the task can be image segmentation, and the output generated by the neural network can assign each pixel of the input image to a class from a set of classes.

[0022] As another example, if the input to the neural network is an Internet resource (e.g., a web page), a document, or a part of a document or features extracted from an Internet resource, a document, or a part of a document, then the task can be to classify the resource or document, i.e., the output generated by the neural network for a given Internet resource, document, or part of a document can be a score for each of a set of topics, where each score represents an estimated likelihood that the Internet resource, document, or part of a document is about that topic.

[0023] As another example, if the input to the neural network is features of an impression context of a particular advertisement, then the output generated by the neural network can be a score representing an estimated likelihood that the particular advertisement will be clicked.

[0024] As another example, if the input to the neural network is features of a personalized recommendation for a user, such as features characterizing the context of the recommendation, such as features characterizing actions previously taken by the user, then the output generated by the neural network can be a score for each of a set of content items, where each score represents an estimated likelihood that the user will respond positively to the recommended content item.

[0025] As another example, if the input to the neural network is a sequence of text in one language, then the output generated by the neural network can be a score for each of a set of text segments in another language, where each score represents an estimated likelihood that the text segment in the other language is a correct translation of the input text into the other language.

[0026] As another example, the task can be an audio processing task. For example, if the input to the neural network is a sequence representing an uttered speech, the output generated by the neural network can be a score for each of a set of text segments, where each score represents an estimated likelihood that the text segment is the correct transcription of the speech. As another example, the task can be a keyword identification task, where, if the input to the neural network is a sequence representing an uttered speech, the output generated by the neural network can indicate whether a particular word or phrase ("hot word") was uttered in the speech. As another example, if the input to the neural network is a sequence representing an uttered speech, the output generated by the neural network can identify the natural language used to utter the speech.

[0027] As another example, the task can be a natural language processing or understanding task, such as, for example, an entailment task, a paraphrasing task, a text similarity task, a sentiment task, a sentence completion task, and a grammar task, etc., which act on some sequence of natural language text.

[0028] As another example, the task can be a text-to-speech task, where the input is text in a natural language or features of text in a natural language, and the network output is a spectrogram or other data defining the audio of the text being spoken in the natural language.

[0029] As another example, the task can be a health prediction task, where the input is electronic health record data of a patient, and the output is a prediction related to the future health of the patient, such as, for example, a predicted treatment that should be prescribed for the patient, the likelihood of an adverse health event occurring for the patient, or a predicted diagnosis for the patient.

[0030] As another example, the task can be an agent control task, where the input is an observation characterizing the state of an environment, and the output defines an action that the agent is to perform in response to the observation. The agent can be, for example, a real-world or simulated robot, a control system for an industrial facility, or a control system for controlling different types of agents.

[0031] Figure 1 An example neural architecture search system 100 is shown. The neural architecture search system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.

[0032] The neural architecture search system 100 is a system that obtains training data 102 for training a neural network to perform a machine learning task and architecture data 104 that defines a plurality of network operators, and uses the training data 102 and the architecture data 104 to determine an optimal neural network architecture for performing the machine learning task, and trains a neural network having the optimal neural network architecture to determine training values of the parameters of the neural network.

[0033] The training data 102 can include a plurality of training examples and corresponding target outputs for each training example. The target output of a given training example is the output that a trained neural network should generate by processing the given training example. In some embodiments, the system 100 divides (e.g., randomly splits) the received training data 102 into a training subset, a validation subset, and an optional test subset.

[0034] The CIFAR-10 dataset, which consists of 60,000 training examples paired with target output classifications selected from ten possible classifications, is an example of such training data. CIFAR-1000 is a related dataset where the classifications are one of 1000 possible classes. Another example of suitable training data is the ImageNet dataset, which consists of over 14 million images paired with target output classifications selected from over 20,000 possible classes. Some or all of these images are also paired with bounding box data that specify the boundaries of regions where an object belonging to one of the possible classes is present.

[0035] The architecture of a neural network generally defines the number of layers in the neural network, the operations performed by each layer, and the connectivity between the layers in the neural network, i.e., which layers in the neural network receive input from which other layers.

[0036] In some embodiments, the architecture data 104 includes data specifying a set of candidate neural network architecture components. Each candidate architecture component can be in the form of a neural network unit or a neural network block. The architecture of the neural network generated by the system 100 based on the architecture data 104, such as the training architecture 122 generated during the search process or the final architecture 150 generated at the end of the search process, can be in the form of a tower. A tower is a neural network that includes a sequence of neural network units, a sequence of neural network blocks, or both, where each unit (or block) after the first unit (or block) in the sequence receives input from one or more earlier units (or blocks) in the sequence, receives network input, or both. For example, each unit can consist of multiple blocks, and each block receives input from one or more previous units and one or more previous blocks within the same unit.

[0037] In some such embodiments, the architecture generated by system 100 from architecture data 104 can have multiple neural network units, where each unit can have multiple neural network blocks, and where each block can be a directed graph for arranging multiple neural network layers. Each neural network layer in a block can be configured to receive an input tensor from a previous layer and generate an output tensor for the input tensor, and the output tensor is fed as an input to the next neural network layer. The multiple neural network layers included in each block can be of different types, that is, they can perform different types of operations on input tensors of different sizes to generate output tensors of different sizes.

[0038] Although different architectures can include different numbers of units (or blocks), the sequence of units (or blocks) in any given candidate includes at least one unit (or block) and at most a fixed maximum number of units (or blocks). In addition to a sequence of one or more neural network units or blocks, each tower can optionally include one or more predetermined components, for example, one or more input layers before the first block in the sequence, one or more intermediate pooling layers, one or more output layers after the last block in the sequence, or a combination thereof.

[0039] In some embodiments, architecture data 104 can be initialized or otherwise derived from one or more baseline network architectures or search spaces. Examples of such baseline network architectures or search spaces are described in more detail below: Gabriel Bender et al. Understanding and simplifying one-shot architecture search; International Conference on Machine Learning, pages 549–558, 2018, Bichen Wu et al. Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 10734–10742, 2019, and Ariel Gordon et al. Morphnet: Fast & simple resource-constrained structure learning of deep networks; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1586–1595, 2018.

[0040] In any of the above embodiments, the architecture data 104 includes data defining a plurality of network operators, each network operator being configured to receive an operator input and generate a corresponding operator output based on processing the operator input according to the current value of a parameter associated with the network operator. For example, an operator can be a neural network layer. For example, each operator can be a convolutional layer, a fully connected layer, a normalization layer, a pooling layer, or an activation layer that applies a series of operations (e.g., transformations) to a layer input to generate a plurality of layer outputs. As another example, an operator can be a component of a neural network layer. For example, each operator can be a filter in a convolutional layer or a neuron in a fully connected layer. As yet another example, an operator can be a combination of two or more neural network layers, or a combination of two or more neural network layer components, as described above.

[0041] Figure 4A An example illustration of a search space is shown. The example search space includes a plurality of neural network units, which in turn include a plurality of neural network blocks, and the neural network blocks in turn include a plurality of operators. Each operator includes a plurality of neural network layers, which include one or more of the following: a 1×1 convolutional layer (“1×1”), which effectively acts as a concatenation aggregator; a k×k depthwise convolutional layer (“k×k DW”); a batch normalization layer (“BN”); and a rectified linear unit activation layer (“ReLU”). An example architecture generated from the search space (as depicted on the left hand side of Figure 4A can be in the form of a tower including a sequence of neural network units, where each unit after the first unit in the sequence receives input from one or more earlier units in the sequence, receives a network input, or both.

[0042] Figure 4B Another example illustration of a search space is shown. The example search space includes a plurality of neural network blocks, and the neural network blocks in turn include a plurality of operators. Each operator includes a plurality of neural network layers, which include one or more of the following: a 1×1 convolutional layer (“1×1”), which effectively acts as a concatenation aggregator; a 3×3 depthwise convolutional layer (“3×3 DW”); a 5×5 depthwise convolutional layer (“5×5 DW”); a batch normalization layer (“BN”); and a rectified linear unit activation layer (“ReLU”). Dashed lines indicate additional skip connections between different operators included in the same block. An example architecture generated from the search space (as shown on the left hand side of Figure 4B can be in the form of a tower that includes a sequence of neural network blocks, where each block after the first block in the sequence receives input from an earlier block in the sequence.

[0043] In some embodiments, the neural architecture search system 100 can additionally receive as input data one or more specified (e.g., user-specified) resource constraints that identify how much computational resources a neural network can consume when performing inference. For example, the resource constraints can specify a target amount of computational resources to be used by the neural network with the final architecture. The target resource usage data specifies: (i) a target memory size that indicates the maximum memory size allowed for creating the final architecture, i.e., the maximum memory that can be occupied by the parameters and architecture data of the final architecture; and (ii) a target number of operations (e.g., floating-point operations per second (FLOPS)) that indicates the maximum number of operations that the neural network with the final architecture can perform to execute a particular machine learning task. As another example, the resource constraints can specify a target runtime latency of the neural network when performing a task and when deployed on one or more computing devices. Thus, the output of the neural architecture search system 100 can be further associated with the specific technical details of the hardware that the neural network is intended to act on.

[0044] The neural architecture search system 100 can receive the training data 102, the architecture data 104, additional input data, or a combination thereof in any of a variety of ways. For example, the system 100 can receive data as an upload from a remote user of the system via a data communication network, e.g., using an application programming interface (API) provided by the system 100. As another example, the system 100 can receive input from a user that specifies which data that the system 100 has maintained should be used as the training data 102 and the architecture data 104.

[0045] To determine the final architecture, the neural architecture search system 100 uses a search process that is repeatedly executed using a variable allocation engine 110, an architecture generation engine 120, and a training engine 130.

[0046] The variable allocation engine 110 can allocate utilization variables 112 to each of a plurality of network operators defined in the architecture data 104. Each utilization variable 112 is associated with a value indicating the likelihood that the network operator is utilized in the final architecture of the neural network.

[0047] The utilization variable allocation engine 110 can determine the associated value of each utilization variable 112 according to the corresponding probability parameterized by one or more tunable parameters ("distribution parameters") that can be maintained by the engine 110. For example, the utilization variable allocation engine 110 can model the utilization variable as a Bernoulli random variable with continuous relaxation. A Bernoulli random variable without continuous relaxation refers to a discrete variable whose value is 1 (probability p) or 0 (probability 1 - p), and the continuous relaxation technique will be further described below. For each network operator, the allocated utilization variable 112 modeled in this way can effectively be regarded as a binary mask. For example, when the value is 1, the network operator is included in the architecture, and when the value is 0, the network operator is pruned from the architecture.

[0048] Figure 5 An example illustration of inserting a mask corresponding to a utilization variable into a network operator is shown. As depicted, the binary masking layer corresponding to the utilization variable allocated to a network operator (e.g., "op 1" at the top) is inserted after this network operator in the architecture. The number next to the edge indicates the number of non - zero channels in the masking layer.

[0049] In response to sampling a fully - zero masking layer inserted after a network operator, the network operator (e.g., "op 2" at the bottom) can be deselected, that is, the network operator is pruned from the current architecture. Thus, deselecting a network operator is equivalent to inserting a zero masking layer after an operator that is not in the selected set of network operators.

[0050] Part of a component of a network operator (e.g., "op 3" at the bottom) can be deselected in response to sampling a masking layer inserted after the network operator that has some zero channels but not all zero channels. As Figure 5 depicted at the bottom, if the network operator is a component of a convolutional neural network layer, that is, a filter of a convolutional neural network layer, in response to sampling a masking layer with 5 zero channels, 5 out of a total of 8 filters of the convolutional neural network layer can be deselected. This is equivalent to modifying the width of the convolutional neural network layer.

[0051] Therefore, the neural architecture search system 100 can act on any of the multiple search spaces in the following way: using the utilization variable allocation engine 110 to allocate utilization variables to different network operators, that is, inserting masks into different neural network layers and different components of neural network layers. This helps the system 100 to perform a fine - grained search over a larger architecture space.

[0052] In particular, system 100 is capable of modeling the space of probability values for utilization variables as a continuous space, and utilization variable allocation engine 110 is capable of modeling the corresponding likelihoods for each network operator utilized in the final architecture as continuous rather than discrete likelihoods. This allows the system to have fine-grained control over the search process. For example, neural architecture search system 100 uses utilization variable allocation engine 110 to assign utilization variables with a given numerical value (e.g., 1) to each network operator according to probabilities determined from a continuous distribution (e.g., Logistic-Sigmoid distribution or a specific distribution) over a set of probability values for utilization variables ranging from 0 to 1, inclusive of both ends.

[0053] Architecture generation engine 120 is capable of generating a training candidate architecture 122 for a neural network based on the values of utilization variables 112 assigned to each of a plurality of network operators. To generate a new training candidate architecture 122 at the start of the search process (or to update an existing training candidate architecture 122 during the search process), architecture generation engine 120 is capable of selecting a set of network operators from the plurality of network operators defined by architecture data 104 according to the current values of the utilization variables assigned to the plurality of network operators. For example, architecture generation engine 120 can select network operators that have been assigned utilization variables whose values fall within a specific range, e.g., greater than 0.9, e.g., equal to 1. Then, the training neural network architecture 122 can be determined by using engine 120 as a combination of the selected set of network operators. For example, architecture generation engine 120 can do this by using an addition aggregator configurable to combine outputs of fixed-shape operators or a 1×1 convolutional layer (i.e., a concatenation aggregator) configurable to combine outputs of variable-shape operators.

[0054] In some embodiments, during the search process, system 100 maintains, e.g., at a memory device accessible to system 100, a set of distribution parameters used by utilization variable allocation engine 110 to generate utilization variables 112 and a set of parameters for the neural network. The set of parameters for the neural network in turn consists of different subsets of parameters associated with different operators of the neural network.

[0055] For a training architecture 122 generated by using architecture generation engine 120 and from architecture data 104 and utilization variables 112, training engine 130 trains an instance of the neural network with training architecture 122 on training data 102 to iteratively update the values of the set of parameters of the neural network and additionally adjust the values of the utilization variables generated by using utilization variable allocation engine 110. Specifically, during training, training engine 130 jointly optimizes two objectives - a computational cost objective and a task execution objective.

[0056] A computational cost target can be derived from user input to the system or from some default computational cost targets associated with deploying a neural network on one or more computing devices to perform a specific machine learning task. For example, the computational cost target can include one or more of the following: (i) the target memory size of the final architecture, (ii) the target floating point operations per second (FLOPS) of the neural network with the final architecture when performing a specific machine learning task, or (iii) the target runtime latency of the neural network when performing a specific machine learning task.

[0057] The task execution target can evaluate the execution metric of the neural network during a training iteration, which measures the execution of the trained neural network for a specific machine learning task. For example, the execution metric can be the loss of the trained neural network on a validation dataset, or the result of some other metric of model accuracy calculated on the validation dataset.

[0058] The training engine 130 determines the update by computing the gradient of the objective function, which evaluates the above two objectives with respect to the network parameters. To adjust the values of the distribution parameters that sequentially define the distribution of the variables, the training engine 130 can backpropagate the determined gradient into the distribution parameters through the variables.

[0059] During the search process, the system 100 can repeatedly use the architecture generation engine 120 to update the trained neural network architecture 122 based on the updated values of the variables of exploitation assigned to each of the plurality of network operators. This allows the system to continuously and adaptively update the training architecture of the neural network to increase the diversity of the search process. At the same time, the set of network parameters associated with the updated training architecture can be updated based on the gradient of the objective function computed by the training engine 130.

[0060] After the search process has terminated, for example, after a specified number of iterations have been performed or after the gradient of the objective function has converged to a specified value, the neural network search system 100 can then output the final architecture data 150 of the neural network. For example, the neural network search system 100 can output data specifying the final neural network architecture 150 to the user who submitted the training data 102. For example, the architecture data can specify the neural network operators that are part of the neural network, the connectivity between neural network operations, and the operations performed by the neural network operators.

[0061] In some embodiments, instead of or in addition to outputting the architecture data 150, the system 100 instantiates an instance of a neural network having the determined architecture and trained parameters. For example, the trained parameters are either trained from scratch with parameter values generated as a result of a search process by the system after determining the final architecture, or are generated by fine-tuning the parameter values generated as a result of the search process. Then, the system 100 uses the trained neural network to process requests received from a user (e.g., via an API provided by the system). That is, the system 100 is capable of receiving an input to be processed, processing the input using the trained neural network, and providing an output generated by the trained neural network or data derived from the generated output in response to the received input.

[0062] Figure 2 FIG. 4 is a flow chart of an example process 200 for searching for a neural network architecture. For convenience, process 200 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed neural architecture search system, such as Figure 1 the neural architecture search system 100, can perform process 200.

[0063] The system receives training data (202) for training a neural network to perform a neural network task. The training data includes a plurality of training examples, and for each training example, a corresponding target output that should be generated by the neural network to perform a particular task.

[0064] The system receives architecture data (204) that defines a plurality of network operators. Generally, when used as part of the architecture of a neural network, each network operator is configured to receive an operator input and generate a corresponding operator output based on processing the operator input according to the current values of the parameters associated with the network operator. For example, an operator can include a neural network layer, a component of a neural network layer, or a combination of two or more neural network layers, or a combination of two or more components of a neural network layer.

[0065] The system assigns an exploitation variable (206) to each of the plurality of network operators that indicates the likelihood of the network operator being utilized in the neural network. The value associated with the exploitation variable can be determined according to a corresponding probability, which in turn can be parameterized by one or more distribution parameters. For example, a continuous probability distribution (e.g., a Logistic-Sigmoid distribution or a specific distribution) can be used to model the exploitation variable as a Bernoulli random variable with continuous relaxation.

[0066] The system generates a neural network (208) for performing a neural network task by jointly updating current values ​​of utilization variables and current values ​​of parameters of the neural network by repeatedly performing the following two steps: (i) sampling a selected set of network operators from a plurality of network operators and according to corresponding current values ​​of the utilization variables, and (ii) training a neural network having an architecture defined by the selected set of network operators on training data to perform the neural network task, as described in more detail below.

[0067] Figure 3 is a flow chart of an example process 300 for training a neural network having an architecture defined by a selected set of network operators. For convenience, process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a suitably programmed neural architecture search system, such as Figure 1 The neural architecture search system 100 is capable of performing process 300.

[0068] In general, the system can repeatedly perform process 300 to generate a neural network having an optimal architecture for performing a neural network task.

[0069] The system samples a selected set of network operators from a plurality of network operators and based on respective current values ​​of utilization variables (302).

[0070] The system trains a neural network having an architecture defined by a selected set of network operators on the training data to perform a neural network task (304). Figure 1 Sampling a selected set of network operators and generating a neural network having an architecture defined by the selected set of network operators is described in more detail, but in brief, this involves using a suitable aggregator (e.g., an additive aggregator or a cascade aggregator) to generate a combination of network operators that have been assigned utilization variables having current values ​​within a particular range (e.g., equal to 1).

[0071] During training, the system jointly updates current values ​​of the utilization variables and current values ​​of the neural network parameters to optimize an objective function that simultaneously evaluates a measure of the computational cost of the neural network and a measure of the neural network's performance of a neural network task associated with the training data.

[0072] To this end, the system computes an objective function (306) that evaluates (i) a measure of the computational cost of the neural network and (ii) a measure of the neural network's performance of the neural network task associated with the training data.

[0073] For example, the objective function can be evaluated as

[0074]

[0075] Among them, is the loss term for the measurement task execution objective, is the loss term for the measurement calculation cost objective, and λ is a regularization factor that scales differently depending on the exact type of the calculation cost objective of interest (e.g., FLOPS), and ⊙ represents the element-wise product between a set of one or more parameters w of the neural network with the training architecture and a binary mask m corresponding to the utilization variable. In this example,

[0076] Here, can be the supervised loss used in the conventional machine learning training for the task, that is, the loss of the training output determined relative to the associated target output included in the training data, where the training output is generated by the neural network processing the training samples according to the current values of the network parameters.

[0077] And it can be evaluated in the following form

[0078]

[0079] Among them, and represent the per-channel binary masks applied to the input of matrix A and the output of matrix B respectively, and π represents the distribution parameter. Matrix A (or B) can be the weight matrix representing the values of the parameters associated with the network operator (e.g., the convolutional layer or the fully connected layer of the neural network).

[0080] The system adjusts the current values of the utilization variable and the current values of the neural network parameters (308) based on the determined gradient of the objective function. The system can do this by calculating the gradient of the objective function with respect to the neural network parameters and then backpropagating the determined gradient of the objective function into one or more distribution parameters through the utilization variable.

[0081] In some cases, m i can be modeled as an independent Bernoulli variable m i ~ Bern(π i ), and the system can use black-box methods (e.g., perturbation or logarithmic derivative methods) to determine the estimate of the gradient with respect to the distribution parameter π.

[0082] In other cases, m i can be modeled as a continuous sample from the Logistic-Sigmoid distribution instead where respectively, l ~ Logistic(0,1), and when τ → 0, with (1 - π i , πi ) has a probability close to (0, 1). Decomposing the logical terms into parameterless components allows the system to backpropagate the gradient of the calculation of the objective function through the mask and learn the value of the distribution parameter π using traditional parameter update techniques such as gradient descent-based techniques.

[0083] The system can then use appropriate update rules, such as stochastic gradient descent update rules, Adam update rules, rmsProp update rules, to apply the gradient to adjust the exploitation variables and network parameters.

[0084] This specification uses the term "configured" in relation to systems and computer program components. For a system of one or more computers configured to perform particular operations or actions, it means that the system has software, firmware, hardware, or a combination thereof installed on it that, in operation, cause the system to perform the operations or actions. For one or more computer programs configured to perform particular operations or actions, it means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operations or actions.

[0085] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus.

[0086] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, equipment, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus can also be or further include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). In addition to the hardware, the apparatus can optionally include code that creates an execution environment for the computer program, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0087] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for a computing environment. The program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data, for example, in one or more scripts in a markup language document, in a single file dedicated to the program being discussed, or in multiple coordinated files, for example, files that hold one or more modules, subroutines, or portions of code. A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0088] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or structured at all, and it can be stored on a storage device in one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed in a different way.

[0089] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers at one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.

[0090] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.

[0091] A computer suitable for executing a computer program can be based on a general-purpose or special-purpose microprocessor or both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for executing or implementing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from, or transfer data to, or both, from such one or more mass storage devices. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.

[0092] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0093] To provide interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser in response to a request received from the web browser on the user's device. Moreover, a computer can interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smart phone running a messaging application) and, in turn, receiving a response message from the user.

[0094] The data processing apparatus for implementing a machine learning model may further include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive parts of machine learning training or production (i.e., inference) workloads.

[0095] A machine learning model can be implemented and deployed using a machine learning framework (e.g., TensorFlow framework, Microsoft Cognitive Toolkit framework, Apache Singa framework, or Apache MXNet framework).

[0096] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server, or includes middleware components, such as an application server, or includes frontend components, such as a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or the computing system can include any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, e.g., a communication network. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), e.g., the Internet.

[0097] The computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship between the client and the server results from computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to a user device, e.g., for the purpose of displaying data to a user interacting with the device that can be a client and receiving user input from the user. Data generated at the user device (e.g., the result of a user interaction) can be received at the server from the user device.

[0098] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or the scope of what is claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although the features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be removed from the combination, and the claimed combination can be directed to a sub-combination or variation of a sub-combination.

[0099] Similarly, although the operations are depicted in the drawings in a particular order and recited in the claims, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0100] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes depicted in the figures do not necessarily need the particular order shown or sequential order to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method for determining an architecture for a neural network, comprising: Receiving training data for training a neural network to perform a neural network task, the training data including a plurality of training examples and corresponding target outputs for each of the plurality of training examples; Receiving architecture data defining a plurality of network operators; Assigning to each of the plurality of network operators an exploitation variable indicative of the likelihood that the network operator is utilized in the neural network; Generating an optimized neural network for performing the neural network task, including repeatedly performing the following operations: Sampling a selected set of network operators from the plurality of network operators and according to the respective current values of the exploitation variables; And Training the neural network having an architecture defined by the selected set of network operators on the training data to perform the neural network task, wherein the training includes: Calculating an objective function that evaluates (i) a measure of the computational cost of the neural network and (ii) a measure of the performance of the neural network for the neural network task associated with the training data; and Adjusting the respective current values of the exploitation variables and the respective current values of the parameters of the neural network based on the determined gradient of the objective function.

2. The method according to claim 1, wherein, The architecture data is initialized from one or more predetermined neural network architectures.

3. The method according to claim 1, further comprising: Removing redundant network operators from the plurality of network operators.

4. The method according to claim 1, wherein, The plurality of network operators includes neural network layers.

5. The method according to claim 4, wherein, The neural network layer includes at least one of the following: a convolutional layer, a fully connected layer, a normalization layer, and an activation layer.

6. The method according to claim 4, wherein, The plurality of operators further includes filters in a convolutional layer or neurons in a fully connected layer.

7. The method according to claim 1, wherein The measure of the computational cost of the neural network includes at least one of the following: size, floating point operations per second (FLOPS), and latency.

8. The method according to claim 1, wherein, Generating the optimized neural network for performing the neural network task further includes: Inserting zero masking layers after each operator that is not one of the selected set of network operators.

9. The method according to claim 1, wherein Generating the optimized neural network for performing the neural network task further includes: Combining the respective outputs of the selected set of network operators using a concat aggregator.

10. The method according to claim 1, wherein, Each exploitation variable is defined by one or more distribution parameters.

11. The method according to claim 10, further comprising: Calculating the determined gradient of the objective function with respect to the one or more distribution parameters.

12. The method according to claim 10, wherein Adjusting the respective current values of the exploitation variables includes: Backpropagating the determined gradient of the objective function into the one or more distribution parameters through the exploitation variables.

13. The method according to any one of claims 10 to 12, wherein The one or more distribution parameters define a binary specific distribution.

14. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 to 13.

15. A non-transitory computer-readable storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 13.