Single Search Of Architecture On Embedded Devices
Through the All-Search (OSFA) process, using weight sharing to search for neural network architecture suitable for low-resource platforms, the problems of high computing costs and mismatch with hardware in the existing technology are solved, and the effect of efficient deployment of neural networks on low-resource platforms is achieved.
Patent Information
- Application Number
- CN202380074245.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-01
- Filing Date
- 2023-08-25
- Publication Date
- 2025-06-06
AI Technical Summary
Deploying advanced neural networks on embedded devices is difficult, and the computational cost of existing neural architecture search and compression algorithms is mainly aimed at convolutional neural networks, or does not match hardware behavior.
A fast neural architecture search process, called Search All at One Time (OSFA), is proposed, using weight sharing searches for convolution, loop, multi-layer perceptron (MLP) or transformer architectures performed on low resource platforms, network accuracy, runtime delay and peak memory usage are jointly optimized for hardware efficiency.
A neural network that maintains or exceeds the accuracy of baseline model on low resource platforms is realized, reducing the number of GPU hours required to find a good architecture, and adapting to the low-cost general adaptation of low-resource platforms.
Smart Images

Figure CN120112918A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. patent application No. 18 / 363,487, filed on August 1, 2023, and entitled “SINGLE SEARCH FOR ARCHITECTURES ON EMBEDDED DEVICES,” which claims the benefit of U.S. provisional patent application No. 63 / 420,511, filed on October 28, 2022, and entitled “SINGLE SEARCH FOR ARCHITECTURES ON EMBEDDED DEVICES,” the disclosures of which are expressly incorporated by reference in their entirety. Technical Field
[0003] Aspects of the present disclosure generally relate to neural network architecture search, and more particularly to one-shot neural architecture search (NAS) for hardware-efficient architectures that run on devices with limited resources. Background Art
[0004] An artificial neural network may include an interconnected group of artificial neurons (e.g., a neuron model). An artificial neural network may be a computing device or represented as a method to be performed by a computing device. A convolutional neural network (CNN) is a type of feedforward artificial neural network. A convolutional neural network may include a collection of neurons, each of which has a receptive field and collectively spells out an input space. Convolutional neural networks such as deep convolutional neural networks (DCNs) have numerous applications. Other types of neural networks include recurrent neural networks, multi-level perceptrons (MLPs), and transformers, among others. These neural network architectures are used in a variety of technologies, such as image recognition, speech recognition, acoustic scene classification, keyword detection, autonomous driving, and other classification tasks.
[0005] Deploying advanced neural networks on embedded devices is difficult but important for the widespread use of deep learning. Existing neural architecture search and compression algorithms can be computationally expensive, focus on convolutional neural networks, or do not match the behavior of the hardware. Summary of the invention
[0006] A processor-implemented method for neural architecture search (NAS) begins by generating an over-parameterized hypernetwork having a plurality of layers. The hypernetwork has a plurality of operator types. Each of the layers includes a maximum hyperkernel corresponding to a search space. The method also includes performing a gradient descent to evolve the maximum hyperkernel into a small kernel corresponding to the search space so as to generate a series of kernel encodings. The method also includes identifying a subset of kernel encodings from the series of kernel encodings for each layer of the hypernetwork based on the gradient descent. The method determines a set of candidate architectures based on the subset of kernel encodings, each of the candidate architectures having a different model size. The method selects a target model from the set of architectures based on meeting hardware specifications, and then applies the target model.
[0007] Other aspects of the present disclosure relate to a device. The device has a memory and one or more processors coupled to the memory. The processor is configured to generate an over-parameterized hypernetwork having multiple layers. The hypernetwork includes multiple operator types. Each of the multiple layers includes a maximum hyperkernel corresponding to a search space. The processor is also configured to perform gradient descent to evolve the maximum hyperkernel into a small kernel corresponding to the search space to generate a series of kernel encodings. The processor is also configured to identify a kernel encoding subset from the series of kernel encodings for each layer of the hypernetwork based on the gradient descent. The processor is also configured to determine a set of candidate architectures based on the kernel encoding subsets. Each candidate architecture in the set of candidate architectures has a different model size. The processor is also configured to select a target model from the set of architectures based on meeting hardware specifications. The processor is also configured to apply the target model.
[0008] Other aspects of the present disclosure relate to an apparatus. The apparatus includes a component for generating an over-parameterized hypernetwork having a plurality of layers. The hypernetwork includes a plurality of operator types. Each of the plurality of layers includes a maximum hyperkernel corresponding to a search space. The apparatus also includes a component for performing gradient descent to evolve the maximum hyperkernel into a small kernel corresponding to the search space to generate a series of kernel encodings. The apparatus also includes a component for identifying a subset of kernel encodings from the series of kernel encodings for each layer of the hypernetwork based on the gradient descent. The apparatus also includes a component for determining a set of candidate architectures based on the subset of kernel encodings. Each candidate architecture in the set of candidate architectures has a different model size. The apparatus also includes a component for selecting a target model from the set of architectures based on satisfying hardware specifications. The apparatus also includes a component for applying the target model.
[0009] Additional features and advantages of the present disclosure will be described below. It will be appreciated by those skilled in the art that the present disclosure can be easily used as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. It will also be appreciated by those skilled in the art that such equivalent constructions do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features considered to be characteristic of the present disclosure, both in terms of its organization and method of operation, together with further objects and advantages, will be better understood when the following description is considered in conjunction with the accompanying drawings. However, it is to be expressly understood that each of the accompanying drawings is provided for illustration and description purposes only and is not intended to be a definition of limitations of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The features, nature and advantages of the present disclosure will become more apparent when the detailed description set forth below is read in conjunction with the accompanying drawings, in which like reference numerals are correspondingly identified throughout.
[0011] Figure 1 An example implementation of a neural network using a system on a chip (SOC) including a general purpose processor in accordance with certain aspects of the present disclosure is illustrated.
[0012] Figure 2A , Figure 2B and Figure 2C is a diagram illustrating a neural network according to aspects of the present disclosure.
[0013] Figure 2D is a diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.
[0014] Figure 3 is a block diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.
[0015] Figure 4 is a diagram illustrating the overall workflow of a neural network architecture search according to aspects of the present disclosure.
[0016] Figure 5 is a diagram illustrating a heterogeneous operator-level search space according to aspects of the present disclosure.
[0017] Figure 6 is a diagram illustrating node partitioning with architectural co-design according to aspects of the present disclosure.
[0018] Figure 7 is a diagram illustrating a weight sharing technique in a hyperkernel according to aspects of the present disclosure.
[0019] Figure 8 is a diagram illustrating progressive growth according to aspects of the present disclosure.
[0020] Fig. 9 is a diagram illustrating progressive zoom-out according to aspects of the present disclosure.
[0021] Fig.10 is a block diagram illustrating knowledge distillation in neural architecture search (NAS) according to aspects of the present disclosure.
[0022] Fig.11 is a flow chart illustrating an example processor-implemented method of searching a neural network architecture in accordance with aspects of the present disclosure. DETAILED DESCRIPTION
[0023] The specific embodiments described below in conjunction with the accompanying drawings are intended as descriptions of various configurations and are not intended to represent the only configurations in which the described concepts may be practiced. In order to provide a thorough understanding of the various concepts, the detailed description includes specific details. However, it is apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, in order to avoid obscuring such concepts, well-known structures and components are shown in block diagram form.
[0024] Based on the teachings, those skilled in the art will recognize that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether the aspect is implemented independently of any other aspect of the present disclosure or implemented in combination with any other aspect. For example, a device or a method may be implemented using any number of aspects described. In addition, the scope of the present disclosure is intended to cover such devices or methods that are practiced using other structures, functionality, or structures and functionality that are supplementary to or different from the various aspects of the present disclosure described. It should be understood that any aspect of the present disclosure disclosed may be embodied by one or more elements of the claims.
[0025] The word “exemplary” is used to mean “serving as an example, instance, or illustration.” Any aspect described as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.
[0026] Although specific aspects are described, numerous variations and permutations of these aspects fall within the scope of the present disclosure. Although some benefits and advantages of preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses or purposes. On the contrary, various aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks and protocols, some of which are illustrated as examples in the accompanying drawings and the following description of preferred aspects. The specific embodiments and drawings are merely illustrative of the present disclosure and are not limiting, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0027] Deploying advanced neural networks on embedded devices is difficult but important for the widespread use of deep learning. Existing neural architecture search and compression algorithms can be computationally expensive, focus on convolutional neural networks, or do not match the behavior of the hardware.
[0028] Aspects of the present disclosure present a fast neural architecture search process, called One-Time Search All (OSFA), to create a set of common network architectures on low-resource hardware platforms. The new search process exploits weight sharing to search for convolutional, recurrent, multi-layer perceptron (MPL), or transformer architectures for execution on low-resource platforms. Network accuracy, runtime latency, and peak memory usage are jointly optimized for hardware efficiency. Using the new search process, accuracy is maintained or exceeded compared to baseline models, enabling low-cost universal adaptation of neural networks to low-resource platforms. The new search process reduces the search cost in terms of graphics processing unit (GPU) hours required to find a good architecture.
[0029] Figure 1 An example implementation of a system on a chip (SOC) 100 is illustrated, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for neural architecture search. Variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), delays, frequency bin information, and task information may be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or may be distributed across multiple blocks. Instructions executed at CPU 102 may be loaded from a program memory associated with CPU 102, or may be loaded from memory block 118.
[0030] The SOC 100 may also include additional processing blocks tailored for specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that may, for example, detect and recognize gestures. In one specific implementation, the NPU 108 is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which may include a global positioning system.
[0031] SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general processor 102 may include code for generating an over-parameterized super network with multiple layers. The super network has multiple operator types. Each of these layers includes a maximum super kernel corresponding to the search space. The general processor 102 may also include code for performing gradient descent to evolve the maximum super kernel into a small kernel corresponding to the search space so as to generate a series of kernel encodings. The general processor 102 may include code for identifying a kernel encoding subset from the series of kernel encodings for each layer of the super network based on the gradient descent. The general processor 102 may also include code for determining a set of candidate architectures based on the kernel encoding subset, each of which has a different model size. The general processor 102 may include code for selecting a target model from the set of architectures based on meeting hardware specifications and then applying the target model.
[0032] Deep learning architectures can perform object recognition tasks by learning to represent inputs at successively higher levels of abstraction in each layer, thereby building useful feature representations of the input data. In this way, deep learning solves the main bottleneck of traditional machine learning. Before the advent of deep learning, machine learning methods for object recognition problems may rely heavily on human-designed features, possibly in conjunction with shallow classifiers. A shallow classifier can be a two-class linear classifier, for example, in which the weighted sum of feature vector components can be compared to a threshold to predict which class the input belongs to. Human-designed features can be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, although deep learning architectures can learn to represent features similar to those that human engineers may design, they require training. In addition, deep networks can learn to represent and recognize new types of features that humans may not have considered.
[0033] A deep learning architecture can learn a hierarchy of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power in specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.
[0034] Deep learning architectures perform particularly well when applied to problems that have a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0035] Neural networks can be designed to have a variety of connection patterns. In a feedforward network, information is passed from a lower layer to a higher layer, where each neuron in a given layer communicates with a neuron in a higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have loops or feedback (also known as top-down) connections. In a loop connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. The loop architecture can help identify patterns that span more than one input data block in the input data blocks that are sequentially delivered to the neural network. The connection from a neuron in a given layer to a neuron in a lower layer is called a feedback (or top-down) connection. When the recognition of a high-level concept can assist in discerning specific low-level features of the input, a network with many feedback connections may be helpful.
[0036] The connections between neural network layers can be fully connected or partially connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In a fully connected neural network 202, a neuron in a first layer may communicate its output to every neuron in a second layer, such that every neuron in the second layer will receive input from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, neurons in a first layer may be connected to a limited number of neurons in a second layer. More generally, the locally connected layers of the locally connected neural network 204 may be configured such that each neuron in the layer will have the same or similar connection pattern, but the connection strengths may have different values (e.g., 210, 212, 214, and 216). The connection patterns of the local connections may produce spatially different receptive fields in higher layers because higher layer neurons in a given region may receive inputs that are tuned through training to characteristics of a limited portion of the total input to the network.
[0037] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 may be configured such that the connection strengths associated with the inputs of each neuron in the second layer are shared (e.g., 208). Convolutional neural networks may be well suited for problems where the spatial location of the inputs is meaningful.
[0038] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2DA detailed example of a DCN 200 designed to recognize visual features from an image 226 input from an image capture device 230 (such as a vehicle-mounted camera) is illustrated. The DCN 200 of the current example can be trained to identify traffic signs and numbers provided on traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic lights.
[0039] The DCN 200 may be trained by supervised learning. During training, an image (such as an image 226 of a speed limit sign) may be presented to the DCN 200, and a forward pass may then be calculated to produce an output 222. The DCN 200 may include a feature extraction portion and a classification portion. Upon receiving the image 226, a convolutional layer 232 may apply a convolutional kernel (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 may be a 5x5 kernel that generates a 28x28 feature map. In this example, because four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to the image 226 at the convolutional layer 232. A convolutional kernel may also be referred to as a filter or a convolutional filter.
[0040] The first set of feature maps 218 may be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 may be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more sets of subsequent feature maps (not shown).
[0041] exist Figure 2D In the example of , the second set of feature maps 220 are convolved to generate a first feature vector 224. In addition, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 may include a number corresponding to a possible feature of the image 226, such as "sign", "60", and "100". A softmax function (not shown) may convert the numbers in the second feature vector 228 into probabilities. Therefore, the output 222 of the DCN 200 may be a probability that the image 226 includes one or more features.
[0042] In this example, the probability of "logo" and "60" in output 222 is higher than the probability of other numbers (such as "30", "40", "50", "70", "80", "90", and "100") in output 222. Before training, the output 222 generated by DCN 200 may be incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the ground truth value of image 226 (e.g., "logo" and "60"). The weights of DCN 200 can then be adjusted so that the output 222 of DCN 200 is closer to the target output.
[0043] To adjust the weights, the learning algorithm may calculate a gradient vector for the weights. The gradient may indicate the amount by which the error will increase or decrease if the weights are adjusted. At the top layer, the gradient may correspond directly to the value of the weights connecting the activated neurons in the penultimate layer and the neurons in the output layer. In lower layers, the gradient may depend on the value of the weights and the calculated error gradients of the higher layers. The weights may then be adjusted to reduce the error. This way of adjusting weights may be referred to as "backward propagation" because it involves a "backward pass" through the neural network.
[0044] In practice, the error gradient of the weights may be calculated over a small number of examples so that the calculated gradient is close to the true error gradient. This approximation method may be referred to as stochastic gradient descent. Stochastic gradient descent may be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, a new image (e.g., a speed limit sign of image 226) may be presented to DCN 200, and a forward pass through DCN 200 may produce output 222, which may be considered an inference or prediction of DCN 200.
[0045] Deep belief network (DBN) is a probabilistic model including multiple layers of hidden nodes. DBN can be used to extract hierarchical representations of training data sets. DBN can be obtained by stacking the layers of restricted Boltzmann machine (RBM). RBM is a type of artificial neural network that can learn probability distributions through a set of inputs. Because RBM can learn probability distributions without information about the category to which each input should be classified, RBM is generally used for unsupervised learning. Using a supervised and unsupervised hybrid paradigm, the bottom RBM of DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the input and target class from the previous layer) and can be used as a classifier.
[0046] A DCN is a network of convolutional networks, configured with additional pooling and normalization layers. DCN has achieved state-of-the-art performance on many tasks. DCN can be trained using supervised learning, where both the input and output targets are known for many examples and are used to modify the weights of the network using a gradient descent method.
[0047] The DCN may be a feed-forward network. In addition, as described above, the connections from a neuron in the first layer of the DCN to a set of neurons in the next higher layer are shared across the neurons in the first layer. The feed-forward and shared connections of the DCN can be used for fast processing. For example, the computational burden of the DCN may be much smaller than that of a similarly sized neural network that includes loops or feedback connections.
[0048] The processing of each layer of the convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green and blue channels of a color image, the convolutional network trained on the input can be considered to be three-dimensional, with two spatial dimensions along the axis of the image and the third dimension capturing color information. The output of the convolutional connection can be considered to form a feature map in a subsequent layer, where each element in the feature map (e.g., 220) receives input from a certain range of neurons in the previous layer (e.g., feature map 218) and from each channel in the multiple channels. The values in the feature map can be further processed with nonlinearity (such as correction, maximum value (max)(0,x)). The values from adjacent neurons can be further pooled, which corresponds to downsampling and can provide additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied by lateral inhibition between neurons in the feature map.
[0049] Figure 3 is a block diagram illustrating a DCN 350. The DCN 350 may include multiple different types of layers based on connection and weight sharing. Figure 3 As shown, DCN 350 includes convolution blocks 354A, 354B. Each of convolution blocks 354A, 354B may be configured with a convolution layer (CONV) 356, a normalization layer (LNorm) 358, and a maximum pooling layer (MAX POOL) 360.
[0050] Although only two of the convolution blocks 354A, 354B are shown, the present disclosure is not limited thereto, and any number of convolution blocks 354A, 354B may be included in the DCN 350 according to design preferences.
[0051] The convolution layer 356 may include one or more convolution filters that may be applied to the input data to generate a feature map. The normalization layer 358 may normalize the output of the convolution filter. For example, the normalization layer 358 may provide whitening or lateral suppression. The maximum pooling layer 360 may provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.
[0052] For example, a parallel filter bank of a deep convolutional network may be loaded on SOC 100 (e.g., Figure 1 ) on the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. In addition, the DCN 350 can access other processing blocks that may exist on the SOC 100, such as the sensor processor 114 and the navigation module 120 dedicated to sensors and navigation, respectively.
[0053] The DCN 350 may also include one or more fully connected layers 362 (FC1 and FC2). The DCN 350 may also include a logistic regression (LR) layer 364. Between each layer 356, 358, 360, 362, 364 of the DCN 350 are weights (not shown) to be updated. The output of each layer in the layers (e.g., 356, 358, 360, 362, 364) may be used as an input to a subsequent layer in the layers (e.g., 356, 358, 360, 362, 364) in the DCN 350 to learn a hierarchical feature representation from the input data 352 (e.g., images, audio, video, sensor data, and / or other input data) supplied at the first convolutional block in the convolutional block 354A. The output of the DCN 350 is a classification score 366 for the input data 352. The classification score 366 may be a set of probabilities, where each probability is a probability that the input data includes a feature from a set of features.
[0054] Deploying advanced neural networks on embedded devices is difficult but important for the widespread use of deep learning. Existing neural network architecture search and compression algorithms can be computationally expensive, focus on convolutional neural networks, or may not match the behavior of the hardware.
[0055] Aspects of the present disclosure present a fast neural network architecture search process, called One-Time Search All (OSFA), to create a set of common network architectures for low-resource hardware platforms. The new search process exploits weight sharing to search for convolutional, recurrent, multi-layer perceptron (MLP), or transformer architectures for execution on low-resource platforms. Network accuracy, runtime latency, and peak memory usage are jointly optimized for hardware efficiency. Using the new search process, accuracy is maintained or exceeded compared to baseline models, enabling low-cost universal adaptation of neural networks to low-resource platforms.
[0056] Neural architecture search (NAS) has revolutionized the design of networks, bringing state-of-the-art (SOTA) performance to targeted use cases. These techniques attempt to find superior neural architectures at the expense of significant search costs by extensively exploring the search space using sampling techniques such as reinforcement learning and evolutionary search. To reduce the search cost, a family of NAS techniques, known as one-shot differentiable NAS, attempts to reduce the search cost by training a single over-parameterized hypernetwork that encodes all possible architectures in the search space.
[0057] At the same time, there is an increasing demand to deploy neural networks on embedded hardware platforms without relying on centralized computation. These platforms have specific hard resource constraints such as model memory, peak memory, bandwidth, and power consumption, as well as expectations for reasonable inference latency. This need to deploy on a unique target platform poses a challenge as the NAS processing will need to be rerun for each platform. Existing one-shot NAS techniques address this problem by selecting a subnetwork from a well-trained “one-shot” (OFA) network, but use a more costly evolutionary search. There is little work on searching for a set of hardware-efficient architectures using the more efficient family of one-shot differentiable NAS with weight sharing, especially when the base architecture is extended beyond convolutional networks for tasks such as recurrent neural networks (RNNs) for keyword spotting.
[0058] To address this disconnect, aspects of the present disclosure introduce One-Stop Search All (OSFA) NAS, a one-shot differentiable NAS technique for designing a set of neural network architectures for embedded platforms in just a few graphics processing unit (GPU) hours. Utilizing a weight sharing mechanism, OSFA can search for convolutional, recurrent, and multi-layer perceptron (MLP) network architectures on low-resource platforms. Network accuracy, runtime latency, and peak memory usage are jointly optimized for hardware efficiency. Figure 4 An example of the overall workflow of OSFA is shown in .
[0059] Figure 4 is a diagram illustrating the overall workflow of a neural network architecture search according to various aspects of the present disclosure. Figure 4In the example of , in the leftmost part (a) before neural architecture search (NAS), a predefined backbone 402 is converted into a hyperkernel parameterization, where all searchable operator configurations are analyzed on the device for real metrics. The predefined backbone 402 is the input for NAS. During the OSFA search, a lookup table (LUT) 406 is used to quickly estimate the hardware-aware loss. A search space 404 is defined for different network types, such as convolution (Conv), gated recurrent unit (GRU), and multi-level perceptron (MLP). The skip operation (SkipOP) process shown in the leftmost part (a) will be described later. Once a network architecture is selected (see the rightmost part (c)), the selected network architecture is fully trained and deployed for downstream tasks.
[0060] Construct an over-parameterized hypernetwork 408 with all possible architecture candidates. Only the maximum hyperkernel of each layer needs to be constructed to encode all candidate choices in the search space 404, allowing different candidate choices in each layer to share hyperkernel weights. The NAS problem is solved by finding which kernel weight subset to use in each layer. To this end, two different search strategies can be adopted: progressive growth and progressive reduction to encode architectural decisions during the search, as seen in the middle part (b). In the middle part (b), different pattern types represent different configurations. For example, the pattern in the hypernetwork 408 can represent a 7x7 configuration, while boxes 412 and 414 can represent 5x5 and 3x3 configurations, respectively, and the pattern in box 416 can represent a long short-term memory (LSTM) layer. Instead of targeting only one final architecture, the hypernetwork 408 is guided to evolve from the maximum hyperkernel encoding to the minimum hyperkernel encoding through gradient descent. A set of network architecture decisions are made along the search trajectory, resulting in many different architectures with various model sizes. Finally, a target model 410 that fits the hardware requirements is sampled for full training, as shown in the rightmost part (c). The target model 410 is then quantized and deployed on the target device.
[0061] Another challenge of deploying the searched model on small hardware devices with tight memory budgets arises when memory usage from large activation tensors becomes a bottleneck. For example, a model with 2.3 million parameters may require more than 4.8MB of peak memory usage. To address this issue, in the case of depth-first scheduling, the convolutional node activations can be split into multiple tiles to reduce the peak memory usage during runtime. Manually designing node splits layer by layer and case by case is inefficient and difficult. Therefore, aspects of the present disclosure further expand the one-shot differentiable NAS search space to automate the process of node splitting. This architectural co-design significantly reduces peak memory usage and allows models to be deployed on small hardware devices without activation bottlenecks.
[0062] Aspects of the present disclosure are directed to a general one-shot differentiable NAS algorithm with hyperkernel weight sharing for convolutional, recurrent, MLP, and transformer architectures. Furthermore, a progressive growing and shrinking search strategy is introduced to discover a set of hardware efficient architectures within a few GPU hours. Activation tensor partitioning is exploited to augment the search space to accommodate large, otherwise irreducible intermediate outputs within strict peak memory constraints on embedded devices.
[0063] For the search space, a macroarchitecture backbone is given as a starting point. The backbone consists of many different layers, such as convolution, depthwise separable convolution, RNN, and fully connected layers. Figure 5 As illustrated, the search space is predefined for each operation type. By doing so, only micro-architectures within each layer are searched.
[0064] Figure 5 is a diagram illustrating a heterogeneous operator-level search space according to various aspects of the present disclosure. Figure 5 In the example of FIG. 5 , the macro architecture 502 includes a convolutional layer 504, a GRU 506, and a fully connected layer 508. For the convolutional layer 504, the expansion ratio and the kernel size are searched. For the ring-based GRU 506 and the column-based fully connected layer 508, the hidden unit size is searched. An example of a heterogeneous operator-level search space for OSFA with convolutional, GRU, and fully connected layers will now be described.
[0065] As seen in portion 510, convolutional layer 504 may be composed of k h ×k w ,i,n,parameterization, respectively represents the spatial filter size (height and width), input channels and output channels. If the baseline architecture uses a square kernel, then let k h =k w = k. For the Moving Inverted Bottleneck (MIB) block and depthwise separable convolutions, the search is performed on the inverted bottleneck expansion ratio e instead of the fixed number of output channels. When multiple channels are defined, the skip operation indicates which channels can be skipped. More details on kernel sharing are given in Figure 7 to describe.
[0066] As seen in section 512, the RNN layer (corresponding to GRU 506) can be parameterized by the input feature size m and the hidden unit size k of the kernel / recurrent kernel. The search is performed on the hidden unit size k of the kernel / recurrent kernel. For a bidirectional layer with multiple gates, the forward kernel f is searched for each gate i. i and the backward kernel b i The different modes for the hyperkernel and hyperloop kernel seen in section 512 represent how weights are shared for the GRU 506 and refer to Figure 7 Describe in more detail.
[0067] As seen in portion 514, the MLP layer (corresponding to the fully connected layer 508) may be parameterized by the input feature size m and by the hidden unit size k. The search is performed on the hidden unit size k.
[0068] For the transformer, the Bidirectional Encoder Representations from Transformers (BERT) encoder model can be selected as the backbone architecture for the search. The model consists of a projection scaled dot product attention layer and a two-layer feed-forward network (FFN). The search is performed on the MLP hidden unit size k in the projected attention (allowing the dot product to be computed in a lower dimension) and the MLP expanded hidden dimension in the FFN. In addition, a special skip operation is included in the search space, which skips the layer and feeds the input directly to the output.
[0069] As mentioned above, activation node splitting is introduced. Many convolutional networks applied to tasks such as computer vision create large intermediate activations. To ensure that the peak memory from these activations fits on the device, convolution node splitting is incorporated into the search space and searched jointly with the kernel weights. Figure 6 As shown, the idea is to split each input tensor into multiple smaller tensors.
[0070] Figure 6 is a diagram illustrating node partitioning with architectural co-design according to aspects of the present disclosure. Figure 6 In the example of , the input tensor 602 can be split into multiple smaller tensors (e.g., nodes 604, 606, see the leftmost part (a)). The symbols {(h,"w"),h∈{1,2,…}",w∈"{1,2,…}"} represent the corresponding number of height and width splits. The split selection can be encoded into the architecture as a path selection problem, as shown in the rightmost part (b). In Figure 6 In the example of the solution selected, the (1,2) option is selected (corresponding to the architecture parameter α), where the two nodes 604, 606 are side by side. The other options (single node (1,1), sequential nodes (2,1), or four nodes (2,2)) are not selected.
[0071] We now discuss weight and architecture co-design. Aspects of the present disclosure relate to weight sharing mechanisms. Under a defined search space, a parameterized hyperkernel is constructed for each operation in the graph, where different candidates are subsets of shared weights. The NAS problem is reduced to finding which subset of kernel weights to use in each layer, such as Figure 7 shown.
[0072] Figure 7 is a diagram illustrating a weight sharing technique in a hyperkernel according to aspects of the present disclosure. Figure 7In the example of , the maximum super kernel based on the search space is represented as the outer ring 702. The less important weights in the outer ring 702 can be pruned. The middle ring 704 and the center ring 706 correspond to smaller subsets of the outer ring 702. Similarly, in the column-based solution, the outer column 712 corresponds to the maximum super kernel, and the middle column 714 and the center column 716 correspond to smaller subsets of the outer column 712.
[0073] For 2D spatial operations, the kernels are ring-based, such as convolutional and recurrent kernels. For linear projection layers, the kernels are column-based, such as those found in RNN inputs, fully connected layers, or feed-forward networks.
[0074] As an illustration of the general approach, Figure 7 , where k_3 represents the maximum kernel size shown in the outer loop 702 in the search space. Taking the gated recurrent unit (GRU) as an example, two super kernels are shown in the outer loop 702 and the outer column 712: the recurrent kernel R and the input kernel W are constructed for each layer. Figure 7 As shown, the weights of R_(k_1×k_1) of the middle ring 704 and R_(k_2×k_2) of the center ring 706 can be regarded as the inner part of the weights of R_(k_3×k_3) of the outer ring 702. The outer ring 702, represented as the weight of R_(k_3×k_3\k_2×k_2), does not contribute to the output of R_(k_2×k_2) of the middle ring 704, but only contributes to R_(k_3×k_3) of the outer ring 702. Similarly, for column-based kernel sharing, the kernel W of each gate is an m×k_3 matrix, where m is the input feature size. Unlike R, the weights in W are shared in column-based kernel sharing. The loop kernels Rn or input kernels Wn of the three gates will be connected, respectively, resulting in a loop kernel R of shape k_3×3·k_3 and an input kernel W of shape m×3·k_3. Note that all three gates share the same kernel selection. In the case of bidirectional RNNs, the forward and backward units are searched separately. For convolutional kernels, the weights are expanded to be shared in the output channel dimension.
[0075] Now refer to Figure 8 and Fig. 9 Discussion of progressive growing and shrinking. To encode NAS decisions into the hyperkernel, two separate search strategies are introduced: progressive growing and progressive shrinking.
[0076] Figure 8 is a diagram illustrating progressive growing according to aspects of the present disclosure. For progressive growing, the hyperkernel is defined as:
[0077]
[0078] in is an indicator function that encodes the architectural choice of using kernel k i . This indicator is approximated by a sigmoid function in the backward pass for differentiation.
[0079] For example, if and but Indication k 2 has been chosen as the kernel size. The selected kernel is chosen to grow outward from the smallest inner core (e.g., the center ring 806 for ring-based progressive growth and the left column 816 for column-based progressive growth). In this case, if The kernel stops from Growth, with indicators Not relevant.
[0080] In contrast, in progressive shrinking, the hyperkernel is defined as:
[0081]
[0082] Fig. 9 is a diagram illustrating a progressive reduction according to various aspects of the present disclosure. Fig. 9 As seen in the example of , the selected kernel selection is always reduced from the largest superkernel 902, 912 to its inner core (eg, the center ring 906 and the left column 916 in the ring-based progressive reduction and the column-based progressive reduction).
[0083] For example, if and but Indication k 1 has been selected as the kernel choice. If Further shrinking to 0, R k becomes 0, which means that the operation will be skipped completely. For example, if the outer loop of the maximum hyperkernel 902 is selected, such as The kernel stops Zoom out, with internal indicators such as Not relevant.
[0084] In general, progressive growing prefers smaller kernel sizes because it only grows when the outer loop becomes important, allowing faster convergence during the search. In contrast, progressive shrinking prefers larger kernel sizes during the initial search phase because it only shrinks when the outer loop becomes less important, resulting in a more extensive exploration of the search space.
[0085] Now we will discuss the weight decision. In the formula, the indicator function determines the kernel selection. The decision is made by comparing the corresponding norm of the weight with a trainable threshold t. Take the following as an example:
[0086]
[0087] The threshold t k=k3 is learned during search time to control the kernel decision. For a convolutional neural network (CNN), the kernel also extends the indicator through a control channel according to Equation 3 With t e Shared in the output channel dimension. In the transformer architecture, column-based kernel sharing is used and the norm of the subkernel is further divided by the norm of the superkernel to accommodate its large dimensionality. In an RNN with multiple gates in each layer (such as a GRU), each layer makes a decision using only the average kernel norm from each gate, as used in Equation 3.
[0088] For architecture encoding, the search space is enlarged to encode the node partitioning of the activation tensor. The architecture parameter α encodes the node partitioning choice for the convolution operation, as Figure 8 As shown in the column-based progressive growth of . Let h = {1, 2, ...}, w = {1, 2, ...} denote the number of splits in each dimension, allowing the creation of |h| × |w| paths for each convolution in the hypernetwork. Each path is associated with a learnable parameter a i,j , where i∈h,j∈w are associated. We formulate the split selection as a path selection problem. Let {(i,j):i∈h,j∈w} be the set of candidate paths, where each candidate path represents a split applied to the input tensor x. To make the search differentiable, we relax the path selection to a softmax over all possible splits:
[0089]
[0090] Among them i,j (x,w) represents the output memory usage of the split from (i,j) in the height and width dimensions. At the end of the search, the partition with the highest a is selected i,j Path.
[0091] The total memory usage during inference consists of three parts: the working set buffers for input and output activation tensors, f s , metadata for graphical representation f m and a persistent memory f for weight storage p , as calculated by Formula 5:
[0092]
[0093] Since CNNs are built on increasing receptive fields, multiple layers of sequential convolutional activations can be split by backpropagating the split selection to earlier convolutional blocks to ensure that they share the same number of splits, adding overlapping tensor data to each split when the receptive field requires it. Figure 6 The depth-first schedule shown, however, the amount of memory required to store the intermediate tensors decreases significantly as the number of partitions increases. This provides a tradeoff from ensuring that the redundant tensor data of the receptive fields remain equal. Therefore, architectural co-design search can jointly search parameters and partitions.
[0094] To search for hardware efficient networks, the model performance of the search architecture including loss objectives, inference latency on the target hardware, and peak memory usage is jointly optimized. Therefore, the multi-objective loss function is defined as:
[0095] l(w|t k ,a i,j )=l acc +λ 1 ·l ms +λ 2 ·l mem , (6)
[0096] Among them l acc is the standard loss for the task (e.g., cross entropy), l ms is the inference latency (in milliseconds) of running the model on the device, and l mem Defined as the maximum working set memory required to store the input and output tensors on all layers, as well as the overall metadata and persistent memory. The total runtime latency is further defined as the sum of the runtime of each layer across the network.
[0097]
[0098] where t k ,t e is the tunable threshold parameter, and a is the reference Figure 6 Architecture parameters discussed. Although the overall loss function is described with respect to accuracy, memory, and latency, other hardware resources such as power consumption may also be considered.
[0099] To avoid having to reanalyze the model's flaws on the hardware each time, operations with various configurations in kernel space are pre-analyzed with associated runtime delays and stored in a lookup table (LUT) (see Figure 4 ). Layer runtimes for kernel selection and architecture space can be retrieved directly from the LUT, allowing fast and accurate metrics.
[0100] The network storage loss term l in Formula 6mem Defined as the maximum working set memory required to store input and output tensors along with overall metadata and persistent memory across all layers, it can be further bounded given a memory budget:
[0101]
[0102] l mem =max{l′ mem -m budget ,0} (9)
[0103] where the layer-by-layer memory usage parameterized by the weight and architecture encoding is calculated based on Equations 4 and 5, corresponds to temporary storage for storing activity data, Corresponding to the metadata storage, corresponds to the persistent memory used to store activations, and L corresponds to the number of layers in the network.
[0104] The loss hyperparameter λ in Equation 6 1 and λ 2 Affects the trade-off between model accuracy, runtime latency, and memory usage. For example, a larger loss hyperparameter λ 1 Values favor higher runtime efficiency, resulting in smaller model architectures after NAS. In progressive shrinking, this happens by shrinking the model architecture from the largest hyperkernel to the smallest hyperkernel. During the search, decisions are made to progressively shrink the architecture, converging to the smallest architecture in the search space by minimizing the overall loss. This one-shot search trajectory results in a set of model architectures with different model sizes and varying inference latency. Any architecture that fits the hardware requirements is then fully trained in a second phase for better accuracy.
[0105] Since the sub-kernels directly correspond to the super-kernels, knowledge distillation (KD) can be applied layer by layer during and after the search phase. Fig.10 is a block diagram illustrating knowledge distillation in neural architecture search (NAS) according to aspects of the present disclosure. Fig.10 In the example, the input runs through each layer twice to compute the superkernel and subkernel (masked superkernel) activations. In addition to the output losses from the supernetwork (teacher) and subnetwork (student), the block-wise mean squared error (MSE) loss between the supernetwork and subnetwork also leads to the cross entropy l defined in Equation 6 acc :
[0106]
[0107] in corresponds to the teacher hypernetwork cross entropy, and Corresponds to the cross entropy of the student subnetwork.
[0108] The disclosed techniques maintain accuracy while reducing memory and latency compared to baselines. Using convolutional networks, activation node segmentation reduces peak memory for tasks such as image classification. Furthermore, a strategy of progressive growth and shrinking allows the discovery of search paths for models where the desired trade-off between latency and memory can be chosen if both are analyzed. On a proxy parameter count metric such as BERT, knowledge distillation using kernel substructures ensures that the searched architecture maintains test loss while being able to outperform a manually compressed baseline.
[0109] Fig.11 1 is a flow chart illustrating an example processor-implemented method 1100 for searching a neural network architecture in accordance with aspects of the present disclosure. Fig.11 As shown, in some aspects, the processor-implemented method 1100 may include generating an over-parameterized hypernetwork having a plurality of layers. The hypernetwork includes a plurality of operator types. Each of the layers includes a maximum hyperkernel corresponding to a search space (block 1102).
[0110] In some aspects, the processor-implemented method 1100 may include performing gradient descent to evolve the largest super kernel into a small kernel corresponding to the search space to generate a series of kernel encodings (block 1104). In some aspects, the processor-implemented method 1100 may include identifying a subset of kernel encodings from the series of kernel encodings for each layer of the super network based on the gradient descent (block 1106). In some aspects, the processor-implemented method 1100 may include determining a set of candidate architectures based on the subset of kernel encodings. Each candidate architecture has a different model size (block 1108). In some aspects, the processor-implemented method 1100 may include selecting a target model from the set of architectures based on satisfying hardware specifications (block 1110). In some aspects, the processor-implemented method 400 may include applying the target model (block 1112).
[0111] Example aspects
[0112] Aspect 1: A processor-implemented method, comprising: generating an over-parameterized hypernetwork having multiple layers, the hypernetwork comprising multiple operator types, each of the multiple layers comprising a maximum hyperkernel corresponding to a search space; performing gradient descent to evolve the maximum hyperkernel into a small kernel corresponding to the search space to generate a series of kernel encodings; based on the gradient descent, for each layer of the hypernetwork, identifying a subset of kernel encodings from the series of kernel encodings; determining a set of candidate architectures based on the subset of kernel encodings, each candidate architecture in the set of candidate architectures having a different model size; selecting a target model from the set of architectures based on satisfying hardware specifications; and applying the target model.
[0113] Aspect 2: The method according to aspect 1 further comprises quantizing the target model.
[0114] Aspect 3: The method according to Aspect 1 or 2 further includes splitting the convolution node activations of the target model by splitting each input tensor into multiple smaller tensors, and the multiple smaller tensors are incorporated into the set of candidate architectures.
[0115] Aspect 4: A method according to any one of the preceding aspects, wherein the multiple operation types include at least one convolutional neural network, at least one recurrent neural network, at least one transformer, and at least one multilayer perceptron (MLP).
[0116] Aspect 5: The method according to any one of the preceding aspects, wherein the super network is based on a specific use case, and a different super network is constructed for each different use case.
[0117] Aspect 6: The method according to any of the preceding aspects, wherein identifying the kernel encoding subset comprises gradually increasing the selected kernel selection from the smallest inner core of the largest super kernel to the largest super kernel.
[0118] Aspect 7: The method according to any one of aspects 1 to 5, wherein identifying the kernel encoding subset comprises gradually reducing the selected kernel selection from a largest super kernel to an inner core of the largest super kernel.
[0119] Aspect 8: The method according to any one of the preceding aspects, wherein applying the target model comprises training the target model.
[0120] Aspect 9: A method according to any one of the preceding aspects, wherein the method comprises a one-shot neural architecture search (NAS).
[0121] Aspect 10: The method according to aspect 9, wherein the one-time NAS is based on memory usage of a device to which the target model is applied, runtime latency of the model running on the device, and model accuracy.
[0122] Aspect 11: A device comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory, the at least one processor being configured to: generate an over-parameterized hypernetwork having multiple layers, the hypernetwork comprising multiple operator types, each of the multiple layers comprising a maximum hyperkernel corresponding to a search space; perform gradient descent to evolve the maximum hyperkernel into a small kernel corresponding to the search space to generate a series of kernel encodings; based on the gradient descent, for each layer of the hypernetwork, identify a kernel encoding subset from the series of kernel encodings; determine a set of candidate architectures based on the kernel encoding subset, each candidate architecture in the set of candidate architectures having a different model size; select a target model from the set of architectures based on meeting hardware specifications; and apply the target model.
[0123] Aspect 12: The apparatus according to Aspect 11, wherein the at least one processor is further configured to quantize the target model.
[0124] Aspect 13: An apparatus according to Aspect 11 or 12, wherein the at least one processor is further configured to split the convolution node activations of the target model by splitting each input tensor into multiple smaller tensors, and the multiple smaller tensors are incorporated into the set of candidate architectures.
[0125] Aspect 14: An apparatus according to any one of Aspects 11 to 13, wherein the multiple operation types include at least one convolutional neural network, at least one recurrent neural network, at least one transformer, and at least one multilayer perceptron (MLP).
[0126] Aspect 15: The apparatus according to any one of Aspects 11 to 14, wherein the super network is based on a specific use case, and a different super network is constructed for each different use case.
[0127] Aspect 16: The apparatus according to any one of aspects 11 to 15, wherein the at least one processor is further configured to gradually increase the selected kernel selection from the smallest inner core of the largest super kernel to the largest super kernel.
[0128] Aspect 17: The apparatus according to any one of aspects 11 to 15, wherein the at least one processor is further configured to gradually reduce the selected kernel selection from the largest super kernel to the internal cores of the largest super kernel.
[0129] Aspect 18: The apparatus according to any one of Aspects 11 to 17, wherein the at least one processor is further configured to train the target model.
[0130] Aspect 19: An apparatus according to any one of aspects 11 to 18, wherein the apparatus is configured to perform a one-shot neural architecture search (NAS).
[0131] Aspect 20: An apparatus according to any one of aspects 11 to 19, wherein the one-time NAS is based on memory usage of a device to which the target model is applied, runtime latency of the model running on the device, and model accuracy.
[0132] Aspect 21: A device comprising: a component for generating an over-parameterized hypernetwork having multiple layers, the hypernetwork including multiple operator types, each of the multiple layers including a maximum hyperkernel corresponding to a search space; a component for performing gradient descent to evolve the maximum hyperkernel into a small kernel corresponding to the search space to generate a series of kernel encodings; a component for identifying a subset of kernel encodings from the series of kernel encodings for each layer of the hypernetwork based on the gradient descent; a component for determining a set of candidate architectures based on the kernel encoding subsets, each candidate architecture in the set of candidate architectures having a different model size; a component for selecting a target model from the set of architectures based on satisfying hardware specifications; and a component for applying the target model.
[0133] Aspect 22: The apparatus according to Aspect 21, further comprising a component for quantizing the target model.
[0134] Aspect 23: The apparatus according to Aspect 21 or 22 further includes a component for splitting the convolution node activations of the target model by splitting each input tensor into multiple smaller tensors, and the multiple smaller tensors are incorporated into the set of candidate architectures.
[0135] Aspect 24: An apparatus according to any one of Aspects 21 to 23, wherein the multiple operation types include at least one convolutional neural network, at least one recurrent neural network, at least one transformer, and at least one multilayer perceptron (MLP).
[0136] Aspect 25: An apparatus according to any one of Aspects 21 to 24, wherein the super network is based on a specific use case, and a different super network is constructed for each different use case.
[0137] Aspect 26: The apparatus according to any one of aspects 21 to 25, wherein the means for identifying the kernel encoding subset includes means for gradually increasing the selected kernel selection from the smallest inner core of the largest super kernel to the largest super kernel.
[0138] Aspect 27: The apparatus according to any one of aspects 21 to 25, wherein the means for identifying the kernel encoding subset includes means for gradually reducing the selected kernel selection from the largest super kernel to an inner core of the largest super kernel.
[0139] Aspect 28: The apparatus according to any one of Aspects 21 to 27, wherein the means for applying the target model comprises means for training the target model.
[0140] Aspect 29: An apparatus according to any one of aspects 21 to 28, wherein the apparatus is configured to perform a one-shot neural architecture search (NAS).
[0141] Aspect 30: An apparatus according to any one of aspects 21 to 29, wherein the one-time NAS is based on memory usage of a device to which the target model is applied, runtime latency of the model running on the device, and model accuracy.
[0142] The various operations of the above methods may be performed by any suitable components capable of performing the corresponding functions. These components may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. In general, where operations are illustrated in the accompanying drawings, these operations may have corresponding paired components plus function components with similar numbers.
[0143] As used, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or another data structure), ascertaining, etc. Additionally, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determining" may include resolving, selecting, choosing, establishing, etc.
[0144] As used, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc.
[0145] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described. A general purpose processor may be a microprocessor, but in an alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.
[0146] The steps or algorithms of the methods described in conjunction with the present disclosure may be directly embodied in hardware, software modules executed by a processor, or a combination of the two. The software module may reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software module may include a single instruction, perhaps multiple instructions, and may be distributed over several different code segments, distributed between different programs, and distributed across multiple storage media. The storage medium may be coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integral with the processor.
[0147] The disclosed methods include one or more steps or actions for implementing the described methods. The steps and / or actions of the methods may be interchangeable with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0148] The functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link various circuits together, including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect a network adapter, etc., to the processing system via the bus. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (e.g., a keypad, a display, a mouse, a joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and will not be described further.
[0149] The processor may be responsible for managing the bus and general processing, including executing software stored on the machine-readable medium. The processor may be implemented using one or more general-purpose processors and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuits that can execute software. Software should be broadly interpreted as meaning instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or other. By way of example, a machine-readable medium may include a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a disk, an optical disk, a hard disk drive, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging materials.
[0150] In hardware implementation, machine-readable media can be part of a processing system separate from a processor. However, as will be readily appreciated by those skilled in the art, machine-readable media or any part thereof can be outside a processing system. By way of example, machine-readable media can include a transmission line, a carrier modulated by data, and / or a computer product separate from a device, all of which can be accessed by a processor through a bus interface. Alternatively or in addition, machine-readable media or any part thereof can be integrated into a processor, such as with a cache and / or a general register stack. Although the various components discussed can be described as having specific locations, such as local components, they can also be configured in various ways, such as certain components being configured as a part of a distributed computing system.
[0151] The processing system can be configured as a general processing system having one or more microprocessors providing processor functionality and an external memory providing at least a portion of a machine-readable medium, all of which are linked together with other support circuits through an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron model and neural system model. As another alternative, the processing system may be implemented with an application specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, support circuits, and at least a portion of a machine-readable medium integrated in a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuits, or circuits capable of performing the various functionalities described throughout this disclosure. Those skilled in the art will recognize how to best implement the functionality of the processing system depending on the specific application and the overall design constraints imposed on the entire system.
[0152] The machine-readable medium may include multiple software modules. These software modules include instructions that cause the processing system to perform various functions when executed by the processor. The software module may include a sending module and a receiving module. Each software module may reside in a single storage device or be distributed across multiple storage devices. By way of example, when a triggering event occurs, the software module may be loaded from a hard drive into a RAM. During the execution of the software module, the processor may load some of the instructions into a cache to increase access speed. Then one or more cache lines may be loaded into a general register stack for execution by the processor. When the functionality of a software module is mentioned below, it will be understood that such functionality is implemented by the processor when executing instructions from the software module. In addition, it should be understood that various aspects of the present disclosure produce improvements in the functionality of processors, computers, machines, or other systems that implement such aspects.
[0153] If implemented in software, each function may be stored on or sent through a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any media that facilitates the transfer of computer programs from one place to another. Storage media can be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, disk storage devices or other magnetic storage devices, or any other media that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is also appropriately referred to as a computer-readable medium. For example, if the software is sent from a website, server, or other remote source using a coaxial cable, optical cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwaves, the coaxial cable, optical cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. As used, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and optical disks. Optical disks, where magnetic disks typically reproduce data magnetically, and optical disks reproduce data optically with lasers. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0154] Thus, some aspects may include a computer program product for performing the operations presented. For example, such a computer program product may include a computer-readable medium having stored (and / or encoded) thereon instructions that can be executed by one or more processors to perform the described operations. For some aspects, the computer program product may include packaging materials.
[0155] In addition, it should be understood that the modules and / or other appropriate components for performing the described methods and techniques can be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such a device can be coupled to a server to facilitate the transmission of the components for performing the described methods. Alternatively, the various methods described can be provided via a storage component (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or a floppy disk) so that once the storage component is coupled to or provided to the device, the user terminal and / or base station can obtain the various methods. In addition, any other suitable technology suitable for providing the described methods and techniques to the device can be used.
[0156] It is to be understood that the claims are not limited to the precise configuration and components illustrated above. Various modifications, changes and variations may be made in the arrangement, operation and details of the methods and apparatus described above without departing from the scope of the claims.
Claims
1. A processor-implemented method, include: generating an over-parameterized hypernetwork having a plurality of layers, the hypernetwork comprising a plurality of operator types, each of the plurality of layers comprising a maximum hyperkernel corresponding to the search space; performing gradient descent to evolve the maximum superkernel into small kernels corresponding to the search space to generate a series of kernel codes; Based on the gradient descent, for each layer of the super network, identifying a subset of kernel codes from the series of kernel codes; determining a set of candidate architectures based on the kernel encoding subset, each candidate architecture in the set of candidate architectures having a different model size; selecting a target model from the set of architectures based on satisfying the hardware specification; as well as The target model is applied. The method according to claim 1 , further comprising quantizing the target model.
3. The method of claim 1 , further comprising splitting convolutional node activations of the target model by splitting each input tensor into a plurality of smaller tensors, the plurality of smaller tensors being incorporated into the set of candidate architectures.
4. The method of claim 1, wherein the plurality of operation types comprises at least one convolutional neural network, at least one recurrent neural network, at least one transformer and / or at least one multi-layer perceptron (MLP).
5. The method according to claim 1, wherein the super network is based on a specific use case, and a different super network is constructed for each different use case.
6. The method of claim 1, wherein identifying the kernel encoding subset comprises gradually increasing the selected kernel selection from a smallest inner core of the largest superkernel to the largest superkernel.
7. The method of claim 1, wherein identifying the kernel encoding subset comprises gradually reducing the selected kernel selection from the largest superkernel to an inner core of the largest superkernel. The method of claim 1 , wherein applying the target model comprises training the target model.
9. The method of claim 1, wherein the method comprises a one-shot neural architecture search (NAS).
10. The method of claim 9, wherein the one-time NAS is based on at least one of: memory usage of a device to which the target model is applied, runtime latency of the model running on the device, model accuracy, and power consumption of the device.
11. A device, include: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory, the at least one processor configured to: generating an over-parameterized hypernetwork having a plurality of layers, the hypernetwork comprising a plurality of operator types, each of the plurality of layers comprising a maximum hyperkernel corresponding to the search space; performing gradient descent to evolve the maximum superkernel into small kernels corresponding to the search space to generate a series of kernel codes; Based on the gradient descent, for each layer of the super network, identifying a subset of kernel codes from the series of kernel codes; determining a set of candidate architectures based on the kernel encoding subset, each candidate architecture in the set of candidate architectures having a different model size; selecting a target model from the set of architectures based on satisfying the hardware specification; as well as The target model is applied.
12. The apparatus of claim 11, wherein the at least one processor is further configured to quantize the target model.
13. The apparatus of claim 11, wherein the at least one processor is further configured to split convolutional node activations of the target model by splitting each input tensor into a plurality of smaller tensors, the plurality of smaller tensors being incorporated into the set of candidate architectures.
14. The apparatus of claim 11, wherein the plurality of operation types comprises at least one convolutional neural network, at least one recurrent neural network, at least one transformer and / or at least one multi-layer perceptron (MLP).
15. The apparatus according to claim 11, wherein the super network is based on a specific use case, and a different super network is constructed for each different use case.
16. The apparatus of claim 11, wherein the at least one processor is further configured to gradually increase the selected kernel selection from a smallest inner core of the largest super kernel to the largest super kernel.
17. The apparatus of claim 11, wherein the at least one processor is further configured to gradually reduce the selected kernel selection from the largest super kernel to an internal core of the largest super kernel.
18. The apparatus of claim 11, wherein the at least one processor is further configured to train the target model.
19. The apparatus of claim 11, wherein the apparatus is configured to perform a one-shot neural architecture search (NAS).
20. The apparatus of claim 19, wherein the one-time NAS is based on at least one of: memory usage of a device to which the target model is applied, runtime latency of the model running on the device, model accuracy, and power consumption of the device.
21. A device, include: means for generating an over-parameterized hypernetwork having a plurality of layers, the hypernetwork comprising a plurality of operator types, each of the plurality of layers comprising a maximal hyperkernel corresponding to a search space; means for performing gradient descent to evolve the maximum superkernel into small kernels corresponding to the search space to generate a series of kernel encodings; means for identifying a subset of kernel codes from the series of kernel codes for each layer of the supernetwork based on the gradient descent; means for determining a set of candidate architectures based on the kernel encoding subset, each candidate architecture in the set of candidate architectures having a different model size; means for selecting a target model from the set of architectures based on satisfying hardware specifications; and A component for applying the target model.
22. The apparatus of claim 21, further comprising means for quantizing the target model.
23. The apparatus of claim 21, further comprising means for splitting convolutional node activations of the target model by splitting each input tensor into a plurality of smaller tensors, the plurality of smaller tensors being incorporated into the set of candidate architectures.
24. The apparatus of claim 21, wherein the plurality of operation types comprises at least one convolutional neural network, at least one recurrent neural network, at least one transformer and / or at least one multi-layer perceptron (MLP).
25. The apparatus according to claim 21, wherein the super network is based on a specific use case, and a different super network is constructed for each different use case.
26. The apparatus of claim 21, wherein the means for identifying the kernel encoding subset comprises means for gradually increasing the selected kernel selection from a smallest inner core of the largest superkernel to the largest superkernel.
27. The apparatus of claim 21, wherein the means for identifying the kernel encoding subset comprises means for gradually reducing the selected kernel selection from the largest superkernel to an inner core of the largest superkernel.
28. The apparatus of claim 21, wherein the means for applying the target model comprises means for training the target model.
29. The apparatus of claim 21, wherein the apparatus is configured to perform a one-shot neural architecture search (NAS).
30. The apparatus of claim 29, wherein the one-time NAS is based on at least one of: memory usage of a device to which the target model is applied, runtime latency of the model running on the device, model accuracy, and power consumption of the device.