Neural Architecture Search with Dynamic Time Budgets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional neural architecture search methods are limited in optimizing complex neural networks for machine learning tasks, as they often rely on fixed computational steps and do not adequately consider the efficiency of training and inference processes, leading to suboptimal performance and computational inefficiencies.

Innovation Solution

The system employs a novel approach to neural architecture search by optimizing complex layer blocks that can process input token sequences in parallel, comparing architectures based on training time and processing efficiency, and using a constrained search space to achieve more complex and efficient neural architectures, resulting in faster training convergence and improved generalization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional neural architecture search methods are used with fixed computational steps, then the search process is simple to implement, but the training efficiency and inference efficiency are suboptimal

Engineering Contradiction:
Improvetraining efficiencyVSAvoidarchitecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transforms the fixed computational steps into dynamic training time budgets. Each architecture is evaluated based on actual training time consumed rather than fixed step counts, allowing the system to adapt the evaluation process to the specific characteristics of each neural network architecture. This dynamic approach enables more accurate comparison of training efficiency across different architectures while maintaining a manageable search process.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the evaluation parameter from fixed computational steps to time-based metrics (training time and inference time). This parameter transformation allows the system to optimize for real-world performance metrics that directly reflect efficiency, resolving the contradiction between simple implementation and optimal efficiency by redefining how architectures are compared.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If complex layer blocks with diverse layers are used, then the neural network achieves better performance and efficiency, but the architecture becomes more difficult to optimize

Engineering Contradiction:
Improveinference efficiencyVSAvoidoptimization difficulty
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the complex neural network into standardized layer blocks that can be independently configured and evaluated. Each layer block contains diverse layers (attention layers, feed-forward layers, normalization layers) but maintains a consistent interface and evaluation protocol. This segmentation allows the optimization system to manage complexity by evaluating blocks independently while still achieving complex overall architectures, resolving the contradiction between performance and optimization difficulty.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses replicated layer blocks throughout the network architecture. Instead of optimizing each layer individually, the system optimizes a single layer block template and replicates it multiple times. This copying approach dramatically reduces optimization difficulty while maintaining the ability to achieve complex performance characteristics through the repeated application of the optimized block pattern.

Inventive Principle:
Principle #26Copying

3Loss of time

If architecture evaluation is based on fixed computational steps, then the comparison process is straightforward, but the computational efficiency is not adequately captured

Engineering Contradiction:
Improvetraining timeVSAvoidevaluation simplicity
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The patent replaces the mechanical counting of computational steps with time-based measurement systems. Instead of tracking discrete operations, the system measures actual wall-clock time for training and inference processes. This substitution provides more accurate capture of computational efficiency while the automated timing infrastructure maintains evaluation simplicity, resolving the contradiction between time capture accuracy and evaluation ease.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240112027A1Neural network architecture search over complex block architectures
Publication Date: 2024.04.04 GOOGLE LLC
  • US20240112027A1 patent drawing
  • US20240112027A1 patent drawing
  • US20240112027A1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for performing neural architecture search for machine learning models. In one aspect, a method comprises receiving training data for a machine learning, generating a plurality of candidate neural networks for performing the machine learning task, wherein each candidate neural network comprises a plurality of instances of a layer block composed of a plurality of layers, for each candidate neural network, selecting a respective type for each of the plurality of layers from a set of layer types that comprises, training the candidate neural network and evaluating performance scores for the trained candidate neural networks as applied to the machine learning task, and determining a final neural network for performing the machine learning task based at least on the performance scores for the candidate neural networks.