Methods, systems, and apparatus for jointly determining
neural network architecture and hardware accelerator architecture, including computer programs encoded on a computer storage medium. In one aspect, a method includes: generating a batch of one or more output sequences using a controller policy, each output sequence in the batch defining a corresponding architecture of a sub-neural network and a corresponding architecture of a hardware accelerator; for each output sequence in the batch: training a corresponding instance of the sub-neural network having the architecture defined by the output sequence; evaluating the
network performance of the trained instance of the sub-neural network; and evaluating the accelerator performance of a corresponding instance of the hardware accelerator having the architecture defined by the output sequence to determine an accelerator performance metric for the instance of the hardware accelerator; and adjusting the controller policy using the
network performance metric and the accelerator performance metric.