This application belongs to the field of
computer technology and provides a task execution method, compilation method,
chip, and device for an
artificial intelligence model. The application first requires compiling the tasks to be processed by the
artificial intelligence model according to a preset compilation method, resulting in multiple sub-tasks. This allows each task to be processed to receive loading and scheduling from the NPU at a moderately granular scheduling unit. Specifically, this application uses the sub-tasks as the basic scheduling unit of the NPU, thereby achieving
resource scheduling at a moderate
granularity that balances
scheduling complexity and task switching flexibility. When encountering urgent tasks, the task switching cost of the NPU is significantly reduced; thus, in the event of sudden changes in the target application running on the
computer device, the response efficiency and
resource utilization of the NPU can be effectively improved.