用于机器学习工作负载的任务调度
By optimizing task scheduling and resource allocation based on NUMA topology in a distributed system and utilizing a shared hardware bus for data communication, the performance bottleneck caused by non-local memory access and data communication in distributed computing systems is solved, thereby improving computing efficiency.
CN114503077BActive Publication Date: 2026-07-17GOOGLE LLC
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2020-09-08
- Publication Date
- 2026-07-17
AI Technical Summary
Technical Problem
In existing distributed computing systems, bandwidth-intensive operations caused by non-local memory access and data communication result in performance bottlenecks, especially when operating across remote resources.
Method used
By using a resource-group-based Non-Uniform Memory Access (NUMA) topology, task scheduling and resource allocation are optimized, data communication is carried out using a shared hardware bus, non-local memory access and data communication are reduced, and resource location utilization is optimized.
Benefits of technology
It reduces the time spent processing machine learning workloads, lowers bandwidth requirements, and improves the performance of computing systems.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN114503077B_ABST
Abstract
描述了用于调度ML工作负载的任务的方法、系统和装置,包括在计算机存储介质上编码的计算机程序。系统接收请求以执行工作负载,并且基于请求来确定资源需求以执行工作负载。系统包括多个主机并且每个主机包括多个加速器。系统基于资源需求和每个主机的加速器来确定被分配来执行工作负载的任务的主机的数量。针对所述数量的主机中的每个主机,系统基于主机的存储器存取拓扑来生成任务规格。规格指定将在主机处使用包括多个加速器的主机的资源执行的任务。系统将任务规格提供给主机,并且在每个主机执行针对主机在任务规格中指定的所分配任务时执行工作负载。
Need to check novelty before this filing date? Find Prior Art