一种深度学习训练资源的自适应分配方法及架构

By monitoring and adaptively adjusting resource configuration in real time under the K8S cluster architecture, the problem of unreasonable resource allocation in deep learning training is solved, achieving efficient resource utilization and automated retry of training, thereby improving training efficiency.

CN116401046BActive Publication Date: 2026-07-17CHINA NANHU ACAD OF ELECTRONICS & INFORMATION TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA NANHU ACAD OF ELECTRONICS & INFORMATION TECH
Filing Date
2023-03-14
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

During deep learning training, existing technologies cannot effectively and dynamically adjust hardware resource allocation, leading to resource waste and training failures, which affects efficiency.

Method used

An adaptive resource allocation method and architecture are adopted. Through the K8S cluster architecture, resource usage is monitored in real time. The resource configuration is dynamically generated using decision recommendation algorithm and adaptive resource adjustment algorithm. When training fails, the resource configuration is automatically adjusted to retry training.

Benefits of technology

It improves resource utilization and training efficiency, reduces manual intervention, automates resource allocation, and ensures training success.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116401046B_ABST
    Figure CN116401046B_ABST
Patent Text Reader

Abstract

本发明公开了一种深度学习训练资源的自适应分配方法及架构,方法包括:根据用户输入或决策推荐算法确定训练的资源配置;实时记录训练中的资源监控数据,并监控训练结果;若训练失败,则根据记录的资源监控数据得到资源结果数据,并利用自适应资源调节算法针对训练失败的资源结果数据调整训练的资源配置,重新进行训练;若训练成功,则结束训练,并根据实时的资源监控数据得到资源结果数据并记录。本发明具有动态生成推荐配置与失败自动重试、自适应优化资源分配方案的功能。
Need to check novelty before this filing date? Find Prior Art