Device and method for adaptively adapting domestic AI accelerator cards to cloud platforms

Through the adaptive adaptation mechanism of the script tool set and the multi-dimensional feature weight algorithm, the problem of rapid integration and deployment of domestically produced AI accelerator cards on the cloud platform was solved, the flexibility and scalability of the cloud platform was improved, the operating performance was optimized, and the application of domestically produced AI technology was promoted.

CN119512835BInactive Publication Date: 2025-09-09SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411536138.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-09-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

AI accelerator cards produced by different manufacturers differ in hardware architecture, interface protocols, and performance parameters, making adaptation and management of cloud platforms difficult. It is also difficult to quickly integrate and deploy domestic AI accelerator cards, and unable to meet diverse AI computing needs.

Method used

By adopting a script tool set, resource management module, adaptation task management module and data storage module, and through modularization and version control strategies, the automatic identification, driver installation, status monitoring and task scheduling of domestic AI accelerator cards are realized. The multi-dimensional feature weight algorithm is combined to optimize resource utilization and form an adaptive adaptation mechanism.

Benefits of technology

It has achieved rapid access and deployment of domestic AI acceleration cards on the cloud platform, improved the flexibility and scalability of the cloud platform, optimized operating performance, reduced deployment time and costs caused by compatibility issues, and promoted the application and development of domestic AI technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119512835B_ABST
    Figure CN119512835B_ABST
Patent Text Reader

Abstract

The present invention discloses a device and method for adaptively adapting a cloud platform to a domestically produced AI accelerator card, belonging to the technical fields of cloud computing and artificial intelligence. The technical problem to be solved by the present invention is how to realize the rapid integration and deployment of domestically produced AI accelerator cards on the cloud platform, improve the flexibility and scalability of the cloud platform, and thus better meet the diversified AI computing needs. The technical solution adopted is: the device includes a script tool set, a resource management module, an adaptation task management module, a data storage module, and an adaptation analysis module; wherein the script tool set integrates script tools for identifying, perceiving, driving, and monitoring physical machine accessories of the CPU and domestically produced AI accelerator cards, and is used to automatically realize hardware identification, driver installation, and hardware status monitoring functions according to the characteristics of the hardware; the resource management module is used to realize platform access of the AI ​​server and management of the domestically produced AI accelerator card, and to monitor the status and load of the node and the AI ​​accelerator card in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of cloud computing and artificial intelligence technology, and specifically to a device and method for adaptively adapting a cloud platform to a domestically produced AI accelerator card. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, AI accelerator cards have become key devices for improving AI computing performance. However, AI accelerator cards produced by different manufacturers vary in hardware architecture, interface protocols, and performance parameters, posing challenges to cloud platform adaptation and management.

[0003] With the rapid development of AI technology, the demand for AI computing resources in different application scenarios is becoming increasingly diverse. Therefore, how to quickly integrate and deploy domestic AI accelerator cards on cloud platforms, improve the flexibility and scalability of cloud platforms, and better meet diverse AI computing needs is a pressing technical challenge. Summary of the Invention

[0004] The technical task of the present invention is to provide a device and method for adaptively adapting domestic AI accelerator cards on a cloud platform to solve the problem of how to quickly integrate and deploy domestic AI accelerator cards on the cloud platform, improve the flexibility and scalability of the cloud platform, and better meet the diverse AI computing needs.

[0005] The technical task of the present invention is achieved in the following manner: a device for adaptively adapting a cloud platform to a domestic AI accelerator card, the device comprising a script tool set, a resource management module, an adaptation task management module, a data storage module, and an adaptation analysis module;

[0006] Among them, the script tool set integrates the recognition, perception, driving and monitoring script tools of physical machine accessories such as CPU and domestic AI accelerator cards, which are used to automatically realize hardware recognition, driver installation and hardware status monitoring functions based on hardware characteristics;

[0007] The resource management module is used to implement platform access for AI servers and manage domestic AI accelerator cards, and to monitor the status and load of nodes and AI accelerator cards in real time.

[0008] The adaptation task management module provides an adaptation management page, allowing adaptation personnel to formulate adaptation tasks based on the characteristics of different AI accelerator cards and schedule and execute tasks based on the node load.

[0009] The data storage module is used to save the adaptation results and the log of the adaptation process to facilitate the tracing of the adaptation process;

[0010] The adaptation analysis module is used to analyze the adaptation results and logs, and form the final adaptation conclusion based on the performance indicator characteristics of the domestic AI accelerator card itself.

[0011] As a preferred option, the script tool set develops plug-in script set management technology, adopts modularization and version control strategies, builds a unified script repository, and combines metadata tagging technology to achieve accurate classification and rapid retrieval of script functions, thereby realizing intelligent management of script sets and ensuring that each script unit is independently maintainable and easy to reuse.

[0012] The script tool set is an adapter with a visual script management interface, which enables script management of driver installation and monitoring collection under different CPU architectures of domestic AI accelerator cards; script management includes installation package management and tool set management;

[0013] Among them, installation package management refers to the management of the installation packages of mainstream domestic AI accelerator cards such as Cambricon, Ascend, and Tianshu Zhixin, as well as the addition, deletion, modification, and query management of the installation packages. It also supports the management of the accelerator card's own software stack, including the driver installation packages, monitoring tool packages, and underlying dependency packages for training and inference of mainstream domestic AI accelerator cards. The management of installation packages includes the fields of the operating system, version, CPU architecture, installation package, installation package checksum, and size supported by the installation package. For example, the Cambricon AI accelerator card supports drivers for mainstream AI accelerator cards such as Cambricon 370-X8 and 370-S4, Cambricon NeuWareSDK, Cambricon PyTorch deep learning framework, and Cambricon's commonly used Deepspeed, Flash Attention, Transformers, PEFT, and other library tools and libraries.

[0014] Toolset management provides a management interface for commonly used script tool sets for domestic AI accelerator cards, enabling the addition, deletion, modification, and query functions of the toolset; the toolset integrates common tool scripts for the installation, monitoring and acquisition, adaptation testing, and fault detection of software packages for commonly used AI accelerator cards such as Ascend, Cambricon, and Tianshu Zhixin; because AI accelerator cards can be used on different CPU architectures and different operating systems, toolset management includes adapted operating systems, CPU architectures, installation package scripts, and verification methods; adaptation scripts include supported AI accelerator card manufacturers, AI accelerator card types, AI accelerator card adaptation items and performance benchmark values, physical machine CPU architectures, operating system types, adaptation scripts, script protocols, and return values.

[0015] As a preferred option, the resource management module provides adapters with a resource interface for domestic AI accelerator cards, enabling automatic discovery, rapid identification, and topology awareness of AI accelerator cards. It also implements node AI accelerator card driver installation and node expansion initialization based on a script tool set, enabling domestic AI accelerator card resources to be quickly connected to the cloud platform.

[0016] The resource management module has computing power registration and node pooling functions;

[0017] Among them, when registering computing power, the registration of computing power resources in the cloud platform is realized based on the plug-in extension mechanism. For various CPUs and domestic AI acceleration card resources on the computing power cluster nodes, a standardized interface is provided to enable the computing power cluster to identify the computing power resources on the nodes, thereby realizing the management and scheduling of computing power resources; and based on the system PXE automatic installation technology, a system automatic installation framework for heterogeneous resources is developed to solve the ability of adaptive installation of multi-architecture operating systems, realize automatic matching of corresponding system images according to different CPU architectures, and complete the installation of node operating systems; the cloud platform scans the device information on the node PCIE, combines the manufacturer and device feature code of the device, and quickly identifies the manufacturer and model information of domestic AI acceleration cards and network cards. Based on the capabilities of the script tool set, the software stack of domestic AI acceleration card accessory drivers and toolkits is installed, and the automatic configuration and initialization of domestic AI acceleration cards for GPU and NPU are realized;

[0018] Node pooling means that after the CPU and domestic AI accelerator card on the node are initialized, the node is labeled accordingly based on the node's CPU architecture and AI accelerator card type, and similar resources are divided into a unified resource pool. For example, if the node server uses a Hygon CPU and a Cambrian 370-X8 AI accelerator card, the node can be added to the resource pool labeled Hygon_Cambican370X8. This allows for unified management and scheduling of similar nodes.

[0019] As a preferred option, the adaptation task management module provides an adaptation management page, allowing adaptation personnel to formulate adaptation tasks based on the characteristics of different AI accelerator cards, schedule corresponding nodes to perform adaptation tasks based on node load, and monitor the execution status of tasks based on monitoring indicators of the node's CPU, memory, and domestic accelerator card utilization to ensure smooth execution and completion of tasks; the details are as follows:

[0020] Generate adaptation tasks: The adaptation task generation module provides a visual management interface for managing the adaptation test task items of domestic AI accelerator cards; among them, the test task items include key task items and indicator task items; key task items are crucial to the overall adaptation. If the indicator value of the overall adaptation is lower than the standard value, the entire adaptation process will be judged as failed; indicator task items are given different weights according to their importance, with weight levels ranging from 1 to 5; the adaptation results will be calculated based on the indicator values ​​and weights of all task items; the adapter adds adaptation tasks according to the characteristics and uses of the domestic AI accelerator card, and sets the adaptation task item list and its weight, and Adjust task items as needed. For conventional domestic AI accelerator card adaptation, this includes driver installation and testing, operating system compatibility testing, toolkit verification, software stack verification, monitoring service verification, training and inference framework verification, benchmark performance testing, common model performance testing, and service product function testing. Driver test results and software stack verification are used as key indicators, and the weights of INT8, FP16, and FP32 benchmark performance indicators are set to 1, 3, and 2, respectively, reflecting the emphasis on AI performance. After the adaptation task is maintained, the adaptation personnel initiate the corresponding adaptation task, which is then scheduled and executed by the task scheduling center.

[0021] Task Scheduling: Design a task scheduler to schedule each adaptation task and its task items, ensuring that each task item is dispatched to the optimal node for execution. Dynamically adjust task scheduling based on node load and task monitoring data to ensure smooth task execution. The core of task scheduling is the task scheduling algorithm. From a resource perspective, when selecting execution nodes for adaptation tasks, a balanced scheduling strategy is implemented based on the node's computing power (such as CPU, AI accelerator card, memory, etc.) and its current load to maximize resource utilization. At the same time, a task scheduling algorithm combining "resource pool filtering + multi-dimensional feature weighting" is designed. The core goal is to select nodes with relatively low current loads and computing power that meet task requirements from eligible nodes to execute tasks.

[0022] Task execution: After the task is scheduled, the task manager will send it to the corresponding node for execution according to the protocol of the task item. After the task is scheduled, the task manager obtains the access method of the scheduling node through the node registration information and sends the adaptation task item to the corresponding node, such as sending the task execution script to the node through SSH remote call. The execution node starts the adaptation task by executing the task script. If an installation package is required, it will automatically download and execute it according to the script content and return the process identifier of the task. After the task is executed, the corresponding adaptation performance and verification result data are returned according to the return value of the script. The relevant log information generated during the task execution process is automatically synchronized to the data storage module for storage.

[0023] Task monitoring: To ensure the normal execution of adaptation tasks, a comprehensive task status monitoring system is built. Based on the monitoring data of CPU utilization, domestic AI accelerator card utilization, and memory utilization on the nodes provided by the cloud platform, if the resource utilization of a node exceeds the limit, an alarm will be issued through email, SMS, etc., and task scheduling for the node will be stopped. The node will be rescheduled after the node utilization returns to normal levels; each adaptation task will be assigned a unique ID (including adaptation item ID and adaptation task ID) to achieve accurate identity identification and tracking, and real-time tracking of task execution progress and status. Once the task execution time exceeds the threshold, automatic intervention will be carried out to ensure the effective utilization of system resources and the controllability of tasks.

[0024] More optimally, the task scheduling algorithm of "resource pool filtering + multi-dimensional feature weighting" is as follows:

[0025] First, resource pools are filtered based on the task item requirements to determine the scope of task execution nodes. Furthermore, based on the node CPU architecture and domestic AI accelerator card type requirements for any task item, resource pools are filtered based on their tag information to select the ones that meet the requirements. For example, if a task requires a Hygon CPU architecture and a Cambrian 370-X8 AI accelerator card, the task scheduler will traverse the resource pools and select the Hygon_Cambican370X8 resource pool that contains both of these tags, thereby determining the scope of task execution nodes.

[0026] Then, based on the multi-dimensional feature weight evaluation, the comprehensive available computing power of each node is calculated and the optimal node is selected for execution; the details are as follows:

[0027] ① Define the feature set: Determine the dimensions or features used for evaluation. The features to be evaluated are closely related to the requirements of task execution. Specifically, for domestic AI accelerator card adaptation, AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth occupancy are used;

[0028] ② Weight allocation: Based on the computing resource requirements of task execution, a weight value is assigned to each feature to distinguish the importance of each resource type (for example, if AI is computing-intensive, the weight of the AI ​​accelerator card is increased). The weight value is adjusted according to the needs of the specific task. Specifically for AI performance adaptation, the weights of AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization are set to 4, 3, 2, and 1, respectively;

[0029] ③ Comprehensive quantitative evaluation: Calculate the comprehensive score of each candidate node based on the weight and quantitative value. The node with the highest comprehensive score is scheduled to execute the adaptation task;

[0030] The resource availability calculation formula for the CPU, AI accelerator card, and memory on a node is as follows: ;

[0031] Where j represents the node ID; i represents the resource type ID; resource weight * current node availability can be used to get the resource scheduling score , the formula is as follows:

[0032] ;

[0033] By Calculated resource scheduling scores for each type By summing up, we can calculate the overall computing power score of each node at present. , the formula is as follows: ;

[0034] By comparing the overall computing power scores of each node, the optimal computing power node can be screened out, thereby achieving balanced scheduling based on computing power quantification. As an optimization, the data storage module designs a unified management framework for the adaptation result set, fully exploring the reuse value of adaptation records of multiple models;

[0035] The adaptation results are divided into three forms: structured data, log information, and business file data;

[0036] Among them, the performance of the adaptation task and the indicative structured data of the adaptation conclusion are directly saved to the public MySQL database for centralized management and quick access;

[0037] Adapt the generated log information and implement unified management through the log management component Kibana;

[0038] Business file data generated during the adaptation process is directly stored in MinIO, allowing the business team to easily query and analyze data.

[0039] The data storage module also introduces a global environment variable table (ENV table) to comprehensively record the environment configuration information of the adaptation task, enhancing configurability and flexibility.

[0040] Preferably, the adaptation analysis module analyzes the adaptation results based on the archived data of the adaptation task, and the adaptation analysis module has the function of automatically generating an adaptation report and comparing adaptation conclusions;

[0041] The automatic generation of an adaptation report automatically calculates adaptation indicators and forms an adaptation conclusion report. AI accelerator cards have many indicators, and each indicator is compared with the benchmark indicator, which is computationally intensive and prone to errors. Therefore, a calculation rule of "key item rejection + weight evaluation" is designed to automatically calculate adaptation indicators based on the rule.

[0042] According to the calculation rules of "key item rejection + weight evaluation", the overall adaptation evaluation score is automatically calculated to form an adaptation analysis report;

[0043] By comparing the adaptation data and analyzing the reports of domestic AI accelerator cards in training tasks of different complete machines, operating systems and large models, we can comprehensively evaluate the stability of domestic AI accelerator cards, the compatibility of complete machines and the applicable scenarios, so as to facilitate the subsequent formulation of delivery plans based on business needs.

[0044] More preferably, the calculation rules of "key item rejection + weight evaluation" are as follows:

[0045] Traverse each key task item in the adaptation task and compare the actual indicators of the corresponding task items with the benchmark values. If the actual indicators are lower than the benchmark value requirements, the overall adaptation fails and the subsequent weight evaluation can be skipped. If the actual indicators meet or exceed the benchmark value requirements, compare other key items and if they all pass the weight evaluation;

[0046] Perform comprehensive quantitative evaluation based on the weight of the task items. Specifically, use i to represent any task item. First, according to the formula Calculate the actual value of task item i indicator and benchmark values Deviation If the actual value is greater than the reference value, the positive deviation result is greater than 0, otherwise the negative deviation result is less than 0; then according to the formula The formula The calculated deviation of each indicator is combined with the task weight By summing up, we can calculate the overall score of each task item of the current adaptation task .

[0047] A method for adaptively adapting a cloud platform to a domestic AI accelerator card, specifically as follows:

[0048] Based on the characteristics and software stack of domestic AI accelerator cards, driver installation, software stack initialization, and monitoring services are managed through scripts and integrated into the script tool set to achieve unified management of tools.

[0049] The cloud platform capabilities based on the script tool set enable automatic discovery, rapid identification, and topology awareness of domestic AI accelerator cards. Initialization tools are used to install the domestic AI accelerator card driver on the node and initialize node expansion, allowing domestic AI accelerator card resources to be quickly connected to the cloud platform.

[0050] Generate adaptation tasks: Based on the characteristics and adaptation requirements of domestic AI accelerator cards, analyze the adaptation verification content to be performed, and based on the capabilities of the scripting toolset, develop an adaptation task list to form specific adaptation tasks;

[0051] Adaptation task scheduling: Design a task scheduler to complete the scheduling of each adaptation task and its task items, ensuring that each task item is scheduled to the optimal node for execution. Dynamically adjust task scheduling based on node load and task monitoring data to ensure smooth task execution. Task items are clearly divided into two categories based on their adaptation importance: key task items and indicator task items. Key task items are crucial to the overall adaptation. Once their indicator value falls below the standard value, the entire adaptation process will be judged as failed. Indicator task items are assigned different weights based on their importance, with weight levels ranging from 1 to 5. The adaptation results will be calculated based on the indicator values ​​and weights of all task items.

[0052] Adaptation task execution: The task manager obtains the access method of the scheduling node through the node registration information, sends the adaptation task item to the corresponding scheduling node, completes the execution of the adaptation task, and records the adaptation log and returns the adaptation result. The data storage module saves the adaptation result and log;

[0053] Adaptive task monitoring: Based on the node CPU utilization, AI accelerator card utilization, and memory utilization monitoring data provided by the cloud platform, task nodes and task status are monitored, and task execution progress and status are tracked in real time. Once the task execution time exceeds the threshold, automatic intervention will be carried out to ensure the effective use of system resources and the controllability of tasks.

[0054] Adaptation result analysis: Automatically obtain adaptation indicators, automatically calculate the overall adaptation comprehensive evaluation score based on the indicator weights, and form an adaptation conclusion report; and based on the comparison with historical adaptation reports, analyze the compatibility, stability and applicable scenarios of domestic AI accelerator cards.

[0055] As a preferred option, the script tool set adopts a modularization and version control strategy. By building a unified script repository and combining it with metadata tagging technology, it can achieve accurate classification and rapid retrieval of script functions. The script tool set includes the script's CPU architecture, operating system version, script protocol, and, if the installation package is involved, also includes the operating system and version supported by the installation package, CPU architecture, installation package content, installation package verification code, and size fields.

[0056] During the adaptation task execution process, the task scheduling algorithm of "resource pool filtering + multi-dimensional feature weighting" is adopted, as follows:

[0057] First, resource pools are filtered based on the task item requirements to determine the scope of task execution nodes. Furthermore, based on the node CPU architecture and domestic AI accelerator card type requirements for any task item, resource pools are filtered based on their tag information to select the ones that meet the requirements. For example, if a task requires a Hygon CPU architecture and a Cambrian 370-X8 AI accelerator card, the task scheduler will traverse the resource pools and select the Hygon_Cambican370X8 resource pool that contains both of these tags, thereby determining the scope of task execution nodes.

[0058] Then, based on the multi-dimensional feature weight evaluation, the comprehensive available computing power of each node is calculated and the optimal node is selected for execution; the details are as follows:

[0059] ① Define the feature set: Determine the dimensions or features used for evaluation. The features to be evaluated are closely related to the requirements of task execution. Specifically, for domestic AI accelerator card adaptation, AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth occupancy are used;

[0060] ② Weight allocation: Based on the computing resource requirements of task execution, a weight value is assigned to each feature to distinguish the importance of each resource type (for example, if AI is computing-intensive, the weight of the AI ​​accelerator card is increased). The weight value is adjusted according to the needs of the specific task. Specifically for AI performance adaptation, the weights of AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization are set to 4, 3, 2, and 1, respectively;

[0061] ③ Comprehensive quantitative evaluation: Calculate the comprehensive score of each candidate node based on the weight and quantitative value. The node with the highest comprehensive score is scheduled to execute the adaptation task;

[0062] Among them, the resource availability of CPU, AI accelerator card and memory on the node The calculation formula is as follows:

[0063] ;

[0064] Wherein, j represents the node identifier; i represents the resource type identifier;

[0065] Resource weight * current node availability can be used to get the resource scheduling score , the formula is as follows:

[0066] ;

[0067] By Calculated resource scheduling scores for each type By summing up, we can calculate the overall computing power score of each node at present. , the formula is as follows:

[0068] ;

[0069] By comparing the overall computing power scores of each node, the optimal computing power node can be screened out, thus achieving balanced scheduling based on computing power quantification. By comparing the overall computing power scores of each node, the optimal computing power node can be screened out, thus achieving balanced scheduling based on computing power quantification.

[0070] The device and method for adaptively adapting a cloud platform to a domestically produced AI accelerator card of the present invention have the following advantages:

[0071] (1) This invention uses Openstack for cloud platform resource management and Cyborg for AI accelerator card management in a cloud computing platform. This platform aims to enable rapid integration and deployment of domestic AI accelerator cards on the cloud platform, reducing deployment time and costs due to compatibility issues, while improving the flexibility and efficiency of the cloud platform and optimizing the operating performance of domestic AI accelerator cards on the cloud platform.

[0072] (2) The present invention comprises a script tool set, a resource management module, an adaptation task management module, a data storage module, and an adaptation analysis module. Through the coordination between these modules, the rapid discovery, access, and adaptation of domestic AI accelerator cards are achieved, which greatly shortens the adaptation cycle of AI accelerator cards, enables rapid access to heterogeneous computing power, promotes the development and application of domestic AI technology, better meets diverse AI computing needs, and effectively enhances the competitiveness and market adaptability of cloud platforms.

[0073] (3) Through the adaptive adaptation mechanism, the present invention enables cloud platforms to quickly integrate and deploy domestic AI accelerator cards, reducing deployment time and costs due to compatibility issues. It can also quickly adjust and optimize its own service capabilities and resource allocation, thereby improving the flexibility and scalability of the cloud platform, optimizing the operating performance of domestic AI accelerator cards, promoting the development and application of domestic AI technologies, improving resource utilization and reducing operating costs, and enhancing the competitiveness and market adaptability of the cloud platform to better meet diverse AI computing needs.

[0074] (4) The script tool set of the present invention integrates script tools for identification, perception, driving, and monitoring of physical machine accessories such as CPUs and domestic AI accelerator cards, and can automatically implement functions such as hardware identification, driver installation, and hardware status monitoring based on hardware characteristics;

[0075] (5) The resource management module of the present invention enables platform access to AI servers and management of domestic AI accelerator cards, and monitors the status and load of nodes and AI accelerator cards in real time;

[0076] (6) The adaptation task management module of the present invention can provide an adaptation management page, allowing adaptation personnel to formulate adaptation tasks based on the characteristics of different AI accelerator cards and schedule the execution of tasks based on the node load, etc.

[0077] (7) The data storage module of the present invention saves the adaptation results and the log of the adaptation process to facilitate tracing of the adaptation process;

[0078] (8) The adaptation analysis module of the present invention analyzes the adaptation results and logs, and forms a final adaptation conclusion based on the performance indicators and other characteristics of the AI ​​accelerator card itself;

[0079] (9) The present invention can be used for the system development of adaptive adaptation engines for heterogeneous computing power, completing the rapid adaptation of heterogeneous computing power accessories such as GPUs and NPUs, and continuously and rapidly updating drivers to be compatible with more servers, enabling cloud platforms to quickly integrate and deploy domestic AI accelerator cards, reducing deployment time and costs caused by compatibility issues, and improving the flexibility and scalability of cloud platforms;

[0080] (10) The present invention solves the limitation of only a few technical experts executing the adaptation process through a modular adaptation script tool set and knowledge base, reduces the difficulty of compatibility adaptation testing, effectively shortens the communication costs and equipment occupation costs of multiple people participating in the test, and reduces the deployment time and cost caused by compatibility issues. It can effectively speed up the adaptation and introduction of domestic AI accelerator cards, improve the compatibility and stability of domestic AI accelerator cards, promote the development and application of domestic AI technology, enhance the competitiveness and market adaptability of cloud platforms, and better meet diverse AI computing needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The present invention will be further described below with reference to the accompanying drawings.

[0082] Attachment Figure 1 A schematic diagram of the structure of a device for adaptively adapting a domestic AI accelerator card to a cloud platform;

[0083] Attachment Figure 2 A flowchart for the execution of the adaptation task. DETAILED DESCRIPTION

[0084] The following is a detailed description of the device and method for adaptively adapting a cloud platform to a domestic AI accelerator card according to the present invention with reference to the accompanying drawings and specific embodiments.

[0085] As attached Figure 1 As shown, this embodiment provides a device for adaptively adapting a cloud platform to a domestic AI accelerator card, the device including a script tool set, a resource management module, an adaptation task management module, a data storage module, and an adaptation analysis module;

[0086] Among them, the script tool set integrates the recognition, perception, driving and monitoring script tools of physical machine accessories such as CPU and domestic AI accelerator cards, which are used to automatically realize hardware recognition, driver installation and hardware status monitoring functions based on hardware characteristics;

[0087] The resource management module is used to implement platform access for AI servers and manage domestic AI accelerator cards, and to monitor the status and load of nodes and AI accelerator cards in real time.

[0088] The adaptation task management module provides an adaptation management page, allowing adaptation personnel to formulate adaptation tasks based on the characteristics of different AI accelerator cards and schedule and execute tasks based on the node load.

[0089] The data storage module is used to save the adaptation results and the log of the adaptation process to facilitate the tracing of the adaptation process;

[0090] The adaptation analysis module is used to analyze the adaptation results and logs, and form the final adaptation conclusion based on the performance indicator characteristics of the domestic AI accelerator card itself.

[0091] As a preferred option, the script tool set develops plug-in script set management technology, adopts modularization and version control strategies, builds a unified script repository, and combines metadata tagging technology to achieve accurate classification and rapid retrieval of script functions, thereby realizing intelligent management of script sets and ensuring that each script unit is independently maintainable and easy to reuse.

[0092] The script tool set is an adapter with a visual script management interface, which enables script management of driver installation and monitoring collection under different CPU architectures of domestic AI accelerator cards; script management includes installation package management and tool set management;

[0093] Among them, installation package management refers to the management of the installation packages of mainstream domestic AI accelerator cards such as Cambricon, Ascend, and Tianshu Zhixin, as well as the addition, deletion, modification, and query management of the installation packages. It also supports the management of the accelerator card's own software stack, including the driver installation packages, monitoring tool packages, and underlying dependency packages for training and inference of mainstream domestic AI accelerator cards. The management of installation packages includes the fields of the operating system, version, CPU architecture, installation package, installation package checksum, and size supported by the installation package. For example, the Cambricon AI accelerator card supports drivers for mainstream AI accelerator cards such as Cambricon 370-X8 and 370-S4, Cambricon NeuWareSDK, Cambricon PyTorch deep learning framework, and Cambricon's commonly used Deepspeed, Flash Attention, Transformers, PEFT, and other library tools and libraries.

[0094] Toolset management provides a management interface for commonly used script tool sets for domestic AI accelerator cards, enabling the addition, deletion, modification, and query functions of the toolset; the toolset integrates common tool scripts for the installation, monitoring and acquisition, adaptation testing, and fault detection of software packages for commonly used AI accelerator cards such as Ascend, Cambricon, and Tianshu Zhixin; because AI accelerator cards can be used on different CPU architectures and different operating systems, toolset management includes adapted operating systems, CPU architectures, installation package scripts, and verification methods; adaptation scripts include supported AI accelerator card manufacturers, AI accelerator card types, AI accelerator card adaptation items and performance benchmark values, physical machine CPU architectures, operating system types, adaptation scripts, script protocols, and return values.

[0095] The resource management module in this embodiment provides adapters with a resource interface for domestic AI accelerator cards, enabling automatic discovery, rapid identification, and topology awareness of AI accelerator cards. It also implements node AI accelerator card driver installation and node expansion initialization based on a script tool set, enabling rapid access to domestic AI accelerator card resources on the cloud platform.

[0096] The resource management module has computing power registration and node pooling functions;

[0097] Among them, when registering computing power, the registration of computing power resources in the cloud platform is realized based on the plug-in extension mechanism. For various CPUs and domestic AI acceleration card resources on the computing power cluster nodes, a standardized interface is provided to enable the computing power cluster to identify the computing power resources on the nodes, thereby realizing the management and scheduling of computing power resources; and based on the system PXE automatic installation technology, a system automatic installation framework for heterogeneous resources is developed to solve the ability of adaptive installation of multi-architecture operating systems, realize automatic matching of corresponding system images according to different CPU architectures, and complete the installation of node operating systems; the cloud platform scans the device information on the node PCIE, combines the manufacturer and device feature code of the device, and quickly identifies the manufacturer and model information of domestic AI acceleration cards and network cards. Based on the capabilities of the script tool set, the software stack of domestic AI acceleration card accessory drivers and toolkits is installed, and the automatic configuration and initialization of domestic AI acceleration cards for GPU and NPU are realized;

[0098] Node pooling means that after the CPU and domestic AI accelerator card on the node are initialized, the node is labeled accordingly based on the node's CPU architecture and AI accelerator card type, and similar resources are divided into a unified resource pool. For example, if the node server uses a Hygon CPU and a Cambrian 370-X8 AI accelerator card, the node can be added to the resource pool labeled Hygon_Cambican370X8. This allows for unified management and scheduling of similar nodes.

[0099] As attached Figure 2As shown, the adaptation task management module in this embodiment provides an adaptation management page, allowing adaptation personnel to formulate adaptation tasks according to the characteristics of different AI accelerator cards, schedule corresponding nodes to perform adaptation tasks according to the node load, and monitor the execution status of tasks based on the monitoring indicators of the node's CPU, memory, and domestic accelerator card utilization to ensure smooth execution and completion of tasks; the details are as follows:

[0100] (1) Generate adaptation tasks: The adaptation task generation module provides a visual management interface for managing the adaptation test task items of domestic AI accelerator cards; among them, the test task items include key task items and indicator task items; key task items are crucial to the overall adaptation. If the indicator value of the overall adaptation is lower than the standard value, the entire adaptation process will be judged as failed; indicator task items are given different weights according to their importance, with weight levels ranging from 1 to 5; the adaptation results will be calculated based on the indicator values ​​and weights of all task items; the adapter adds adaptation tasks according to the characteristics and uses of the domestic AI accelerator card, and sets the list of adaptation task items and their weights , and adjust the task items as needed; for conventional domestic AI accelerator card adaptation, including driver installation and testing, operating system compatibility testing, toolkit verification, software stack verification, monitoring service verification, training and inference framework verification, benchmark performance testing, common model performance testing and service product function testing, and take driver test results and software stack verification as key indicators, and set the weights of INT8, FP16 and FP32 benchmark performance indicators to 1, 3 and 2, reflecting the emphasis on AI performance; after the adaptation task is maintained, the adaptation personnel start the corresponding adaptation task, and the task scheduling center schedules and executes the adaptation task;

[0101] (2) Task scheduling: Design a task scheduler to complete the scheduling of each adaptation task and its task items, and schedule each task item to the optimal node for execution. Dynamically adjust the task scheduling based on the node load and task monitoring data to ensure the smooth execution of the task. The core of task scheduling is the task scheduling algorithm. From the resource level, when selecting the execution node for the adaptation task, a balanced scheduling strategy is implemented based on the node's computing power (such as CPU, AI accelerator card, memory, etc.) and its current load to maximize resource utilization. At the same time, a task scheduling algorithm of "resource pool filtering + multi-dimensional feature weighting" is designed. The core goal is to select nodes with relatively low current load and computing power that meet the task requirements from qualified nodes to execute the task.

[0102] (3) Task execution: After the task is scheduled, the task manager sends it to the corresponding node for execution according to the protocol of the task item; after the task scheduling is completed, the task manager obtains the access method of the scheduling node through the node registration information, and sends the adaptation task item to the corresponding node, such as sending the task execution script to the node through SSH remote call; the execution node starts the adaptation task by executing the task script. If an installation package is required, it will automatically download and execute it according to the script content and return the process identifier of the task; after the task is completed, the corresponding adaptation performance and verification result data are returned according to the return value requirements of the script; the relevant log information generated during the task execution process is automatically synchronized to the data storage module for storage;

[0103] (4) Task monitoring: In order to achieve the normal execution of the adaptation task, a comprehensive task status monitoring system is built. Based on the monitoring data of CPU utilization, domestic AI acceleration card utilization and memory utilization on the node provided by the cloud platform, if the resource utilization of a node exceeds the limit, an alarm will be issued through email, SMS, etc., and the task scheduling of the node will be stopped. The node will be rescheduled after the node utilization returns to normal levels; each adaptation task will be assigned a unique ID (including adaptation item ID and adaptation task ID) to achieve accurate identity identification and tracking, and real-time tracking of task execution progress and status. Once the task execution time exceeds the threshold, automatic intervention will be carried out to ensure the effective utilization of system resources and the controllability of the task.

[0104] The task scheduling algorithm of "resource pool filtering + multi-dimensional feature weighting" in this embodiment is as follows:

[0105] First, resource pools are filtered based on the task item requirements to determine the scope of task execution nodes. Furthermore, based on the node CPU architecture and domestic AI accelerator card type requirements for any task item, resource pools are filtered based on their tag information to select the ones that meet the requirements. For example, if a task requires a Hygon CPU architecture and a Cambrian 370-X8 AI accelerator card, the task scheduler will traverse the resource pools and select the Hygon_Cambican370X8 resource pool that contains both of these tags, thereby determining the scope of task execution nodes.

[0106] Then, based on the multi-dimensional feature weight evaluation, the comprehensive available computing power of each node is calculated and the optimal node is selected for execution; the details are as follows:

[0107] ① Define the feature set: Determine the dimensions or features used for evaluation. The features to be evaluated are closely related to the requirements of task execution. Specifically, for domestic AI accelerator card adaptation, AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth occupancy are used;

[0108] ② Weight allocation: Based on the computing resource requirements of task execution, a weight value is assigned to each feature to distinguish the importance of each resource type (for example, if AI is computing-intensive, the weight of the AI ​​accelerator card is increased). The weight value is adjusted according to the needs of the specific task. Specifically for AI performance adaptation, the weights of AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization are set to 4, 3, 2, and 1, respectively;

[0109] ③ Comprehensive quantitative evaluation: Calculate the comprehensive score of each candidate node based on the weight and quantitative value. The node with the highest comprehensive score is scheduled to execute the adaptation task;

[0110] The resource availability calculation formula for the CPU, AI accelerator card, and memory on a node is as follows: ;

[0111] Where j represents the node ID; i represents the resource type ID; resource weight * current node availability can be used to get the resource scheduling score , the formula is as follows:

[0112] ;

[0113] By Calculated resource scheduling scores for each type By summing up, we can calculate the overall computing power score of each node at present. , the formula is as follows: ;

[0114] By comparing the overall computing power scores of each node, the optimal computing power node can be screened out, thereby achieving balanced scheduling based on computing power quantification. As an optimization, the data storage module designs a unified management framework for the adaptation result set, fully exploring the reuse value of adaptation records of multiple models;

[0115] The adaptation results are divided into three forms: structured data, log information, and business file data;

[0116] Among them, the performance of the adaptation task and the indicative structured data of the adaptation conclusion are directly saved to the public MySQL database for centralized management and quick access;

[0117] Adapt the generated log information and implement unified management through the log management component Kibana;

[0118] Business file data generated during the adaptation process is directly stored in MinIO, allowing the business team to easily query and analyze data.

[0119] The data storage module also comprehensively records the environment configuration information of the adaptation task by introducing a global environment variable table (ENV table), enhancing configurability and flexibility. The adaptation analysis module in this embodiment analyzes the adaptation results based on the archived data of the adaptation task. The adaptation analysis module has the function of automatically generating an adaptation report and comparing the adaptation conclusions;

[0120] The automatic generation of an adaptation report automatically calculates adaptation indicators and forms an adaptation conclusion report. AI accelerator cards have many indicators, and each indicator is compared with the benchmark indicator, which is computationally intensive and prone to errors. Therefore, a calculation rule of "key item rejection + weight evaluation" is designed to automatically calculate adaptation indicators based on the rule.

[0121] According to the calculation rules of "key item rejection + weight evaluation", the overall adaptation evaluation score is automatically calculated to form an adaptation analysis report;

[0122] By comparing the adaptation data and analyzing the reports of domestic AI accelerator cards in training tasks of different complete machines, operating systems and large models, we can comprehensively evaluate the stability of domestic AI accelerator cards, the compatibility of complete machines and the applicable scenarios, so as to facilitate the subsequent formulation of delivery plans based on business needs.

[0123] The calculation rules of "key item rejection + weight evaluation" in this embodiment are as follows:

[0124] ① Traverse each key task item in the adaptation task and compare the actual indicators of the corresponding task items with the benchmark values. If the actual indicators are lower than the benchmark value requirements, the overall adaptation fails and the subsequent weight evaluation can be skipped. If the actual indicators meet or exceed the benchmark value requirements, compare other key items and if they all pass the weight evaluation;

[0125] ② Carry out comprehensive quantitative evaluation based on the weight of the task items. Specifically: use i to represent any task item, first according to the formula Calculate the actual value of task item i indicator and benchmark values Deviation If the actual value is greater than the reference value, the positive deviation result is greater than 0, otherwise the negative deviation result is less than 0; then according to the formula The formula The calculated deviation of each indicator is combined with the task weight By summing up, we can calculate the overall score of each task item of the current adaptation task .

[0126] Example 2:

[0127] This embodiment provides a method for adaptively adapting a cloud platform to a domestic AI accelerator card. The method is as follows:

[0128] S1. Based on the characteristics and software stack of domestic AI accelerator cards, driver installation, software stack initialization, and monitoring services are managed through scripts and integrated into the script tool set to achieve unified management of tools.

[0129] S2. Cloud platform capabilities based on script toolsets enable automatic discovery, rapid identification, and topology awareness of domestic AI accelerator cards. Initialization tools are used to install domestic AI accelerator card drivers and initialize node expansion, enabling rapid access to cloud platform resources.

[0130] S3. Generate adaptation tasks: Based on the characteristics and adaptation requirements of domestic AI accelerator cards, analyze the adaptation verification content to be performed, and based on the capabilities of the scripting toolset, develop an adaptation task list to form specific adaptation tasks;

[0131] S4. Adaptation Task Scheduling: Design a task scheduler to schedule each adaptation task and its task items, ensuring that each task item is dispatched to the optimal node for execution. Dynamically adjust task scheduling based on node load and task monitoring data to ensure smooth task execution. Task items are clearly divided into two categories based on their importance to adaptation: critical task items and indicator task items. Critical task items are crucial to the overall adaptation. Once their indicator value falls below the standard value, the entire adaptation process will be deemed unsuccessful. Indicator task items are assigned different weights based on their importance, ranging from 1 to 5. The adaptation results will be calculated based on the indicator values ​​and weights of all task items.

[0132] S5. Adaptation task execution: The task manager obtains the access method of the scheduling node through the node registration information, sends the adaptation task item to the corresponding scheduling node, completes the execution of the adaptation task, and records the adaptation log and returns the adaptation result. The data storage module saves the adaptation result and log;

[0133] S6. Adaptive Task Monitoring: Based on the CPU utilization, AI accelerator card utilization, and memory utilization monitoring data on the node provided by the cloud platform, task nodes and task status are monitored, and task execution progress and status are tracked in real time. Once the task execution time exceeds the threshold, automatic intervention will be carried out to ensure the effective use of system resources and the controllability of tasks.

[0134] S7. Adaptation result analysis: Automatically obtain adaptation indicators, automatically calculate the overall adaptation comprehensive evaluation score based on the indicator weights, and form an adaptation conclusion report; and based on the comparison with historical adaptation reports, analyze the compatibility, stability and applicable scenarios of domestic AI accelerator cards.

[0135] The script toolset in this embodiment adopts a modular and version control strategy. By building a unified script repository and combining metadata tag technology, it realizes accurate classification and rapid retrieval of script functions. Among them, the script toolset includes the script's CPU architecture, operating system version, script protocol, and if it involves the installation package, it also includes the operating system and version supported by the installation package, CPU architecture, installation package content, installation package verification code and size fields.

[0136] In the adaptation task execution process in step S5 of this embodiment, the task scheduling algorithm of "resource pool filtering + multi-dimensional feature weight" is adopted as follows:

[0137] First, resource pools are filtered based on the task item requirements to determine the scope of task execution nodes. Furthermore, based on the node CPU architecture and domestic AI accelerator card type requirements for any task item, resource pools are filtered based on their tag information to select the ones that meet the requirements. For example, if a task requires a Hygon CPU architecture and a Cambrian 370-X8 AI accelerator card, the task scheduler will traverse the resource pools and select the Hygon_Cambican370X8 resource pool that contains both of these tags, thereby determining the scope of task execution nodes.

[0138] Then, based on the multi-dimensional feature weight evaluation, the comprehensive available computing power of each node is calculated and the optimal node is selected for execution; the details are as follows:

[0139] ① Define the feature set: Determine the dimensions or features used for evaluation. The features to be evaluated are closely related to the requirements of task execution. Specifically, for domestic AI accelerator card adaptation, AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth occupancy are used;

[0140] ② Weight allocation: Based on the computing resource requirements of task execution, a weight value is assigned to each feature to distinguish the importance of each resource type (for example, if AI is computing-intensive, the weight of the AI ​​accelerator card is increased). The weight value is adjusted according to the needs of the specific task. Specifically for AI performance adaptation, the weights of AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization are set to 4, 3, 2, and 1, respectively;

[0141] ③ Comprehensive quantitative evaluation: Calculate the comprehensive score of each candidate node based on the weight and quantitative value. The node with the highest comprehensive score is scheduled to execute the adaptation task;

[0142] The resource availability calculation formula for the CPU, AI accelerator card, and memory on a node is as follows: ;

[0143] Where j represents the node ID; i represents the resource type ID; resource weight * current node availability can be used to get the resource scheduling score , the formula is as follows:

[0144] ;

[0145] By Calculated resource scheduling scores for each type By summing up, we can calculate the overall computing power score of each node at present. , the formula is as follows:

[0146] ;

[0147] By comparing the overall computing power scores of each node, the optimal computing power node can be screened out, thereby achieving balanced scheduling based on computing power quantification.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A device for adaptively adapting a cloud platform to a domestic AI accelerator card, characterized in that: The device includes a script tool set, a resource management module, an adaptation task management module, a data storage module and an adaptation analysis module; Among them, the script tool set integrates the recognition, perception, driving and monitoring script tools of physical machine accessories such as CPU and domestic AI accelerator cards, which are used to automatically realize hardware recognition, driver installation and hardware status monitoring functions based on hardware characteristics; The resource management module is used to implement platform access for AI servers and manage domestic AI accelerator cards, and to monitor the status and load of nodes and AI accelerator cards in real time. The adaptation task management module is used to provide an adaptation management page, so that the adaptation personnel can formulate adaptation tasks according to the characteristics of different AI accelerator cards, and schedule and execute tasks according to the load of the node; among them, the adaptation task management module generates adaptation tasks specifically as follows: The adaptation task generation module provides a visual management interface for managing the adaptation test task items of domestic AI accelerator cards; among them, the test task items include key task items and indicator task items; key task items are crucial to the overall adaptation. If the indicator value of the overall adaptation is lower than the standard value, the entire adaptation process will be judged as failed; indicator task items are given different weights according to their importance, and the weights, etc. The level is 1 to 5; the adaptation result will be calculated based on the index values ​​and weights of all task items; the adaptation personnel will add adaptation tasks according to the characteristics and uses of domestic AI accelerator cards, set the adaptation task item list and its weights, and adjust the task items as needed; the scheduling and execution of tasks according to the node load are specifically: design a task scheduler to complete the scheduling of each adaptation task and its task items, realize the scheduling of each task item to the optimal node for execution, and dynamically adjust the task scheduling according to the node load and task monitoring data to ensure the smooth execution of the task; the core of task scheduling is the task scheduling algorithm; from the resource level, select the execution node for the adaptation task At each point in time, a balanced scheduling strategy is executed according to the computing power of the node and its current load to maximize resource utilization; at the same time, a task scheduling algorithm of "resource pool filtering + multi-dimensional feature weighting" is designed. The core goal is to select nodes with relatively low current load and computing power that meet the task requirements from qualified nodes to execute tasks; the task scheduling algorithm of "resource pool filtering + multi-dimensional feature weighting" is specifically as follows: first, the resource pool is filtered according to the requirements of the task item to determine the scope of the task execution node; and according to the requirements of any task item for the node CPU architecture and the type of domestic AI accelerator card, the resource pool is filtered with reference to the label information of the resource pool. Filter to select the resource pool that meets the requirements; then, based on the multi-dimensional feature weight evaluation, calculate the comprehensive available computing power of each node and select the optimal node for execution; specifically: ① Define the feature set: determine the dimensions or features used for evaluation, and the evaluated features are closely related to the requirements of task execution; ② Weight allocation: Based on the focus of task execution on computing power resources, assign a weight value to each feature to distinguish the importance of each resource type, and the weight value is adjusted according to the needs of the specific task; ③ Comprehensive quantitative evaluation: Calculate the comprehensive score of each candidate node based on the weight and quantitative value. The node with the highest comprehensive score is scheduled to execute the adaptation task; The data storage module is used to save the adaptation results and the log of the adaptation process to facilitate the tracing of the adaptation process; The adaptation analysis module is used to analyze the adaptation results and logs, and form the final adaptation conclusion based on the performance indicator characteristics of the domestic AI accelerator card itself.

2. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 1, characterized in that: The script tool set develops plug-in script set management technology, adopts modularization and version control strategies, builds a unified script repository, and combines metadata tagging technology to achieve accurate classification and rapid retrieval of script functions; The script tool set is an adapter with a visual script management interface, which enables script management of driver installation and monitoring collection under different CPU architectures of domestic AI accelerator cards; script management includes installation package management and tool set management; Installation package management refers to the management of domestic AI accelerator card installation packages, as well as the addition, deletion, modification, and query management of installation packages. It also supports the management of the accelerator card's own software stack, including the domestic AI accelerator card driver installation package, monitoring toolkit, and underlying dependency packages for training and inference. Installation package management includes fields such as the operating system, version, CPU architecture, installation package, installation package checksum, and size supported by the installation package. Toolset management provides a management interface for commonly used script tool sets for domestic AI accelerator cards, enabling the addition, deletion, modification, and query functions of the toolset; the toolset integrates commonly used tool scripts for AI accelerator card software package installation, monitoring and acquisition, adaptation testing, and fault detection; toolset management includes adapted operating systems, CPU architectures, installation package scripts, and verification methods; adaptation scripts include supported AI accelerator card manufacturers, AI accelerator card types, AI accelerator card adaptation items and performance benchmark values, physical machine CPU architectures, operating system types, adaptation scripts, script protocols, and return values.

3. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 1, characterized in that: The resource management module provides adapters with a resource interface for domestic AI accelerator cards, enabling automatic discovery, rapid identification, and topology awareness of AI accelerator cards. It also implements node AI accelerator card driver installation and node expansion initialization based on a script tool set, enabling rapid access to domestic AI accelerator card resources on the cloud platform. The resource management module has computing power registration and node pooling functions; Among them, when registering computing power, the registration of computing power resources in the cloud platform is realized based on the plug-in extension mechanism. For various CPUs and domestic AI acceleration card resources on the computing power cluster nodes, a standardized interface is provided to enable the computing power cluster to identify the computing power resources on the nodes, thereby realizing the management and scheduling of computing power resources; and based on the system PXE automatic installation technology, a system automatic installation framework for heterogeneous resources is developed to solve the ability of adaptive installation of multi-architecture operating systems, realize automatic matching of corresponding system images according to different CPU architectures, and complete the installation of node operating systems; the cloud platform scans the device information on the node PCIE, combines the manufacturer and device feature code of the device, and quickly identifies the manufacturer and model information of domestic AI acceleration cards and network cards. Based on the capabilities of the script tool set, the software stack of domestic AI acceleration card accessory drivers and toolkits is installed, and the automatic configuration and initialization of domestic AI acceleration cards for GPU and NPU are realized; Node pooling means that after the CPU and domestic AI accelerator card on the node are initialized, corresponding identifiers are added to the node according to the node's CPU architecture and AI accelerator card type, and similar resources are divided into a unified resource pool.

4. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 1, characterized in that: The adaptation task management module provides an adaptation management page, allowing adaptation personnel to formulate adaptation tasks based on the characteristics of different AI accelerator cards, schedule corresponding nodes to perform adaptation tasks based on node load, and monitor the execution status of tasks based on monitoring indicators of the node's CPU, memory, and domestic accelerator card utilization to ensure smooth task execution and completion. The adaptation task management module generates adaptation tasks for conventional domestic AI accelerator cards, including driver installation and testing, operating system compatibility testing, toolkit verification, software stack verification, monitoring service verification, training and inference framework verification, benchmark performance testing, common model performance testing, and service product function testing. Driver test results and software stack verification are used as key indicators, and the weights of INT8, FP16, and FP32 benchmark performance indicators are set to 1, 3, and 2, respectively, reflecting the emphasis on AI performance. After the adaptation task is maintained, the adaptation personnel start the corresponding adaptation task, and the task scheduling center schedules and executes the adaptation task. After the task is scheduled, the task manager sends it to the corresponding node for execution according to the task item protocol; After the task scheduling is completed, the task manager obtains the access method of the scheduling node through the node registration information and sends the adaptation task item to the corresponding node; the execution node starts the adaptation task by executing the task script; After the task is completed, the corresponding adaptation performance and verification result data are returned according to the return value requirements of the script; The relevant log information generated during the task execution process is automatically synchronized to the data storage module for storage; When the adaptation task management module performs task monitoring, it is based on the monitoring data of CPU utilization, domestic AI acceleration card utilization and memory utilization on the node provided by the cloud platform; each adaptation task will be assigned a unique ID to achieve accurate identification and tracking, and track the task execution progress and status in real time. Once the task execution time exceeds the threshold, it will automatically intervene to ensure the effective utilization of system resources and the controllability of the task.

5. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 4, characterized in that: Determine the dimensions or features to be used for evaluation. When the evaluation features are closely related to the requirements of task execution, specifically for domestic AI accelerator card adaptation, use AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization; Based on the computing resource requirements of task execution, a weight value is assigned to each feature to distinguish the importance of each resource type. The weight value is adjusted according to the needs of specific tasks. Specifically for AI performance adaptation, the weights of AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization are set to 4, 3, 2, and 1, respectively. The comprehensive score of each candidate node is calculated based on the weight and quantization value. When the node with the highest comprehensive score is scheduled to execute the adaptation task, the resource availability of the CPU, AI accelerator card and memory on the node is calculated. The calculation formula is as follows: ; Wherein, j represents the node identifier; i represents the resource type identifier; The resource scheduling score can be obtained by multiplying the resource weight by the current node's availability. The formula is as follows: ; By Calculated resource scheduling scores for each type By summing up, we can calculate the overall computing power score of each node at present. , the formula is as follows: ; By comparing the overall computing power scores of each node, the optimal computing power node can be screened, thereby achieving balanced scheduling based on computing power quantification.

6. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 1, characterized in that: The data storage module is designed to adapt to a unified management framework for result sets, fully tapping the reuse value of adaptation records for multiple models; The adaptation results are divided into three forms: structured data, log information, and business file data; Among them, the performance of the adaptation task and the indicative structured data of the adaptation conclusion are directly saved to the public MySQL database for centralized management and quick access; Adapt the generated log information and implement unified management through the log management component Kibana; Business file data generated during the adaptation process is directly stored in MinIO, allowing the business team to easily query and analyze data. The data storage module also comprehensively records the environment configuration information of the adaptation task by introducing a global environment variable table.

7. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 1, characterized in that: The adaptation analysis module analyzes the adaptation results based on the archived data of the adaptation task. The adaptation analysis module has the function of automatically generating an adaptation report and comparing the adaptation conclusions. The automatic generation of the adaptation report automatically calculates the adaptation indicators and forms an adaptation conclusion report. A calculation rule of "key item rejection + weight evaluation" is designed to automatically calculate the adaptation indicators according to the rule. According to the calculation rules of "key item rejection + weight evaluation", the overall adaptation evaluation score is automatically calculated to form an adaptation analysis report; By comparing the adaptation data of domestic AI accelerator cards in training tasks of different complete machines, operating systems and large models and analyzing the reports, we comprehensively evaluate the stability of domestic AI accelerator cards, the compatibility of complete machines and the applicable scenarios, so as to facilitate the subsequent formulation of delivery plans based on business needs.

8. The device for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 7, characterized in that: The calculation rules for "key item rejection + weighted evaluation" are as follows: Traverse each key task item in the adaptation task and compare the actual indicators of the corresponding task items with the benchmark values. If the actual indicators are lower than the benchmark value requirements, the overall adaptation fails and the subsequent weight evaluation can be skipped; If the actual indicators meet or exceed the benchmark value requirements, other key items will be compared and all will pass the weighted assessment; Perform comprehensive quantitative evaluation based on the weight of the task items. Specifically, use i to represent any task item. First, according to the formula Calculate the actual value of task item i indicator and benchmark values Deviation If the actual value is greater than the reference value, the positive deviation result is greater than 0, otherwise the negative deviation result is less than 0; then according to the formula The formula The calculated deviation of each indicator is combined with the task weight By summing up, we can calculate the overall score of each task item of the current adaptation task .

9. A method for adaptively adapting a cloud platform to a domestic AI accelerator card, characterized in that: The method is as follows: Based on the characteristics and software stack of domestic AI accelerator cards, driver installation, software stack initialization, and monitoring services are managed through scripts and integrated into the script tool set to achieve unified management of tools. The cloud platform capabilities based on the script tool set enable automatic discovery, rapid identification, and topology awareness of domestic AI accelerator cards. Initialization tools are used to install the domestic AI accelerator card driver on the node and initialize node expansion, allowing domestic AI accelerator card resources to be quickly connected to the cloud platform. Generate adaptation tasks: Based on the characteristics and adaptation requirements of domestic AI accelerator cards, analyze the adaptation verification content to be performed, and based on the capabilities of the scripting toolset, develop an adaptation task list to form specific adaptation tasks; Adaptation task scheduling: Design a task scheduler to complete the scheduling of each adaptation task and its task items, realize the scheduling of each task item to the optimal node for execution, and dynamically adjust the task scheduling according to the node load and task monitoring data to ensure the smooth execution of the task; among them, task items are clearly divided into two categories according to the importance of adaptation: key task items and indicator task items; key task items are crucial to the overall adaptation. Once their indicator values ​​are lower than the standard values, the entire adaptation process will be judged as failed; indicator task items are given different weights according to their importance, with weight levels ranging from 1 to 5; the adaptation results will be calculated based on the indicator values ​​and weights of all task items; the task scheduling algorithm of "resource pool filtering + multi-dimensional feature weight" is adopted during the execution of the adaptation task, specifically: first, the resource pool is filtered according to the requirements of the task item to determine the task execution The scope of nodes; and according to the requirements of any task item for the node CPU architecture and the type of domestic AI accelerator card, the resource pool is filtered with reference to the label information of the resource pool to screen out the resource pool that meets the requirements; then, based on the multi-dimensional feature weight evaluation, the comprehensive available computing power of each node is calculated, and the optimal node is selected for execution; specifically: ① Define the feature set: determine the dimensions or features used for evaluation, and the evaluated features are closely related to the requirements of task execution; ② Weight allocation: according to the focus of the task execution on the demand for computing power resources, a weight value is assigned to each feature to distinguish the importance of each resource type, and the weight value is adjusted according to the needs of the specific task; ③ Comprehensive quantitative evaluation: calculate the comprehensive score of each candidate node based on the weight and quantitative value, and the node with the highest comprehensive score is scheduled to execute the adaptation task; Adaptation task execution: The task manager obtains the access method of the scheduling node through the node registration information, sends the adaptation task item to the corresponding scheduling node, completes the execution of the adaptation task, and records the adaptation log and returns the adaptation result. The data storage module saves the adaptation result and log; Adaptive task monitoring: Based on the node CPU utilization, AI accelerator card utilization, and memory utilization monitoring data provided by the cloud platform, task nodes and task status are monitored, and task execution progress and status are tracked in real time. Once the task execution time exceeds the threshold, automatic intervention will be carried out to ensure the effective use of system resources and the controllability of tasks. Adaptation result analysis: Automatically obtain adaptation indicators, automatically calculate the overall adaptation comprehensive evaluation score based on the indicator weights, and form an adaptation conclusion report; and based on the comparison with historical adaptation reports, analyze the compatibility, stability and applicable scenarios of domestic AI accelerator cards.

10. The method for adaptively adapting a cloud platform to a domestic AI accelerator card according to claim 9, characterized in that: The script tool set adopts a modular and version control strategy. By building a unified script repository and combining it with metadata tagging technology, it can achieve accurate classification and rapid retrieval of script functions. The script tool set includes the script's CPU architecture, operating system version, script protocol, and, for installation packages, also includes the operating system and version supported by the installation package, CPU architecture, installation package content, installation package verification code, and size fields. Determine the dimensions or features to be used for evaluation. When the evaluation features are closely related to the requirements of task execution, specifically for domestic AI accelerator card adaptation, use AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization; Based on the computing resource requirements of task execution, a weight value is assigned to each feature to distinguish the importance of each resource type. The weight value is adjusted according to the needs of specific tasks. Specifically for AI performance adaptation, the weights of AI accelerator card utilization, CPU utilization, memory utilization, and network bandwidth utilization are set to 4, 3, 2, and 1, respectively. The comprehensive score of each candidate node is calculated based on the weight and quantization value. When the node with the highest comprehensive score is scheduled to execute the adaptation task, the resource availability of the CPU, AI accelerator card and memory on the node is calculated. The calculation formula is as follows: ; Wherein, j represents the node identifier; i represents the resource type identifier; The resource scheduling score can be obtained by multiplying the resource weight by the current node's availability. The formula is as follows: ; By Calculated resource scheduling scores for each type By summing up, we can calculate the overall computing power score of each node at present. , the formula is as follows: ; By comparing the overall computing power scores of each node, the optimal computing power node can be screened out, thereby achieving balanced scheduling based on computing power quantification.

Citation Information

Patent Citations

  • Method for scheduling heterogeneous AI accelerator card resources in Kubernetes cluster

    CN117632472A

  • Cloud resource automatic allocation system

    CN118363765A