Hybrid heterogeneous computing power management method and device for intelligent computing center
By determining task-specific device requirements and scheduling devices within intelligent computing centers, the method optimizes the management of mixed heterogeneous computing resources, addressing inefficiencies in resource allocation and enhancing utilization rates.
Patent Information
- Application Number
- CN202510386633.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
How to realize the unified management and efficient utilization of hybrid heterogeneous computing power in the intelligent computing center, and solve the problem of low computing resource utilization caused by differences in different regions and computing resources.
By receiving target tasks, the required equipment information is determined based on the hybrid heterogeneous computing power of the intelligent computing center, and the target equipment is scheduled to perform tasks in the equipment collection, and the equipment attributes, parameters and label information are used for screening and scheduling, so as to realize the management of hybrid heterogeneous computing power.
It realizes unified management of different regions and computing resources, improves the utilization rate of computing power resources, and improves the overall computing performance and efficiency of the computing center.
Smart Images

Figure CN120315871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly to a method and device for managing hybrid heterogeneous computing power in an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] Since the emergence of intelligent computing centers, how to manage the hybrid heterogeneous computing power of intelligent computing centers is an urgent problem to be solved. Otherwise, it is difficult to uniformly manage computing power resources for different regions and different computing resource differences, resulting in very low utilization rate of computing power resources. Summary of the Invention
[0008] The present invention provides a method and device for managing hybrid heterogeneous computing power in an intelligent computing center, which is used to solve the problem of managing the hybrid heterogeneous computing power of an intelligent computing center.
[0009] To solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for managing hybrid heterogeneous computing power in an intelligent computing center, including:
[0011] Step S1: Receive a target task;
[0012] Step S2: Determine the required device information for completing the target task based on the hybrid heterogeneous computing power of the intelligent computing center;
[0013] Step S3: Schedule the target device to execute the target task in the managed device set based on the hybrid heterogeneous computing power of the intelligent computing center and the required device information.
[0014] Optionally, the required device information includes first device information; Step S3 includes:
[0015] Step S31: Based on the hybrid heterogeneous computing power of the intelligent computing center, filter the devices in the device set that do not meet the first condition to obtain at least one device that meets the first condition, where the first condition is a condition determined based on the first device information, and the at least one device includes the target device;
[0016] Step S32: Schedule the target device to execute the target task.
[0017] Optionally, the first device information includes at least one of the following:
[0018] Device attributes, device parameters, regions to which the device belongs, device functions, operating states of the device, pre-specified devices, the number of instances allowed to run on the device, concurrent input / output I / O capabilities of the device, non-uniform memory access (NUMA) topology of the device, pre-set parameters;
[0019] Among them, the device parameters include parameters of at least one of a central processing unit (CPU), a random access memory (RAM), and a disk capacity.
[0020] Optionally, the required device information includes second device information; after Step S31 and before Step S32, the method further includes:
[0021] Step S33: Based on the hybrid heterogeneous computing power of the intelligent computing center, determine the target device among at least two devices that meet the first condition according to the weight of the second device information.
[0022] Optionally, Step S33 includes:
[0023] Step S331: When the second device information includes at least two metrics, based on the hybrid heterogeneous computing power of the intelligent computing center, weight the evaluation results corresponding to the at least two metrics of each device among the at least two devices to obtain the weighted result of each device;
[0024] Step S332: Among the at least two devices, determine the device with the largest weighted result as the target device.
[0025] Optionally, the second device information includes at least one of the device's RAM, available disk capacity, available virtual central processing unit (vCPU), workload, number of running instances on the same server, and number of scheduling failure boot attempts.
[0026] Optionally, the method further includes:
[0027] Step S4: Add tag information to each device in the device set, where the tag information is device information;
[0028] The step S3 includes:
[0029] Step S34: Based on the hybrid heterogeneous computing power of the intelligent computing center, according to the required device information and the tag information, schedule a target device in the managed device set to execute the target task.
[0030] In a second aspect, the present invention provides a hybrid heterogeneous computing power management device for an intelligent computing center, including:
[0031] A receiving module, configured to receive a target task;
[0032] A first determination module, configured to determine the required device information for completing the target task based on the hybrid heterogeneous computing power of the intelligent computing center;
[0033] A scheduling module, configured to schedule a target device in the managed device set to execute the target task according to the hybrid heterogeneous computing power of the intelligent computing center and the required device information.
[0034] In a third aspect, the present invention provides a server, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the hybrid heterogeneous computing power management method for the intelligent computing center as described in the first aspect above are implemented.
[0035] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the hybrid heterogeneous computing power management method for the intelligent computing center as described in the first aspect above are implemented.
[0036] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the hybrid heterogeneous computing power management method for the intelligent computing center as described in the first aspect above are implemented.
[0037] In the present invention, based on the hybrid heterogeneous computing power of the intelligent computing center, the required device information corresponding to the target task is determined, and the target device is called from the managed device set according to the required device information to execute the target task, which can realize the management of the hybrid heterogeneous computing power of the intelligent computing center. For different regions and different computing resource differences, it is convenient to uniformly manage the computing power resources and improve the utilization rate of computing power resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0039] Figure 1 is a schematic flowchart of the method for managing the hybrid heterogeneous computing power of the intelligent computing center of the present invention;
[0040] Figure 2 is the test scheme of the hybrid heterogeneous computing power management platform of the present invention;
[0041] Figure 3 is a schematic structural diagram of the device for managing the hybrid heterogeneous computing power of the intelligent computing center of the present invention;
[0042] Figure 4 is a schematic structural diagram of a server of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The technical solutions in the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0044] First, the technical terms related to the present invention will be briefly described below.
[0045] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0046] The "Computational Power (CP)" described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing power of a data center and includes general computing power, supercomputing power, and intelligent computing power. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS). The larger the value, the stronger the comprehensive computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。
[0047] The "Network Power (NP)" described in the present invention refers to: the manifestation of the data transmission ability of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.
[0048] The "Storage Power (SP)" described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage ability of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0049] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying power, and data storage power, and can realize centralized computing, storage, transmission, and application of information.
[0050] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0051] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and supercomputing power.
[0052] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0053] The "intelligent computing power" described in the present invention refers to a computing platform that is deployed on a large scale for various artificial intelligence innovation applications based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing (NLP), machine vision, and so on.
[0054] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0055] The "intelligent computing center" described in the present invention refers to a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0056] The "intelligent computing center" described in the present invention includes but is not limited to intelligent computing centers.
[0057] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0058] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, which has computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0059] The "supercomputing center" described in the present invention refers to: namely, a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0060] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0061] The "hybrid heterogeneous computing power" described in the present invention refers to: combining multiple different architectures and different types of computing units (such as CPUs, GPUs, FPGAs, etc.) to form a collaborative computing system. It realizes higher computing performance and energy efficiency ratio by integrating the advantages of different types of devices.
[0062] The "unified management" described in the present invention refers to: uniformly managing and scheduling at least two or more computing resources to achieve centralized resource management and improve resource utilization.
[0063] See Figure 1 , Figure 1 which is a method for unified management of hybrid heterogeneous computing power in an intelligent computing center provided by the present invention, including:
[0064] Step S1: Receive a target task;
[0065] Step S2: Based on the hybrid heterogeneous computing power of the intelligent computing center, determine the required device information for completing the target task;
[0066] Step S3: Based on the hybrid heterogeneous computing power of the intelligent computing center and the required device information, schedule target devices in the unified managed device set to execute the target task.
[0067] Among them, the target task can be a task to be processed, such as tasks like virtual machine instance creation, container deployment, batch processing jobs, etc. The unified managed device set can include physical devices and / or virtual devices.
[0068] Based on the hybrid heterogeneous computing power resources of the intelligent computing center, parse the target task to obtain the computing resources required to complete the target task and the required device information.
[0069] The device information may include the attribute information of the device (such as the architecture type x86), device parameters (such as the number of CPU cores, the capacity of Random Access Memory (RAM), disk space, etc.), regional restrictions, etc., and may also include resource utilization metrics (such as available Virtual Central Processing Unit (vCPU), remaining RAM), load status (Input / Output (I / O) concurrency, number of running instances), the number of historical scheduling failures, etc.
[0070] Before performing scheduling, tags can be added to the set of managed devices according to the device information. In this way, when receiving a target task, according to the required device information and the added tags, the target device is screened and allocated from the set of managed devices to execute the target task.
[0071] For example, when obtaining a task for large model training, the task is parsed to determine that the task has a high demand for the CPU, and the target device corresponding to the high-performance CPU type is scheduled to execute.
[0072] Another example is that if the device information required by the task is 8-core CPU + 32GB RAM, the device parameters need to meet the conditions of CPU cores ≥ 8 and RAM ≥ 32GB.
[0073] The above method of the present invention can be applied to various architectures such as, for example, bare metal architecture, heterogeneous computing architecture, computing optimization architecture, graphics computing optimization architecture, etc. The above architectures can adopt technologies such as passthrough technology, virtualization technology, etc.
[0074] Before performing task scheduling, tags can be defined for the images corresponding to the above architectures. Thus, the corresponding devices are matched according to the tags of the images.
[0075] Optionally, in some embodiments, the required device information includes first device information; the step S3 includes:
[0076] Step S31: Based on the hybrid heterogeneous computing power of the intelligent computing center, filter the devices in the device set that do not meet the first condition to obtain at least one device that meets the first condition. The first condition is a condition determined based on the first device information, and the at least one device includes the target device;
[0077] Step S32: Schedule the target device to execute the target task.
[0078] Among them, the first condition includes a condition for the first device information. For example, when the first device information includes the core utilization rate of the device CPU, the first condition may be that the CPU core utilization rate is greater than a preset value; when the first device information includes the device function, the first condition may be that the device has function A.
[0079] By setting the first condition, devices that do not meet the first condition are filtered out, leaving at least one device that meets the first condition.
[0080] When there is 1 device that meets the first condition, schedule this device to execute the target task.
[0081] When there are more than two devices that meet the first condition, a target device can be further screened and determined from these devices according to a screening mechanism.
[0082] When screening, the devices can also be aggregated first. For example, aggregated by region or by specific characteristics. When screening, the aggregated devices can be screened according to the device information, which can improve the screening efficiency.
[0083] Physical hosts with certain characteristics are grouped into a cluster of hosts. A physical host can be added to a set of multiple aggregated hosts to cooperate with the flavor (virtual machine template) to control the virtual machine to select hosts with different architectures or different configurations for the physical host. Host aggregation groups can be created separately. When creating a cloud host, specify the host aggregation group to match multi-architecture and multi-category host devices.
[0084] For example, the following code can be used in the front end to classify the managed platform and managed devices:
[0085] export const flavorArchitectures={
[0086] x86_architecture:t('X86 Architecture'),
[0087] heterogeneous_computing:t('Heterogeneous Computing'),
[0088] bare_metal:t('Bare Metal'),
[0089] arm_architecture:t('ARM Architecture'),
[0090] custom:t('Custom'),
[0091] all: t('All Flavors'),
[0092] };
[0093] export const x86CategoryList = {
[0094] general_purpose: t('General Purpose'),
[0095] compute_optimized: t('Compute Optimized'),
[0096] memory_optimized: t('Memory Optimized'),
[0097] big_data: t('Big Data'),
[0098] local_ssd: t('Local SSD'),
[0099] high_clock_speed: t('High Clock Speed'),
[0100] };
[0101] Optionally, in some embodiments, the first device information includes at least one of the following:
[0102] device attributes, device parameters, region to which the device belongs, device functions, operating status of the device, pre-specified devices, number of instances allowed to run on the device, concurrent input / output I / O capabilities of the device, non-uniform memory access (NUMA) topology of the device, pre-set parameters;
[0103] Among them, the device parameters include parameters of at least one of a central processing unit (CPU), random access memory (RAM), and disk capacity.
[0104] Device attributes may include inherent characteristics at the hardware or software level of the device. For example, the model of the device, operating system, CPU architecture, etc. For example, physical host architectures are distinguished by tags or architecture type x86_64.
[0105] When filtering devices according to device attributes, for example, matching the attributes required by the image (such as a host with specified attributes defined on the image, then filtering out hosts other than the above attributes).
[0106] Device parameters may include CPU, RAM, disk capacity, etc. Among them, the parameters of the CPU may include the number of cores, clock frequency, etc.
[0107] When filtering devices according to device parameters, for example, filtering devices with cores that do not meet the requirements based on CPU core utilization; or selecting devices with sufficient RAM according to RAM.
[0108] For another example, filter the host CPU allocation ratio (cpu_allocation_ratio) by CPU core number and each aggregation. If the value for each aggregation is not found, it will fallback to the global default cpu_allocation_ratio. If multiple values are found for a host (meaning the host is in two different aggregations with different ratio settings), the minimum value can be used.
[0109] For yet another example, confirm whether the device host has the physical PCI devices (such as GPUs) required by the instance.
[0110] The disk capacity represents the size of the storage space.
[0111] When selecting a target device according to the disk capacity, for example: filter the hosts by the disk allocation ratio, and only those with enough disk space to host the instance. The disk allocation ratio of the virtual disk to the physical disk is defaulted to 1.0. The total disk size allowed for allocation is the physical disk multiplied by this ratio.
[0112] The region to which the device belongs can be a region determined according to the geographical location or the network address. For example, restrict the instance to run in the specified availability zone; when multiple regions are specified, any one of the regions can be selected.
[0113] The device functions, for example, support specific hardware or topology requirements. For example, check whether the functions provided by the device computing service meet any additional specifications associated with the instance type. A host with the specified instance type can be created.
[0114] The running status of the device, including the running status, enabled status, non-overloaded status, isolated status, etc. For example, exclude hosts in the resource - insufficient status.
[0115] The pre - specified devices, that is, the specific devices or device sets pre - set. For example, check whether the aggregation metadata meets any additional specifications associated with the instance type. A host with the specified instance type can be created.
[0116] The number of instances allowed to run on the device, that is, the maximum number of instances allowed to run pre - set, and determine whether the conditions of the first device information are met based on the maximum number of instances.
[0117] The concurrent input / output I / O capacity of the device. According to the device concurrent capacity, determine whether the device resource requirements for completing the target task are met.
[0118] The non-uniform memory access (NUMA) topology of the device, for example, filters the corresponding devices according to the topology requirements of the instance request.
[0119] Other preset parameters, which can set the conditions for other parameters of the device according to the task type, etc.
[0120] For example, deploy the instance and the specified instance group on the same host;
[0121] Create a server group with an "anti-affinity" policy through the server group API. When booting a new server, provide a scheduler hint. When a server is scheduled, anti-affinity will be enforced among all servers in the group.
[0122] The maximum number of I / O-intensive instances allowed to run on this host.
[0123] Filter the hosts (i.e., devices) by setting the I / O per aggregation. If the value for each aggregation is not found, it will fallback to the global default value. If multiple values are found for a host (indicating that the device host is in two or more different aggregations with different maximum I / O settings), the minimum value can be used.
[0124] Filter the hosts by the number of instances set with the maximum number of instances per host per aggregation. If the value for each aggregation is not found, it will fallback to the global default value. If multiple values are found for a host (indicating that the host is in two or more different aggregations with different maximum instance number settings per host), the minimum value will be used.
[0125] Filter the compute nodes (devices) by the number of running instances.
[0126] Filter the host memory allocation ratio setting by RAM per aggregation. If the value for each aggregation is not found, it will fallback to the global default value. If multiple values are found for a host (indicating that the host is in two different aggregations with different ratio settings), the minimum value will be used.
[0127] Allow new instances on hosts within the same IP block.
[0128] Filter the hosts by the concurrent I / O on the host, and filter out hosts with excessive concurrent I / O.
[0129] Filter the hosts that have attempted to be scheduled, and only pass the hosts that have not been attempted before.
[0130] Do not filter and indicate all available host devices. When specifying multiple key values, the multiple key values can be separated by commas, such as "value1, value2". If not specified, it means passing all hosts.
[0131] Based on the above first device information, corresponding first conditions can be set for screening, so as to obtain devices that meet the first conditions to execute the target task.
[0132] Optionally, in some embodiments, the required device information includes second device information; after step S31 and before step S32, the method further includes:
[0133] Step S33: Based on the hybrid heterogeneous computing power of the intelligent computing center, determine the target device among at least two devices that meet the first conditions according to the weights of the second device information.
[0134] Among them, the second device information can also include information such as the attributes and parameters of the device. The second device information can include one or more metrics, and corresponding weights can be set for each metric.
[0135] Filter out some devices according to the first conditions to obtain at least two devices that meet the first conditions. Among these at least two devices, further sort according to the weights, so as to select the device with the largest or smallest weight.
[0136] For example, when the second device information includes the disk capacity of the device, the weights of each device can be obtained among at least two devices that meet the first conditions according to the disk capacity, so as to select the device with the largest weight as the target device.
[0137] Optionally, in some embodiments, step S33 includes:
[0138] Step S331: When the second device information includes at least two metrics, based on the hybrid heterogeneous computing power of the intelligent computing center, weight the evaluation results corresponding to the at least two metrics of each device among the at least two devices to obtain the weighted result of each device;
[0139] Step S332: Among the at least two devices, determine the device with the largest weighted result as the target device.
[0140] Among them, two metrics of the device, for example, the RAM and disk capacity of the device.
[0141] When the second device information includes at least two metrics, the evaluation results of each metric are weighted according to the weights to obtain the weighted result of each device.
[0142] For example, the evaluation result of the first device for the metric RAM is 90, and the evaluation result for the metric disk capacity is 40 points; the evaluation result of the second device for the metric RAM is 60, and the evaluation result for the metric disk capacity is 80 points; the weight of the metric RAM is 0.8, and the weight of the disk capacity is 0.2. Weighting according to the above weights, the weighted results of the two devices are 80 and 64 respectively.
[0143] Determine the device with the largest weighted result as the target device, indicating that the comprehensive performance of the target device is better.
[0144] Optionally, in some embodiments, the second device information includes at least one of the device's RAM, available disk capacity, available virtual central processing unit vCPU, workload, number of running instances on the same server, and number of scheduling failure boot attempts.
[0145] According to any one of the above second device information, multiple devices that meet the first condition can be sorted according to their weights, and thus the device with the largest or smallest sort can be selected.
[0146] As some alternative embodiments, the weights can be determined according to the following second device information:
[0147] Calculate the weight according to the available RAM on the computing node (i.e., the device) and sort by weight. If the memory weight multiplier of the filtering scheduler is negative, then the host with the least available RAM (for stacked hosts rather than dispersed hosts).
[0148] The hosts are weighted and sorted according to the available disk space, and the one with the largest weight wins. If the multiplier is negative, then the host with less available disk space will win (for stacked hosts rather than dispersed hosts).
[0149] Calculate the weight according to the available vCPU on the computing node and sort by the combined weight. If the CPU weight multiplier of the filtering scheduler is negative, then the host with the least available CPU (for stacked hosts rather than dispersed hosts).
[0150] Calculate the weight according to the workload of the computing node host. By default, it is best to select a host with a light workload for the PCI weight evaluator: Calculate the weight according to the number of PCI devices on the host and the number of PCI devices requested by the instance.
[0151] Calculate the weight according to the number of instances running on the same server group. The largest weight defines the preferred host for the new instance.
[0152] Calculate the weight according to the number of instances running on the same server group as a negative value. The largest weight defines the preferred host for the new instance.
[0153] Calculate the weight of the host according to the number of recent failed boot attempts of the device. Preferentially call the device with fewer failed boot attempts.
[0154] Calculate the weight according to various metrics of the computing node host.
[0155] According to the above second device information, determine the weight or weighted value of each device according to one or more items, so as to obtain a ranking according to the weight size or weighted result size, and select the most qualified call.
[0156] Based on the above method, the most qualified device can be further screened to execute the target task, thereby improving the resource utilization rate.
[0157] Optionally, in some embodiments, the method further includes:
[0158] Step S4: Add tag information to each device in the device set, where the tag information is device information;
[0159] The step S3 includes:
[0160] Step S34: Based on the hybrid heterogeneous computing power of the intelligent computing center, according to the required device information and the tag information, schedule the target device in the managed device set to execute the target task.
[0161] Add tag information to each device in the device set.
[0162] For example, according to the type of CPU, classify the devices in the managed device set as follows, and add corresponding tags to the devices in each category:
[0163] The first category is general purpose: balanced CPU, memory and storage resources, suitable for a variety of application scenarios.
[0164] The second category is compute optimized: high CPU performance, usually with more CPU cores and higher clock speeds, suitable for tasks that require a large amount of computing resources but have low requirements for memory and storage.
[0165] The third category is memory optimized: a large amount of memory resources, suitable for tasks that require high memory bandwidth and low latency, and suitable for application programs that need to frequently access and process a large amount of data.
[0166] The fourth category is high clock speed: CPUs with high clock speeds, which can process more instructions per unit time, suitable for tasks that require fast processing and response.
[0167] When receiving a target task, according to the tags of each device, it is determined whether the corresponding device requirement information conditions are met, so as to facilitate quickly filtering or selecting the corresponding device.
[0168] See Figure 2 , Figure 2 is a test solution for the hybrid heterogeneous computing power management platform provided by the present invention. The test purpose is to verify that the computing power resources (computing power platform) support the management of hybrid heterogeneous computing power. Through the test, it is shown that the hybrid heterogeneous computing power management platform can support the x86 computing resource architecture type and the GPU computing resource type.
[0169] See Figure 3 , Figure 3 is a hybrid heterogeneous computing power management device 200 of an intelligent computing center provided by the present invention. The device includes:
[0170] A receiving module 301, configured to receive a target task;
[0171] A first determination module 302, configured to determine the required device information for completing the target task based on the hybrid heterogeneous computing power of the intelligent computing center;
[0172] A scheduling module 303, configured to schedule a target device to execute the target task in the managed device set according to the hybrid heterogeneous computing power of the intelligent computing center and the required device information.
[0173] Optionally, the required device information includes first device information; the scheduling module 303 includes:
[0174] A filtering sub-module, configured to filter devices in the device set that do not meet the first condition based on the hybrid heterogeneous computing power of the intelligent computing center, to obtain at least one device that meets the first condition, where the first condition is a condition determined based on the first device information, and the at least one device includes the target device;
[0175] A scheduling sub-module, configured to schedule the target device to execute the target task.
[0176] Optionally, the first device information includes at least one of the following:
[0177] Device attributes, device parameters, regions to which the device belongs, device functions, operating states of the device, pre-specified devices, the number of instances allowed to run on the device, concurrent input / output I / O capabilities of the device, non-uniform memory access NUMA topology of the device, pre-set parameters;
[0178] Among them, the device parameters include parameters of at least one of a central processing unit CPU, a random access memory RAM, and a disk capacity.
[0179] Optionally, the required device information includes second device information; the apparatus further includes:
[0180] A second determination module, configured to determine the target device from at least two devices that meet the first condition according to the weight of the second device information based on the hybrid heterogeneous computing power of the intelligent computing center.
[0181] Optionally, the second determination module includes:
[0182] A calculation sub-module, configured to, when the second device information includes at least two metrics, perform weighting on the evaluation results corresponding to the at least two metrics of each device among the at least two devices based on the hybrid heterogeneous computing power of the intelligent computing center to obtain a weighted result for each device;
[0183] A determination sub-module, configured to determine, among the at least two devices, the device with the largest weighted result as the target device.
[0184] Optionally, the second device information includes at least one of the device's RAM, available disk capacity, available virtual central processing unit (vCPU), workload, number of running instances on the same server, and number of scheduling failure boot attempts.
[0185] Optionally, the apparatus further includes:
[0186] An adding module, configured to add tag information to each device in the device set, where the tag information is device information;
[0187] The scheduling module 303 is specifically configured to:
[0188] Based on the hybrid heterogeneous computing power of the intelligent computing center, schedule a target device to execute the target task in the managed device set according to the required device information and the tag information.
[0189] The hybrid heterogeneous computing power management apparatus of the intelligent computing center provided by the present invention can implement each process of the above-described embodiments of the hybrid heterogeneous computing power management method of the intelligent computing center, with the technical features corresponding one by one and achieving the same technical effects. To avoid repetition, details are not described herein again.
[0190] It should be noted that the hybrid heterogeneous computing power management apparatus of the intelligent computing center in the present invention may be a device, or a component, integrated circuit, or chip in an electronic device.
[0191] Please refer to Figure 4, the present invention also provides a server 110, including a processor 111, a memory 112, and a computer program stored on the memory 112 and executable on the processor 111. When the computer program is executed by the processor 111, it implements each process of the above-mentioned method embodiment for hybrid heterogeneous computing power management in the intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0192] The present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, it implements each process of the above-mentioned method embodiment for hybrid heterogeneous computing power management in the intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium can be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0193] The embodiment of the present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above-mentioned Figure 1 method embodiment for hybrid heterogeneous computing power management in the intelligent computing center shown and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0194] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0196] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.
Claims
1. A method for hybrid heterogeneous computing power management in an intelligent computing center, characterized in that Including: Step S1: Receive a target task; Step S2: Based on the hybrid heterogeneous computing power of the intelligent computing center, determine the required device information for completing the target task; Step S3: Based on the hybrid heterogeneous computing power of the intelligent computing center and the required device information, schedule a target device in the managed device set to execute the target task.
2. The method according to claim 1, wherein The required device information includes first device information; Step S3 includes: Step S31: Based on the hybrid heterogeneous computing power of the intelligent computing center, filter the devices in the device set that do not meet the first condition to obtain at least one device that meets the first condition. The first condition is a condition determined based on the first device information, and the at least one device includes the target device; Step S32: Schedule the target device to execute the target task.
3. The method according to claim 2, characterized in that, The first device information includes at least one of the following: Device attributes, device parameters, the area to which the device belongs, device functions, the operating status of the device, pre-specified devices, the number of instances allowed to run on the device, the concurrent input / output I / O capabilities of the device, the non-uniform memory access (NUMA) topology of the device, pre-set parameters; Among them, the device parameters include parameters of at least one of the central processing unit (CPU), random access memory (RAM), and disk capacity.
4. The method according to claim 2, wherein The required device information includes second device information; after Step S31 and before Step S32, the method further includes: Step S33: Based on the hybrid heterogeneous computing power of the intelligent computing center, determine the target device among at least two devices that meet the first condition according to the weight of the second device information.
5. The method according to claim 4, characterized in that Step S33 includes: Step S331: When the second device information includes at least two metrics, based on the hybrid heterogeneous computing power of the intelligent computing center, weight the evaluation results corresponding to the at least two metrics of each device among the at least two devices to obtain the weighted result of each device; Step S332: Among the at least two devices, determine the device with the largest weighted result as the target device.
6. The method according to claim 4, characterized in that, The second device information includes at least one of the RAM of the device, available disk capacity, available virtual central processing unit (vCPU), workload, number of running instances on the same server, and number of scheduling failure retry attempts.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: Step S4: Add tag information to each device in the device set, and the tag information is device information; Step S3 includes: Step S34: Based on the hybrid heterogeneous computing power of the intelligent computing center, according to the required device information and the tag information, schedule a target device in the managed device set to execute the target task.
8. A hybrid heterogeneous computing power management device for an intelligent computing center, characterized in that, Including: A receiving module for receiving a target task; A first determination module for determining the required device information for completing the target task based on the hybrid heterogeneous computing power of the intelligent computing center; A scheduling module for scheduling a target device in the managed device set to execute the target task based on the hybrid heterogeneous computing power of the intelligent computing center and the required device information.
9. A server, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor, the program, when executed by the processor, implementing the steps of the method for hybrid heterogeneous computing power management of the intelligent computing center according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the method for hybrid heterogeneous computing power management of the intelligent computing center according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, it implements the steps of the method for hybrid heterogeneous computing power management of the intelligent computing center according to any one of claims 1 to 7.