Computing force load balancing method and device of intelligent computing center

By obtaining the computing node status data of the intelligent computing center and reasonably allocating tasks based on task type and node capability evaluation information, the problem of unbalanced load of computing nodes is solved, and efficient utilization of computing power resources and stable system operation is achieved.

CN120448110APending Publication Date: 2025-08-08DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510517706.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the intelligent computing center, unreasonable allocation of computing tasks leads to unbalanced load of computing nodes, which may lead to excessive or low individual nodes, affect business operations or waste of resources. How to achieve computing power load balancing has become an urgent problem.

Method used

By obtaining the operating status data of the computing node, determining the load capacity evaluation information based on the type of computing task and the node status, and assigning the tasks to the most suitable target computing node, monitoring and adjusting the task allocation in real time to ensure load balancing.

Benefits of technology

The full and rational use of computing power resources of each computing node is achieved, the concentration of computing tasks is avoided, and the task processing efficiency and system stability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448110A_ABST
    Figure CN120448110A_ABST
Patent Text Reader

Abstract

The invention provides a computing power load balancing method and device for an intelligent computing center, and relates to the technical field of computing power infrastructure, and the method comprises the steps: obtaining the operation state data of a plurality of computing nodes; receiving a calculation task; determining load capacity evaluation information of the plurality of computing nodes according to the task type of the computing task and the running state data of the plurality of computing nodes; and distributing the calculation task to a target calculation node in the plurality of calculation nodes according to the load capacity evaluation information of the plurality of calculation nodes. Thus, it is avoided that calculation tasks are concentrated on a few calculation nodes, calculation power resources of all the calculation nodes are fully and reasonably utilized, and therefore load balancing of calculation power is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a computing power load balancing method and device for an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.

[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] As the core hub for massive data processing and complex computing tasks, the Intelligent Computing Center is responsible for providing powerful computing power support for numerous enterprises, research institutions, and various intelligent applications. When a large number of computing tasks appear in the Intelligent Computing Center, if the computing tasks are not allocated properly or the computing tasks are unavailable, it may cause the load on a single computing node to be too high or too low, or the task may be scheduled to an unavailable computing task, resulting in scheduling failure. If the computing node is overloaded, its processing capacity will reach its limit, causing the computing task processing speed to slow down and affecting the normal operation of the business; if the computing node is underloaded, computing resources will be idle, resulting in resource waste; if the computing node is unavailable, it will be automatically scheduled to other available computing nodes. Therefore, since the emergence of intelligent computing centers, how to achieve computing power load balancing has become a technical problem that needs to be solved urgently. Summary of the Invention

[0008] The present invention provides a computing power load balancing method and device for an intelligent computing center to solve the problem of how to achieve computing power load balancing.

[0009] To solve the above problems, the present invention is achieved as follows:

[0010] In a first aspect, the present invention provides a method for balancing computing power load in an intelligent computing center, comprising:

[0011] Step S1: Obtaining the operating status data of multiple computing nodes;

[0012] Step S2: receiving a computing task;

[0013] Step S3: determining load capacity evaluation information of the plurality of computing nodes according to the task type of the computing task and the operating status data of the plurality of computing nodes;

[0014] Step S4: Allocate the computing task to a target computing node among the multiple computing nodes according to the load capacity evaluation information of the multiple computing nodes.

[0015] In one embodiment, step S3 includes:

[0016] Step S31: determining the resource requirement of the computing task according to the task type of the computing task;

[0017] Step S32: determining the remaining amount of resources of the plurality of computing nodes according to the operating status data of the plurality of computing nodes;

[0018] Step S33: Determine load capacity evaluation information of the multiple computing nodes based on the resource requirements of the computing task and the remaining resources of the multiple computing nodes.

[0019] In one embodiment, step S33 includes:

[0020] Step S331: Obtain performance data of the multiple computing nodes;

[0021] Step S332: Calculate the weights of the plurality of computing nodes according to the performance data of the plurality of computing nodes;

[0022] Step S333: Determine load capacity evaluation information of the multiple computing nodes according to the resource demand of the computing task, the remaining resources of the multiple computing nodes, and the weights of the multiple computing nodes.

[0023] In one embodiment, the load capacity evaluation information of the plurality of computing nodes includes load capacity evaluation values of the plurality of computing nodes, and step S4 includes:

[0024] Step S41: sorting the load capacity evaluation values of the plurality of computing nodes in descending order;

[0025] Step S42: Select computing nodes whose load capacity evaluation values rank in the top N and are greater than a first preset threshold as target computing nodes, where N is a positive integer.

[0026] In one embodiment, after step S4, the method further includes:

[0027] Step S5: During the execution of the computing task, monitoring the load value of the target computing node;

[0028] Step S6: When the load value of the target computing node exceeds a second preset threshold, part or all of the computing task is transferred to other computing nodes among the multiple computing nodes except the target computing node.

[0029] In one embodiment, the method further comprises:

[0030] Step S7: setting a network bandwidth limit and a network flow limit for the target computing node according to the resource demand of the computing task and the resource usage of the target computing node;

[0031] Step S8: monitor the network traffic and network traffic of the target computing node in real time;

[0032] Step S9: When the network traffic of the target computing node exceeds the network bandwidth limit and / or the network traffic exceeds the network traffic limit, take flow limiting measures.

[0033] In a second aspect, the present invention further provides a computing power load balancing device for an intelligent computing center, comprising:

[0034] A first acquisition module is used to obtain the operating status data of multiple computing nodes;

[0035] A first receiving module, configured to receive a computing task;

[0036] A first determining module is configured to determine load capacity evaluation information of the plurality of computing nodes according to a task type of the computing task and the operating status data of the plurality of computing nodes;

[0037] The first allocation module is configured to allocate the computing task to a target computing node among the multiple computing nodes according to load capacity evaluation information of the multiple computing nodes.

[0038] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for computing power load balancing of the intelligent computing center as described in the first aspect above are implemented.

[0039] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the computing power load balancing method of the intelligent computing center as described in the first aspect above.

[0040] In a fifth aspect, the present invention further provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps in the computing power load balancing method of the intelligent computing center as described in the first aspect above.

[0041] In the present invention, the operating status data of multiple computing nodes is obtained; a computing task is received; load capacity assessment information of the multiple computing nodes is determined based on the task type of the computing task and the operating status data of the multiple computing nodes; and the computing task is allocated to a target computing node among the multiple computing nodes based on the load capacity assessment information of the multiple computing nodes. This avoids the concentration of computing tasks on a small number of computing nodes, allows the computing power resources of each computing node to be fully and reasonably utilized, and thus achieves computing load balancing. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0043] Figure 1 This is a flow chart of a computing power load balancing method for an intelligent computing center provided by the present invention;

[0044] Figure 2 This is a structural diagram of a computing power load balancing device for an intelligent computing center provided by the present invention;

[0045] Figure 3 This is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0048] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0049] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0050] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.

[0051] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.

[0052] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0053] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0054] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on central processing unit (CPU) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0055] The "intelligent computing power" mentioned in the present invention refers to: a computing platform based on large-scale deployment of special chips such as graphics processing units (GPUs), field programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs) for various innovative artificial intelligence applications, such as natural language processing and machine vision.

[0056] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.

[0057] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0058] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".

[0059] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0060] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0061] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0062] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0063] The "computing power load balancing" mentioned in the present invention refers to a technology that evenly distributes computing power tasks to multiple computing nodes to avoid excessive or low load on a single computing node, thereby improving the overall computing power resource utilization efficiency and ensuring stable and efficient operation of the system.

[0064] In the prior art, when a large number of computing tasks appear in an intelligent computing center, if the computing tasks are not allocated reasonably or the computing tasks are unavailable, it may cause the load of a single computing node to be too high or too low, or it may be scheduled to an unavailable computing task, resulting in scheduling failure. If the load of the computing node is too high, its processing capacity reaches its limit, which will cause the processing speed of the computing task to slow down, affecting the normal operation of the business; if the load of the computing node is too low, it will cause the computing resources to be idle, resulting in resource waste; if the computing node is unavailable, it will be automatically scheduled to other available computing nodes. Therefore, since the emergence of intelligent computing centers, how to achieve computing power load balancing of each computing node has become a technical problem that needs to be solved urgently. In order to achieve computing power load balancing of the intelligent computing center, in the present invention, the operating status data of multiple computing nodes are obtained; computing tasks are received; according to the task type of the computing task and the operating status data of the multiple computing nodes, the load capacity evaluation information of the multiple computing nodes is determined; according to the load capacity evaluation information of the multiple computing nodes, the computing task is allocated to the target computing node among the multiple computing nodes. In this way, computing tasks are avoided from being concentrated in a few computing nodes, and the computing power resources of each computing node are fully and reasonably utilized, thereby achieving load balancing of computing power.

[0065] For details, see Figure 1 , Figure 1 This is a flow chart of a computing power load balancing method for an intelligent computing center provided by the present invention. Figure 1 As shown, the following steps are included:

[0066] Step S1: Obtaining the operating status data of multiple computing nodes;

[0067] In this step, computing tasks within an intelligent computing center are typically distributed across multiple compute nodes. Each compute node is an entity with independent computing capabilities, and may be a server, a virtual machine, or a computing unit. Obtaining operational status data for these compute nodes is essential for the subsequent rational allocation of computing tasks. This data may include CPU utilization, memory usage, and disk I / O status.

[0068] Step S2: receiving a computing task;

[0069] In this step, the Intelligent Computing Center receives external computing tasks. These tasks come from a wide range of sources, including user-submitted data analysis requests, simulation tasks in scientific research projects, and batch data processing requirements in enterprise business processes.

[0070] Step S3: determining load capacity evaluation information of the plurality of computing nodes according to the task type of the computing task and the operating status data of the plurality of computing nodes;

[0071] In this step, different types of computing tasks have very different requirements for computing resources. For example, data mining tasks often require a large amount of CPU computing power and memory to process and analyze data, while graphics rendering tasks rely more on the powerful computing power of the GPU to quickly generate images.

[0072] Combined with the previously acquired compute node operating status data, we can evaluate each compute node's load capacity for the current computing task. For a specific computing task, if a compute node has low CPU utilization, sufficient memory, and hardware resources that match the task type (such as a high-performance GPU for graphics rendering), then the node's load capacity is strong. Conversely, if the node's resources are already heavily utilized and lack the necessary hardware support, then its load capacity is weak. Through this comprehensive evaluation, a load capacity assessment can be generated for each compute node, which intuitively reflects the node's ability to handle the current task.

[0073] Step S4: Allocate the computing task to a target computing node among the multiple computing nodes according to the load capacity evaluation information of the multiple computing nodes.

[0074] In this step, after obtaining load capacity assessment information for each compute node, the Intelligent Computing Center's load balancer uses this information to assign computing tasks to the most appropriate target compute nodes. This assignment is based on ensuring efficient and stable task execution. If a task requires high CPU power, the system prioritizes nodes with low CPU utilization and high computing power. If the task involves storing and processing large amounts of data, the system selects nodes with ample memory and good disk I / O performance.

[0075] In the above embodiment, the operating status data of multiple computing nodes is obtained; a computing task is received; load capacity assessment information of the multiple computing nodes is determined based on the task type of the computing task and the operating status data of the multiple computing nodes; and the computing task is allocated to a target computing node among the multiple computing nodes based on the load capacity assessment information of the multiple computing nodes. This avoids the concentration of computing tasks on a small number of computing nodes, allows the computing power resources of each computing node to be fully and reasonably utilized, and thus achieves computing load balancing.

[0076] In one embodiment, step S3 includes:

[0077] Step S31: determining the resource requirement of the computing task according to the task type of the computing task;

[0078] Step S32: determining the remaining amount of resources of the plurality of computing nodes according to the operating status data of the plurality of computing nodes;

[0079] Step S33: Determine load capacity evaluation information of the multiple computing nodes based on the resource requirements of the computing task and the remaining resources of the multiple computing nodes.

[0080] In the above embodiments, different types of computing tasks have significantly different requirements for computing resources. Computing tasks are diverse, such as data mining, machine learning training, graphics rendering, video encoding, etc. Each task type has its own unique resource requirements.

[0081] Take data mining tasks, for example. These typically require processing large amounts of data, placing high demands on CPU computing power and memory storage capacity. Data mining requires sorting, filtering, and analyzing massive amounts of data, requiring powerful CPU computing power to ensure processing speed, as well as sufficient memory to store intermediate results and data. Graphics rendering tasks, on the other hand, rely heavily on the parallel computing capabilities of the GPU. Graphics rendering requires processing large amounts of graphic data, such as lighting calculations, texture mapping, and polygon rendering. The GPU's parallel architecture efficiently completes these tasks, placing extremely high demands on GPU performance.

[0082] By clarifying the type of computing task, we can more accurately estimate the specific amount of CPU, memory, GPU, disk I / O and other resources required to complete the task based on past experience and the characteristics of that type of task, that is, determine the resource requirements of the computing task.

[0083] Each compute node has its own hardware resource configuration, including the number of CPU cores, memory capacity, GPU performance, and disk space. During operation, these nodes continuously process various tasks, consuming a certain amount of resources. By obtaining operational status data for multiple compute nodes, we can understand the current resource usage of each node. For example, by monitoring CPU utilization, we can determine the proportion of tasks currently being processed by the CPU to its total processing capacity; by viewing memory usage, we can determine the occupied and remaining available memory space. By subtracting the currently used resources from each node's total resources, we can determine the node's remaining resources, including remaining CPU computing power, memory space, GPU processing power, and available disk space.

[0084] After obtaining the resource requirements of the computing task and the remaining resources of each computing node, the load capacity of each computing node can be evaluated.

[0085] For a specific computing task, if a computing node's remaining resources can meet the task's resource requirements with some margin, then the node has a strong load capacity. For example, if a computing task requires 2GB of memory and 20% of CPU computing power, and a node has 3GB of memory and 30% of CPU computing power remaining, then the node is capable of handling this task and has some resources left to handle other possible tasks.

[0086] Conversely, if a node's remaining resources cannot meet the task's resource requirements, or can only barely meet them, then the node's load capacity is weak. By comprehensively comparing the remaining resources of all nodes with the task's resource requirements, a load capacity assessment can be generated for each node. This information can be expressed in numerical values, levels, or other formats, intuitively reflecting each node's ability to handle the current computing task.

[0087] In this embodiment, by separately determining the resource requirements of a computing task and the remaining resources of a computing node, each node's task processing capability can be more accurately assessed. When allocating tasks, tasks can be accurately assigned to the most suitable nodes, avoiding assigning tasks to nodes with insufficient resources, thereby improving the success rate and efficiency of task processing.

[0088] In one embodiment, step S33 includes:

[0089] Step S331: Obtain performance data of the multiple computing nodes;

[0090] Step S332: Calculate the weights of the plurality of computing nodes according to the performance data of the plurality of computing nodes;

[0091] Step S333: Determine load capacity evaluation information of the multiple computing nodes according to the resource demand of the computing task, the remaining resources of the multiple computing nodes, and the weights of the multiple computing nodes.

[0092] In the above embodiments, the performance of each computing node varies. Performance data can reflect the node's ability to process computing tasks. This performance data includes multiple aspects, such as the CPU's main frequency, number of cores, cache size, etc.

[0093] To more accurately measure the importance and processing power of each computing node in the entire system, its weight needs to be calculated based on performance data. Weight is a relative value that represents the degree of advantage of the node in processing tasks. There are various methods for calculating weights. For example, a weighted average can be used to assign different weight coefficients to different performance indicators. Then, a comprehensive score can be calculated based on the performance data of each node, and the comprehensive score is converted into a weight. Assuming that the weight coefficients for CPU performance, memory performance, GPU performance, and disk performance are 0.4, 0.2, 0.2, and 0.2, respectively, the comprehensive score of each node is obtained by weighted calculation of the corresponding performance indicators of each node. The higher the score, the greater the weight of the node, which means that it has a greater advantage in processing tasks.

[0094] When determining a compute node's load capacity assessment information, multiple factors need to be considered. The resource requirements of a computing task specify the quantity of various resources required to complete the task; the remaining resources of a computing node reflect the resources currently available for the node to process the task; and the weight of the computing node reflects the node's inherent processing power advantage. Combining these three factors allows for a more comprehensive and accurate assessment of each node's ability to handle specific computing tasks. For example, for a task requiring high CPU computing power, a node may not have the highest remaining resources, but its CPU performance and weight may be high, so it may have a higher load capacity for handling this task. Through comprehensive calculations, a load capacity assessment value can be generated for each node. The higher the value, the greater the node's ability to handle the current task.

[0095] In one embodiment, the load capacity evaluation information of the plurality of computing nodes includes load capacity evaluation values of the plurality of computing nodes, and step S4 includes:

[0096] Step S41: sorting the load capacity evaluation values of the plurality of computing nodes in descending order;

[0097] Step S42: Select computing nodes whose load capacity evaluation values rank in the top N and are greater than a first preset threshold as target computing nodes, where N is a positive integer.

[0098] In the above embodiment, the load capacity assessment values of multiple computing nodes are sorted from largest to smallest to clearly understand the order in which each computing node is most capable of handling the current computing task. This sorting method allows for intuitive identification of nodes with greater processing power, providing a clear basis for subsequent target computing node selection.

[0099] The computing nodes whose load capacity evaluation values ​​rank in the top N and are greater than the first preset threshold are selected as target computing nodes. The first preset threshold here is a standard value set according to the actual situation of the system and the task requirements. Only computing nodes with a load capacity evaluation value greater than this threshold are considered to have sufficient capacity to process the task, avoiding the selection of nodes that are ranked high but have insufficient actual processing capacity. At the same time, the top N are selected to determine the specific number of target computing nodes. The value of N is determined based on factors such as the scale of the task and the resource situation of the system to ensure that the selected nodes can not only meet the task processing requirements but also make rational use of system resources.

[0100] In one embodiment, after step S4, the method further includes:

[0101] Step S5: During the execution of the computing task, monitoring the load value of the target computing node;

[0102] Step S6: When the load value of the target computing node exceeds a second preset threshold, part or all of the computing task is transferred to other computing nodes among the multiple computing nodes except the target computing node.

[0103] In the above embodiment, after a computing task is assigned to a target computing node and starts to be executed, in order to ensure that the task can be successfully completed and the stable operation of the entire computing system is guaranteed, the load value of the target computing node needs to be continuously monitored.

[0104] A compute node's load is a comprehensive indicator that reflects the system resources used by the compute node when processing tasks. It typically includes CPU utilization, memory utilization, disk I / O busyness, and network bandwidth utilization. By monitoring these indicators in real time, we can dynamically understand the operating status of the target compute node. For example, persistently high CPU utilization may indicate that the target compute node is processing a large number of computing tasks and is operating under high load. If memory utilization approaches saturation, this may lead to frequent data exchange, affecting task processing speed.

[0105] The second preset threshold is a critical value set in advance based on system performance and task characteristics. When the target computing node's load exceeds this threshold, it indicates that the target computing node has approached or reached its processing capacity limit. Continuing to execute tasks on this target computing node may cause the computing task to slow down or even fail, increasing the risk of system crashes.

[0106] At this point, the computing task needs to be adjusted. If the task is divisible, such as some large data processing tasks, parts of the task can be moved to other less-loaded computing nodes for continued execution. If the task is indivisible or the target computing node is overloaded, severely impacting task execution, the entire task needs to be moved to another suitable computing node. This ensures that the task can continue to be processed on appropriate computing resources and avoids the adverse effects of node overload.

[0107] In one embodiment, the method further comprises:

[0108] Step S7: setting a network bandwidth limit and a network flow limit for the target computing node according to the resource demand of the computing task and the resource usage of the target computing node;

[0109] Step S8: monitor the network traffic and network traffic of the target computing node in real time;

[0110] Step S9: When the network traffic of the target computing node exceeds the network bandwidth limit and / or the network traffic exceeds the network traffic limit, take flow limiting measures.

[0111] In the above embodiment, network bandwidth limits and network traffic limits are set for the target computing node based on the resource requirements of the computing task and the resource usage of the target computing node. The resource requirements of the computing task determine the approximate range of network resources required to complete the task, while the resource usage of the target computing node reflects its current resource occupancy. Taking these two factors into consideration, network resources can be reasonably allocated to the target computing node, and appropriate network bandwidth limits (i.e., limiting the network transmission rate that the node can use) and network traffic limits (i.e., limiting the amount of data that the node can transmit within a certain period of time) can be set to ensure the rational allocation and effective utilization of network resources.

[0112] Monitor the target compute node's network traffic and bandwidth usage in real time. When the target compute node's network traffic exceeds the network bandwidth limit and / or the network traffic exceeds the network traffic limit, implement throttling measures. This means that if a node's network usage exceeds a pre-defined reasonable range, it will be subject to throttling, such as reducing the network transmission rate or suspending some network transmission tasks, to prevent excessive network resource usage and impacting the normal operation of other compute nodes or tasks.

[0113] By setting limits and taking current limiting measures, we can avoid network congestion caused by excessive network resource occupation by a single computing node, thereby ensuring the stability and reliability of the entire network environment and ensuring that other computing tasks can carry out network communication normally.

[0114] See Figure 2 , Figure 2 This is a structural diagram of a computing power load balancing device for an intelligent computing center provided by the present invention. Figure 2 As shown, the computing power load balancing device 200 of the intelligent computing center includes:

[0115] A first acquisition module 201 is used to acquire the operating status data of multiple computing nodes;

[0116] A first receiving module 202, configured to receive a computing task;

[0117] A first determining module 203 is configured to determine load capacity evaluation information of the plurality of computing nodes according to the task type of the computing task and the operating status data of the plurality of computing nodes;

[0118] The first allocation module 204 is configured to allocate the computing task to a target computing node among the multiple computing nodes according to the load capacity evaluation information of the multiple computing nodes.

[0119] In one embodiment, the first determining module 203 includes:

[0120] A first determining unit, configured to determine a resource requirement of the computing task according to a task type of the computing task;

[0121] A second determining unit, configured to determine remaining amounts of resources of the plurality of computing nodes according to the operating status data of the plurality of computing nodes;

[0122] The third determining unit is configured to determine load capacity evaluation information of the plurality of computing nodes according to the resource demand of the computing task and the remaining resources of the plurality of computing nodes.

[0123] In one embodiment, the third determining unit includes:

[0124] A first acquisition subunit, configured to acquire performance data of the plurality of computing nodes;

[0125] A first calculation subunit, configured to calculate weights of the plurality of computing nodes according to the performance data of the plurality of computing nodes;

[0126] The first determining subunit is configured to determine load capacity evaluation information of the plurality of computing nodes according to the resource demand of the computing task, the remaining resources of the plurality of computing nodes, and the weights of the plurality of computing nodes.

[0127] In one embodiment, the first allocation module 204 includes:

[0128] A first sorting unit, configured to sort the load capacity evaluation information of the plurality of computing nodes in descending order;

[0129] The first selection unit is used to select a computing node whose load capacity evaluation information ranks in the top N and is greater than a first preset threshold as a target computing node, where N is a positive integer.

[0130] In one embodiment, the apparatus further comprises:

[0131] A first monitoring module is used to monitor the load value of the target computing node during the execution of the computing task;

[0132] The first transfer module is configured to transfer part or all of the computing task to other computing nodes among the multiple computing nodes except the target computing node when the load value of the target computing node exceeds a second preset threshold.

[0133] In one embodiment, the apparatus further comprises:

[0134] A first setting module is used to set a network bandwidth limit and a network flow limit for the target computing node according to the resource demand of the computing task and the resource usage of the target computing node;

[0135] a second monitoring module for monitoring the network traffic and network traffic of the target computing node in real time;

[0136] The first flow limiting module is configured to take flow limiting measures when the network flow of the target computing node exceeds the network bandwidth limit and / or the network flow exceeds the network flow limit.

[0137] The computing power load balancing device of the intelligent computing center provided by the present invention is capable of implementing the various processes of each embodiment of the computing power load balancing method of the above-mentioned intelligent computing center. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0138] It should be noted that the computing power load balancing device of the intelligent computing center in the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0139] The present invention also provides an electronic device, see Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device includes a memory 301, a processor 302, and a program or instruction stored in the memory 301. When the program or instruction is executed by the processor 302, Figure 1 Any steps in the corresponding embodiment of the computing power load balancing method of the intelligent computing center and the same beneficial effects are achieved will not be repeated here.

[0140] The processor 302 may be a CPU, an ASIC, an FPGA or a GPU.

[0141] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned embodiment of the method for implementing computing power load balancing in an intelligent computing center can be accomplished through hardware associated with program instructions, and the program can be stored in a readable medium.

[0142] The present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above Figure 1 Any step in the corresponding embodiment of the method for balancing computing power load of an intelligent computing center can achieve the same technical effect. To avoid repetition, it is not repeated here. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0143] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the implementation method of the computing power load balancing method of the corresponding intelligent computing center can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0144] The terms "first", "second" and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. In addition, the terms "comprise" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. In addition, "and / or" is used in this application to represent at least one of the connected objects, for example A and / or B and / or C, which means comprising seven situations including single A, single B, single C, and both A and B exist, both B and C exist, both A and C exist, and both A, B and C exist.

[0145] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0146] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of each embodiment of the present application.

[0147] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A computing power load balancing method for an intelligent computing center, characterized in that: include: Step S1: Obtaining the operating status data of multiple computing nodes; Step S2: receiving a computing task; Step S3: determining load capacity evaluation information of the plurality of computing nodes according to the task type of the computing task and the operating status data of the plurality of computing nodes; Step S4: Allocate the computing task to a target computing node among the multiple computing nodes according to the load capacity evaluation information of the multiple computing nodes.

2. The method according to claim 1, wherein The step S3 comprises: Step S31: determining the resource requirement of the computing task according to the task type of the computing task; Step S32: determining the remaining amount of resources of the plurality of computing nodes according to the operating status data of the plurality of computing nodes; Step S33: Determine load capacity evaluation information of the multiple computing nodes based on the resource requirements of the computing task and the remaining resources of the multiple computing nodes.

3. The method according to claim 2, wherein The step S33 includes: Step S331: Obtain performance data of the multiple computing nodes; Step S332: Calculate the weights of the plurality of computing nodes according to the performance data of the plurality of computing nodes; Step S333: Determine load capacity evaluation information of the multiple computing nodes according to the resource demand of the computing task, the remaining resources of the multiple computing nodes, and the weights of the multiple computing nodes.

4. The method according to claim 1, wherein The load capacity evaluation information of the plurality of computing nodes includes load capacity evaluation values of the plurality of computing nodes; the step S4 includes: Step S41: sorting the load capacity evaluation values of the plurality of computing nodes in descending order; Step S42: Select computing nodes whose load capacity evaluation values rank in the top N and are greater than a first preset threshold as target computing nodes, where N is a positive integer.

5. The method according to any one of claims 1 to 4, characterized in that After step S4, the method further includes: Step S5: During the execution of the computing task, monitoring the load value of the target computing node; Step S6: When the load value of the target computing node exceeds a second preset threshold, part or all of the computing task is transferred to other computing nodes among the multiple computing nodes except the target computing node.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Step S7: setting a network bandwidth limit and a network flow limit for the target computing node according to the resource demand of the computing task and the resource usage of the target computing node; Step S8: monitor the network traffic and network traffic of the target computing node in real time; Step S9: When the network traffic of the target computing node exceeds the network bandwidth limit and / or the network traffic exceeds the network traffic limit, take flow limiting measures.

7. A computing power load balancing device for an intelligent computing center, characterized in that: include: A first acquisition module is used to obtain the operating status data of multiple computing nodes; A first receiving module, configured to receive a computing task; A first determining module, configured to determine load capacity evaluation information of the plurality of computing nodes according to a task type of the computing task and the operating status data of the plurality of computing nodes; The first allocation module is configured to allocate the computing task to a target computing node among the multiple computing nodes according to load capacity evaluation information of the multiple computing nodes.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for balancing computing power load of an intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the computing power load balancing method of the intelligent computing center according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the computing power load balancing method of the intelligent computing center as described in any one of claims 1 to 6.