Computing power node load balancing method and device, equipment and storage medium

Through the prediction model, estimate the task resource requirements and real-time load filtering of the best nodes, combined with multi-dimensional feature matching to optimize task allocation, the problem that traditional load balancing strategies cannot perceive load and high computational complexity in real time is solved, and more efficient load balancing and cross-cloud adaptability are achieved.

CN120104352AInactive Publication Date: 2025-06-06SHENZHEN XUNCE TECH CO LTD

Patent Information

Application Number
CN202510590659.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional load balancing strategies cannot perceive server load status in real time, resulting in uneven task allocation, high computational complexity, difficulty in adapting across cloud environments, and failure to fully consider the performance differences of heterogeneous nodes, resulting in wasted computing power or node overload.

Method used

By reading the task feature information of the computing power task to be calculated, using the preset prediction model to predict the computing power resource, dynamically filtering the best nodes with real-time load, and optimizing task allocation through multi-dimensional feature matching.

Benefits of technology

It improves load balancing efficiency, reduces system computing overhead, enhances cross-cloud environment adaptability, avoids waste of computing power and node overload, and achieves more accurate resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104352A_ABST
    Figure CN120104352A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computing power load balancing, and discloses a computing power node load balancing method and device, equipment and a storage medium. The computing power node load balancing method comprises the steps that a target computing power task is read, task feature information is extracted, and the computing power resource quantity needed by the task is calculated through a prediction model; calculating the load of each node in real time, and screening candidate nodes meeting the resource margin; and evaluating the matching degree of the candidate nodes and the tasks, and selecting an optimal node to perform task allocation. According to the method, the task resource demand is accurately predicted through the prediction model, the optimal node is dynamically screened in combination with the real-time load, the problems that a traditional strategy cannot sense the instantaneous load, the calculation complexity is high and cross-cloud adaptation is difficult are solved, the load balancing efficiency is improved, the system calculation overhead is reduced, and meanwhile the cross-cloud environment adaptability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computing power load balancing technology, and in particular to a computing power node load balancing method, device, equipment and storage medium. Background Art

[0002] In a distributed computing environment, load balancing of computing nodes is a key technology to ensure efficient operation of the system. Traditional load balancing strategies, such as polling and weighted polling, are simple and easy to implement, but have obvious limitations. These strategies cannot perceive the load status of the server in real time, such as instantaneous fluctuations in CPU and memory, resulting in uneven task distribution and affecting the overall performance of the system. In addition, dynamic load balancing algorithms, such as the least connection and shortest response time algorithms, can allocate tasks based on real-time monitoring data, but their computational complexity is high, which increases the computational overhead of the system. At the same time, the threshold setting of these algorithms often relies on manual experience, lacks adaptive capabilities, and is difficult to cope with complex and changing computing environments. In terms of cloud environment adaptation, the automatic expansion and contraction strategy relies on the application programming interface provided by the cloud platform, but in cross-cloud or hybrid cloud scenarios, due to the differences in application programming interfaces of different cloud platforms, it is difficult to unify the strategy, the resource pool is seriously fragmented, and the resource utilization efficiency is reduced. In addition, with the increase in the heterogeneity of computing nodes, such as the introduction of acceleration nodes such as GPU and NPU, traditional load balancing algorithms fail to fully consider the performance differences of these nodes, resulting in the coexistence of the risk of computing power waste or node overload.

[0003] In view of the above problems, the existing technology needs to be improved urgently. Summary of the invention

[0004] The main purpose of the present invention is to provide a computing power node load balancing method, device, equipment and storage medium, which has the advantages of improving load balancing efficiency, reducing system computing overhead and enhancing adaptability across cloud environments.

[0005] A first aspect of the present invention provides a computing power node load balancing method, the computing power node load balancing method comprising: Read the computing task to be calculated, and extract the task feature information of the computing task; Input the task feature information into a preset prediction model to predict the amount of computing power resources, and output the amount of computing power resources required to calculate the computing power task; Calculate the real-time computing load of each computing node; According to the real-time computing load of each computing node and the amount of computing resources required for the computing task, a number of candidate computing nodes for calculating the computing task are screened out; Calculate the matching score between each candidate computing power node and the computing power task; According to the matching score, the best computing node is selected from each of the candidate computing nodes, and the computing task is assigned to the best computing node.

[0006] In a first implementation of the first aspect of the present invention, before reading the computing power task to be calculated, the method further includes: Obtain the task information corresponding to each historical computing task and the amount of historical computing resources used to complete each historical computing task; According to the order of task completion time, feature extraction is performed on the task information of each historical computing task to obtain a task feature vector, and the historical computing resources used to complete each historical computing task are vectorized to obtain the corresponding historical computing resources vector; The task feature vector and the historical computing power resource vector are used as inputs of the initial prediction model, and the prediction deviation value of the historical computing power resource corresponding to each historical computing power task is trained and outputted. Based on the prediction deviation value, the weight matrix and bias of the input layer and hidden layer of the initial prediction model are adjusted to obtain a trained prediction model.

[0007] In a second implementation of the first aspect of the present invention, the prediction model includes an input layer, multiple hidden layers, and an output layer in sequence; the output of a previous layer of the multiple hidden layers is used as the input of a subsequent layer; the hidden layer includes multiple operation nodes arranged in sequence, each of the operation nodes includes a weight matrix, a bias, and an addition operation node and a primary operation node for connecting the operation nodes in the previous layer; Among them, the weight matrix is ​​used to perform dot multiplication with the vector matrix output by a corresponding operation node in the previous layer to obtain a result vector matrix; the bias is used to offset the elements in the result vector matrix; the addition operation node is used to accumulate the bias and each element in the result vector matrix to obtain an accumulated vector matrix; the primary operation node is used to perform linear or nonlinear operations on each element in the accumulated vector matrix output by the addition operation node to obtain an output vector matrix.

[0008] In a third implementation of the first aspect of the present invention, calculating the real-time computing load of each computing node includes: Detect resource usage of each computing node; Calculate the benchmark resource consumption rate of each computing power node based on the resource utilization rate of each computing power node; Calculate the computing performance evaluation value of each computing power node based on the benchmark resource consumption rate of each computing power node; Get the current traffic information of each computing power node; Calculate the performance coefficient of each computing power node based on the current traffic information of each computing power node; The real-time load of each computing power node is calculated based on the performance evaluation value and performance coefficient of each computing power node.

[0009] In a fourth implementation of the first aspect of the present invention, screening out a number of candidate computing nodes for calculating the computing task according to the real-time computing load of each computing node and the amount of computing resources required for the computing task includes: Convert the real-time computing load of each computing node into the corresponding computing resources, and calculate the real-time remaining computing resources of each computing node; Determine respectively whether there is a target computing power node whose real-time remaining computing power resources are greater than the computing power resources required by the computing power task; If there are several target computing nodes whose real-time remaining computing power resources are greater than the computing power resources required by the computing power task, then the real-time load fluctuation and computing power load stability of each target computing power node within a preset time period are calculated; According to the real-time load fluctuation and computing load stability of each target computing node within a preset time period, a number of candidate computing nodes for calculating the computing task are screened out.

[0010] In a fifth implementation of the first aspect of the present invention, calculating the matching score between each candidate computing power node and the computing power task includes: Respectively extracting node computing power resource characteristics of each of the candidate computing power nodes and computing power resource characteristics required for the computing power task; The similarity between the node computing power resource characteristics of each candidate computing power node and the computing power resource characteristics required for the computing power task; The feature similarities corresponding to each candidate computing power node are weighted and summed up respectively to obtain a matching score between each candidate computing power node and the computing power task.

[0011] In a sixth implementation of the first aspect of the present invention, the computing power node load balancing method further includes: When the best computing power node calculates the computing power task, splitting the computing power task into multiple sub-computing power tasks; Determine the resource consumption rate of each sub-computing task according to the task type of each sub-computing task; Monitor the running time of each sub-computing task; According to the running time and resource consumption rate of each sub-computing task, the amount of computing resources used to complete the computing task is statistically obtained.

[0012] A second aspect of the present invention further provides a computing power node load balancing device, the computing power node load balancing device comprising: A reading module, used to read the computing task to be calculated and extract task feature information of the computing task; A prediction module, used to input the task feature information into a preset prediction model to predict the amount of computing power resources, and output the amount of computing power resources required to calculate the computing power task; The first computing module is used to calculate the real-time computing load of each computing node; A screening module, used to screen out a number of candidate computing nodes for calculating the computing task according to the real-time computing load of each computing node and the amount of computing resources required for the computing task; A second calculation module is used to calculate the matching score between each candidate computing power node and the computing power task; The allocation module is used to select the best computing power node from each of the candidate computing power nodes according to the matching score, and allocate the computing power task to the best computing power node.

[0013] In a first implementation of the second aspect of the present invention, the computing power node load balancing device further includes: The training module is used to obtain the task information corresponding to each historical computing power task and the historical computing power resources used to complete each historical computing power task; extract the features of the task information of each historical computing power task in the order of task completion time to obtain the task feature vector, and vectorize the historical computing power resources used to complete each historical computing power task to obtain the corresponding historical computing power resource vector; use the task feature vector and the historical computing power resource vector as the input of the initial prediction model, train and output the prediction deviation value of the historical computing power resource corresponding to each historical computing power task, and adjust the weight matrix and bias of the input layer and hidden layer of the initial prediction model based on the prediction deviation value to obtain a trained prediction model.

[0014] In a second implementation of the second aspect of the present invention, the prediction model includes an input layer, multiple hidden layers, and an output layer in sequence; the output of a previous layer of the multiple hidden layers is used as the input of a subsequent layer; the hidden layer includes multiple operation nodes arranged in sequence, each of the operation nodes includes a weight matrix, a bias, and an addition operation node and a primary operation node for connecting the operation nodes in the previous layer; Among them, the weight matrix is ​​used to perform dot multiplication with the vector matrix output by a corresponding operation node in the previous layer to obtain a result vector matrix; the bias is used to offset the elements in the result vector matrix; the addition operation node is used to accumulate the bias and each element in the result vector matrix to obtain an accumulated vector matrix; the primary operation node is used to perform linear or nonlinear operations on each element in the accumulated vector matrix output by the addition operation node to obtain an output vector matrix.

[0015] In a third implementation of the second aspect of the present invention, the first calculation module is specifically configured to: Detect resource usage of each computing node; Calculate the benchmark resource consumption rate of each computing power node based on the resource utilization rate of each computing power node; Calculate the computing performance evaluation value of each computing power node based on the benchmark resource consumption rate of each computing power node; Get the current traffic information of each computing power node; Calculate the performance coefficient of each computing power node based on the current traffic information of each computing power node; The real-time load of each computing power node is calculated based on the performance evaluation value and performance coefficient of each computing power node.

[0016] In a fourth implementation of the second aspect of the present invention, the screening module is specifically used for: Convert the real-time computing load of each computing node into the corresponding computing resources, and calculate the real-time remaining computing resources of each computing node; Determine respectively whether there is a target computing power node whose real-time remaining computing power resources are greater than the computing power resources required by the computing power task; If there are several target computing nodes whose real-time remaining computing power resources are greater than the computing power resources required by the computing power task, then the real-time load fluctuation and computing power load stability of each target computing power node within a preset time period are calculated; According to the real-time load fluctuation and computing load stability of each target computing node within a preset time period, a number of candidate computing nodes for calculating the computing task are screened out.

[0017] In a fifth implementation manner of the second aspect of the present invention, the second calculation module is specifically configured to: Respectively extracting node computing power resource characteristics of each of the candidate computing power nodes and computing power resource characteristics required for the computing power task; The similarity between the node computing power resource characteristics of each candidate computing power node and the computing power resource characteristics required for the computing power task; The feature similarities corresponding to each candidate computing power node are weighted and summed up respectively to obtain a matching score between each candidate computing power node and the computing power task.

[0018] In a sixth implementation of the second aspect of the present invention, the computing power node load balancing device further includes: The statistical module is used to split the computing task into multiple sub-computing tasks when the optimal computing node calculates the computing task; determine the resource consumption rate of each sub-computing task according to the task type of each sub-computing task; monitor the running time of each sub-computing task; and obtain the amount of computing resources used to complete the computing task according to the running time and resource consumption rate of each sub-computing task.

[0019] The third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned computing power node load balancing method.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned computing power node load balancing method.

[0021] From the above, it can be seen that the computing power node load balancing method, device, computer equipment and storage medium provided by the present application accurately predict the task resource requirements through the prediction model, and dynamically select the best nodes based on the real-time load. It solves the problems that traditional strategies cannot perceive the instantaneous load, have high computational complexity and difficulty in cross-cloud adaptation. It has the advantages of improving load balancing efficiency, reducing system computing overhead and enhancing adaptability across cloud environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A schematic diagram of a flow chart of an embodiment of a method for load balancing computing power nodes in an embodiment of the present invention; Figure 2 A schematic diagram of a flow chart of an embodiment of training a prediction model in an embodiment of the present invention; Figure 3 A schematic diagram of a flow chart of an embodiment of calculating real-time computing load in an embodiment of the present invention; Figure 4 A schematic diagram of a flow chart of an embodiment of selecting candidate computing nodes in an embodiment of the present invention; Figure 5 A schematic diagram of a flow chart of an embodiment of scoring the matching degree between candidate computing nodes and computing tasks in an embodiment of the present invention; Figure 6 A schematic diagram of a flow chart of an embodiment of the amount of computing resources used by a centralized computing task in an embodiment of the present invention; Figure 7 The present invention is a schematic diagram of the functional modules of a computing power node load balancing device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0024] Embodiment 1: In the prior art, traditional load balancing methods use polling or weighted polling strategies, which cannot perceive the dynamic load status of the server in real time, such as instantaneous fluctuations in CPU or memory usage, resulting in uneven resource allocation. Dynamic algorithms such as the minimum number of connections rely on real-time monitoring, which increases computing overhead and threshold setting relies on manual experience. In a cloud environment, cross-platform resource scheduling strategies are difficult to unify, and resource pool fragmentation reduces efficiency. At the same time, the performance differences of heterogeneous nodes such as GPU-accelerated servers have not been fully optimized, and there is a risk of computing power waste or overload.

[0025] In order to solve the above problems, a balancing method is needed that can dynamically perceive node load, reduce computational complexity and adapt to heterogeneous environments. Traditional static strategies cannot cope with real-time load changes, and dynamic algorithm monitoring overhead is too large. Therefore, this embodiment introduces a prediction model to estimate task resource requirements, selects candidate nodes based on the real-time node load, and then optimizes task allocation through multi-dimensional feature matching, which can reduce complexity while ensuring efficiency.

[0026] like Figure 1 As shown, this application proposes the following technical solutions: 101. Read the computing task to be calculated, and extract task feature information of the computing task; In this embodiment, the task characteristic information refers to the resource type, computing intensity and priority attributes required for task execution. Specifically, data packet parsing technology can be used to extract core indicators in the task parameters (pre-set).

[0027] 102. Input the task feature information into a preset prediction model to predict the amount of computing power resources, and output the amount of computing power resources required to calculate the computing power task; The prediction model refers to a neural network trained based on historical task data. It predicts resource consumption by inputting task feature vectors, and a multi-layer perceptron can be used to achieve nonlinear mapping.

[0028] 103. Calculate the real-time computing load of each computing node; Real-time computing load refers to the ratio of the node's currently available resources to the total resources. It is calculated by monitoring the CPU, memory, and network bandwidth usage. Specifically, resource probes can be used to collect data in real time.

[0029] 104. Screening out a number of candidate computing nodes for calculating the computing task according to the real-time computing load of each computing node and the amount of computing resources required for the computing task; Candidate computing node screening refers to selecting nodes with smaller load fluctuations from computing nodes that meet resource margins. Specifically, the load variance value of each computing node within the time window can be calculated.

[0030] 105. Calculate the matching score between each candidate computing power node and the computing power task; The matching score refers to the weighted value of the similarity between the task requirement characteristics and the node resource characteristics. Specifically, the cosine similarity algorithm can be used to calculate the correlation between multi-dimensional feature vectors.

[0031] 106. According to the matching score, select the best computing node from the candidate computing nodes, and assign the computing task to the best computing node.

[0032] In this embodiment, when a new computing task arrives, the system parses the task parameters to generate a feature vector, which is input into the trained prediction model to obtain the estimated resource amount. At the same time, each computing node reports real-time load data through a resource probe to calculate the remaining available resources. The system selects nodes whose remaining resources are greater than the task requirements as candidate sets, and excludes nodes whose instantaneous load fluctuations are too large. For each candidate node, its hardware configuration, current load distribution and other features are extracted, and the similarity is calculated with the task requirement features, and a matching score is generated in combination with the weight coefficient. Finally, the node with the highest score is selected to perform the task to achieve optimal resource utilization.

[0033] Compared with the prior art, the traditional polling strategy only allocates in a fixed order and cannot be adjusted according to the real-time load. This embodiment uses a prediction model to accurately estimate resource requirements to avoid over- or under-allocation. Dynamic load calculation combined with historical fluctuation analysis can screen nodes with higher stability and reduce the risk of task execution interruption. The multi-dimensional feature matching mechanism takes into account both node hardware performance and task requirement characteristics, achieves precise adaptation in heterogeneous environments, and improves resource utilization compared to a single indicator selection strategy.

[0034] Through the above technical solutions, this application effectively solves the problem of insufficient adaptability of traditional load balancing methods in dynamic environments. Resource demand prediction reduces the need for manual intervention, real-time load screening ensures node availability, and feature matching mechanism improves heterogeneous resource utilization. While ensuring efficiency, the task allocation process reduces performance fluctuations caused by load mutations and realizes dynamic load balancing in distributed computing environments.

[0035] Embodiment 2: In this embodiment, a new machine learning model is further proposed to be used as a training model of the prediction model (ie, an initial prediction model).

[0036] In this embodiment, the prediction model includes an input layer, multiple hidden layers and an output layer in sequence, the output of the previous layer in the multiple hidden layers is used as the input of the next layer, the hidden layer includes multiple operation nodes arranged in sequence, each operation node includes a weight matrix, a bias, and an addition operation node and a primary operation node for connecting each operation node in the previous layer, wherein the weight matrix is ​​used to multiply the vector matrix output by the corresponding operation node in the previous layer to obtain a result vector matrix, the bias is used to offset the elements in the result vector matrix, the addition operation node is used to accumulate the bias and each element in the result vector matrix to obtain an accumulated vector matrix, and the primary operation node is used to perform linear or nonlinear operations on each element in the accumulated vector matrix to obtain an output vector matrix.

[0037] Among them, the weight matrix refers to the set of parameters used to perform linear transformation on the feature data output by the previous layer. Specifically, it can be implemented by a randomly initialized floating-point matrix, and the input features are mapped to the new feature space through point multiplication operations. The bias refers to the offset parameter used to adjust the result of the linear transformation. Specifically, it can be implemented by a vector that matches the dimension of the weight matrix, and is used to correct the impact of feature distribution deviation on the prediction result. The addition operation node refers to the calculation unit that performs the superposition of matrix elements and biases. Specifically, it can be implemented by an element-by-element adder, which is used to integrate the intermediate results after linear transformation and bias adjustment. The primary operation node refers to the processing unit that applies a function transformation to the accumulated result. Specifically, it can be implemented by a ReLU activation function or a Sigmoid function, and the model's ability to express complex correlation relationships is enhanced through nonlinear mapping.

[0038] Specifically, the input layer receives the task feature vector and passes it to the first hidden layer. The operation node of each hidden layer performs the dot product operation between the weight matrix and the input vector to generate a preliminary feature mapping result. The bias is superimposed on each element in the dot product result to form an adjusted intermediate vector. The addition operation node adds the bias to the dot product result element by element to ensure that the feature transformation process retains the distribution characteristics of the original data. The primary operation node applies an activation function to the accumulated result, such as using the ReLU function to return negative values ​​to zero to filter invalid features, or using the Sigmoid function to compress the output to the 0-1 range to enhance feature discrimination. Multiple hidden layers perform the above operations in sequence, and the nonlinear correlation between task characteristics and historical resource quantities is converted into quantifiable prediction results through layer-by-layer abstraction. Finally, the output layer generates a predicted value of computing power resources.

[0039] Compared with the prior art, traditional prediction models mostly use a single hidden layer structure or a linear regression model, which cannot effectively capture the high-order nonlinear relationship between task characteristics and resource requirements. In the prior art, the activation function only acts on the final output layer, resulting in the lack of nonlinear adjustment capabilities of the intermediate feature transformation. This embodiment uses a multi-layer hidden cascade structure and layer-by-layer nonlinear activation to enable the model to gradually extract deep patterns in task characteristics. For example, it identifies the feature combination of CPU-intensive tasks in the first hidden layer, and associates the dynamic relationship between memory consumption and task concurrency in subsequent hidden layers, thereby significantly improving the prediction accuracy.

[0040] Through the above technical solution, this application solves the problem of insufficient feature expression ability of the prediction model due to its single structure, and effectively captures the complex relationship between task characteristics and computing resource requirements through multi-layer nonlinear transformations. The synergy between the weight matrix and the bias in the operation node reduces the impact of feature distribution deviation on the prediction results, and the layer-by-layer application of the activation function enhances the model's ability to fit the implicit rules in historical data. This structure enables the prediction model to adapt to the differences in resource characteristics of heterogeneous computing nodes. For example, it automatically identifies the relationship between video memory occupancy and computing core utilization for GPU acceleration tasks, and then accurately predicts resource requirements under different hardware environments.

[0041] Embodiment 3: like Figure 2 As shown, in this embodiment, the prediction model is further trained based on the model of the above embodiment 2, and the specific training process is as follows: 201. Obtain task information corresponding to each historical computing task and the amount of historical computing resources used to complete each historical computing task; In this embodiment, the task information corresponding to the historical computing power task refers to the metadata set generated during the task execution process, which can be implemented by task type identification, input data scale, computational complexity level, dependency library version number, and processor instruction set characteristics to characterize the resource allocation requirements during task execution.

[0042] The historical computing power resource volume refers to the quantitative indicators of the actual hardware resources consumed. It can be implemented by CPU core occupancy time, memory peak usage, GPU memory occupancy rate, and network bandwidth occupancy percentage to reflect the actual resource consumption of task execution.

[0043] 202. Extract features from task information of each historical computing task in order of task completion time to obtain a task feature vector, and vectorize the amount of historical computing resources used to complete each historical computing task to obtain a corresponding historical computing resource amount vector; The task feature vector refers to the process of converting unstructured task information into a numerical vector. Specifically, it can be achieved by using word embedding technology to encode text descriptions, or by normalizing and concatenating numerical parameters to establish a machine-processable feature expression.

[0044] The historical computing power resource quantity vector refers to the process of integrating multi-dimensional resource consumption indicators into a unified vector. Specifically, it can be achieved by adopting a multi-channel time series data alignment method, and the monitoring data of different sampling frequencies are interpolated to form equal-length vectors to maintain the spatiotemporal correlation of resource consumption characteristics.

[0045] The prediction deviation value refers to the difference measure between the model prediction result and the actual consumption. Specifically, it can be achieved by calculating the sum of the squares of the differences in each dimension between the predicted resource quantity vector and the actual resource quantity vector using the mean square error function, which is used to guide the iterative optimization of the model parameters.

[0046] 203. Use the task feature vector and the historical computing power resource vector as inputs of the initial prediction model, train and output the prediction deviation value of the historical computing power resource corresponding to each historical computing power task, and adjust the weight matrix and bias of the input layer and hidden layer of the initial prediction model based on the prediction deviation value to obtain a trained prediction model.

[0047] In this embodiment, the acquired historical task information is parsed into a combination of structured data and unstructured data, the task type identifier and computational complexity level are converted into binary vectors through unique hot encoding, and the input data scale is compressed by logarithmic transformation to form a normalized scalar. The discrete monitoring indicators in the historical computing power resource quantity are converted into time series fragments, and the resource consumption curve of fixed length is intercepted through a sliding window. The task feature vector and the resource quantity vector are aligned according to the task completion timestamp to form a training sample pair with time series association.

[0048] The number of neurons in the input layer of the initial prediction model is set to the dimension value of the task feature vector, and the hidden layer activation function uses a leaky linear rectifier unit to alleviate the gradient vanishing problem. During the training process, after the forward propagation generates the predicted resource quantity vector, the back propagation algorithm calculates the gradient change of each hidden layer weight matrix according to the predicted deviation value, and the bias dynamically adjusts the compensation value according to the error distribution. After multiple rounds of iterative training, the connection weights between the hidden layer neurons form a nonlinear mapping relationship between task characteristics and resource requirements, and the model automatically learns the dependence pattern of different task types on computing resources.

[0049] Compared with the prior art, the traditional dynamic resource prediction method relies on manually setting resource consumption thresholds, and operation and maintenance personnel are required to manually adjust the prediction parameters based on experience, which results in response lag and subjective error problems. However, this embodiment constructs an automated training mechanism so that the prediction model can autonomously mine the correlation between features and resource consumption from historical task execution data, eliminating the uncertainty caused by manual experience intervention. Compared with the static threshold setting method, the dynamic adjustment mechanism of the weight matrix in this embodiment can adapt to changes in the distribution of task features in different periods, and the adaptive compensation function of the bias can effectively handle prediction deviations caused by system environment fluctuations. In addition, the time-sequential data processing method retains the continuity characteristics of task execution, enabling the model to capture the evolution trend of resource consumption patterns.

[0050] Through the above technical solutions, this application realizes the self-optimization capability of the computing resource prediction model, significantly reduces the workload of manual parameter adjustment, and improves the accuracy of resource demand prediction. The trained prediction model can accurately identify the implicit relationship between task characteristics and resource consumption, avoiding resource allocation overload or idle problems caused by insufficient manual experience. The model parameters are continuously optimized through historical data drive, and can adapt to changes in task types and hardware environment upgrades, ensuring that distributed data centers can achieve accurate prediction of computing resources at different operation stages.

[0051] Embodiment 4: like Figure 3 As shown, in this embodiment, the real-time computing load of each computing node is calculated in the following manner, including: 301. Detect the resource utilization rate of each computing power node; In this embodiment, resource utilization refers to the proportion of hardware resources occupied by the computing power node per unit time. It can be achieved by real-time collection of the CPU utilization, memory occupancy, disk I / O rate and GPU computing unit load rate of the monitoring node to reflect the node's immediate resource consumption status.

[0052] 302. Calculate the benchmark resource consumption rate of each computing node according to the resource utilization rate of each computing node; The benchmark resource consumption rate refers to the average resource consumption level of a computing power node under a stable operating state. It can be achieved by using a sliding window algorithm to perform weighted average calculation on historical data of resource utilization, in order to eliminate the impact of instantaneous fluctuations on load assessment.

[0053] 303. Calculate the computing performance evaluation value of each computing node according to the benchmark resource consumption rate of each computing node; The computing performance evaluation value refers to a quantitative evaluation index of the hardware processing capability of a computing node. It can be achieved by performing normalized weighted calculations on the node's processor main frequency, number of cores, cache capacity, and accelerator type. It is used to distinguish differences in computing capabilities among heterogeneous nodes.

[0054] 304. Obtain the current traffic information of each computing power node; Current traffic information refers to the data throughput status of the computing power node in network communication. It can be achieved by collecting real-time statistics of the inbound and outbound bandwidth occupancy rates, data packet transmission delays, and packet loss rates of the network interface to reflect the impact of network transmission on the efficiency of computing task execution.

[0055] 305. Calculate the performance coefficient of each computing power node according to the current traffic information of each computing power node; The performance coefficient refers to the degree of match between the network transmission efficiency of the computing power node and the computing task requirements. It can be achieved by normalizing the difference between the bandwidth occupancy rate in the traffic information and the data transmission rate required by the task. It is used to evaluate the dynamic impact of the network status on the load.

[0056] 306. Calculate the real-time load of each computing power node based on the performance evaluation value and performance coefficient of each computing power node.

[0057] Real-time load refers to the overall load level after integrating the hardware resource processing capacity and network transmission efficiency. It can be achieved by performing linear weighted summation and normalization on the computing performance evaluation value and the performance coefficient to generate multi-dimensional dynamic load evaluation results.

[0058] In this embodiment, when calculating the real-time load of the computing power node, the operating data of the CPU, memory, disk and accelerator card are first periodically collected through the resource monitoring module to generate a multi-dimensional resource utilization index. These indicators are input into the sliding window calculation unit, and the historical data is exponentially weighted averaged according to the preset time window length to eliminate the instantaneous peak interference and generate a benchmark resource consumption rate. Subsequently, a performance evaluation matrix is ​​established according to the hardware configuration parameters of the node, and the parameters such as the processor main frequency and the number of cores are normalized, and then matrix multiplication is performed with the benchmark resource consumption rate to generate a computing performance evaluation value. At the same time, the real-time traffic data of the node is collected through the network probe, the matching degree of the bandwidth occupancy rate and the task required bandwidth is extracted, and the performance performance coefficient is constructed in combination with the transmission delay. Finally, the computing performance evaluation value and the performance performance coefficient are input into the weighted calculation unit, and a linear combination is performed according to the preset hardware weight and network weight to generate a real-time load value reflecting the comprehensive load state of the node.

[0059] Compared with existing technologies, traditional dynamic load algorithms only rely on a single indicator of CPU or memory usage for load evaluation, which cannot effectively distinguish the differences in hardware processing capabilities of heterogeneous nodes, and does not consider the impact of network transmission on task execution efficiency. This solution constructs a dual hardware evaluation system of benchmark resource consumption rate and computing performance evaluation value, combined with the performance coefficient generated by network traffic, to achieve multi-dimensional coupling calculation of hardware resources and network status. This coupling mechanism breaks through the limitation of traditional algorithms that only focus on a single resource indicator, so that the load evaluation results can simultaneously reflect the computing power and communication efficiency of the node.

[0060] Through the above technical solutions, this application effectively solves the evaluation deviation problem caused by the single monitoring dimension of the traditional dynamic load algorithm, and improves the load evaluation accuracy in complex heterogeneous environments through multi-dimensional fusion calculation of hardware resources and network status. At the same time, the sliding window algorithm and normalization processing method are adopted to avoid the reliance on manually set static thresholds and reduce the complexity of algorithm implementation. Through the automated dynamic weight allocation mechanism, the real-time load status of the node is accurately quantified, providing a reliable basis for subsequent task scheduling.

[0061] Embodiment 5: like Figure 4 As shown, in this embodiment, the candidate computing power nodes are selected in the following manner, including: 401. Convert the real-time computing load of each computing node into the corresponding computing resource amount, and calculate the real-time remaining computing resource amount of each computing node; Converting real-time computing load into computing resources means normalizing the CPU, memory, and storage resources occupied by the processes currently running on the node. Specifically, the weighted summation method can be used to map multi-dimensional resource consumption into a value of uniform dimension.

[0062] The real-time remaining computing power resources are obtained by subtracting the real-time computing power load from the total available resources. This can be achieved by obtaining the node hardware configuration data through the resource monitoring interface to evaluate the node's current ability to carry new tasks.

[0063] 402. Determine whether there is a target computing node whose real-time remaining computing power resources are greater than the computing power resources required by the computing task; 403. If there are several target computing nodes whose real-time remaining computing power resources are greater than the computing power resources required by the computing task, calculate the real-time load fluctuation and computing power load stability of each target computing power node within a preset time period; The real-time load fluctuation within a preset time period refers to the standard deviation of the resource utilization rate of the node within the historical time window. Specifically, a sliding window algorithm can be used to calculate the resource change amplitude within the last N seconds.

[0064] The stability of computing load is an indicator that reflects the ability of a node to operate continuously and stably. It can be achieved by counting the proportion of time that resource utilization is within the safety threshold within a preset time period, and used to predict the future load trend of the node.

[0065] 404. According to the real-time load fluctuation and computing load stability of each target computing node within a preset time period, select a number of candidate computing nodes for calculating the computing task.

[0066] In this embodiment, the screening process of candidate computing power nodes is divided into two stages. First, a preliminary screening is performed based on the real-time remaining computing power resources to filter out a set of nodes that meet the minimum resource requirements of the task, ensuring the basic feasibility of task allocation. Then, for the nodes after the preliminary screening, a stability assessment model is constructed by calculating the load fluctuation amplitude and resource stability index in the past time period. For example, within a 5-minute time window, if the CPU utilization standard deviation of a node is less than 3% and the memory occupancy rate remains in the range of 40%-60% for more than 80% of the time, it is judged to have high stability. Finally, a set of candidate nodes that meet the current resource requirements and have stable service capabilities are screened out to provide high-quality input for subsequent matching scores.

[0067] Compared with the existing technology, the traditional load balancing algorithm only relies on the node load status at the current moment to make decisions, which is prone to misjudgment due to instantaneous load fluctuations. For example, the minimum number of connections algorithm only counts the current number of active connections and cannot identify nodes that are about to release a large amount of resources. However, this solution can identify whether the node is in a stable operating state by introducing load fluctuation monitoring in the time dimension, avoiding assigning tasks to nodes that are about to be overloaded. The dynamic threshold setting in the existing technology relies on manual experience adjustment. This solution automatically evaluates the node's continuous service capability by quantifying the load stability index, reducing the need for manual intervention.

[0068] Through the above technical solutions, the present application can effectively improve the accuracy and stability of candidate node screening. In dynamic load scenarios, allocation errors caused by abnormal instantaneous node load are avoided through dual evaluation of real-time remaining resources and historical load fluctuations. For example, in a cloud computing environment, when a node's CPU occupancy rate temporarily soars due to a sudden task, this embodiment can accurately determine whether the node is suitable for receiving new tasks by analyzing its load stability over the past 5 minutes. Through an automated screening mechanism, the empirical dependence on manually setting load thresholds is reduced, and the balance of resource allocation and overall system utilization are improved.

[0069] Embodiment 6: like Figure 5 As shown, this embodiment specifically adopts the following method to score the matching degree between the candidate computing power nodes and the computing power tasks, including: 501. Extracting node computing resource characteristics of each candidate computing node and computing resource characteristics required for the computing task respectively; The node computing resource characteristics refer to a set of quantitative indicators that reflect the hardware performance of the computing node. Specifically, they can be represented by a combination of processor core number, memory capacity, and accelerator type to describe the heterogeneous computing capabilities of the node.

[0070] The computing resource characteristics required for a task refer to a set of parameters that represent the resource conditions required to achieve the target computing task. Specifically, they can be represented by a combination of thread concurrency requirements, video memory occupancy, and floating-point operation intensity to describe the degree of dependence of the task on heterogeneous resources.

[0071] 502. Determine the similarity between the node computing resource characteristics of each candidate computing node and the computing resource characteristics required for the computing task; Feature similarity refers to the degree of proximity between the features of two types of resources in multidimensional space. It can be calculated using the cosine similarity algorithm to quantify the degree of fit between node capabilities and task requirements.

[0072] 503. Perform weighted summation on the feature similarities corresponding to each of the candidate computing nodes to obtain a matching score between each of the candidate computing nodes and the computing task.

[0073] In this embodiment, the weighted sum refers to a linear combination after weights are assigned to multi-dimensional similarities, which can be specifically implemented by a dynamic weight adjustment mechanism to balance the differences in the impact of different resource dimensions on task execution.

[0074] In this embodiment, in a distributed data center environment, the heterogeneity of candidate nodes makes it difficult for traditional load balancing strategies to accurately match task requirements. By establishing a vectorized model of node resource characteristics and task requirement characteristics, heterogeneous indicators such as GPU acceleration capability and memory bandwidth are incorporated into the feature space. The calculation of feature similarity can be converted into an angle measure of a high-dimensional vector, which effectively identifies candidate nodes with specific hardware advantages. The weighting coefficient is dynamically adjusted according to the task type. For example, image processing tasks give GPU performance a higher weight, while data analysis tasks focus on memory capacity weight. This multi-dimensional dynamic matching mechanism breaks through the limitations of a single CPU utilization indicator, allowing task allocation to adapt to the dedicated acceleration capabilities of nodes with different computing power.

[0075] Compared with existing technologies, traditional methods only allocate tasks based on single indicators such as CPU utilization or memory usage, and cannot adapt to dedicated computing scenarios of heterogeneous accelerators such as GPU / NPU. This solution can accurately identify nodes with dedicated acceleration capabilities required for tasks by constructing a multi-dimensional feature space for similarity matching. Compared with the static threshold screening mechanism, dynamic weight adjustment can optimize resource adaptation strategies according to task types and avoid waste of computing power due to human experience errors.

[0076] Through the above technical solutions, this application effectively solves the problem of feature mismatch between heterogeneous computing nodes and complex task requirements. In GPU-intensive task scenarios, nodes with sufficient video memory and parallel computing capabilities can be accurately screened to avoid task interruptions caused by insufficient video memory. In mixed load scenarios, the dynamic weight mechanism can balance demand conflicts in different resource dimensions and reduce the risk of performance degradation due to resource contention. This solution can adjust resource evaluation dimensions according to real-time task requirements and improve the overall utilization of heterogeneous computing clusters.

[0077] Embodiment 7: like Figure 6 As shown, in this embodiment, the following methods are specifically used to count the amount of computing resources used to complete the computing task, specifically including: 601. When the best computing power node calculates the computing power task, split the computing power task into multiple sub-computing power tasks; Task splitting refers to decomposing a single task into multiple independently executable sub-units based on computational logic or data dependencies. This can be achieved by using a workflow engine to divide task boundaries and establish a communication mechanism between sub-tasks. This operation refines the granularity of resource monitoring from the task level to the sub-task level.

[0078] 602. Determine the resource consumption rate of each sub-computing task according to the task type of each sub-computing task; Resource consumption rate refers to the proportion of CPU, memory, and storage resources occupied by a specific type of task per unit time. This can be achieved by establishing a mapping table between task types and historical resource consumption data. This parameter reflects the differences in characteristics of different computing units.

[0079] 603. Monitor the running time of each sub-computing task; Runtime monitoring refers to recording the time interval from the start to the termination of a subtask. This can be achieved by collecting task execution logs using a distributed tracing system. This data is used to quantify the time-varying characteristics of resource consumption.

[0080] 604. According to the running time and resource consumption rate of each sub-computing task, the amount of computing resources used to complete the computing task is statistically obtained.

[0081] In this embodiment, after the target computing power task is assigned to the optimal computing power node, the task scheduler first parses its calculation flow chart to deconstruct the subtask sequence. For each subtask, the preset resource consumption rate parameter table is matched according to the data processing type involved, such as the GPU memory consumption coefficient associated with image recognition tasks, and the CPU core occupancy coefficient associated with numerical calculation tasks. During the operation process, the monitoring agent records the actual execution time of each subtask in real time, and multiplies the resource consumption rate by the runtime to obtain the resource consumption of each subtask. Finally, by accumulating the resource consumption of all subtasks, the total resource consumption evaluation value of the target computing power task is formed, which is used as feedback data for the subsequent task scheduling strategy optimization.

[0082] Compared with existing technologies, traditional resource monitoring methods only focus on resource usage at the overall task level and cannot identify differences in resource requirements of different computing units within a task. This solution, through task decoupling and hierarchical monitoring mechanisms, can capture the actual resource consumption characteristics of heterogeneous computing units and eliminate resource allocation errors caused by task complexity estimation bias.

[0083] Through the above technical solution, this application realizes dynamic and fine-grained monitoring of the execution process of computing tasks, effectively solving the problem of insufficient accuracy in resource utilization assessment. By establishing a task decomposition and type association mechanism, it is possible to accurately identify the resource consumption characteristics of different computing links; through the association calculation of running time and consumption rate, a closed-loop feedback data is formed to support the optimization of resource allocation strategy, avoiding the risk of node overload caused by global load estimation deviation.

[0084] Embodiment 8: The present invention also provides a computing power node load balancing device. Figure 7 As shown, in this embodiment, the computing power node load balancing device includes: The reading module 701 is used to read the computing task to be calculated and extract the task feature information of the computing task; Prediction module 702, used for inputting the task feature information into a preset prediction model to predict the amount of computing power resources, and outputting the amount of computing power resources required for calculating the computing power task; The first calculation module 703 is used to calculate the real-time computing load of each computing node; A screening module 704 is used to screen out a number of candidate computing nodes for calculating the computing task according to the real-time computing load of each computing node and the amount of computing resources required for the computing task; The second calculation module 705 is used to calculate the matching score between each candidate computing power node and the computing power task; The allocation module 706 is used to select the best computing node from each of the candidate computing nodes according to the matching score, and allocate the computing task to the best computing node.

[0085] Compared with existing technologies, traditional load balancing devices rely on static strategies or single indicators to allocate tasks, such as round-robin scheduling based only on the current number of connections or response time of the node, which cannot effectively perceive the instantaneous load fluctuations of the node or the characteristics of heterogeneous resources. However, this device quantifies the task resource requirements through a dynamic prediction model, combines real-time load monitoring with a multi-dimensional feature matching mechanism, and can achieve accurate scheduling of heterogeneous nodes in a cross-cloud environment. For example, in a hybrid cloud scenario, the device can simultaneously access public cloud virtual machines and local GPU servers, eliminate the impact of resource fragmentation through a unified load assessment model, and avoid the problem of threshold deviation set by manual experience.

[0086] Through the above technical solutions, this application can dynamically perceive the real-time load status of nodes and quantify the task resource requirements, solving the uneven distribution problem caused by static thresholds in traditional strategies. The resource utilization of heterogeneous nodes is optimized through the feature matching scoring mechanism, for example, deep learning tasks are automatically assigned to nodes equipped with NPU to reduce GPU resource waste. Furthermore, the modular design supports cross-cloud platform deployment, effectively integrates the dispersed resource pools in the hybrid cloud environment, and improves the efficiency of large-scale task scheduling.

[0087] The present invention also provides a computer device, which includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the computing power node load balancing method in the above-mentioned embodiments.

[0088] Compared with the existing technology, traditional load balancing devices rely on fixed polling strategies or manually set weights and cannot dynamically adapt to changes in node status. This solution integrates the prediction model with the real-time monitoring module to enable the device to automatically sense node resource fluctuations. For example, when a node suddenly has a high load, the system automatically removes it from the candidate set. In addition, the existing technology requires manual configuration of each platform interface in a multi-cloud environment, while this solution uses a unified resource abstraction layer to encapsulate the application programming interface calls of heterogeneous cloud platforms into a standard instruction set to achieve transparent scheduling of cross-cloud resources.

[0089] Through the above technical solutions, the present application solves the hysteresis problem of static strategies and realizes dynamic decision-making based on real-time load fluctuations; reduces the subjective error of manually set thresholds and automatically quantifies resource requirements through prediction models; overcomes the fragmentation problem of cross-cloud resource pools and establishes a unified resource scheduling interface; optimizes the utilization of heterogeneous nodes, automatically matches accelerated computing units according to task characteristics, and avoids resource waste caused by general-purpose nodes processing dedicated tasks.

[0090] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the computing power node load balancing method.

[0091] Compared with the existing technology, the existing load balancing solution relies on static configuration files and cannot respond to node status fluctuations in real time. This embodiment implements a fully automated decision-making process by solidifying dynamic prediction and evaluation algorithms through instructions. The existing technology requires manual maintenance of scheduling strategies for different cloud platforms. This embodiment can shield the differences in underlying platforms through a unified instruction set. The existing algorithm does not fully consider the characteristics of heterogeneous computing units. This embodiment optimizes the resource matching mechanism through feature similarity calculation.

[0092] Through the above technical solution, this application solves the problem that traditional methods cannot perceive node load changes in real time, and improves resource allocation accuracy through dynamic prediction models and real-time monitoring; reduces the running overhead of dynamic algorithms, solidifies the core logic through preset instructions to avoid repeated calculations, and can achieve unified deployment of resource scheduling strategies across cloud environments, eliminate the impact of resource pool fragmentation, optimize the computing power utilization of heterogeneous nodes, and adapt to the computing characteristics of different hardware acceleration units through feature matching mechanisms.

[0093] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0094] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0095] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A computing node load balancing method, characterized in that: The computing power node load balancing method includes: Read the computing task to be calculated, and extract the task feature information of the computing task; Input the task feature information into a preset prediction model to predict the amount of computing power resources, and output the amount of computing power resources required to calculate the computing power task; Calculate the real-time computing load of each computing node; According to the real-time computing load of each computing node and the amount of computing resources required for the computing task, a number of candidate computing nodes for calculating the computing task are screened out; Calculate the matching score between each candidate computing power node and the computing power task; According to the matching score, the best computing node is selected from each of the candidate computing nodes, and the computing task is assigned to the best computing node.

2. The computing node load balancing method according to claim 1, characterized in that: Before reading the computing task to be calculated, the method further includes: Obtain the task information corresponding to each historical computing task and the amount of historical computing resources used to complete each historical computing task; According to the order of task completion time, feature extraction is performed on the task information of each historical computing task to obtain a task feature vector, and the historical computing resources used to complete each historical computing task are vectorized to obtain the corresponding historical computing resources vector; The task feature vector and the historical computing power resource vector are used as inputs of the initial prediction model, and the prediction deviation value of the historical computing power resource corresponding to each historical computing power task is trained and outputted. Based on the prediction deviation value, the weight matrix and bias of the input layer and hidden layer of the initial prediction model are adjusted to obtain a trained prediction model.

3. The computing node load balancing method according to claim 2, characterized in that: The prediction model includes an input layer, multiple hidden layers and an output layer in sequence; the output of the previous layer of the multiple hidden layers is used as the input of the next layer; the hidden layer includes multiple operation nodes arranged in sequence, each of which includes a weight matrix, a bias, and an addition operation node and a primary operation node for connecting the operation nodes in the previous layer; Among them, the weight matrix is ​​used to perform dot multiplication with the vector matrix output by a corresponding operation node in the previous layer to obtain a result vector matrix; the bias is used to offset the elements in the result vector matrix; the addition operation node is used to accumulate the bias and each element in the result vector matrix to obtain an accumulated vector matrix; the primary operation node is used to perform linear or nonlinear operations on each element in the accumulated vector matrix output by the addition operation node to obtain an output vector matrix.

4. The computing power node load balancing method according to claim 1, characterized in that: The calculation of the real-time computing load of each computing node includes: Detect resource usage of each computing node; Calculate the benchmark resource consumption rate of each computing power node based on the resource utilization rate of each computing power node; Calculate the computing performance evaluation value of each computing power node based on the benchmark resource consumption rate of each computing power node; Get the current traffic information of each computing power node; Calculate the performance coefficient of each computing power node based on the current traffic information of each computing power node; The real-time load of each computing power node is calculated based on the performance evaluation value and performance coefficient of each computing power node.

5. The computing node load balancing method according to claim 1, characterized in that: The candidate computing nodes for calculating the computing task are selected based on the real-time computing load of each computing node and the amount of computing resources required for the computing task, including: Convert the real-time computing load of each computing node into the corresponding computing resources, and calculate the real-time remaining computing resources of each computing node; Determine respectively whether there is a target computing power node whose real-time remaining computing power resources are greater than the computing power resources required by the computing power task; If there are several target computing nodes whose real-time remaining computing power resources are greater than the computing power resources required by the computing power task, then the real-time load fluctuation and computing power load stability of each target computing power node within a preset time period are calculated; According to the real-time load fluctuation and computing load stability of each target computing node within a preset time period, a number of candidate computing nodes for calculating the computing task are screened out.

6. The computing node load balancing method according to claim 1, characterized in that: The calculating of the matching score between each candidate computing power node and the computing power task includes: Respectively extracting node computing power resource characteristics of each of the candidate computing power nodes and computing power resource characteristics required for the computing power task; The feature similarity between the node computing power resource characteristics of each candidate computing power node and the feature of the computing power resource required for the computing power task; The feature similarities corresponding to each candidate computing power node are weighted and summed up respectively to obtain a matching score between each candidate computing power node and the computing power task.

7. The computing node load balancing method according to any one of claims 1 to 6, characterized in that: The computing power node load balancing method also includes: When the best computing power node calculates the computing power task, splitting the computing power task into multiple sub-computing power tasks; Determine the resource consumption rate of each sub-computing task according to the task type of each sub-computing task; Monitor the running time of each sub-computing task; According to the running time and resource consumption rate of each sub-computing task, the amount of computing resources used to complete the computing task is statistically obtained.

8. A computing node load balancing device, characterized in that: The computing power node load balancing device comprises: A reading module, used to read the computing task to be calculated and extract task feature information of the computing task; A prediction module, used to input the task feature information into a preset prediction model to predict the amount of computing power resources, and output the amount of computing power resources required to calculate the computing power task; The first computing module is used to calculate the real-time computing load of each computing node; A screening module, used to screen out a number of candidate computing nodes for calculating the computing task according to the real-time computing load of each computing node and the amount of computing resources required for the computing task; A second calculation module is used to calculate the matching score between each candidate computing power node and the computing power task; The allocation module is used to select the best computing power node from each of the candidate computing power nodes according to the matching score, and allocate the computing power task to the best computing power node.

9. A computer device, characterized in that: The computer device includes: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the computing power node load balancing method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the computing power node load balancing method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Calculation power network collaborative service scheduling method and system based on service intention

    CN117241393A

  • Resource allocation method based on computing power network and related equipment

    CN117834560A

  • Heterogeneous hardware computing power scheduling method and device, equipment and medium

    CN118626263A

  • Task management method and project management system for scientific research platform

    CN119476899A

  • Dynamic scheduling method for cloud computing resource pool

    CN119537025A

Cited By

  • Thread block configuration method and device, readable storage medium and program product

    CN120315751A

  • Task allocation method, device and equipment, storage medium and computer program product

    CN120762926A

  • Cross-platform shared data synchronization optimization system

    CN120785901A

  • Gradient scheduling method and system based on cloud platform

    CN120849059A

  • Data processing method and device, equipment and medium

    CN120872553A