Calculation power scheduling method and device
By identifying the latency sensitivity of the jobs to be processed and analyzing green energy data, the scheduling scores of candidate computing power clusters are determined, which solves the problem of low green energy power utilization and achieves more efficient computing power scheduling and cost optimization.
Patent Information
- Application Number
- CN202511314613.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-15
AI Technical Summary
The computing power scheduling method in the existing technology has a low utilization rate of green energy electricity, resulting in high operating costs of data centers and poor scheduling stability and flexibility.
By identifying the latency sensitivity of the pending jobs, candidate computing power clusters are determined, including candidate computing power clusters based on matching green and grid power resource requirements. Based on the current green energy data and current cluster load information of each candidate computing power, the scheduling score of the candidate computing power cluster is determined, and the candidate computing power cluster with the highest scheduling score is used as the target computing power cluster, and the pending jobs are distributed to the target computing power cluster for execution.
It improves the utilization rate of green energy and electricity, reduces the operating costs of data centers, and improves the stability and flexibility of computing power scheduling.
Smart Images

Figure CN120803680A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a computing power scheduling method and device. BACKGROUND
[0002] With the rapid development of big data and artificial intelligence technology, computing power resources have become the core resources to support complex computing tasks and intelligent applications, and data centers, as key computing hubs, integrate large-scale servers, GPU clusters, and high-speed interconnection components, etc. Among them, green data centers based on green energy power have greater advantages in reducing operating costs compared to traditional data centers based on only grid power. However, the computing power scheduling method in the related art has a low utilization rate of green energy power, resulting in a still high operating cost of the data center, and the stability and flexibility of the computing power scheduling are also poor. SUMMARY
[0003] In order to solve the problems of the prior art, the embodiments of the present application provide a computing power scheduling method and device. The technical solution is as follows: In one aspect, a computing power scheduling method is provided, the method comprising: obtaining a to-be-processed job; identifying the time delay sensitivity of the to-be-processed job to obtain a time delay sensitivity identification result of the to-be-processed job; in a case where the time delay sensitivity identification result indicates that the to-be-processed job is not time delay sensitive, determining a candidate computing power cluster matching the resource requirement of the to-be-processed job from a plurality of computing power clusters; the plurality of computing power clusters include a first type of computing power cluster based on green energy power and a second type of computing power cluster based on grid power; determining a scheduling score of each candidate computing power cluster according to current green energy data and current cluster load information of each candidate computing power cluster, and taking the candidate computing power cluster with the highest scheduling score as a target computing power cluster; distributing the to-be-processed job to the target computing power cluster, so that the target computing power cluster executes the to-be-processed job.
[0004] In some exemplary embodiments, the identifying the time delay sensitivity of the to-be-processed job to obtain the time delay sensitivity identification result of the to-be-processed job comprises: extracting metadata of the to-be-processed job to determine the field content of the time delay sensitivity identification field in the metadata; in a case where the field content is not empty, determining the time delay sensitivity identification result of the to-be-processed job based on the field content of the time delay sensitivity identification field; In a case where the field content is empty, based on a job type and a job estimated execution duration in the metadata, a time delay sensitivity identification result of the to-be-processed job is determined.
[0005] In some example embodiments, the determining, according to current green energy data and current cluster load information of each of the candidate computing power clusters, a scheduling score of each of the candidate computing power clusters comprises: For each of the candidate computing power clusters, based on current green energy data of the candidate computing power cluster, a green energy availability score corresponding to the candidate computing power cluster is determined; the current green energy data comprises current power generation and predicted power generation of green energy corresponding to the candidate computing power cluster; the green energy availability score represents a power supply capability of the green energy corresponding to the candidate computing power cluster; Based on current cluster load information of the candidate computing power cluster, a cluster load score corresponding to the candidate computing power cluster is determined; the cluster load score represents an idle degree of the candidate computing power cluster; Based on weighted sum processing of the green energy availability score and the cluster load score corresponding to the candidate computing power cluster, a scheduling score of the candidate computing power cluster is obtained.
[0006] In some example embodiments, the weighted sum processing of the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster comprises: Based on a proportion of current output power of green energy corresponding to the candidate computing power cluster in a rated demand power of the candidate computing power cluster, an operation cost score of the candidate computing power cluster is determined; the operation cost score represents an electricity cost of running the to-be-processed job in the candidate computing power cluster; Based on weighted sum processing of the green energy availability score, the cluster load score and the operation cost score corresponding to the candidate computing power cluster, a scheduling score of the candidate computing power cluster is obtained.
[0007] In some example embodiments, the weighted sum processing of the green energy availability score, the cluster load score and the operation cost score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster comprises: According to a storage location of computing data required by the to-be-processed job, a data transmission cost is determined; Based on the data transmission cost, a data affinity score corresponding to the candidate computing power cluster is determined; the data affinity score represents a distance between the computing data required by the to-be-processed job and the candidate computing power cluster; The green energy availability score, the cluster load score, the operation cost score, and the data affinity score corresponding to the candidate computing power cluster are weighted and summed to obtain a scheduling score of the candidate computing power cluster.
[0008] In some example embodiments, the weighting and summing of the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster comprises: obtaining a set of weight coefficients, the set of weight coefficients including weight coefficients corresponding to each of the items in the weighting and summing, weighting and summing the green energy availability score and the cluster load score corresponding to the candidate computing power cluster based on the set of weight coefficients to obtain the scheduling score of the candidate computing power cluster.
[0009] In some example embodiments, the method further comprises: in a case where the time delay sensitivity identification result indicates that the to-be-processed job is time delay sensitive, increasing a first weight coefficient in the set of weight coefficients and decreasing a second weight coefficient in the set of weight coefficients to obtain an adjusted set of weight coefficients, the first weight coefficient being a weight coefficient corresponding to the cluster load score item, and the second weight coefficient being a weight coefficient corresponding to the green energy availability score item; weighting and summing the green energy availability score and the cluster load score corresponding to the candidate computing power cluster based on the adjusted set of weight coefficients to obtain the scheduling score of the candidate computing power cluster, and selecting a candidate computing power cluster with the highest scheduling score as a target computing power cluster.
[0010] In some example embodiments, after the to-be-processed job is distributed to the target computing power cluster, the method further comprises: detecting green energy data and cluster load information of each computing power cluster; updating current green energy data and current cluster load information of each computing power cluster based on the detection results.
[0011] In another aspect, a computing power scheduling device is provided, the device comprising: a job obtaining module configured to obtain a to-be-processed job; a time delay sensitivity identification module configured to identify the time delay sensitivity of the to-be-processed job to obtain a time delay sensitivity identification result of the to-be-processed job; The candidate computing power cluster determination module is configured to determine, from a plurality of computing power clusters, a candidate computing power cluster matching a resource requirement of the to-be-processed job, in a case where the time delay sensitivity identification result indicates that the to-be-processed job is non-time delay sensitive, wherein the plurality of computing power clusters include a first type of computing power cluster based on green energy power and a second type of computing power cluster based on grid mains power. The cluster scheduling score determination module is configured to determine, according to current green energy data and current cluster load information of each of the candidate computing power clusters, a scheduling score of each of the candidate computing power clusters, and take the candidate computing power cluster with the highest scheduling score as a target computing power cluster. The job distribution module is configured to distribute the to-be-processed job to the target computing power cluster, so that the target computing power cluster executes the to-be-processed job.
[0012] In some example embodiments, the time delay sensitivity identification module includes: The field content determination module is configured to extract metadata of the to-be-processed job, and determine field content of a time delay sensitivity identification field in the metadata. The first determination module is configured to, in a case where the field content is not empty, determine a time delay sensitivity identification result of the to-be-processed job based on the field content of the time delay sensitivity identification field. The second determination module is configured to, in a case where the field content is empty, determine a time delay sensitivity identification result of the to-be-processed job based on a job type and a job estimated execution duration in the metadata.
[0013] In some example embodiments, the cluster scheduling score determination module includes: The green energy availability score determination module is configured to, for each of the candidate computing power clusters, determine a green energy availability score corresponding to the candidate computing power cluster based on current green energy data of the candidate computing power cluster, wherein the current green energy data includes a current power generation amount and a predicted power generation amount of a green energy corresponding to the candidate computing power cluster, and the green energy availability score represents a power supply capability of the green energy corresponding to the candidate computing power cluster. The cluster load score determination module is configured to determine a cluster load score corresponding to the candidate computing power cluster based on current cluster load information of the candidate computing power cluster, wherein the cluster load score represents an idle degree of the candidate computing power cluster. The weighting processing module is configured to perform weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster, to obtain a scheduling score of the candidate computing power cluster.
[0014] In some example embodiments, the weighting processing module includes: The operation cost score determination module is configured to determine an operation cost score of the candidate computing power cluster based on a proportion of a current output power of green energy corresponding to the candidate computing power cluster in a rated demand power of the candidate computing power cluster; and the operation cost score represents an electricity cost of running the to-be-processed job on the candidate computing power cluster. The weighting processing first submodule is configured to perform weighted summation processing on the green energy availability score, the cluster load score and the operation cost score corresponding to the candidate computing power cluster to obtain a scheduling score of the candidate computing power cluster.
[0015] In some example embodiments, the weighting processing first submodule includes: The data transmission cost determination module is configured to determine a data transmission cost according to a storage location of the calculation data required by the to-be-processed job. The data affinity score determination module is configured to determine a data affinity score corresponding to the candidate computing power cluster based on the data transmission cost; and the data affinity score represents a distance between the calculation data required by the to-be-processed job and the candidate computing power cluster. The weighting processing second submodule is configured to perform weighted summation processing on the green energy availability score, the cluster load score, the operation cost score and the data affinity score corresponding to the candidate computing power cluster to obtain a scheduling score of the candidate computing power cluster.
[0016] In some example embodiments, the weighting processing module is specifically configured to: obtain a set of weight coefficients; the set of weight coefficients includes weight coefficients corresponding to respective items in the weighted summation processing; perform weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster based on the set of weight coefficients to obtain a scheduling score of the candidate computing power cluster; and select a candidate computing power cluster with the highest scheduling score as the target computing power cluster.
[0017] In some example embodiments, the apparatus further includes: The weight coefficient adjustment module is configured to, in a case where the time delay sensitivity identification result indicates that the to-be-processed job is time delay sensitive, increase a first weight coefficient in the set of weight coefficients and decrease a second weight coefficient in the set of weight coefficients to obtain an adjusted set of weight coefficients; the first weight coefficient is a weight coefficient corresponding to the cluster load score item, and the second weight coefficient is a weight coefficient corresponding to the green energy availability score item. The summation processing module is configured to perform weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster based on the adjusted set of weight coefficients to obtain a scheduling score of the candidate computing power cluster; and select a candidate computing power cluster with the highest scheduling score as the target computing power cluster.
[0018] In some exemplary embodiments, the apparatus further comprises: a detection module configured to detect green energy data and cluster load information of each of the computing power clusters; an updating module configured to update current green energy data and current cluster load information of each of the computing power clusters based on the detection results.
[0019] In another aspect, an electronic device is provided, comprising a processor and a memory, the memory having stored therein at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the computing power scheduling method of any of the above aspects.
[0020] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium having stored therein at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the computing power scheduling method of any of the above aspects.
[0021] In another aspect, a computer program product or computer program is provided, the computer program product or computer program comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the computing power scheduling method of any of the above aspects.
[0022] The embodiments of the present application can effectively transfer a large amount of computing load to computing power clusters with abundant green energy power, improve the utilization rate of green energy power and the on-site consumption rate of green energy, effectively reduce the operating cost of the data center, and adapt to the dynamic time-varying nature of green energy power, thereby improving the stability and flexibility of computing power scheduling. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the description of the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0024] Figure 1 is a system architecture diagram of a joint scheduling provided by an embodiment of the present application; Figure 2 is a flow diagram of a computing power scheduling method provided by an embodiment of the present application; Figure 3 is a flow diagram of another computing power scheduling method provided by an embodiment of the present application; Figure 4 is a flow diagram of another computing power scheduling method provided by an embodiment of the present application; Figure 5 is a structural block diagram of a computing power scheduling device provided by an embodiment of the present application; Figure 6 is a hardware structural block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely in the following description with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of protection of the present application.
[0026] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0028] It can be understood that in the specific embodiments of the present application, data related to user information and the like are involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions.
[0029] Please refer to Figure 1 , which is a system architecture diagram of joint scheduling provided by an embodiment of the present application, including a joint scheduling node 110 and multiple data centers 120 distributed in different geographical locations (such as region a, region b, region c, and region d). The joint scheduling node 110 and each data center 120 can communicate through network connection. The multiple data centers 120 include a green data center 121 and a traditional data center 122, wherein the green data center 121 uses green energy power (such as wind power, hydroelectric power, and photovoltaic power) and grid commercial power for power supply, and the traditional data center 122 only uses grid commercial power for power supply. For example, Figure 1 As shown in the figure, the green data center 121 is deployed in region a and region b, and the traditional data center 122 is deployed in region c and region d.
[0030] The green data center 121 and the traditional data center 122 can each include multiple computing power clusters. The cluster types of the multiple computing power clusters can include but are not limited to Kubernetes clusters and Slurm clusters. The multiple computing power clusters can be heterogeneous computing power clusters. The heterogeneous computing power in the heterogeneous computing power clusters can include but is not limited to graphics processing units (GPUs), neural network processing units (NPUs), tensor processing units (TPUs), FPGA (Field Programmable Gate Array) hardware, etc. It should be noted that the computing power clusters in the embodiments of the present application include first-type computing power clusters and second-type computing power clusters. The power supply sources of the first-type computing power clusters include green energy power, and the power supply sources of the second-type computing power clusters are only grid commercial power. It can be understood that the power supply sources of the first-type computing power clusters can also include grid commercial power in addition to green energy power.
[0031] Specifically, in each data center, including the green data center 121 and the traditional data center 122, an energy monitoring agent is deployed, which is used to detect the green energy data of each computing power cluster in the data center to which it belongs at a preset time interval, and send the detected green energy data of each computing power cluster to the joint scheduling node 110, so that the joint scheduling node 110 can obtain the current green energy data of each computing power cluster, and provide data basis for scheduling decision. The time interval of green energy data detection, i.e. the preset time interval, can be set based on actual needs. Generally, the shorter the preset time interval is set, the stronger the real-time detection is, and real-time green energy data detection can be achieved. In the embodiment of the present application, the green energy data of the computing power cluster can represent the green energy supply state of the computing power cluster, and specifically can include the current power generation, i.e. the current output power (kW) of green energy, and the predicted power generation, which is the power generation curve of a future period of time (such as the next 6 hours) predicted based on the current environmental information (such as light intensity, wind speed) and the historical power generation data model. The historical power generation data model is a trained neural network model used to predict the power generation curve of a future period of time. In actual application, the current environmental information and the power generation data of a recent period of time (such as the last week) can be input into the historical power generation data model to predict the power generation curve of the future period of time.
[0032] In each computing power cluster, a cluster load monitoring agent is deployed, such as Prometheus Node Exporter, kube-state-metrics in a Kubernetes cluster, and SlurmExporter for Prometheus in a Slurm cluster. The cluster load monitoring agent is used to detect the cluster load information of each computing power cluster in the data center to which it belongs at a preset time interval, and send the detected cluster load information of each computing power cluster to the joint scheduling node 110, so that the joint scheduling node 110 can obtain the current cluster load information of each computing power cluster, and provide data basis for scheduling decision. The preset time interval can be set based on actual needs. Generally, the shorter the preset time interval is set, the stronger the real-time detection is, and real-time cluster load information detection can be achieved. In the embodiment of the present application, the cluster load information of the computing power cluster represents the load state of the computing power cluster, and specifically can include resource utilization rate, such as real-time utilization rate of CPU, memory, and GPU, resource allocation rate, i.e. the ratio of the amount of resources allocated to existing jobs to the total amount of resources, job queue state, such as the number of jobs waiting for execution and the total waiting time, and node state, such as whether each computing node is available.
[0033] The joint scheduling node 110 stores current green energy data and current cluster load information of each computing power cluster. It can be understood that the current green energy data and the current cluster load information are updated according to the detection results of each energy monitoring agent and the detection results of each cluster load monitoring agent.
[0034] It should be noted that the nodes / servers involved in the embodiments of the present application can be independent physical servers, or server clusters or distributed systems formed by multiple physical servers, or cloud servers providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal devices involved in the embodiments of the present application include, but are not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, and the like.
[0035] In an exemplary embodiment, the joint scheduling node 110 and each computing power cluster can be a node device in a blockchain system, capable of sharing the acquired and generated information to other node devices in the blockchain system, realizing information sharing between multiple node devices. The multiple node devices in the blockchain system can be configured with the same blockchain, which is composed of multiple blocks, and the adjacent blocks have an association relationship, so that when the data in any block is tampered with, it can be detected through the next block, thereby avoiding the data in the blockchain being tampered with, and ensuring the security and reliability of the data in the blockchain.
[0036] Please refer to Figure 2 , which is a flowchart of a computing power scheduling method provided by an embodiment of the present application. The method can be applied to the joint scheduling node in Figure 1 It should be noted that the present specification provides method operation steps as described in the embodiments or flowcharts, but more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many execution orders, and does not represent the only execution order. In actual system or product execution, the method order can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, as shown in Figure 2 , the computing power scheduling method of the embodiments of the present application can include: S201, acquiring a to-be-processed job.
[0037] Specifically, a terminal device can be used to show a user a job configuration interface. The terminal device can generate a scheduling job based on job parameters configured in the job configuration interface in response to a job submission instruction triggered based on the job configuration interface, and send the scheduling job to a joint scheduling node. Thus, the joint scheduling node obtains the scheduling job submitted by the user and processes the scheduling job as a to-be-processed job.
[0038] The to-be-processed job contains metadata generated based on user configuration. The metadata can include field contents of a plurality of preset fields. The plurality of preset fields can include a cluster type field, a resource requirement field, a running time estimation field, a time delay sensitivity identification field, and a job type label field. The field contents of the preset fields can be configured by the user through the job configuration interface. Specifically, the cluster type field is used to specify the cluster type corresponding to the job, such as a Kubernetes cluster job or a Slurm cluster job. The resource requirement field is used to configure resource requirement information, such as the number of CPU cores, the size of memory, the type and quantity of GPUs, and the like. The running time estimation field is used to configure the estimated execution time length of the job. The time delay sensitivity identification field is used to configure whether the job is time delay sensitive. Usually, it is a Boolean type label (true / false) displayed and set by the user. The job type label field is used to configure the job type, such as model training (training), online inference (inference), data batch processing (batch-processing), and interactive application (interactive).
[0039] S203, identifying the time delay sensitivity of the to-be-processed job to obtain a time delay sensitivity identification result of the to-be-processed job.
[0040] The time delay sensitivity identification result of the to-be-processed job indicates whether the to-be-processed job is time delay sensitive. Usually, the to-be-processed job that is time delay sensitive has high sensitivity to running time delay and needs to be executed immediately. For example, online inference and interactive application are “time delay sensitive” jobs and must be executed immediately. The to-be-processed job that is not time delay sensitive has low sensitivity to running time delay and can be delayed, suspended, or migrated. For example, AI model training, data batch processing, and scientific simulation are usually “non-time delay sensitive” jobs.
[0041] In some example embodiments, the step S203 can comprise: extracting metadata of the to-be-processed job, determining field content of a time-sensitive identification field in the metadata; in a case where the field content is not empty, determining the time-sensitive identification result of the to-be-processed job based on the field content of the time-sensitive identification field; and in a case where the field content is empty, determining the time-sensitive identification result of the to-be-processed job based on a job type and a job estimated execution duration in the metadata.
[0042] Specifically, if the field content of the time-sensitive identification field is not empty, it indicates that the user has set whether the scheduled job is time-sensitive, and at this time, explicit identification can be performed based on the field content, i.e., determining the time-sensitive identification result of the to-be-scheduled job according to the field content set by the user. For example, if the user sets the field content of the time-sensitive identification field as false, it can be determined that the time-sensitive identification result of the to-be-scheduled job is non-time-sensitive, and vice versa, if the user sets the field content of the time-sensitive identification field as true, it can be determined that the time-sensitive identification result of the to-be-scheduled job is time-sensitive.
[0043] If the field content of the time-sensitive identification field is empty, it indicates that the user has not set whether the scheduled job is time-sensitive, and at this time, implicit identification can be performed according to the job type and the job estimated execution duration in the metadata. Specifically, when the job type is a preset first job type, or the job estimated execution duration is less than a preset execution duration threshold, it is determined that the time-sensitive identification result of the to-be-scheduled job is time-sensitive, wherein the preset first job type can be set based on actual experience, for example, the preset first job type is online inference and interactive application, and the preset execution duration threshold can also be set based on actual experience, for example, the preset execution duration threshold is 30 minutes. When the job type is a preset second job type or the job estimated execution duration is greater than the preset execution duration threshold, it is determined that the time-sensitive identification result of the to-be-scheduled job is non-time-sensitive, wherein the preset second job type can be set based on actual experience, for example, the preset second job type is model training and batch processing.
[0044] The above-mentioned embodiments combine explicit identification and implicit identification of time sensitivity based on the metadata of the to-be-processed job, thereby accurately identifying the sensitivity of the to-be-processed job to runtime delay, and providing a basis for subsequent more fine-grained algorithm scheduling based on the time-sensitive identification result.
[0045] S205, in the case where the time delay sensitivity identification result indicates that the to-be-processed job is a non-time delay sensitive type, determining a candidate computing power cluster matching the resource requirement of the to-be-processed job from a plurality of computing power clusters.
[0046] The plurality of computing power clusters include a first type of computing power cluster based on green energy power and a second type of computing power cluster based on grid mains.
[0047] Specifically, the resource requirement of the to-be-processed job can be obtained from the metadata of the to-be-processed job, and the joint scheduling center can filter out the computing power cluster matching the resource requirement of the to-be-processed job as the candidate computing power cluster based on the current cluster load information of each computing power cluster. It can be understood that the candidate computing power cluster can include the first type of computing power cluster and / or the second type of computing power cluster.
[0048] S207, determining a scheduling score of each candidate computing power cluster according to the current green energy data and the current cluster load information of each candidate computing power cluster, and selecting the candidate computing power cluster with the highest scheduling score as the target computing power cluster.
[0049] Specifically, the joint scheduling node can traverse each candidate computing power cluster. In the traversal process, a scheduling score is calculated for each candidate computing power cluster based on the current green energy data and the current cluster load information of the candidate computing power cluster. The scheduling score can make the candidate computing power cluster with abundant green energy power and low cluster load obtain a high score. After the traversal ends, the candidate computing power cluster with the highest scheduling score is selected as the target computing power cluster, so that the "non-time delay sensitive type" scheduling job can be preferentially scheduled to the green computing power cluster with the highest scheduling score, thereby maximizing the utilization rate of green energy.
[0050] In some exemplary embodiments, for each candidate computing power cluster, the above step S207 can include, when implemented: determining a green energy availability score corresponding to the candidate computing power cluster based on the current green energy data of the candidate computing power cluster; determining a cluster load score corresponding to the candidate computing power cluster based on the current cluster load information of the candidate computing power cluster; and performing weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster.
[0051] The current green energy data includes the current power generation and the predicted power generation of the green energy corresponding to the candidate computing power cluster, and the green energy availability score represents the power supply capability of the green energy corresponding to the candidate computing power cluster; and the cluster load score represents the idle degree of the candidate computing power cluster.
[0052] In a specific implementation, the green energy availability score of the candidate computing cluster can be obtained by using the following formula (1): (1) wherein, represents the green energy availability score of the candidate computing cluster; represents the current power generation of the green energy corresponding to the candidate computing cluster, i.e., the currently available green energy power; represents the average green energy power in the estimated execution time of the to-be-processed job based on the corresponding predicted power generation; represents the total consumed power when the candidate computing cluster is running at full capacity, which can be obtained by summing the power consumption of each computing device in the candidate computing cluster; represents a decay factor (0 to 1) for adjusting the trust degree of the future prediction.
[0053] The cluster load score of the candidate computing cluster can be obtained by using the following formula (2): (2) wherein, represents the cluster load score of the candidate computing cluster; represents the amount of resources (such as CPU occupancy, GPU occupancy, memory temporary occupancy, and network bandwidth occupancy) allocated to the candidate computing cluster; represents the total amount of resources of the candidate computing cluster.
[0054] When performing weighted summation processing, the green energy availability score and the cluster load score can be respectively set with a weight coefficient, and then the green energy availability score and the cluster load score are weighted and summed based on the respective weight coefficients to obtain the scheduling score of the corresponding candidate computing cluster. Through the scheduling score, the candidate computing cluster with sufficient green energy power and low cluster load can obtain a high score, which can guide the scheduling of the job flow to the candidate computing cluster and can avoid sending the job to the computing cluster that is about to be saturated, thereby improving the job startup speed and the running stability.
[0055] It can be understood that when the green energy power supply is insufficient, tends to 0, at which time the score is extremely low, thereby effectively preventing the newly scheduled job from being distributed to the candidate computing cluster, and achieving dynamic adjustment of the cluster load and the green energy supply.
[0056] In some example embodiments, when the scheduling score of the candidate computing power cluster is obtained by performing weighted sum processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster, the operation cost score of the candidate computing power cluster can be determined based on the proportion of the current output power of the green energy corresponding to the candidate computing power cluster in the rated demand power of the candidate computing power cluster; and the scheduling score of the candidate computing power cluster can be obtained by performing weighted sum processing on the green energy availability score, the cluster load score and the operation cost score corresponding to the candidate computing power cluster.
[0057] The operation cost score represents the power cost of running the to-be-processed job by the candidate computing power cluster.
[0058] In a specific implementation, the operation cost score of the candidate computing power cluster can be obtained by the following formula (3): (3) In formula (3), represents the operation cost score of the candidate computing power cluster; represents the unit cost of green energy power (usually very low or 0); represents the unit price of the current commercial power; represents the proportion of green energy power of the candidate computing power cluster, = , represents the current power generation of the green energy corresponding to the candidate computing power cluster, i.e., the currently available green energy power, represents the rated demand power of the candidate computing power cluster; represents the highest commercial power price in all computing power clusters, for normalization.
[0059] When performing the weighted sum processing, a weight coefficient can be set for the green energy availability score, the cluster load score and the operation cost score, respectively, and then the green energy availability score, the cluster load score and the operation cost score are weighted and summed based on the respective weight coefficients to obtain the scheduling score of the corresponding candidate computing power cluster. Through the scheduling score, the candidate computing power cluster with abundant green energy power, low cluster load and low operation cost can obtain a high score, which can improve the consumption rate of green energy in the data center while effectively reducing the operation cost of the data center.
[0060] In some example embodiments, when the scheduling score of the candidate computing power cluster is obtained by performing weighted sum processing on the green energy availability score, the cluster load score and the operation cost score corresponding to the candidate computing power cluster, the method can comprise: determining a data transmission cost according to a storage location of the required computing data of the to-be-processed job; determining a data affinity score corresponding to the candidate computing power cluster based on the data transmission cost; and obtaining the scheduling score of the candidate computing power cluster by performing weighted sum processing on the green energy availability score, the cluster load score, the operation cost score and the data affinity score corresponding to the candidate computing power cluster.
[0061] The data affinity score represents a distance between the required computing data of the to-be-processed job and the candidate computing power cluster.
[0062] In a specific implementation, the data affinity score of the candidate computing power cluster can be obtained by the following formula (4): (4) In formula (4), the data affinity score of the candidate computing power cluster is represented by Sdata. The data transmission cost is proportional to the transmission time or the transmission cost.
[0063] Specifically, if a large data set required by the job (such as AI training data) already exists in the storage of a candidate computing power cluster, Sdata=0, and the score of the candidate computing power cluster is 1. If the data needs to be transmitted across clusters or across regions, since the data transmission cost is proportional to the transmission time or the transmission cost, Sdatais attenuated according to the transmission time or the transmission cost.
[0064] When performing the weighted sum processing, a weight coefficient can be set for the green energy availability score, the cluster load score, the operation cost score and the data affinity score, respectively, and then the green energy availability score, the cluster load score, the operation cost score and the data affinity score are weighted and summed based on the respective weight coefficients to obtain the scheduling score of the corresponding candidate computing power cluster. Through the scheduling score, the candidate computing power cluster with abundant green energy power, low cluster load, low operation cost and short distance can obtain a high score, which can improve the consumption rate of green energy in the data center, effectively reduce the operation cost of the data center, reduce the delay and network cost caused by data transmission, and improve the scheduling efficiency.
[0065] An exemplary method for obtaining the scheduling score of the candidate computing power cluster based on the weighted sum processing of the green energy availability score and the cluster load score corresponding to the candidate computing power cluster can include: obtaining a weight coefficient set; the weight coefficient set includes weight coefficients corresponding to each term in the weighted sum processing, and the weighted sum processing of the green energy availability score and the cluster load score corresponding to the candidate computing power cluster is performed based on the weight coefficient set to obtain the scheduling score of the candidate computing power cluster.
[0066] Specifically, the scheduling score of the candidate computing power cluster can be represented by the following formula (5): (5) wherein, represents the weight coefficient of each term, which can be set by an administrator according to actual needs (such as cost priority), and .
[0067] S209, distributing the to-be-processed job to the target computing power cluster, so that the target computing power cluster executes the to-be-processed job.
[0068] Specifically, if the target computing power cluster is a Kubernetes cluster, the joint scheduling node can convert the description file of the to-be-processed job into the job format of the Kubernetes API Server, and then submit it to the target computing power cluster, and if the target computing power cluster is a Slurm cluster, the joint scheduling node can submit the job script of the to-be-processed job to the controller of the target computing power cluster through the sbatch command.
[0069] The technical scheme of the embodiments of the present application can effectively transfer the huge computing load to the computing power cluster with abundant green energy power, improve the utilization rate of green energy power and the on-site consumption rate of green energy, and maximize the use of zero marginal cost green energy, thereby reducing the electricity bill of the data center. At the same time, each degree of green electricity used reduces the dependence on fossil fuel power, thereby significantly reducing the operating cost of the data center.
[0070] In some exemplary embodiments, as Figure 3 shown, the method can further include: S211, in the case where the time delay sensitivity identification result indicates that the to-be-processed job is time delay sensitive, adjusting the first weight coefficient in the weight coefficient set and adjusting the second weight coefficient in the weight coefficient set to obtain an adjusted weight coefficient set.
[0071] The first weight coefficient is a weight coefficient corresponding to the cluster load score item, and the second weight coefficient is a weight coefficient corresponding to the green energy availability score item.
[0072] S213, based on the adjusted weight coefficient set, performing weighted sum processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain a scheduling score of the candidate computing power cluster, and taking the candidate computing power cluster with the highest scheduling score as the target computing power cluster.
[0073] Specifically, the weight coefficient set includes weight coefficients corresponding to each sub-score item used to calculate the scheduling score, and the sum of the weight coefficients corresponding to each sub-score item is 1, wherein each sub-score item includes the green energy availability score and the cluster load score, and of course can also include the operation cost score and the data affinity score.
[0074] In the initial weight coefficient set, the first weight coefficient is less than the second weight coefficient, so as to preferentially schedule the non-time-sensitive scheduling job to the green computing power cluster with the highest current scheduling score to maximize the utilization rate of green energy. When the to-be-processed job is a time-sensitive job, by increasing the first weight coefficient and decreasing the second weight coefficient, the first weight coefficient can be greater than the second weight coefficient, so as to preferentially select the computing power cluster that is currently the most idle and responds the fastest, regardless of whether it is a green computing power cluster, to achieve the goal of "fast execution" for the time-sensitive job. By identifying and scheduling the time-sensitive job, the "green" can be pursued while ensuring that the time-sensitive scheduling job with high priority is responded in time, realizing the dual goals of "green" and "performance".
[0075] In some exemplary embodiments, as shown in Figure 4 The method can further include: S401, detecting green energy data and cluster load information of each computing power cluster.
[0076] S403, updating the current green energy data and the current cluster load information of each computing power cluster based on the detection result.
[0077] Specifically, after submitting the to-be-processed job, the joint scheduling node can detect the green energy data and cluster load information of the corresponding computing power cluster through the energy monitoring agent and cluster load monitoring agent corresponding to each computing power cluster, and update the current green energy data and current cluster load information of each computing power cluster based on the detection result, thereby realizing dynamic feedback, continuously monitoring the green energy state and computing power cluster load information, and when the green computing power cluster's green energy power supply decreases, its scheduling score will automatically decrease, and the joint scheduling node will automatically stop distributing new jobs to it, thereby realizing the coupling of computing power load and green energy power supply, enabling the computing power scheduling to adapt to the changes of green energy power and computing power cluster load, continuously optimizing the matching of computing power demand and green energy power supply, and improving the automation and stability of computing power scheduling.
[0078] Corresponding to the computing power scheduling method provided in the above several embodiments, the embodiment of the present application also provides a computing power scheduling device. Since the computing power scheduling device provided by the embodiment of the present application corresponds to the computing power scheduling method provided by the above several embodiments, the implementation of the foregoing computing power scheduling method is also applicable to the computing power scheduling device provided by the present embodiment, which will not be described in detail in the present embodiment.
[0079] Please refer to Figure 5 , which shows a structure schematic diagram of a computing power scheduling device provided by an embodiment of the present application. The device has the function of implementing the computing power scheduling method in the above method embodiments. The function can be realized by hardware, or the corresponding software can be executed by hardware. As Figure 5 shown, the computing power scheduling device 500 can include: A job acquisition module 510 is configured to acquire a to-be-processed job. A time delay sensitivity identification module 520 is configured to identify the time delay sensitivity of the to-be-processed job to obtain a time delay sensitivity identification result of the to-be-processed job. A candidate computing power cluster determination module 530 is configured to, in a case where the time delay sensitivity identification result indicates that the to-be-processed job is not time delay sensitive, determine a candidate computing power cluster matching the resource demand of the to-be-processed job from a plurality of computing power clusters; the plurality of computing power clusters include a first type of computing power cluster based on green energy power and a second type of computing power cluster based on grid mains. A cluster scheduling score determination module 540 is configured to determine the scheduling score of each candidate computing power cluster according to the current green energy data and current cluster load information of each candidate computing power cluster, and determine the candidate computing power cluster with the highest scheduling score as a target computing power cluster. A job distribution module 550 is configured to distribute the to-be-processed job to the target computing power cluster, so that the target computing power cluster executes the to-be-processed job.
[0080] In some example embodiments, the time sensitivity identification module 520 comprises: a field content determination module configured to extract metadata of the to-be-processed job and determine field content of a time sensitivity identification field in the metadata; a first determination module configured to, in a case where the field content is not empty, determine a time sensitivity identification result of the to-be-processed job based on the field content of the time sensitivity identification field; a second determination module configured to, in a case where the field content is empty, determine the time sensitivity identification result of the to-be-processed job based on a job type and a job estimated execution duration in the metadata.
[0081] In some example embodiments, the cluster scheduling score determination module 540 comprises: a green energy availability score determination module configured to, for each of the candidate computing power clusters, determine a green energy availability score corresponding to the candidate computing power cluster based on current green energy data of the candidate computing power cluster; the current green energy data comprises current and predicted power generation of green energy corresponding to the candidate computing power cluster; the green energy availability score represents power supply capability of the green energy corresponding to the candidate computing power cluster; a cluster load score determination module configured to determine a cluster load score corresponding to the candidate computing power cluster based on current cluster load information of the candidate computing power cluster; the cluster load score represents an idle degree of the candidate computing power cluster; a weighting processing module configured to perform weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain a scheduling score of the candidate computing power cluster.
[0082] In some example embodiments, the weighting processing module comprises: an operation cost score determination module configured to determine an operation cost score of the candidate computing power cluster based on a proportion of current output power of green energy corresponding to the candidate computing power cluster in rated demand power of the candidate computing power cluster; the operation cost score represents power cost of running the to-be-processed job on the candidate computing power cluster; a weighting processing first sub-module configured to perform weighted summation processing on the green energy availability score, the cluster load score and the operation cost score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster.
[0083] In some example embodiments, the weighting processing first sub-module comprises: a data transmission cost determination module configured to determine a data transmission cost according to a storage location of the computation data required by the to-be-processed job; a data affinity score determination module configured to determine a data affinity score corresponding to the candidate computing power cluster based on the data transmission cost, the data affinity score representing a distance between the computation data required by the to-be-processed job and the candidate computing power cluster; a weighting processing second submodule configured to perform weighted summation processing on the green energy availability score, the cluster load score, the operation cost score and the data affinity score corresponding to the candidate computing power cluster to obtain a scheduling score of the candidate computing power cluster.
[0084] In some exemplary embodiments, the weighting processing module is specifically configured to: obtain a set of weight coefficients, the set of weight coefficients including weight coefficients corresponding to respective items in the weighted summation processing; and perform weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster based on the set of weight coefficients to obtain the scheduling score of the candidate computing power cluster.
[0085] In some exemplary embodiments, the apparatus 500 further includes: a weight coefficient adjustment module configured to, in a case where the time delay sensitivity identification result indicates that the to-be-processed job is time delay sensitive, increase a first weight coefficient in the set of weight coefficients and decrease a second weight coefficient in the set of weight coefficients to obtain an adjusted set of weight coefficients, the first weight coefficient being a weight coefficient corresponding to the cluster load score item and the second weight coefficient being a weight coefficient corresponding to the green energy availability score item; a summation processing module configured to perform weighted summation processing on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster based on the adjusted set of weight coefficients to obtain the scheduling score of the candidate computing power cluster.
[0086] In some exemplary embodiments, the apparatus 500 further includes: a detection module configured to detect green energy data and cluster load information of each of the computing power clusters; an updating module configured to update current green energy data and current cluster load information of each of the computing power clusters based on a result of the detection.
[0087] It should be noted that the apparatus provided by the above embodiments, in realizing its functions, is only exemplified by the above division of functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the apparatus and method embodiments provided by the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0088] The electronic device provided in the embodiments of the present application includes a processor and a memory, and the memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement any one of the computing power scheduling methods provided in the embodiments of the present application.
[0089] The memory can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory can also include a memory controller to provide access of the processor to the memory.
[0090] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal, a server or a similar computing device, that is, the above electronic device can include a computer terminal, a server or a similar computing device. Taking the case of running on a server as an example, Figure 6 is the hardware structure block diagram of the server running a computing power scheduling method provided in the embodiments of the present application, like Figure 6As shown, the server 600 can vary greatly in configuration and performance, and can include one or more Central Processing Units (CPU) 610 (which can include, but is not limited to, processing devices such as microprocessors, MCUs, or programmable logic devices, FPGAs, etc.), a memory 630 for storing data, one or more storage media 620 (such as one or more mass storage devices) for storing applications 623 or data 622. The memory 630 and the storage media 620 can be of the temporary or persistent storage variety. The programs stored in the storage media 620 can include one or more modules, each of which can include a series of instructions for operating on the server. Further, the CPU 610 can be configured to communicate with the storage media 620 to execute a series of instructions of the storage media 620 on the server 600. The server 600 can also include one or more power supplies 660, one or more wired or wireless network interfaces 650, one or more input / output interfaces 640, and / or one or more operating systems 621, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0091] The input / output interface 640 can be configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the server 600. In one example, the input / output interface 640 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the input / output interface 640 can be a radio frequency (RF) module configured to communicate with the Internet via a wireless manner.
[0092] Those of ordinary skill in the art can understand that, Figure 6 The structure shown is merely illustrative and does not limit the structure of the electronic device described above. For example, the server 600 can include more or fewer components than those shown, or have a different configuration than that shown. Figure 6 Figure 6 The structure shown is merely illustrative and does not limit the structure of the electronic device described above. For example, the server 600 can include more or fewer components than those shown, or have a different configuration than that shown.
[0093] The embodiments of the present application also provide a computer readable storage medium, which can be arranged in an electronic device to save at least one instruction or at least one program for implementing a computing power scheduling method. The at least one instruction or the at least one program is loaded and executed by the processor to implement any one of the computing power scheduling methods provided by the above method embodiments.
[0094] Embodiments of the present application also provide a computer program product or computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the electronic device perform any one of the computing power scheduling methods provided by the above method embodiments.
[0095] Optionally, in the embodiment, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0096] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
[0097] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0098] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.
[0099] The above is only the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A computing power scheduling method, characterized in that: The method comprises: Get pending jobs; Performing delay sensitivity identification on the pending job to obtain a delay sensitivity identification result of the pending job; If the delay sensitivity identification result indicates that the job to be processed is non-delay sensitive, determining a candidate computing power cluster that matches the resource requirements of the job to be processed from multiple computing power clusters; the multiple computing power clusters include a first type of computing power cluster based on green energy power and a second type of computing power cluster based on grid power; Determine the scheduling score of each candidate computing power cluster based on the current green energy data and current cluster load information of each candidate computing power cluster, and select the candidate computing power cluster with the highest scheduling score as the target computing power cluster; Distribute the pending job to the target computing power cluster so that the target computing power cluster executes the pending job.
2. The method according to claim 1, characterized in that The performing delay sensitivity identification on the pending job to obtain a delay sensitivity identification result of the pending job includes: Extracting metadata of the job to be processed, and determining the field content of the metadata corresponding to the delay sensitivity identification field; When the content of the field is not empty, determining the delay sensitivity identification result of the to-be-processed job based on the field content of the delay sensitivity identification field; When the content of the field is empty, the delay sensitivity identification result of the to-be-processed job is determined based on the job type and the estimated execution duration of the job in the metadata.
3. The method according to claim 1, characterized in that Determining the scheduling score of each candidate computing power cluster according to the current green energy data and current cluster load information of each candidate computing power cluster includes: For each candidate computing power cluster, a green energy availability score corresponding to the candidate computing power cluster is determined based on the current green energy data of the candidate computing power cluster; the current green energy data includes the current power generation and predicted power generation of the green energy corresponding to the candidate computing power cluster; the green energy availability score represents the green energy supply capacity corresponding to the candidate computing power cluster; Determine a cluster load score corresponding to the candidate computing power cluster based on the current cluster load information of the candidate computing power cluster; the cluster load score represents the idleness of the candidate computing power cluster; A weighted summation process is performed based on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster.
4. The method according to claim 3, characterized in that The scheduling score of the candidate computing power cluster obtained by performing weighted summation based on the green energy availability score and cluster load score corresponding to the candidate computing power cluster includes: Determining an operating cost score for the candidate computing power cluster based on a proportion of the current output power of the green energy corresponding to the candidate computing power cluster in the rated required power of the candidate computing power cluster; the operating cost score represents the electricity cost of running the job to be processed in the candidate computing power cluster; A weighted summation process is performed based on the green energy availability score, cluster load score, and operating cost score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster.
5. The method according to claim 4, characterized in that The scheduling score of the candidate computing power cluster obtained by performing weighted summation processing based on the green energy availability score, cluster load score, and operating cost score corresponding to the candidate computing power cluster includes: Determining a data transmission cost based on a storage location of the computational data required for the job to be processed; Determining a data affinity score corresponding to the candidate computing power cluster based on the data transmission cost; the data affinity score represents the distance between the computing data required for the job to be processed and the candidate computing power cluster; A scheduling score of the candidate computing power cluster is obtained by performing weighted summation based on the green energy availability score, cluster load score, operating cost score and data affinity score corresponding to the candidate computing power cluster.
6. The method according to claim 3, characterized in that The scheduling score of the candidate computing power cluster obtained by performing weighted summation based on the green energy availability score and cluster load score corresponding to the candidate computing power cluster includes: Obtaining a weight coefficient set; the weight coefficient set includes weight coefficients corresponding to each item in the weighted summation process; Based on the weight coefficient set, a weighted summation process is performed on the green energy availability score and the cluster load score corresponding to the candidate computing power cluster to obtain the scheduling score of the candidate computing power cluster.
7. The method according to claim 6, characterized in that The method further comprises: If the delay sensitivity identification result indicates that the to-be-processed job is delay-sensitive, increase a first weight coefficient in the weight coefficient set and decrease a second weight coefficient in the weight coefficient set to obtain an adjusted weight coefficient set; the first weight coefficient is a weight coefficient corresponding to the cluster load score item, and the second weight coefficient is a weight coefficient corresponding to the green energy availability score item; Based on the adjusted weight coefficient set, the green energy availability score and cluster load score corresponding to the candidate computing power cluster are weighted and summed to obtain the scheduling score of the candidate computing power cluster, and the candidate computing power cluster with the highest scheduling score is used as the target computing power cluster.
8. The method according to claim 6, characterized in that After distributing the pending job to the target computing power cluster, the method further includes: Detecting green energy data and cluster load information of each computing power cluster; Based on the detection results, the current green energy data and current cluster load information of each computing power cluster are updated.
9. A computing power scheduling device, characterized in that: The device comprises: Job acquisition module, used to obtain pending jobs; A delay sensitivity identification module, configured to perform delay sensitivity identification on the job to be processed and obtain a delay sensitivity identification result of the job to be processed; a candidate computing power cluster determination module, configured to determine, when the delay sensitivity identification result indicates that the pending job is non-delay-sensitive, a candidate computing power cluster that matches the resource requirements of the pending job from a plurality of computing power clusters; the plurality of computing power clusters including a first type of computing power cluster based on green energy power and a second type of computing power cluster based on grid power; A cluster scheduling score determination module is used to determine the scheduling score of each candidate computing power cluster based on the current green energy data and current cluster load information of each candidate computing power cluster, and select the candidate computing power cluster with the highest scheduling score as the target computing power cluster; The job distribution module is used to distribute the pending job to the target computing power cluster so that the target computing power cluster executes the pending job.
Citation Information
Patent Citations
Homemade heterogeneous computing power cluster-oriented job scheduling method and system
CN118519766A
Cross-cluster operation method and device and storage medium
CN119248537A
Cross-department government affair big data business co-processing method based on computing power platform and data fusion
CN120318018A