Resource scheduling method and electronic equipment

Through optical interconnection technology and time slice reallocation, the problems caused by the limitations of electrical signal transmission in traditional computing device resource pooling systems have been solved, flexible deployment of physical computing nodes and improved resource utilization have been achieved, and transmission performance and bit error rate have been optimized.

CN120743567AActive Publication Date: 2025-10-03LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511255569.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-03
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In traditional computing device resource pooling systems, limitations in electrical signal transmission lead to problems such as inflexible physical deployment, low resource scheduling efficiency, difficulty in integrating heterogeneous computing power, and limited transmission quality.

Method used

It uses optical interconnection to connect to the target resource pool, breaks through the distance limitations of traditional cable transmission through optical signal transmission, realizes flexible deployment of physical computing nodes, optimizes resource utilization through time slice reallocation, and supports cross-node resource pooling.

Benefits of technology

Significantly reduce bit error rates and power loss, improve resource utilization, optimize transmission performance, and ensure performance close to that of local physical computing nodes when scaling on a large scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743567A_ABST
    Figure CN120743567A_ABST
Patent Text Reader

Abstract

The invention discloses a resource scheduling method and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: obtaining a creation request sent by a target computing platform; screening a physical computing node from the target resource pool, and sending the creation request to the physical computing node, so that the physical computing node creates a virtual computing node; obtaining a to-be-processed task sent by the target computing platform; acquiring current state information of the physical computing nodes, and distributing a first time slice for each physical computing node; based on the required time slice of the task to be processed and the current state information, distributing a second time slice to the virtual computing node after the first time slice is distributed; and sending the to-be-processed task to the virtual computing node after the second time slice is allocated so as to call the virtual computing node after the second time slice is allocated to execute the to-be-processed task, thereby solving the problems of inflexible physical deployment, low resource scheduling efficiency, difficulty in heterogeneous computing power integration and limited transmission quality, improving the resource utilization efficiency, and improving the resource utilization rate. And the bit error rate and the transmission loss are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a resource scheduling method and electronic equipment. Background Art

[0002] In large-scale AI (Artificial Intelligence) model training and cloud computing scenarios, core issues arise, including low GPU (Graphics Processing Unit) resource utilization, inefficient scheduling, and the physical limitations of traditional electrical interconnect architectures. While existing GPU virtualization technology can partition resources on a single machine, the transmission distance limitations of PCIe (Peripheral Component Interconnect Express) cables make it difficult to build a flexible resource pool across physical devices. Furthermore, electrical signal transmission suffers from significant electromagnetic interference, high bit error rates, and severe power loss.

[0003] It can be seen that how to solve the problems of inflexible physical deployment, low resource scheduling efficiency, difficulty in integrating heterogeneous computing power and limited transmission quality caused by the limitations of electrical signal transmission in traditional computing device resource pooling systems, improve resource utilization efficiency, optimize transmission performance, and reduce bit error rate and transmission loss are issues that technical personnel in this field need to solve. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a resource scheduling method and electronic device that can address the problems of inflexible physical deployment, low resource scheduling efficiency, difficulty integrating heterogeneous computing power, and limited transmission quality caused by electrical signal transmission limitations in traditional computing device resource pooling systems. This method improves resource utilization efficiency, optimizes transmission performance, and reduces bit error rates and transmission losses. The specific solution is as follows: In a first aspect, the present application discloses a resource scheduling method, which is applied to a resource controller, wherein the resource controller is connected to a target resource pool via an optical interconnection, and the target resource pool is configured with a plurality of physical computing nodes; wherein the method comprises: Obtaining a creation request for creating a virtual computing node sent by a target computing platform; Filtering the corresponding physical computing node from the target resource pool according to the creation request, and sending the creation request to the physical computing node so that the physical computing node creates the corresponding virtual computing node based on the creation request; Forward the creation success response returned by the physical computing node to the target computing platform to obtain the pending tasks sent by the target computing platform; Acquire current status information of the physical computing nodes in the target resource pool, and allocate a corresponding first time slice to each physical computing node based on the current status information; Allocate a second time slice to the virtual computing node in each physical computing node after being allocated the first time slice based on the required time slice of the task to be processed and the current state information; The task to be processed is sent to the virtual computing node allocated the second time slice, so as to call the virtual computing node allocated the second time slice to execute the task to be processed.

[0005] In a second aspect, the present application discloses an electronic device, comprising: memory for storing computer programs; A processor is used to implement the steps of the aforementioned resource scheduling method when executing a computer program.

[0006] In a third aspect, the present application discloses a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of the aforementioned resource scheduling method when executed by a processor.

[0007] It can be seen that the present application provides a resource scheduling method, including obtaining a creation request for creating a virtual computing node sent by a target computing platform; filtering out the corresponding physical computing node from the target resource pool according to the creation request, and sending the creation request to the physical computing node, so that the physical computing node creates the corresponding virtual computing node based on the creation request; forwarding the creation success response returned by the physical computing node to the target computing platform to obtain the pending task sent by the target computing platform; obtaining the current status information of the physical computing nodes in the target resource pool, and allocating a corresponding first time slice to each physical computing node based on the current status information; allocating a second time slice to the virtual computing node in each physical computing node after the first time slice is allocated based on the required time slice and current status information of the pending task; sending the pending task to the virtual computing node after the second time slice is allocated, so as to call the virtual computing node after the second time slice is allocated to execute the pending task. The present application is applied to a resource controller, which is connected to a target resource pool via an optical interconnection method. The target resource pool is configured with several physical computing nodes. The resource controller obtains a creation request for creating a virtual computing node sent by a target computing platform, selects the corresponding physical computing node from the target resource pool according to the creation request, and sends the creation request to the physical computing node. The resource controller is connected to the target resource pool via an optical interconnection method, and utilizes the characteristics of long transmission distance and strong anti-electromagnetic interference of optical signals to break through the transmission distance limit of standard cables of traditional high-speed serial computer expansion buses, and realizes flexible deployment of physical computing nodes within a range of hundreds of meters. The physical computing node creates a corresponding virtual computing node based on the creation request, and forwards the creation success response returned by the physical computing node to the target computing platform to obtain the pending tasks sent by the target computing platform, thereby realizing physical decoupling, significantly reducing the bit error rate and power loss, and ensuring close proximity to this during large-scale expansion. The system calculates the performance of the physical computing nodes, obtains the current status information of the physical computing nodes in the target resource pool, allocates a corresponding first time slice to each physical computing node based on the current status information, and allocates a second time slice to the virtual computing node in each physical computing node after the first time slice is allocated based on the required time slice and current status information of the task to be processed. Through time slice allocation, the effect of on-demand allocation of resources is achieved, and resource utilization is maximized. It supports cross-node resource pooling and dynamic time division multiplexing, and sends the task to be processed to the virtual computing node after the second time slice is allocated, so as to call the virtual computing node after the second time slice is allocated to execute the task to be processed. It solves the problems of inflexible physical deployment, low resource scheduling efficiency, difficulty in integrating heterogeneous computing power and limited transmission quality caused by the limitation of electrical signal transmission in the resource pooling system of traditional computing equipment, optimizes transmission performance, and reduces bit error rate and transmission loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0009] Figure 1 This is a flow chart of a resource scheduling method disclosed in this application; Figure 2 This is a transmission flow chart of an optical interconnection method disclosed in this application; Figure 3 A structural diagram of a resource scheduling system disclosed in this application; Figure 4 A specific flow chart for implementing resource scheduling disclosed in this application; Figure 5 This application discloses a schematic diagram of the structure of a resource scheduling device disclosed in this application. DETAILED DESCRIPTION

[0010] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0011] In AI large-scale model training and cloud computing scenarios, core issues such as low GPU resource utilization, insufficient scheduling efficiency, and the physical limitations of traditional electrical interconnection architectures arise. While existing GPU virtualization technology can achieve single-machine resource segmentation, it is difficult to build a flexible resource pool across physical devices due to the transmission distance limitations of PCIe cables. Furthermore, electrical signal transmission suffers from defects such as high electromagnetic interference, high bit error rates, and severe power consumption. It can be seen that in traditional computing device resource pooling systems, how to solve the problems of inflexible physical deployment, low resource scheduling efficiency, difficulty integrating heterogeneous computing power, and limited transmission quality caused by electrical signal transmission limitations, and how to improve resource utilization efficiency, optimize transmission performance, and reduce bit error rates and transmission losses are issues that technical personnel in this field need to address.

[0012] See also Figure 1 As shown, an embodiment of the present invention discloses a resource scheduling method, which is applied to a resource controller, wherein the resource controller is connected to a target resource pool via an optical interconnection, and the target resource pool is configured with a plurality of physical computing nodes; wherein the method may specifically include: Step S11: Obtain a creation request for creating a virtual computing node sent by the target computing platform.

[0013] In this embodiment, a first connection relationship is established between the resource controller and the target computing platform using a cable; and a creation request for creating a virtual computing node sent by a user is obtained from the target computing platform based on the first connection relationship.

[0014] In this application, the target computing platform is connected to the GPU Controller (resource controller) through a cable, and the physical computing nodes are connected to the GPU Controller through optical interconnection (optical fiber and optical module). The target computing platform receives and sends computing tasks to the physical computing nodes in a software-defined manner. The GPU Controller serves as a central device, assuming the function of a bridge for high-speed signal transmission between the general computing platform and the physical computing nodes. It is also used to monitor the target resource pool and schedule GPU resources, forming a unified virtualized resource pool across physical devices.

[0015] Step S12: Filter out corresponding physical computing nodes from the target resource pool according to the creation request, and send the creation request to the physical computing node, so that the physical computing node creates a corresponding virtual computing node based on the creation request.

[0016] In this embodiment, optical fibers and optical modules are used to establish a second connection relationship between the resource controller and the physical computing nodes in the target resource pool through optical interconnection. The corresponding physical computing nodes are screened out from the target resource pool based on the creation request. The electrical signal corresponding to the creation request is split, and the split electrical signal is converted into an optical signal using an optical module. The optical signal is transmitted to the physical computing node using the optical fiber and the second connection relationship, so that the physical computing node can restore the optical signal to obtain the electrical signal corresponding to the creation request before the conversion, so that the physical computing node can create the corresponding virtual computing node based on the creation request.

[0017] In this embodiment, when receiving a creation request for a virtual compute node from the target computing platform, the target computing platform sends a vGPU (virtual compute node) creation request to the GPU Controller. After receiving the request, the GPU Controller selects the best matching physical compute node that meets the parameters based on the scheduling module, creates the virtual compute node in the attached target resource pool, and sends a creation success response to the target computing platform.

[0018] The optical interconnection transmission process in this application is as follows Figure 2As shown, the PCIe signal in the link is converted from an end-to-end electrical signal to an optical signal. This application uses a PCIe Switch that supports splitting PCIe x16 signals into 2x8 + a QSFP-DD (Quad Small Form-factor Pluggable Double Density) optical module for photoelectric conversion to form an optical interconnection link. In actual applications, the PCIe x16 signal is split into two x8 signals through the PCIe Switch and input into the QSFP-DD optical module. The optical module converts the level information of the PCIe electrical signal into the power information of the optical signal and realizes optical signal transmission through optical fiber. Before entering the target resource pool, the optical module at the target resource pool end receives the electrical signal, and restores the optical signal back to the electrical signal and transmits it to the physical computing node.

[0019] In addition, the resource controller may also establish a third connection relationship with the target resource pool by using the optical switch, so as to access resources of the target resource pool based on the third connection relationship, the optical switch, and the optical channel.

[0020] When multiple general-purpose computing hosts are directly connected to an optical switch, each host can access any GPU resource in the target resource pool in parallel through the optical channel. Thanks to the bandwidth advantage of more than 100Gbps and nanosecond latency characteristics of a single optical fiber transmission channel, the pooled system can still maintain computing performance close to that of a local GPU even in large-scale expansion scenarios.

[0021] Step S13: forward the creation success response returned by the physical computing node to the target computing platform to obtain the to-be-processed task sent by the target computing platform.

[0022] In this embodiment, when the physical computing node successfully creates the corresponding virtual computing node, the second connection relationship is used to obtain the creation success response returned by the physical computing node; based on the first connection relationship between the resource controller and the target computing platform, the creation success response is forwarded to the target computing platform to obtain the pending tasks sent by the target computing platform.

[0023] This application breaks through the physical boundary limitations of the electrical interconnection architecture and realizes ultra-long-distance deployment of GPU nodes; improves resource utilization efficiency, supports cross-node resource pooling, and supports dynamic time division multiplexing; optimizes transmission performance, can significantly reduce bit error rate, reduce transmission loss, and ensure computing performance close to that of local GPUs during large-scale expansion.

[0024] Step S14: obtaining current status information of the physical computing nodes in the target resource pool, and allocating a corresponding first time slice to each physical computing node based on the current status information.

[0025] In this embodiment, the target resource pool is monitored and managed in real time; wherein, the real-time monitoring includes physical computing node utilization monitoring and network load monitoring; the target time parameters and load status of the physical computing nodes in the target resource pool monitored in real time are obtained, and the corresponding first time slice is allocated to each physical computing node based on the current status information.

[0026] This application allocates a corresponding first time slice to each physical compute node. Physical compute node time slicing is a virtualization technology that allows multiple workloads or VMs (Virtual Machines) to share a single GPU by dividing processing time into discrete slices. Time slices are dynamically allocated based on the usage of each GPU's time slices, allocating portions of the GPU's compute and memory resources to different tasks or users. A time slice reallocation module is used to achieve on-demand resource allocation. This enables multiple tasks to be executed concurrently on a single GPU, maximizing resource utilization.

[0027] Step S15: allocating a second time slice to the virtual computing nodes in each physical computing node after being allocated the first time slice based on the required time slice of the task to be processed and the current state information.

[0028] In this embodiment, a load balancing algorithm is used to allocate a second time slice to the virtual computing nodes in each physical computing node after being allocated the first time slice based on the required time slice of the task to be processed, the target time parameter, and the load status.

[0029] Specifically, determine a target ratio; the target ratio is the ratio between the required time slice of the task to be processed and the allocated time slice of any virtual computing node in the target time parameter; determine whether the target ratio is greater than a preset threshold; if the target ratio is greater than the preset threshold, calculate the second time slice of any virtual computing node based on the required time slice and the allocated time slice; the second time slice is the remaining time slice of any virtual computing node; use a load balancing algorithm and based on the load status, allocate the second time slice to the other virtual computing nodes in each physical computing node after the first time slice is allocated, except for any virtual computing node.

[0030] In this embodiment, as the GPU usage varies, the GPU load will also change. In response to the load changes, the time slice allocation mechanism ensures that the time slice resources obtained by the vGPU roughly match its load, achieving load balancing between vGPUs. The evaluation criteria for dynamically allocating time slices to the vGPU is to first calculate the ratio between the required time slice of the task to be processed and the allocated time slice of any virtual computing node in the target time parameter, and then determine whether the ratio is greater than a preset threshold. For example, the preset threshold is set to 9 / 10. That is, if the proportion of time slices used by a vGPU is less than 90% of the allocated time slices, the remaining time slices of the vGPU will be allocated to other virtual computing nodes except for any virtual computing node, thereby maximizing the utilization of GPU resources.

[0031] Step S16: sending the task to be processed to the virtual computing node allocated the second time slice, so as to call the virtual computing node allocated the second time slice to execute the task to be processed.

[0032] In this embodiment, after calling the virtual computing node to execute the pending task after the second time slice is allocated, the task execution success response sent by the target resource pool is obtained, and the task execution success response is sent to the target computing platform; when the resource release task sent by the target computing platform is obtained, the resource release task is forwarded to the physical computing node to be released, so that the physical computing node to be released releases resources for the corresponding virtual computing node based on the resource release task.

[0033] In this embodiment, after receiving the creation success response, the target computing platform calls the virtual computing node resources to execute the pending tasks. After processing the pending tasks, the target computing platform returns the execution results to the target resource pool and releases the resources of the created virtual computing node.

[0034] This application proposes a resource scheduling system, the structure of which is as follows: Figure 3As shown in the figure, it consists of a target computing platform, a GPU Controller (resource controller), and a target resource pool. The target computing platform is primarily responsible for initiating requests to create, call, and release GPU resources based on computing and application needs. The GPU Controller primarily manages all GPU resources in the resource pool through the GPU monitoring and GPU scheduling modules. The GPU monitoring module is responsible for monitoring GPU utilization, network load, and status. The GPU scheduling module is configured with a load balancing algorithm to optimally schedule GPU resources. The target resource pool integrates hardware drivers and runtime libraries from major vendors. It uses different processes to execute and respond to vGPU resource call API (Application Program Interface) requests from GPU clients, fulfilling their resource call services and reclaiming and releasing resources after use.

[0035] The specific process of implementing resource scheduling in this application is as follows Figure 4 As shown, the core is to use optical interconnect technology to replace traditional electrical signal transmission and optimize GPU resource virtualization through time-slice reallocation. Optical fiber connects the GPUController and GPU nodes (physical computing nodes). Leveraging the long transmission distance and strong electromagnetic interference resistance of optical signals, this overcomes the distance limitations of traditional PCIe cables and enables flexible deployment of GPU nodes over distances of hundreds of meters. This architecture physically decouples GPU resources from the CPU host and, through the time-slice reallocation module, achieves on-demand resource allocation, maximizing resource utilization.

[0036] In addition, the optical interconnection technology in this application can achieve computing power collaboration across data centers and can be applied to hybrid cloud scenarios, intelligent computing centers, etc. GPU virtual pooling technology solves the problem of fragmented resource utilization and can be applied to large model inference tasks. Its low latency characteristics can be applied to fields such as autonomous driving.

[0037] In addition, if the GPU and optical module run at full power for a long time, it will lead to energy waste. Therefore, this application can also form a dynamic energy efficiency adjustment mechanism through GPU dynamic frequency reduction and optical module low power mode, thereby avoiding energy waste. For example, GPU dynamic frequency reduction: when the GPU utilization rate is lower than 20%, the core frequency is automatically reduced, such as from 1.8GHz to 1.2GHz, and the power consumption is reduced by 40%; optical module low power mode: when the optical link idle time exceeds 5 minutes, the optical module automatically enters "sleep mode", and the power consumption is reduced from 15W to 3W. When there is data transmission, it wakes up within 100 microseconds, which does not affect the low latency requirement. In addition, a visual panel can also be provided to provide users with a vGPU resource monitoring page, which displays the utilization rate, video memory usage, and task progress of the allocated vGPU in real time, and supports fault alarm push, for example, push in the form of SMS, email, enterprise WeChat, etc., so that users can grasp the resource status in real time.

[0038] In this embodiment, a creation request for creating a virtual computing node sent by a target computing platform is obtained; the corresponding physical computing node is filtered out from the target resource pool according to the creation request, and the creation request is sent to the physical computing node, so that the physical computing node creates the corresponding virtual computing node based on the creation request; the creation success response returned by the physical computing node is forwarded to the target computing platform to obtain the pending task sent by the target computing platform; the current status information of the physical computing nodes in the target resource pool is obtained, and a corresponding first time slice is allocated to each physical computing node based on the current status information; based on the required time slice and current status information of the pending task, a second time slice is allocated to the virtual computing node in each physical computing node after the first time slice is allocated; the pending task is sent to the virtual computing node after the second time slice is allocated, so as to call the virtual computing node after the second time slice is allocated to execute the pending task. The present application is applied to a resource controller, which is connected to a target resource pool via an optical interconnection method. The target resource pool is configured with several physical computing nodes. The resource controller obtains a creation request for creating a virtual computing node sent by a target computing platform, selects the corresponding physical computing node from the target resource pool according to the creation request, and sends the creation request to the physical computing node. The resource controller is connected to the target resource pool via an optical interconnection method, and utilizes the characteristics of long transmission distance and strong anti-electromagnetic interference of optical signals to break through the transmission distance limit of standard cables of traditional high-speed serial computer expansion buses, and realizes flexible deployment of physical computing nodes within a range of hundreds of meters. The physical computing node creates a corresponding virtual computing node based on the creation request, and forwards the creation success response returned by the physical computing node to the target computing platform to obtain the pending tasks sent by the target computing platform, thereby realizing physical decoupling, significantly reducing the bit error rate and power loss, and ensuring close proximity to this during large-scale expansion. The system calculates the performance of the physical computing nodes, obtains the current status information of the physical computing nodes in the target resource pool, allocates a corresponding first time slice to each physical computing node based on the current status information, and allocates a second time slice to the virtual computing node in each physical computing node after the first time slice is allocated based on the required time slice and current status information of the task to be processed. Through time slice allocation, the effect of on-demand allocation of resources is achieved, and resource utilization is maximized. It supports cross-node resource pooling and dynamic time division multiplexing, and sends the task to be processed to the virtual computing node after the second time slice is allocated, so as to call the virtual computing node after the second time slice is allocated to execute the task to be processed. It solves the problems of inflexible physical deployment, low resource scheduling efficiency, difficulty in integrating heterogeneous computing power and limited transmission quality caused by the limitation of electrical signal transmission in the resource pooling system of traditional computing equipment, optimizes transmission performance, and reduces bit error rate and transmission loss.

[0039] See also Figure 5As shown, an embodiment of the present invention discloses a resource scheduling device, which is applied to a resource controller. The resource controller is connected to a target resource pool via an optical interconnection. The target resource pool is configured with a plurality of physical computing nodes. The device may specifically include: A creation request acquisition module 11 is used to acquire a creation request for creating a virtual computing node sent by a target computing platform; A creation request forwarding module 12 is configured to select a corresponding physical computing node from a target resource pool according to the creation request and send the creation request to the physical computing node so that the physical computing node creates a corresponding virtual computing node based on the creation request; The pending task acquisition module 13 is used to forward the creation success response returned by the physical computing node to the target computing platform to obtain the pending task sent by the target computing platform; A first time slice allocation module 14 is configured to obtain current status information of the physical computing nodes in the target resource pool and allocate a corresponding first time slice to each physical computing node based on the current status information; A second time slice allocation module 15 is configured to allocate a second time slice to the virtual computing nodes in each physical computing node after being allocated the first time slice based on the required time slice of the task to be processed and the current state information; The pending task sending module 16 is configured to send the pending task to the virtual computing node allocated the second time slice, so as to call the virtual computing node allocated the second time slice to execute the pending task.

[0040] In some specific embodiments, creating a request acquisition module 11 may specifically include: A first connection relationship establishing module, configured to establish a first connection relationship between the resource controller and the target computing platform using a cable; The creation request specific acquisition module is used to obtain a creation request for creating a virtual computing node sent by a user from the target computing platform based on the first connection relationship.

[0041] In some specific embodiments, creating the request forwarding module 12 may specifically include: A second connection relationship establishing module, configured to establish a second connection relationship between the resource controller and the physical computing nodes in the target resource pool by using optical fibers and optical modules and by optical interconnection; A physical computing node screening module is used to screen out corresponding physical computing nodes from a target resource pool based on a creation request; An electrical signal splitting and conversion module, configured to split the electrical signal corresponding to the creation request and convert the split electrical signal into an optical signal using an optical module; The restoration module is used to transmit the optical signal to the physical computing node by using the optical fiber and the second connection relationship, so that the physical computing node restores the optical signal to obtain the electrical signal corresponding to the creation request before conversion.

[0042] In some specific embodiments, the pending task acquisition module 13 may specifically include: A creation success response acquisition module is used to obtain a creation success response returned by the physical computing node using the second connection relationship when the physical computing node successfully creates the corresponding virtual computing node; The creation success response forwarding module is configured to forward the creation success response to the target computing platform based on the first connection relationship between the resource controller and the target computing platform.

[0043] In some specific embodiments, the first time slice allocation module 14 may specifically include: Real-time monitoring and management module, used to monitor and manage the target resource pool in real time; real-time monitoring includes physical computing node utilization monitoring and network load monitoring; The target time parameter and load status acquisition module is used to obtain the target time parameters and load status of the physical computing nodes in the target resource pool monitored in real time.

[0044] In some specific embodiments, the second time slice allocation module 15 may specifically include: The second time slice specific allocation module is used to use the load balancing algorithm and allocate the second time slice to the virtual computing nodes in each physical computing node after the first time slice is allocated based on the required time slice, target time parameter and load status of the task to be processed.

[0045] In some specific embodiments, the second time slice allocation module 15 may specifically include: A target ratio determination module is used to determine a target ratio; the target ratio is the ratio between the time slice required for the task to be processed and the allocated time slice of any virtual computing node in the target time parameter; A judgment module, used to judge whether the target ratio is greater than a preset threshold; A second time slice calculation module is configured to calculate a second time slice of any virtual computing node based on the required time slice and the allocated time slice if the target ratio is greater than a preset threshold; the second time slice is a remaining time slice of any virtual computing node; The module for allocating a second time slice to other virtual computing nodes is used to allocate a second time slice to other virtual computing nodes except any virtual computing node in each physical computing node after being allocated the first time slice, using a load balancing algorithm and based on load status.

[0046] In some specific embodiments, the resource scheduling device may further include: A task execution success response sending module is used to obtain a task execution success response sent by the target resource pool after calling the virtual computing node allocated the second time slice to execute the pending task, and send the task execution success response to the target computing platform; The resource release module is used to forward the resource release task to the physical computing node to be released when it obtains the resource release task sent by the target computing platform, so that the physical computing node to be released can release resources of the corresponding virtual computing node based on the resource release task.

[0047] In some specific embodiments, the resource controller establishes a third connection relationship with the target resource pool using the optical switch, so as to access resources of the target resource pool based on the third connection relationship, the optical switch, and the optical channel.

[0048] Among them, the description of the features in the embodiment corresponding to the resource scheduling device can refer to the relevant description of the embodiment corresponding to the resource scheduling method, and will not be repeated here.

[0049] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above resource scheduling method embodiments.

[0050] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned resource scheduling method embodiments when running.

[0051] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0052] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0053] The above is a detailed introduction to a resource scheduling method and electronic device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only applicable to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A resource scheduling method, characterized in that: Applied to a resource controller, the resource controller is connected to a target resource pool via an optical interconnection, and the target resource pool is configured with a plurality of physical computing nodes; wherein the method comprises: Obtaining a creation request for creating a virtual computing node sent by a target computing platform; Filtering a corresponding physical computing node from the target resource pool according to the creation request, and sending the creation request to the physical computing node, so that the physical computing node creates a corresponding virtual computing node based on the creation request; Forwarding the creation success response returned by the physical computing node to the target computing platform to obtain the pending task sent by the target computing platform; Acquire current status information of the physical computing nodes in the target resource pool, and allocate a corresponding first time slice to each of the physical computing nodes based on the current status information; Allocate a second time slice to the virtual computing node in each of the physical computing nodes after being allocated the first time slice based on the time slice required for the task to be processed and the current state information; The to-be-processed task is sent to the virtual computing node allocated with the second time slice, so as to call the virtual computing node allocated with the second time slice to execute the to-be-processed task.

2. The resource scheduling method according to claim 1, characterized in that: The step of obtaining a creation request for creating a virtual computing node sent by the target computing platform includes: establishing a first connection relationship between the resource controller and the target computing platform using a cable; A creation request for creating a virtual computing node sent by a user is obtained from the target computing platform based on the first connection relationship.

3. The resource scheduling method according to claim 1, characterized in that: The step of selecting a corresponding physical computing node from the target resource pool according to the creation request and sending the creation request to the physical computing node includes: Using optical fibers and optical modules, and through optical interconnection, a second connection relationship is established between the resource controller and the physical computing nodes in the target resource pool; Filtering corresponding physical computing nodes from the target resource pool based on the creation request; Splitting the electrical signal corresponding to the creation request, and converting the split electrical signal into an optical signal using the optical module; The optical signal is transmitted to the physical computing node by using the optical fiber and the second connection relationship, so that the physical computing node restores the optical signal to obtain an electrical signal corresponding to the creation request before conversion.

4. The resource scheduling method according to claim 3, characterized in that: The forwarding the creation success response returned by the physical computing node to the target computing platform includes: When the physical computing node successfully creates the corresponding virtual computing node, using the second connection relationship to obtain a creation success response returned by the physical computing node; Based on the first connection relationship between the resource controller and the target computing platform, the creation success response is forwarded to the target computing platform.

5. The resource scheduling method according to claim 1, characterized in that: The acquiring the current status information of the physical computing node in the target resource pool includes: Real-time monitoring and management of target resource pools; real-time monitoring includes physical computing node utilization monitoring and network load monitoring; The target time parameters and load status of the physical computing nodes in the target resource pool monitored in real time are obtained.

6. The resource scheduling method according to claim 5, characterized in that: The allocating a second time slice to the virtual computing node in each of the physical computing nodes after being allocated the first time slice based on the required time slice of the task to be processed and the current state information includes: A load balancing algorithm is used, and based on the required time slice of the task to be processed, the target time parameter, and the load status, a second time slice is allocated to the virtual computing node in each of the physical computing nodes after being allocated the first time slice.

7. The resource scheduling method according to claim 6, characterized in that: The utilizing a load balancing algorithm and allocating a second time slice to the virtual computing node in each of the physical computing nodes after being allocated the first time slice based on the required time slice of the task to be processed, the target time parameter, and the load status includes: Determine a target ratio; the target ratio is the ratio between the time slice required for the task to be processed and the allocated time slice of any virtual computing node in the target time parameter; Determining whether the target ratio is greater than a preset threshold; If the target ratio is greater than a preset threshold, a second time slice of the any virtual computing node is calculated based on the required time slice and the allocated time slice; the second time slice is a remaining time slice of the any virtual computing node; A load balancing algorithm is used and based on the load status, a second time slice is allocated to the other virtual computing nodes in the physical computing nodes except for the any one virtual computing node after the first time slice is allocated.

8. The resource scheduling method according to any one of claims 1 to 7, characterized in that: Also includes: After calling the virtual computing node allocated the second time slice to execute the pending task, obtaining a task execution success response sent by the target resource pool, and sending the task execution success response to the target computing platform; When a resource release task sent by the target computing platform is obtained, the resource release task is forwarded to the physical computing node to be released, so that the physical computing node to be released releases resources of the corresponding virtual computing node based on the resource release task.

9. The resource scheduling method according to claim 1, wherein: The resource controller establishes a third connection relationship with the target resource pool by using the optical switch, so as to access resources of the target resource pool based on the third connection relationship, the optical switch and the optical channel.

10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the resource scheduling method according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Cloud computing virtual machine allocation and adjustment method based on multiple objectives

    CN106506657A

  • Method and device for realizing virtual GPU and system

    CN108984264A

  • Heterogeneous computing-oriented resource scheduling method, node, system, equipment and medium

    CN114996018A

  • GPU resource scheduling method and device, electronic equipment and storage medium

    CN115564635A

  • Resource management method, device and system and storage medium

    CN115686839A