A computer network system, job scheduling system and method
By shorting the computer with short-circuited jumper in the computer network system, flexible scheduling between computers is achieved, and the problems of low resource utilization efficiency and long task delays in the prior art are solved when computing resources are tight, and resource utilization efficiency and calculation tasks are improved.
Patent Information
- Application Number
- CN202411889884.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In the field of high-performance computing, the prior art cannot effectively manage resource allocation when computing resources are tight, resulting in low resource utilization efficiency and long task execution delays, especially in the ring network topology.
By carrying short-circuit jumpers in computer network systems, flexible scheduling between computers is achieved. When at least one computer equipped with a short-circuited jumper, the adjacent computer without a short-circuited jumper is directly connected through the short-circuited jumper to form a ring network topology and improve resource utilization efficiency.
It realizes flexible scheduling of computing resources, improves resource utilization efficiency, reduces task execution delays, avoids short-term effects, and is suitable for high-concurrency and large-scale computing tasks.
Smart Images

Figure CN119336518B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a computer network system, a job scheduling system and a method. Background Art
[0002] In the field of high-performance computing, computing jobs are mainly compute-intensive and require efficient computing resource support. The scheduling of computing jobs usually depends on the platform's job scheduling system. Through reasonable resource allocation, jobs are assigned to multiple computing nodes in the computing cluster for execution. However, each job has different requirements for computing resources, including the number of CPU cores. When each computing job is assigned the required resources and starts to execute, it cannot be interrupted or migrated. If an interruption occurs, it needs to be re-executed. The main function of the job scheduling system is to allocate reasonable computing resources to jobs in the computing platform according to the job's demand for computing resources, and start the job. However, when the computing platform is heavily loaded, the job scheduling system may not be able to provide sufficient computing resources for all jobs in time. At this time, jobs that have not been allocated resources will be placed in a waiting queue, waiting for the release of computing resources. Although the existing job scheduling mechanism can effectively manage the allocation of computing resources in most cases, in some scenarios, especially when resources are highly scarce, there are still problems such as low resource utilization efficiency and long task execution delays.
[0003] The existing technology currently has no short-circuit jumpers and short-circuit controllers, and the commonly used solution is the traditional bridge network card-based solution, which belongs to the L2 layer data link layer connectivity. Although it can achieve interconnection between different computing nodes, it has inherent delays and bandwidth bottlenecks due to its working principle, especially in the ring network topology. In delay-sensitive application scenarios, the traditional bridge network card connection method cannot meet the needs of efficient data transmission, limiting the performance of the overall computing platform, especially in the execution of high-concurrency, large-scale computing tasks. Summary of the invention
[0004] In response to the deficiencies in the prior art, the present application provides a computer network system, a scheduling system and a scheduling method. In the computer network system, about half of the computers are equipped with short-circuit jumpers. Computers equipped with short-circuit jumpers can all undertake single-node computing tasks, making the entire computer network flexible to schedule. After a single node is occupied, the remaining nodes are not blocked, thereby improving hardware utilization efficiency.
[0005] To achieve the above-mentioned purpose, the present invention provides a computer network system, comprising a plurality of computers without short-circuit jumpers, wherein a computer equipped with a short-circuit jumper is arranged between two adjacent computers without short-circuit jumpers, the computers without short-circuit jumpers and the computers equipped with short-circuit jumpers are both equipped with two central processing units and each central processing unit is connected to a high-speed network card, and the high-speed network cards between the computers without short-circuit jumpers and the computers equipped with short-circuit jumpers form a ring network topology through a wired connection; when the computers equipped with short-circuit jumpers are not short-circuited, the computers that are not short-circuited are connected to the two adjacent computers without short-circuit jumpers through the high-speed network card; when at least one computer equipped with a short-circuit jumper is short-circuited, the two computers without short-circuit jumpers adjacent to the short-circuited computer are directly connected at the physical layer through the short-circuit jumper.
[0006] Furthermore, the short-circuit jumper components include a short-circuit jumper and a short-circuit controller. The short-circuit jumper is a wire connected between two high-speed network cards of a computer and is used to control the connection status of the circuit. The short-circuit controller is a logic circuit used to detect the short-circuit jumper status and related instructions. The related instructions include identifying and changing whether the short-circuit jumper is short-circuited.
[0007] Furthermore, when at least one computer equipped with a short-circuited jumper is short-circuited, the remaining computers without the short-circuited jumper and the computers equipped with the short-circuited jumper that are not short-circuited still form a ring network topology through wired connections.
[0008] The present application also provides a job scheduling system containing a computer network system, including a job scheduling manager, which is used to receive jobs and assign computers to received jobs. The computers are computers without short-circuiting jumpers and computers equipped with short-circuiting jumpers in the computer network system, which are used to calculate the assigned jobs.
[0009] Furthermore, the job scheduling manager includes a job queue, a job scheduler and a node manager;
[0010] The job queue is used to receive job requests and sort multiple job requests according to the sorting principle set by the job scheduling strategy. After obtaining the job sorting status, it initiates a job computing request to the job scheduler. The job computing request includes the required number of computing nodes.
[0011] The job scheduler is used to initiate a request for the number of computing nodes to the node manager, and after the available computing nodes returned by the node manager, control the corresponding number of computers equipped with short-circuit jumpers to be short-circuited or not short-circuited, and allocate the corresponding number of computing nodes according to the job computing request;
[0012] The node manager is used to manage all computers in the computer system. Based on the computing node quantity request sent by the job scheduler, the node manager returns available computing nodes to the job scheduler. The available computing nodes include computers without short-circuiting jumpers and computers equipped with short-circuiting jumpers but not short-circuited.
[0013] The present application also provides a scheduling method, comprising the following steps:
[0014] Step 1: Receive the job submission instruction sent by the client, which includes the computing requirements of the submitted job, including the number of computing nodes;
[0015] Step 2: The job scheduling manager puts the received jobs into the job queue, and the job queue is sorted according to the order in which they enter the queue;
[0016] Step 3: The job scheduling manager determines the number of available computing nodes M, determines the number of computing nodes m required for the first-order job in the job queue, allocates the corresponding number of nodes to the job and starts execution;
[0017] Step 4: Get the required computing quantity of the job in the first order in the update task queue, and search for the corresponding number of target nodes in the remaining (Mm) computing nodes; if there is a match, assign the corresponding target node to the job and start execution;
[0018] Step 5: Until the remaining computing nodes are allocated according to step 4 or the remaining computing nodes cannot meet the computing requirements of the jobs in the job queue, wait for the running jobs to be executed and terminated, and then release the computing resources of the corresponding executing computing nodes, and reallocate the computing nodes to the jobs to be scheduled.
[0019] Furthermore, in steps 3 and 4, the jobs are classified according to the required number of computing nodes to form corresponding node computing tasks. When the job belongs to a single-node computing task, the job scheduler controls at least one computer equipped with a short-circuit jumper in the available computing nodes to short-circuit, and assigns the single-node computing task to the short-circuited computer.
[0020] The present invention adopts the above technical solution to achieve the following technical effects: 1. The high-performance computing network in the present application includes multiple computers without short-circuit jumpers and computers equipped with short-circuit jumpers. The opening and closing of the short-circuit jumpers can realize whether two computers without short-circuit jumpers adjacent to the short-circuited computer are directly connected at the physical layer through the short-circuit jumpers. This solves the problem that only two network cards can be bridged in the prior art, and the bridged network card belongs to the L2 layer connection, which is about one order of magnitude higher than the delay of the present application; it also solves the bottleneck of low efficiency in the ring network for delay-sensitive applications such as high-performance computing and direct memory access.
[0021] The present application installs a computer equipped with a short-circuit jumper between two computers without a short-circuit jumper, so that about half of the computers equipped with the short-circuit jumper can undertake single-node computing tasks after short-circuiting, making the entire computer network scheduling more flexible.
[0022] After the computer equipped with the short-circuiting jumper is short-circuited, the line conditions of the remaining computers are consistent and the delays are similar, which can avoid the short-board effect.
[0023] The present application also provides a scheduling method, which allocates single-node computing tasks to a computer equipped with a short-circuited jumper and short-circuited, and allows the remaining nodes to still form a ring, which can undertake one or more other computing tasks, thereby achieving flexible scheduling of a ring topology computer cluster and achieving the technical effect of non-blocking remaining nodes after a single node is occupied. The method is suitable for high-performance computing tasks such as fluid simulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of a computer network system;
[0025] Figure 2 It is a workflow diagram of a job scheduling system including a computer network system;
[0026] Figure 3 Schematic diagram of the scheduling method.
[0027] Figure numerals: 1. Computer without short-circuit jumper; 2. Computer equipped with short-circuit jumper; 3. High-speed network card; 4. Short-circuit jumper; 5. Short-circuit controller. DETAILED DESCRIPTION
[0028] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0029] Refer to Figure 1 As shown, the present application provides a computer network system, including a high-performance computing network, the high-performance computing network includes a plurality of computers without short-circuit jumpers, and computers equipped with short-circuit jumpers are arranged between adjacent computers without short-circuit jumpers. (The present application uses 4 computers as Figure 1As shown in the figure, the computers are numbered N1, N2, N3 and N4 from top to bottom. The computers without short-circuit jumpers and the computers equipped with short-circuit jumpers are both equipped with two central processing units and each central processing unit is connected to a high-speed network card, and each central processing unit is connected to a network card, so that the CPU load is relatively balanced. The high-speed network cards between the computers without short-circuit jumpers and the computers equipped with short-circuit jumpers form a ring network topology through wired connections, and the wired connections include high-speed copper cables or optical fiber connections. The short-circuit jumper interval enables about half of the computers in the computer network system to undertake single-node computing tasks, making the entire computer network system flexible to schedule; when all short-circuit jumpers are short-circuited, the line conditions between the remaining computers without short-circuit jumpers are consistent, and the delays are similar, which can avoid the short board effect. Interval installation can effectively reduce direct interference between adjacent nodes, reduce signal reflection and crosstalk, optimize signal quality, and improve network reliability; at the same time, interval installation can make full use of the physical space in actual deployment, facilitate maintenance and management, and ensure that the overall network performance is not affected in complex computing network systems. The short-circuit jumper includes a short-circuit jumper and a short-circuit controller. The short-circuit jumper is a wire installed between two high-speed network cards of the computer and is used to control the connection state of the circuit. The short-circuit controller is a logic circuit used to detect the state of the short-circuit jumper and related instructions. The related instructions include identifying and changing whether the short-circuit jumper is short-circuited. When all computers equipped with the short-circuit jumper are not short-circuited, the computers that are not short-circuited are connected to the two adjacent computers without the short-circuit jumper through the high-speed network card, and all computers together form a Ring network topology (i.e., N1-N2-N3-N4-N1); when at least one computer equipped with a shorting jumper is short-circuited (N2 is set to be short-circuited), the two computers without shorting jumpers adjacent to the short-circuited computer are directly connected at the physical layer through the shorting jumper, so that the remaining computers without shorting jumpers and the computers equipped with shorting jumpers and not short-circuited form a ring network topology (i.e., N1-N3-N4-N1), and the short-circuited computer (N2) independently undertakes single-node computing tasks.
[0030] like Figure 2 As shown, the present application also provides a job scheduling system containing a computer network system, including a job scheduling manager, the job scheduling manager is used to receive jobs and allocate computing nodes without short-circuit jumpers and / or computing nodes equipped with short-circuit jumpers (computing nodes are equivalent to computers) to the received jobs, and the computers without short-circuit jumpers and computers equipped with short-circuit jumpers in the computing system are used to calculate the assigned jobs. The specific job scheduling manager includes a job queue, a job scheduler, and a node manager. The job queue is used to receive job requests and sort multiple job requests according to the sorting principle set by the job scheduling strategy, using a first-come, first-served scheduling algorithm.
[0031] The job scheduler is used to initiate a request for the number of computing nodes to the node manager, and after the available computing nodes returned by the node manager, control the corresponding number of computers equipped with short-circuit jumpers to be short-circuited or not, and allocate the corresponding number of computing nodes according to the job computing request.
[0032] The node manager is used to manage all computers in the computer system. Based on the computing node quantity request sent by the job scheduler, the node manager returns available computing nodes to the job scheduler. The available computing nodes include computers without short-circuiting jumpers and computers equipped with short-circuiting jumpers but not short-circuited.
[0033] like Figure 3 As shown, the present application also provides a job scheduling method, comprising the following steps:
[0034] Step 1: Receive the job submission instruction sent by the client, which includes the computing requirements of the submitted job, including the number of computing nodes;
[0035] Step 2: The job scheduling manager puts the received jobs into the job queue, and the job queue is sorted according to the order in which they enter the queue;
[0036] Step 3: The job scheduling manager determines the number of available computing nodes M, determines the number of computing nodes m required for the first-order job in the job queue, allocates the corresponding number of nodes to the job and starts execution;
[0037] Step 4: Get the required computing quantity of the job in the first order in the update task queue, and search for the corresponding number of target nodes in the remaining (Mm) computing nodes; if there is a match, assign the corresponding target node to the job and start execution;
[0038] Step 5: Until more than the computing nodes are allocated according to step 4 or the remaining computing nodes cannot meet the computing requirements of the jobs in the job queue, wait for the running jobs to be completed and terminated, and then release the computing resources of the corresponding executing computing nodes, and reallocate the computing nodes to the jobs to be scheduled.
[0039] Specific jobs are classified into corresponding node computing tasks according to the required number of computing nodes. When the job belongs to a single-node computing task, the job scheduler controls one of the available computing nodes equipped with a short-circuit jumper to short-circuit the computer and assigns the short-circuited computing node to the single-node computing task.
[0040] The present application uses a scheduling method and a scheduling system to achieve the allocation of single-node computing tasks to computers equipped with short-circuited jumpers and short-circuited. The remaining computing nodes without short-circuited jumpers and computing nodes equipped with short-circuited jumpers and not short-circuited can still undertake single-node computing tasks or multi-node computing tasks, thereby achieving the technical effect of a single node occupying the remaining nodes without blocking. For example, all computer nodes (taking 4 computing nodes as an example) are not currently executing work, and the job scheduling manager has received jobs 1, 2, 3, and n. When job 1 is a 4-node computing task, all computers equipped with short-circuited jumpers are not short-circuited and all 4 computing nodes are allocated to job 1. The remaining jobs can only wait for job 1 to be executed, release and recycle the computing resources of the computing nodes, and then allocate the computing nodes to the jobs to be scheduled according to the scheduling strategy;
[0041] When Job 1 is a 3-node computing task, one of the computers equipped with a short-circuit jumper is controlled to be short-circuited, so that the remaining 3 computing nodes form a ring and the 3 computing nodes are assigned to Job 1; when Job 2 is a single-node computing task, the job scheduler will assign the short-circuited computing node to Job 2, but when Job 2 is a multi-node computing task, it can only wait for Job 1 to be executed, release the computing resources of the recycled computing nodes, and then assign the computing nodes to the scheduled Job 2 according to the scheduling strategy.
[0042] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A computer network system, characterized in that: The invention comprises a plurality of computers without short-circuit jumpers, wherein a computer equipped with a short-circuit jumper is arranged between two adjacent computers without short-circuit jumpers, the computers without short-circuit jumpers and the computers equipped with short-circuit jumpers are both equipped with two central processing units and each central processing unit is connected with a high-speed network card, and the high-speed network cards between the computers without short-circuit jumpers and the computers equipped with short-circuit jumpers are connected by wires to form a ring network topology; when the computers equipped with short-circuit jumpers are not short-circuited, the computers not short-circuited are connected to the two adjacent computers without short-circuit jumpers through the high-speed network cards; when at least one computer equipped with a short-circuit jumper is short-circuited, the two adjacent computers without short-circuit jumpers are directly connected at the physical layer through the short-circuit jumper, and the remaining computers without short-circuit jumpers and the computers equipped with short-circuit jumpers that are not short-circuited still form a ring network topology through wired connections; The short-circuit jumper components include a short-circuit jumper and a short-circuit controller. The short-circuit jumper is a wire connected between two high-speed network cards of a computer and is used to control the connection state of the circuit. The short-circuit controller is a logic circuit used to detect the state of the short-circuit jumper and related instructions. The related instructions include identifying and changing whether the short-circuit jumper is short-circuited.
2. A job scheduling system comprising the computer network system according to claim 1, characterized in that: The system comprises a job scheduling manager, which is used for receiving jobs and allocating computers to the received jobs. Computers without short-circuit jumpers and computers equipped with short-circuit jumpers in the computer network system are used for calculating the allocated jobs.
3. The job scheduling system according to claim 2, characterized in that: The job scheduling manager includes a job queue, a job scheduler and a node manager; The job queue is used to receive job requests and sort multiple job requests according to the sorting principle set by the job scheduling strategy. After obtaining the job sorting status, it initiates a job computing request to the job scheduler. The job computing request includes the required number of computing nodes. The job scheduler is used to initiate a request for the number of computing nodes to the node manager, and after the available computing nodes returned by the node manager, control the corresponding number of computers equipped with short-circuit jumpers to be short-circuited or not short-circuited, and allocate the corresponding number of computing nodes according to the job computing request; The node manager is used to manage all computers in the computer system. Based on the computing node quantity request sent by the job scheduler, the node manager returns available computing nodes to the job scheduler. The available computing nodes include computers without short-circuiting jumpers and computers equipped with short-circuiting jumpers but not short-circuited.
4. A scheduling method based on the job scheduling system according to claim 3, characterized in that: The following steps are involved: Step 1: Receive the job submission instruction sent by the client, which includes the computing requirements of the submitted job, including the number of computing nodes; Step 2: The job scheduling manager puts the received jobs into the job queue, and the job queue is sorted according to the order in which they enter the queue; Step 3: The job scheduling manager determines the number of available computing nodes M, determines the number of computing nodes m required for the first-order job in the job queue, allocates the corresponding number of nodes to the job and starts execution; Step 4: Get the required computing quantity of the job in the first order in the update task queue, and search for the corresponding number of target nodes in the remaining Mm computing nodes; if there is a match, assign the corresponding target node to the job and start execution; Step 5: Until the remaining computing nodes are allocated according to step 4 or the remaining computing nodes cannot meet the computing requirements of the jobs in the job queue, wait for the running jobs to be executed and terminated, and then release the computing resources of the corresponding executing computing nodes, and reallocate the computing nodes to the jobs to be scheduled.
5. The scheduling method according to claim 4, characterized in that: In steps 3 and 4, the jobs are classified according to the required number of computing nodes to form corresponding node computing tasks. When the job belongs to a single-node computing task, the job scheduler controls at least one computer equipped with a short-circuit jumper in the available computing nodes to short-circuit, and assigns the single-node computing task to the short-circuited computer.
Citation Information
Patent Citations
Distributed computing multiple application function asynchronous concurrent scheduling method
CN102063336A
Method for achieving high availability of computer operation scheduling system
CN103279386A