Server system, job execution method and apparatus, device, and medium
By introducing extended computing domain and expansion controller in the server system, the problems of low expansion performance and limited number of computing units in the prior art are solved, and efficient and low-cost incremental expansion of computing units are achieved, meeting the growing computing power demand.
Patent Information
- Application Number
- PCT/CN2024/120444
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-09-23
- Publication Date
- 2025-06-05
AI Technical Summary
In the existing server architecture, the expansion performance of computing units is low and the number of expansions is limited, making it difficult to meet the growing computing power demand.
By introducing an extended computing domain into the server system, using the extension controller to communicate with the server, allocate and manage extended computing tasks, the incremental expansion of the computing unit is realized. The extended computing unit in the extended computing domain may include a graphics processor, a field programmable gate array, etc., and communicate through a high-speed serial computer extended bus standard protocol.
The expansion performance and number of computing units are improved, and the incremental expansion of low-cost and weakly coupled computing units are realized, meeting the demand for high computing power.
Smart Images

Figure CN2024120444_05062025_PF_FP_ABST
Abstract
Description
A server system, job execution method, device, equipment and medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311599060.5, and application name “A server system, job execution method, device, equipment and medium”, all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and more specifically, to a server system, a job execution method, an apparatus, a device, and a medium. Background Art
[0004] AI servers require continuous expansion of computing units to meet the growing demand for computing power. In existing server architectures, this expansion is achieved by expanding the number of CPUs (Central Processing Units). However, due to limitations such as processor channel size and server space, existing technologies have limited scalability and the number of expansion units is limited.
[0005] Therefore, how to improve the expansion performance and expansion quantity of computing units is a technical problem that needs to be solved by those skilled in the art.
[0006] Summary of the Invention
[0007] The present application provides a server system, comprising a server and an extended computing domain; the server comprises a processor control domain and a local computing domain, the local computing domain comprising a plurality of local computing units, the processor control domain being connected to the local computing units via a high-speed serial computer expansion bus standard protocol interface, the local computing units being used to perform local computing tasks;
[0008] The extended computing domain includes an extended controller and multiple extended computing units connected to the extended controller. The server is connected to the extended controller through an extended extension line and / or an external communication interface that complies with the high-speed serial computer expansion bus standard protocol. The extended controller is used to communicate with the server to obtain extended computing tasks, and the extended computing units are used to execute extended computing tasks.
[0009] The expansion controller communicates with the server based on an external communication protocol, and the external communication protocol includes a remote direct data access communication protocol and / or an Ethernet protocol.
[0010] The expansion controller includes a control unit and a switching unit. The control unit is used to perform control flow communication with the server, and the switching unit is used to perform data flow communication with the server.
[0011] Among them, the switching unit includes an upstream port, a switching matrix and a downstream port connected in sequence. The switching unit communicates data streams with the processor control domain through the upstream port, and communicates data streams with the extended computing unit through the downstream port. Each downstream port is connected to an extended computing unit, and there is a communication connection between every two downstream ports in the switching matrix.
[0012] The extended computing unit includes any one or a combination of any two of a graphics processor, a field programmable gate array, and a cross-point device.
[0013] The extended computing units communicate with each other based on an internal communication protocol, which includes a high-speed serial computer expansion bus standard protocol and / or a point-to-point transmission protocol of the high-speed serial computer expansion bus standard protocol.
[0014] In the process of communication between the extended controller and the server, the controller of the sending end enters the kernel state to create a communication link between the sending end and the receiving end, copies the data to be sent in the memory of the sending end to the hardware cache, and encapsulates the data to be sent in the hardware cache into a data packet according to the communication protocol between the sending end and the receiving end and sends it to the receiving end through the communication link; in which, if the sending end is a server and the receiving end is an extended computing unit, the controller of the sending end is the processor control domain; if the sending end is an extended computing unit and the receiving end is a server, the controller of the sending end is an extended controller.
[0015] In the process of the extended controller communicating with the processor control domain based on the remote direct data access communication protocol, the controller at the sending end bypasses the kernel state to copy the data to be sent to the hardware cache, and encapsulates the data to be sent in the hardware cache into a data packet based on the remote direct data access communication protocol and sends it to the receiving end; in which, if the sending end is a server and the receiving end is an extended computing unit, the controller at the sending end is the processor control domain; if the sending end is an extended computing unit and the receiving end is a server, the controller at the sending end is an extended controller.
[0016] The memory of the local computing unit communicates with the memory of the extended computing unit by direct memory access.
[0017] Among them, the local computing unit is used to execute computing tasks with delay sensitivity greater than or equal to a first preset value, the extended computing unit is used to execute computing tasks with delay sensitivity less than the first preset value, the communication sparsity between computing tasks executed in the extended computing domain is less than or equal to a second preset value, and the communication sparsity between computing tasks executed in the extended computing domain and computing tasks executed in the local computing domain is greater than the second preset value.
[0018] The processor control domain is used to merge the computing task execution results of the local computing unit and the computing task execution results of the extended computing unit.
[0019] The present application provides a job execution method, which is applied to a server in the above-mentioned server system, including:
[0020] Get the target job and split it into multiple computing tasks;
[0021] Sending computing tasks to local computing units in the server and extended computing units in the extended computing domain in the server system for execution;
[0022] Obtain the execution results of the local computing task of the local computing unit and the execution results of the extended computing task of the extended computing unit;
[0023] The local computing task execution results and the computing task execution results are merged to obtain the execution results of the target job.
[0024] Before sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, the method further includes:
[0025] Computing tasks are divided into local computing tasks and extended computing tasks according to their delay sensitivity and the communication sparsity between different computing tasks.
[0026] Among them, the local computing task is a computing task whose delay sensitivity is greater than or equal to the first preset value, the extended computing task is a computing task whose execution delay sensitivity is less than the first preset value, the communication sparsity between the extended computing tasks is less than or equal to the second preset value, and the communication sparsity between the extended computing tasks and the local computing tasks is greater than the second preset value.
[0027] The step of sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution includes:
[0028] The local computing tasks are sent to the local computing units in the server for execution, and the extended computing tasks are sent to the extended computing units in the extended computing domain in the server system for execution.
[0029] The target job is a model training job, which is split into multiple computing tasks, including:
[0030] Split the model training job into multiple sub-model training jobs;
[0031] Accordingly, the computing task is sent to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, including:
[0032] Sending the sub-model training job to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, thereby obtaining a trained sub-model;
[0033] Accordingly, obtaining the execution results of the local computing task of the local computing unit and the execution results of the extended computing task of the extended computing unit includes:
[0034] Obtain the sub-model trained by the local computing unit and the sub-model trained by the extended computing unit;
[0035] Accordingly, the local computing task execution results and the computing task execution results are merged to obtain the execution results of the target job, including:
[0036] The sub-model trained by the local computing unit and the sub-model trained by the extended computing unit are merged to obtain a trained model.
[0037] The target job is a federated learning job, which is split into multiple computing tasks, including:
[0038] Split the federated learning task into model training tasks and parameter update tasks;
[0039] Accordingly, the computing task is sent to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, including:
[0040] Sending the parameter update task to the local computing unit in the server so that the local computing unit updates the model parameters of the model according to the merged gradient in the parameter update task; wherein the merged gradient is the merged result of the gradients obtained by training multiple extended computing units;
[0041] Sending the model training job to the extended computing unit in the extended computing domain of the server system so that the extended computing unit can use local data to train the model corresponding to the model training job to obtain gradients; wherein the model corresponding to the model training job is the latest model updated by the local computing unit;
[0042] Accordingly, obtaining the execution results of the local computing task of the local computing unit and the execution results of the extended computing task of the extended computing unit includes:
[0043] Get the gradient obtained from the training of the extended computing unit;
[0044] Get the latest model updated by the local computing unit;
[0045] Accordingly, the local computing task execution results and the computing task execution results are merged, including:
[0046] The gradients obtained from training multiple extended computing units are merged to obtain a merged gradient.
[0047] The present application provides a job execution device, which is applied to a server in the above-mentioned server system, and includes:
[0048] The splitting module is used to obtain the target job and split it into multiple computing tasks;
[0049] A sending module, configured to send computing tasks to local computing units in the server and extended computing units in the extended computing domain in the server system for execution;
[0050] An acquisition module is used to obtain the execution results of the local computing task of the local computing unit and the execution results of the extended computing task of the extended computing unit;
[0051] The merging module is used to merge the local computing task execution results and the computing task execution results to obtain the execution results of the target job.
[0052] The present application provides an electronic device, comprising:
[0053] a memory for storing computer-readable instructions;
[0054] A processor is configured to implement the steps of the above-mentioned job execution method when executing computer-readable instructions.
[0055] The present application provides one or more non-volatile computer-readable storage media storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the above-mentioned job execution method.
[0056] Correspondingly, the processor control domain includes any one or a combination of multiple items of a single chip microcomputer, a programmable logic controller, a digital signal processor, a field programmable gate array, and a central processing unit.
[0057] Accordingly, the multiple local computing units include any one or a combination of multiple items including a graphics computing processor, a field programmable gate array, a tensor processing unit accelerator card, a data processor, and a neural network processor.
[0058] Accordingly, the control unit includes any one or a combination of multiple items of an operation control core, an instruction decoder, a clock and timing controller, and a data buffer. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings are used to provide a further understanding of the present disclosure and constitute part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure, but do not constitute a limitation of the present disclosure. In the drawings:
[0060] FIG1 is a structural diagram of a server system in the related art;
[0061] FIG2 is a schematic diagram of a PCIe Switch expansion mode in the related art;
[0062] FIG3 is a structural diagram of a server system according to one or more exemplary embodiments;
[0063] FIG4 is a structural diagram of an expansion controller according to one or more exemplary embodiments;
[0064] FIG5 is a structural diagram of a switching unit according to one or more exemplary embodiments;
[0065] FIG6 is a schematic diagram illustrating a connection method between a switching matrix and a computing unit according to one or more exemplary embodiments;
[0066] FIG7 is a flowchart of a method for executing a job according to one or more exemplary embodiments;
[0067] FIG8 is a schematic diagram showing vertical splitting of a model in distributed training according to one or more exemplary embodiments;
[0068] FIG9 is a schematic diagram of federated learning according to one or more exemplary embodiments;
[0069] FIG10 is a structural diagram of a job execution device according to one or more exemplary embodiments;
[0070] Fig. 11 is a structural diagram of an electronic device according to one or more exemplary embodiments. DETAILED DESCRIPTION
[0071] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. In addition, in the embodiments of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0072] In existing server architectures, computing unit expansion is achieved by expanding the number of CPU channels. As shown in Figure 1, the processor controller connects to memory, and the processor control domain connects to computing units via PCIe interfaces. Due to the limited number of CPU channels, in addition to individual computing units that can be directly connected, PCIe switches can also be used for tree-like cascading expansion. The PCIe switch expansion model is shown in Figure 2. Due to limitations such as processor channel size and server space, existing technologies have low computing unit expansion performance and a limited number of expansions.
[0073] Therefore, the present application adds an extended computing domain on the basis of realizing the expansion of computing units through local computing domains, and realizes the incremental expansion of computing units. The server is connected to the extended computing domain through a PCIe extension line or an external communication interface to realize a weak coupling connection between the server and the extended computing domain, and realizes the weak coupling expansion of computing units. Furthermore, the extended computing domain, as an independent server operation unit, includes an extended controller for controlling the extended computing units in the extended computing domain, realizing exchange with the upstream server and issuing extended computing tasks to the downstream extended computing units, that is, the logic of the extended computing domain is controlled by the extended controller, and there is no need for the processing control domain in the server to control it. The processing channel is only occupied when the extended controller communicates with the server, and the processing channel will not be occupied at other times. The expansion of the computing unit will not be limited by the server processing channel and space, and there is no need to increase the server processing channel, thereby realizing low-cost expansion of the computing unit. It can be seen that the present application realizes low-cost, weakly coupled incremental expansion of the computing unit, and improves the expansion performance and expansion quantity of the computing unit.
[0074] The present application embodiment discloses a server system, including a server and an extended computing domain; the server includes a processor control domain and a local computing domain, the local computing domain includes multiple local computing units, the processor control domain is connected to the local computing units via a high-speed serial computer expansion bus standard protocol interface, and the local computing units are used to perform local computing tasks;
[0075] The extended computing domain includes an extended controller and multiple extended computing units connected to the extended controller. The server is connected to the extended controller through an extended extension line and / or an external communication interface that complies with the high-speed serial computer expansion bus standard protocol. The extended controller is used to communicate with the server to obtain extended computing tasks, and the extended computing units are used to execute extended computing tasks.
[0076] In this embodiment, as shown in FIG3 , the expansion of the computing unit is achieved through the local computing domain, that is, the processor control domain is connected to the computing unit through the PCIe interface, and multiple computing units can also be connected through the PCIe switch to achieve tree-shaped cascade expansion of the computing unit. The CPU and the local computing domain serve as the main control domain, and some tasks that are sensitive to delay requirements and inseparable can be run on the local computing domain. Specifically, the processor control domain includes any one or more combinations of microcontrollers, programmable logic controllers (PLC, Programmable Logic Controller), digital signal processors (DSP, Digital Signal Processor), field programmable gate arrays (FPGA, Field-Programmable Gate Array) and CPUs commonly used in servers. The local computing unit can be a graphics computing processor, a field programmable gate array, a tensor processing unit accelerator card TPU, a DPU (Data Processing Unit, data processor), an NPU (neural network processor), etc.
[0077] On this basis, further expansion of computing units is achieved by extending the computing domain. This serves as an auxiliary or secondary control domain. Tasks that are relatively insensitive to latency requirements and interact infrequently with other tasks can run in the extended computing domain, preventing overall system performance degradation caused by frequent communication.
[0078] The extended computing domain includes an extended controller and multiple extended computing units connected to it. These units can include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and extended processing units (xPUs). The extended controller in the extended computing domain, operating independently of the host system, can be implemented using an FPGA or chip. The extended controller exchanges with upstream servers and distributes extended computing tasks to downstream extended computing units.
[0079] As a feasible implementation manner, the extension controller includes a control unit and a switching unit. The control unit is used to perform control flow communication with the server, and the switching unit is used to perform data flow communication with the server.
[0080] In a specific implementation, as shown in FIG4 , the extension controller includes a control unit and a switching unit. The control unit and the switching unit work in coordination. The control unit is responsible for the control flow, and the switching unit is responsible for the data flow. The control unit includes any one of an operation control core, an instruction decoder, a clock and timing controller, and a data buffer, or is obtained by integrating multiple of them. The control unit obtained by integrating one or more of them can realize functions such as a control module, a configuration module, and a protocol support module. The realization of specific functions can be obtained by manually configuring the control unit. Specifically, the control unit includes a control module, a configuration module, and a protocol support module. The main functions of the control module and the configuration module include: device enumeration detection, configuration space information setting, creation of various resources required for communication, device communication mode, upstream and downstream command parsing, interface connection management, and link training. Resources required for communication include RDMA (Remote Direct Memory Access) control paths, which create queue pairs (QPs), completion queues (CQs), and memory regions (MRs) at the sending and receiving ends. Device communication methods include P2P (Peer-to-Peer), DMA (Direct Memory Access), and GPUDirect. Interface connection management includes the interconnection methods of interface modules in switches. The protocol support module supports internal communication protocols such as the High-Speed Serial Computer Expansion Bus Standard Protocol, the Point-to-Point Transmission Protocol for the High-Speed Serial Computer Expansion Bus Standard Protocol, and the Neural Network Processor Interconnect Protocol, as well as external communication protocols such as the Remote Direct Data Access Communication Protocol and the Ethernet Protocol.
[0081] The expansion controller communicates with the server based on external communication protocols, such as Remote Direct Data Access (RDMA) and Ethernet. Depending on the type of expansion unit, the protocol support module supports expansion of different types of units through protocol switching and supports different modes of protocol parsing through communication with the server. The expansion units communicate with each other based on internal communication protocols, such as PCIe (PCIe Express), PCIe Peer-to-Peer (PCIe Peer-to-Peer), and NVLink (NVIDIA High-Speed Interconnect Protocol). This creates a highly cohesive connection within the expansion domain and a weakly coupled connection outside the expansion domain.
[0082] As a feasible implementation method, the switching unit includes an upstream port, a switching matrix and a downstream port connected in sequence. The switching unit communicates data streams with the processor control domain through the upstream port, and the switching unit communicates data streams with the extended computing unit through the downstream port. Each downstream port is connected to an extended computing unit, and there is a communication connection between every two downstream ports in the switching matrix.
[0083] In a specific implementation, as shown in Figure 5, a switching unit includes upstream ports, downstream ports, and a switching matrix. The upstream interface is used to connect to the server, and the downstream interface is used to connect to the extended computing unit. However, unlike the topological connection method of a PCIe switch, the switching matrix not only forwards communication messages between the upstream interface and the corresponding downstream interface, but also is responsible for the topological connection method between the extended computing units. The switching matrix is a hardware structure on a backplane switch that is used to achieve high-speed point-to-point connections between various line cards. The connection method between the switching matrix and the computing unit is shown in Figure 6. Each downstream port is connected to an extended computing unit, and there is a communication connection between every two downstream ports in the switching matrix. Compared with the tree-like computing unit expansion, the matrix-type computing unit expansion is realized, which improves the expansion performance and expansion quantity of computing units.
[0084] The server and expansion domain can be connected via PCIe extension cables or external communication interfaces. PCIe extension cables are bus-level expansions, unlike traditional models that rely on costly expansion of CPU paths to expand the number of communication channels, or rely on the limited number of fixed channels on a PCIe switch within a limited space. By using an independent control unit to expand the controller, the bus expansion model can overcome spatial limitations. For example, a high-speed, fully interconnected structure using a proprietary lightweight protocol can be implemented using the architecture shown in Figure 6. External communication interfaces can have multiple modes, such as the RDMA communication protocol with low kernel overhead or the Ethernet protocol. Alternatively, communication methods such as the motherboard's RJ45 (Registered Jack 45) network port, which consumes more CPU resources, can be used.
[0085] As a feasible implementation method, during the communication between the extended controller and the server, the controller of the sending end enters the kernel state to create a communication link between the sending end and the receiving end, copies the data to be sent in the memory of the sending end to the hardware cache, and encapsulates the data to be sent in the hardware cache into a data packet and sends it to the receiving end through the communication link according to the communication protocol between the sending end and the receiving end; wherein, if the sending end is a server and the receiving end is an extended computing unit, the controller of the sending end is a processor control domain; if the sending end is an extended computing unit and the receiving end is a server, the controller of the sending end is an extended controller.
[0086] When the server sends data to the extended computing unit, the processor control domain enters kernel mode to establish a communication link between the server and the extended controller and creates the necessary resources for communication. As the receiving end, the extended controller parses the communication protocol, configures information such as the cache address where the data is stored, and notifies the network card and other hardware to prepare to receive the data. The processor control domain configures information such as the cache address and data length for the data to be sent, and notifies the network card and other hardware to send the data. The processor control domain copies the data to be sent from the memory cache to the cache of the network card and other hardware, then encapsulates the data packet according to the protocol and sends it to the extended controller. After receiving the data packet, the extended controller writes it to the memory of the extended computing unit. When the extended computing unit sends data to the server, the extended controller enters kernel mode to establish a communication link between the server and the extended controller and creates the necessary resources for communication. As the receiving end, the processor control domain parses the communication protocol, configures information such as the cache address where the data is stored, and notifies the network card and other hardware to prepare to receive the data. The extended controller configures information such as the cache address and data length for the data to be sent, and notifies the network card and other hardware to send the data. The extended controller copies the data to be sent in the memory of the extended computing unit to the cache of hardware such as the network card, and then encapsulates the data packet according to the protocol and sends it to the processor control domain. After receiving the data packet, the processor control domain writes it into the memory.
[0087] As another feasible implementation method, in the process of the extended controller communicating with the processor control domain based on the remote direct data access communication protocol, the controller at the sending end bypasses the kernel state to copy the data to be sent to the hardware cache, and encapsulates the data to be sent in the hardware cache into a data packet and sends it to the receiving end based on the remote direct data access communication protocol; wherein, if the sending end is a server and the receiving end is an extended computing unit, the controller at the sending end is the processor control domain; if the sending end is an extended computing unit and the receiving end is a server, the controller at the sending end is an extended controller.
[0088] In a specific implementation, the processor control domain and the extended controller can interact through RDMA. This implementation is different from the above implementation in that the controller on the sending end bypasses the kernel state to copy the data to be sent to the hardware cache, and does not require the participation of the sending segment memory.
[0089] As another feasible implementation, direct memory access is used to communicate between the memory of the local computing unit and the memory of the extended computing unit. In a specific implementation, if the extended computing unit is a GPU, GPUdirect RDMA can be used to directly send protocol-encapsulated data packets from the local computing unit memory to the extended computing unit memory, or vice versa.
[0090] As a preferred embodiment, the local computing unit is used to execute computing tasks with delay sensitivity greater than or equal to a first preset value, the extended computing unit is used to execute computing tasks with delay sensitivity less than the first preset value, the communication sparsity between computing tasks executed in the extended computing domain is less than or equal to a second preset value, and the communication sparsity between computing tasks executed in the extended computing domain and computing tasks executed in the local computing domain is greater than the second preset value.
[0091] In specific implementation, for some jobs that require computational splitting, such as distributed machine learning, computing tasks can be divided into local computing tasks and extended computing tasks. Local computing tasks are computing tasks with high delay sensitivity or cannot be split, and extended computing tasks are computing tasks with low delay sensitivity. At the same time, the communication sparsity between computing tasks executed in the extended computing domain is low, that is, the communication is intensive, while the communication sparsity between computing tasks executed in the extended computing domain and computing tasks executed in the local computing domain is high, that is, the communication is sparse, ultimately achieving high cohesion and low coupling to form high-performance computing in the extended computing domain and incremental improvement of system computing performance between computing domains.
[0092] Assume that the percentage of local computing tasks in the local computing domain to the total computing tasks is X, and the percentage of extended computing tasks in the extended computing domain to the total computing tasks is Y, X + Y = 1. The extended computing domain improves performance by N times, but the communication overhead it brings is Z. The total performance speedup ratio is 1 / (X + Y / N + Z). From the above formula, it can be seen that when the communication overhead is not considered, the higher the performance improvement ratio of the extended domain, the higher the system acceleration performance improvement. However, excessive communication overhead will reduce system performance.
[0093] As a feasible implementation, the processor control domain is used to merge the computational task execution results of the local computing unit and the computational task execution results of the extended computing unit. In a specific implementation, after the different computing domains complete the calculation, the processor control domain merges the final calculation results and completes the computation task after the calculation is completed.
[0094] The server system provided by the embodiment of the present application, on the basis of realizing the expansion of the computing unit through the local computing domain, adds the extended computing domain, realizes the incremental expansion of the computing unit. The server is connected to the extended computing domain through a PCIe extension line or an external communication interface, realizes a weak coupling connection between the server and the extended computing domain, and realizes the weak coupling expansion of the computing unit. Further, the extended computing domain, as an independent server operation unit, includes an extended controller for controlling the extended computing unit in the extended computing domain, realizes exchange with the upstream server, and sends extended computing tasks to the downstream extended computing unit, that is, the logic of the extended computing domain is controlled by the extended controller, and there is no need for the processing control domain in the server to control it. The processing channel is only occupied when the extended controller communicates with the server, and the processing channel is not occupied at other times. The expansion of the computing unit will not be limited by the server processing channel and space, and there is no need to increase the server processing channel, thereby realizing low-cost expansion of the computing unit. It can be seen from this that the server system provided by the embodiment of the present application realizes low-cost, weakly coupled incremental expansion of the computing unit, and improves the expansion performance and expansion quantity of the computing unit.
[0095] The present application discloses a method for executing a job. Referring to FIG7 , a flowchart of a method for executing a job according to an exemplary embodiment is shown. As shown in FIG7 , the method includes:
[0096] S101: Obtain a target job and split the target job into multiple computing tasks;
[0097] S102: Sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution;
[0098] This embodiment is applied to the server system provided by the above embodiment. In a specific implementation, the target job is split into multiple computing tasks, which are sent to the local computing units in the local computing domain and the extended computing units in the extended computing domain for execution.
[0099] As a preferred embodiment, before sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, it also includes: dividing the computing task into local computing tasks and extended computing tasks according to the delay sensitivity of the computing task and the communication sparsity between different computing tasks; wherein the local computing task is a computing task with a delay sensitivity greater than or equal to a first preset value, the extended computing task is a computing task with an execution delay sensitivity less than the first preset value, the communication sparsity between the extended computing tasks is less than or equal to a second preset value, and the communication sparsity between the extended computing tasks and the local computing task is greater than the second preset value.
[0100] In specific implementation, computing tasks can be divided into local computing tasks and extended computing tasks. Local computing tasks are computing tasks with high delay sensitivity or cannot be split, and extended computing tasks are computing tasks with low delay sensitivity. At the same time, the communication sparsity between computing tasks executed in the extended computing domain is low, that is, the communication is intensive, while the communication sparsity between computing tasks executed in the extended computing domain and computing tasks executed in the local computing domain is high, that is, the communication is sparse, ultimately achieving high cohesion and low coupling to form high-performance computing in the extended computing domain and incremental improvement of system computing performance between computing domains.
[0101] Furthermore, local computing tasks are sent to local computing units in the server for execution, and extended computing tasks are sent to extended computing units in the extended computing domain in the server system for execution. In specific implementations, local computing tasks and extended computing tasks are mapped to the local computing domain and the extended computing domain, respectively, based on the task division results. In order to provide a high-performance and flexible computing base as much as possible, both the local computing domain and the extended computing domain adopt a high-speed interconnected communication method for computing regions. The domain where the task is located can perform high-performance computing with the highest parallel computing and high-speed communication method based on the hardware resource status, connection topology information, and task type. The difference is that the control of the local computing domain is completed by the CPU, while the control of the extended computing domain is completed by the extended controller on the extended domain.
[0102] S103: Obtaining the local computing task execution result of the local computing unit and the extended computing task execution result of the extended computing unit;
[0103] S104: Merge the local computing task execution result and the computing task execution result to obtain the execution result of the target job.
[0104] In specific implementations, the ideal computing method is to decouple the local computing domain from the extended computing domain. After the calculations of different computing domains are completed, the processor control domain merges the final calculation results, and the computing task is completed after the calculations are processed. In a non-ideal state, there is a state where the local computing domain and the extended computing domain are task-coupled, but the above task division has minimized the coupling as much as possible, that is, reduced the interaction between different computing domains. The interaction process uses different interactive communication methods with the processor control domain according to the different expansion methods. Or, to maximize performance, all available extended communication methods can be used to interact with the processor control domain. Finally, the task collaboration is completed and the processor control domain merges and completes the job.
[0105] The job execution method provided in this application utilizes the server system provided in the above embodiment to execute the target job, thereby improving the job execution efficiency.
[0106] The above embodiments can be applied to error detection, including dual-core lockstep detection and heterogeneous parallel multi-core detection. The core idea of the dual-core lockstep detection technology is to use two identical processor cores in a computer system and let them execute the same instruction sequence at the same time. During the execution process, the two processor cores will compare the execution results with each other. If it is found that the execution results of the two cores are inconsistent, the system will immediately enter the safe mode, stop running and perform fault diagnosis and repair. Heterogeneous parallel multi-core refers to the use of multiple processor cores to perform the same instruction detection in the computer architecture. These processor cores may have different architectures, functions or performance characteristics. Regardless of which of the above error detection methods is used, it can run in parallel in the local computing domain and the extended computing domain according to the error detection requirements.
[0107] The above embodiment can be applied to model training, which includes the following steps:
[0108] Step 1: Split the model training job into multiple sub-model training jobs;
[0109] Step 2: Send the sub-model training job to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, and obtain the trained sub-model;
[0110] Step 3: Obtain the sub-model trained by the local computing unit and the sub-model trained by the extended computing unit;
[0111] Step 4: Merge the sub-model trained by the local computing unit and the sub-model trained by the extended computing unit to obtain a trained model.
[0112] In practice, the entire model is divided into multiple sub-models. Each computing unit uses the same data to train a different sub-model. The processor control domain then merges the trained sub-models from each computing unit to produce the final trained model. For example, as shown in Figure 8, the vertical splitting of the model in distributed training involves training different sub-models in the local computing domain and the extended computing domain. The computational communication between the server and the extended computing domain is the result of the parallel computing units in each model.
[0113] The above embodiment can be applied to federated learning, including the following steps:
[0114] Step 1: Split the federated learning job into model training tasks and parameter update tasks;
[0115] Step 2: Send the model training job to the extended computing unit in the extended computing domain of the server system, so that the extended computing unit can use local data to train the model corresponding to the model training job to obtain gradients. The model corresponding to the model training job is the latest model updated by the local computing unit.
[0116] Step 3: Obtain the gradient obtained by the extended computing unit training, and merge the gradients obtained by the training of multiple extended computing units to obtain a merged gradient;
[0117] Step 4: Send the parameter update task to the local computing unit in the server so that the local computing unit updates the model parameters of the model according to the merged gradient in the parameter update task; where the merged gradient is the merged result of the gradients obtained by training multiple extended computing units;
[0118] Step 5: Obtain the latest model updated by the local computing unit and send the latest model to the extended computing unit so that the extended computing unit can update its own model.
[0119] In the specific implementation, as shown in Figure 9, each extended computing unit in the extended computing domain downloads the latest model from the server, uses local data to train the model to obtain the gradient, encrypts and uploads it to the server, and the server merges the gradients of each extended computing unit to obtain a merged gradient. The local computing unit updates the model parameters according to the merged gradient and returns the updated model to each extended computing unit in the extended computing domain. Each extended computing unit in the extended computing domain updates its own model, and the above process continues until the calculation is completed.
[0120] The following introduces a job execution device provided in an embodiment of the present application. The job execution device described below and the job execution method described above can be referenced to each other.
[0121] Referring to FIG10 , a structural diagram of a job execution device according to an exemplary embodiment is shown. As shown in FIG10 , the device includes:
[0122] A splitting module 100 is used to obtain a target job and split the target job into multiple computing tasks;
[0123] A sending module 200 is used to send the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution;
[0124] An acquisition module 300 is configured to acquire the execution results of the local computing task of the local computing unit and the execution results of the extended computing task of the extended computing unit;
[0125] The merging module 400 is used to merge the local computing task execution result and the computing task execution result to obtain the execution result of the target job.
[0126] The job execution device provided in this application utilizes the server system provided in the above embodiments to execute the target job, thereby improving the job execution efficiency.
[0127] Based on the above embodiment, as a preferred implementation, it further includes:
[0128] The partitioning module is used to divide the computing tasks into local computing tasks and extended computing tasks according to the delay sensitivity of the computing tasks and the communication sparsity between different computing tasks.
[0129] Based on the above embodiments, as a preferred implementation method, the local computing task is a computing task whose delay sensitivity is greater than or equal to a first preset value, the extended computing task is a computing task whose execution delay sensitivity is less than the first preset value, the communication sparsity between the extended computing tasks is less than or equal to a second preset value, and the communication sparsity between the extended computing tasks and the local computing tasks is greater than the second preset value.
[0130] Based on the above embodiments, as a preferred implementation, the sending module 200 is specifically used to: send local computing tasks to the local computing unit in the server for execution, and send extended computing tasks to the extended computing unit in the extended computing domain in the server system for execution.
[0131] On the basis of the above embodiments, as a preferred implementation mode, the target job is a model training job, and the splitting module 100 is specifically used to: split the model training job into multiple sub-model training jobs; accordingly, the sending module 200 is specifically used to: send the sub-model training job to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, to obtain the trained sub-model; accordingly, the acquisition module 300 is specifically used to: obtain the sub-model trained by the local computing unit and the sub-model trained by the extended computing unit; accordingly, the merging module 400 is specifically used to: merge the sub-model trained by the local computing unit and the sub-model trained by the extended computing unit to obtain a trained model.
[0132] Based on the above embodiment, as a preferred implementation method, the target job is a federated learning job, and the splitting module 100 is specifically used to: split the federated learning job into a model training task and a parameter update task; accordingly, the sending module 200 is specifically used to: send the parameter update task to the local computing unit in the server, so that the local computing unit updates the model parameters of the model according to the merged gradient in the parameter update task; wherein the merged gradient is the merged result of the gradients obtained by training multiple extended computing units; send the model training job to the extended computing unit in the extended computing domain in the server system, so that the extended computing unit uses local data to train the model corresponding to the model training job to obtain the gradient; wherein the model corresponding to the model training job is the latest model updated by the local computing unit; accordingly, the acquisition module 300 is specifically used to: obtain the gradient obtained by training the extended computing unit; obtain the latest model updated by the local computing unit; accordingly, the merging module 400 is specifically used to: merge the gradients obtained by training multiple extended computing units to obtain a merged gradient.
[0133] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0134] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present application, the embodiment of the present application further provides an electronic device. FIG11 is a structural diagram of an electronic device according to an exemplary embodiment. As shown in FIG11 , the electronic device includes:
[0135] Communication interface 1, capable of exchanging information with other devices such as network devices;
[0136] The processor 2 is connected to the communication interface 1 to implement information exchange with other devices and is used to execute the job execution method provided by one or more of the above technical solutions when running computer-readable instructions. The computer-readable instructions are stored in the memory 3.
[0137] Of course, in actual applications, the various components in the electronic device are coupled together via bus system 4. It will be understood that bus system 4 is used to enable communication between these components. In addition to a data bus, bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in FIG11 , all of these buses are labeled as bus system 4.
[0138] The memory 3 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer-readable instructions for operating on the electronic device.
[0139] It is understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 3 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.
[0140] The method disclosed in the above-mentioned embodiment of the present application can be applied to processor 2 or implemented by processor 2. Processor 2 may be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above-mentioned method can be completed by the integrated logic circuit of the hardware in processor 2 or instructions in the form of software. The above-mentioned processor 2 can be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 2 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in memory 3. Processor 2 reads the program in memory 3 and completes the steps of the above-mentioned method in combination with its hardware.
[0141] When the processor 2 executes the program, the corresponding processes in each method of the embodiment of the present application are implemented. For the sake of brevity, they are not repeated here.
[0142] In exemplary embodiments, the present application also provides one or more non-volatile computer-readable storage media storing computer-readable instructions, such as a memory 3 storing computer-readable instructions. The computer-readable instructions are executable by a processor 2 to perform the aforementioned method steps. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface mount storage, optical disk, or CD-ROM.
[0143] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, ROM, RAM, disks or optical disks, etc. Various media that can store program codes.
[0144] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for an electronic device (which can be a personal computer, server, network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0145] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A server system, characterized in that: It includes a server and an extended computing domain; the server includes a processor control domain and a local computing domain, the local computing domain includes a plurality of local computing units, the processor control domain is connected to the local computing units via a high-speed serial computer extended bus standard protocol interface, and the local computing units are used to perform local computing tasks; The extended computing domain includes an extended controller and multiple extended computing units connected to the extended controller. The server is connected to the extended controller via an extended extension line and / or an external communication interface that complies with the high-speed serial computer expansion bus standard protocol. The extended controller is used to communicate with the server to obtain extended computing tasks, and the extended computing units are used to execute the extended computing tasks.
2. The server system according to claim 1, characterized in that: The expansion controller communicates with the server based on an external communication protocol, wherein the external communication protocol includes a remote direct data access communication protocol and / or an Ethernet protocol.
3. The server system according to claim 1, characterized in that: The expansion controller includes a control unit and a switching unit. The control unit is used to perform control flow communication with the server, and the switching unit is used to perform data flow communication with the server.
4. The server system according to claim 3, characterized in that: The switching unit includes an upstream port, a switching matrix and a downstream port connected in sequence. The switching unit performs data stream communication with the processor control domain through the upstream port, and the switching unit performs data stream communication with the extended computing unit through the downstream port. Each of the downstream ports is connected to an extended computing unit, and there is a communication connection between every two downstream ports in the switching matrix.
5. The server system according to claim 1, characterized in that: The extended computing unit includes any one or a combination of any one of a graphics processor, a field programmable gate array, and a cross-point device.
6. The server system according to claim 1, characterized in that: The extended computing units communicate with each other based on an internal communication protocol, wherein the internal communication protocol includes a high-speed serial computer expansion bus standard protocol and / or a point-to-point transmission protocol of a high-speed serial computer expansion bus standard protocol.
7. The server system according to claim 1, characterized in that: During the communication between the extended controller and the server, the controller of the sending end enters the kernel state to create a communication link between the sending end and the receiving end, copies the data to be sent in the memory of the sending end to the hardware cache, and encapsulates the data packets to be sent in the hardware cache into data packets and sends them to the receiving end through the communication link according to the communication protocol between the sending end and the receiving end; wherein, if the sending end is the server and the receiving end is the extended computing unit, the controller of the sending end is the processor control domain; if the sending end is the extended computing unit and the receiving end is the server, the controller of the sending end is the extended controller.
8. The server system according to claim 1, characterized in that: In the process of the extended controller communicating with the processor control domain based on the remote direct data access communication protocol, the controller at the sending end bypasses the kernel state to copy the data to be sent to the hardware cache, and encapsulates the data packets to be sent in the hardware cache and sends them to the receiving end based on the remote direct data access communication protocol; wherein, if the sending end is the server and the receiving end is the extended computing unit, the controller at the sending end is the processor control domain; if the sending end is the extended computing unit and the receiving end is the server, the controller at the sending end is the extended controller.
9. The server system according to claim 1, characterized in that: The memory of the local computing unit communicates with the memory of the extended computing unit by means of direct memory access.
10. The server system according to claim 1, characterized in that: The local computing unit is used to execute computing tasks whose delay sensitivity is greater than or equal to a first preset value, the extended computing unit is used to execute computing tasks whose delay sensitivity is less than the first preset value, the communication sparsity between computing tasks executed in the extended computing domain is less than or equal to a second preset value, and the communication sparsity between computing tasks executed in the extended computing domain and computing tasks executed in the local computing domain is greater than the second preset value.
11. The server system according to claim 1, characterized in that: The processor control domain is used to merge the computing task execution result of the local computing unit and the computing task execution result of the extended computing unit.
12. A method for executing a job, characterized in that: Applied to a server in a server system as claimed in any one of claims 1 to 11, the method comprising: Obtain a target job, and split the target job into multiple computing tasks; Sending the computing task to a local computing unit in the server and an extended computing unit in an extended computing domain in the server system for execution; Obtaining a local computing task execution result of the local computing unit and an extended computing task execution result of the extended computing unit; and The local computing task execution result and the computing task execution result are combined to obtain the execution result of the target job.
13. The method for executing a job according to claim 12, characterized in that: Before sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, the method further includes: The computing tasks are divided into local computing tasks and extended computing tasks according to the delay sensitivity of the computing tasks and the communication sparsity between different computing tasks.
14. The method for executing a job according to claim 13, characterized in that: The local computing task is a computing task whose delay sensitivity is greater than or equal to a first preset value, the extended computing task is a computing task whose execution delay sensitivity is less than the first preset value, the communication sparsity between the extended computing tasks is less than or equal to a second preset value, and the communication sparsity between the extended computing task and the local computing task is greater than the second preset value.
15. The method for executing a job according to claim 13, characterized in that: Sending the computing task to a local computing unit in the server and an extended computing unit in an extended computing domain in the server system for execution includes: The local computing task is sent to a local computing unit in the server for execution, and the extended computing task is sent to an extended computing unit in an extended computing domain in the server system for execution.
16. The method for executing a job according to claim 12, characterized in that: The target job is a model training job, and the target job is split into multiple computing tasks, including: Splitting the model training job into multiple sub-model training jobs; Accordingly, the computing task is sent to the local computing unit in the server and the extended computing unit in the server system. Extended compute unit execution in the domain, including: Sending the sub-model training job to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution, to obtain a trained sub-model; Correspondingly, the obtaining of the local computing task execution result of the local computing unit and the extended computing task execution result of the extended computing unit includes: Obtaining the sub-model trained by the local computing unit and the sub-model trained by the extended computing unit; Accordingly, the local computing task execution result and the computing task execution result are combined to obtain the execution result of the target job, including: The sub-model trained by the local computing unit and the sub-model trained by the extended computing unit are merged to obtain a trained model.
17. The method for executing a job according to claim 12, characterized in that: The target job is a federated learning job, which is split into multiple computing tasks, including: Splitting the federated learning job into a model training task and a parameter updating task; Accordingly, sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution includes: Sending the parameter update task to the local computing unit in the server so that the local computing unit updates the model parameters of the model according to the merged gradient in the parameter update task; wherein the merged gradient is a merged result of the gradients obtained by training of the plurality of the extended computing units; and Sending the model training job to an extended computing unit in an extended computing domain in the server system, so that the extended computing unit uses local data to train the model corresponding to the model training job to obtain a gradient; wherein the model corresponding to the model training job is the latest model updated by the local computing unit; Correspondingly, the obtaining of the local computing task execution result of the local computing unit and the extended computing task execution result of the extended computing unit includes: Obtaining a gradient obtained by training the extended computing unit; and Obtaining the latest model updated by the local computing unit; Accordingly, merging the local computing task execution result and the computing task execution result includes: The gradients obtained by training the multiple extended computing units are merged to obtain the merged gradient.
18. A work execution device, characterized in that: Applied to a server in a server system according to any one of claims 1 to 11, the device comprising: A splitting module, used for acquiring a target job and splitting the target job into multiple computing tasks; A sending module, used for sending the computing task to the local computing unit in the server and the extended computing unit in the extended computing domain in the server system for execution; An acquisition module, used to acquire the execution result of the local computing task of the local computing unit and the execution result of the extended computing task of the extended computing unit; The merging module is used to merge the local computing task execution result and the computing task execution result to obtain the execution result of the target job.
19. An electronic device, characterized in that: include: a memory for storing computer readable instructions; A processor, configured to implement the steps of the job execution method according to any one of claims 12 to 17 when executing the computer-readable instructions.
20. One or more non-volatile computer-readable storage media storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the method according to any one of claims 12 to 17.
Citation Information
Patent Citations
Distributed information processing method and device
CN111679860A
Job task processing method and device of computing system, storage medium and processor
CN112486646A
Task execution method and storage device
CN113821311A
Memory extension system and computing node
CN115858146A
Server system, job execution method, apparatus and device, and medium
CN117312215A