Array server computing power cluster method, control module, array server and medium
Through the array server computing power cluster method, the cluster mode is determined using operation information and computing power cluster requirements, data packets are divided and processed, which solves the problem of low computing power utilization efficiency caused by independent operation of processing modules in the array server, and realizes more efficient computing power utilization and data processing.
Patent Information
- Application Number
- CN202310185257.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-02-17
AI Technical Summary
The processing modules in existing array servers run independently, resulting in low computing power utilization efficiency and cannot meet the growth of AI computing power demand.
Through the array server computing power cluster method, the operation information and computing power cluster requirements are obtained, the cluster mode is determined, and the data packet is divided into multiple subpackets and sent to the processing module for processing, and the results are summarized and output.
It improves the computing power utilization rate of the processing module, enhances the efficiency of large-scale data processing, and reduces the computing and storage pressure of a single processing module.
Smart Images

Figure CN116302512B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to an array server computing power clustering method, a control module, an array server, and a medium. Background Art
[0002] The array server contains multiple processing modules to provide computing power. Each processing module specifically includes CPU computing power (Central Processing Unit), GPU computing power (Graphics Processing Unit), NPU computing power (Neural Network Processing Unit), VPU computing power (VedioProcessing Unit), etc. At present, each processing module runs an independent system, and each processing module is used as an individual. It is understandable that there is an upper limit to computing power when used as an individual, and the demand for AI computing power clusters has increased 300,000 times in the past 6 years. A single node can no longer meet the AI computing power requirements. How to make full use of the computing power of the processing modules in the array server is a technical problem that needs to be solved urgently by those skilled in the art.
[0003] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0004] Based on this, in order to solve the above problems, an array server computing power clustering method, control module, array server and medium are proposed, which can efficiently utilize all processing modules in the array server to realize computing power clustering.
[0005] This application solves the technical problem by adopting the following technical solutions:
[0006] The present application provides an array server computing power clustering method, which is applied to a control module in the array server and includes the following steps: obtaining operating information and a data packet including computing power cluster requirements, where the operating information is used to indicate the operating status of processing modules in the array server; determining a cluster mode based on the computing power cluster requirements and the operating information; dividing the data packet into multiple sub-packets according to the cluster mode control and sending them to multiple processing modules for processing; and summarizing and outputting the processing results of the processing modules.
[0007] In an optional embodiment of the present application, the computing power cluster requirement includes a computing power size requirement, and the operating information includes the number of idle modules and the remaining computing power of the processing module; the cluster mode is determined according to the computing power cluster requirement and the operating information, including: if the computing power size requirement is greater than the remaining computing power, then matching the first cluster mode, the first cluster mode is used to divide the data packet into multiple first sub-packets, and the first sub-packets are used to preferentially fill the processing module with the largest remaining computing power; if the computing power size requirement is less than or equal to the remaining computing power, and the number of idle modules is less than a preset number, then matching the second cluster mode, the second cluster mode is used to divide the data packet into multiple second sub-packets, and the second sub-packets are used to average the remaining computing power of each processing module; if the computing power size requirement is less than or equal to the remaining computing power, and the number of idle modules is greater than the preset number, then matching the third cluster mode, the third cluster mode is used to divide the data packet into multiple third sub-packets of equal size, and the third sub-packets are used to preferentially fill the processing module with the largest remaining computing power.
[0008] In an optional embodiment of the present application, the computing power cluster requirement includes at least one computing power capability requirement, which indicates the processing task that the data packet needs to perform, and the computing power capability requirement includes one of the data processing requirement, graphics processing requirement, video processing requirement, and neural network processing requirement; the processing module includes multiple computing power processors, and the computing power processors are used to complete the corresponding processing tasks. The computing power processors include: central processing unit, graphics processing unit, video processor, and neural network processor; the cluster mode is determined according to the computing power cluster requirement and operation information, including: determining the computing power processor to which the data packet will be distributed according to the computing power capability requirement to determine the cluster mode.
[0009] In an optional embodiment of the present application, when the computing power requirement includes a neural network processing requirement, the data packet is divided into multiple sub-packets according to the cluster mode and sent to multiple processing modules for processing, including: controlling the processing module to perform gradient calculation according to the preset model and sub-packets; summarizing the processing results of the processing module and outputting them, including: obtaining the parameter gradient obtained by each processing module; summarizing the parameter gradient through gradient aggregation operation for updating the parameters of the preset model and conducting the next round of training.
[0010] In an optional embodiment of the present application, the computing power cluster requirement also includes a priority, and the operating information includes the remaining computing power of the processing module; when multiple data packets are obtained, the cluster mode is determined according to the computing power cluster requirement and the operating information, including: if the remaining computing power is higher than the idle threshold, then matching the fourth cluster mode, the fourth cluster mode is used to divide the high-priority data packets into the first processing module, and divide the low-priority data packets into the second processing module, the first processing module is the processing module with the largest remaining computing power in the array server, and the second module is the remaining processing modules in the array server except the first processing module; if the remaining computing power is less than or equal to the idle threshold, then matching the fifth cluster mode, the fifth cluster mode is used to divide and process the high-priority data packets first, and then divide the low-priority data packets.
[0011] The present application also provides a control module, comprising a processor and a memory: the processor is configured to execute a computer program stored in the memory to implement the aforementioned method.
[0012] The present application also provides an array server, comprising a control module and at least one processing module; the control module is used to execute the method described above; the processing module is used to receive sub-packets sent by the control module, process the sub-packets and feed back the processing results to the control module.
[0013] In an optional embodiment of the present application, the array server also includes a network module, which is used to establish a local area network to connect the control module and all processing modules in the array server; when multiple processing modules are selected for the computing power cluster, the network module establishes a subnet between the selected processing modules to form a network isolation between the subnet and other processing modules.
[0014] In an optional embodiment of the present application, the processing module includes multiple computing power processors, which are used to complete corresponding processing tasks. The computing power processors include: a central processing unit, a graphics processing unit, a video processor, and a neural network processor.
[0015] The present application also provides a computer-readable storage medium storing a computer program, which implements the aforementioned method when the computer program is executed by a processor.
[0016] The embodiments of the present application have the following beneficial effects:
[0017] This application can divide the data packets into cluster modes according to the needs of the data packets and the specific operation status of the array server, and distribute them to the processing modules for processing, thereby making full use of the computing power of the processing modules, improving the processing efficiency of large-scale data in parallel, and reducing the computing and storage pressure on a single processing module.
[0018] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, which can be implemented in accordance with the contents of the description, and to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the following preferred embodiments are specifically described in detail with reference to the accompanying drawings. It should be understood that the above general description and the detailed description below are only exemplary and explanatory and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] in:
[0021] Figure 1 A schematic diagram of a flow chart of a method for clustering computing power of array servers provided in one embodiment;
[0022] Figure 2 A graph illustrating computing power usage provided in one embodiment;
[0023] Figure 3 A schematic bar chart showing computing power usage in a first cluster mode provided by an embodiment;
[0024] Figure 4 A schematic bar chart showing computing power usage in the second cluster mode provided by one embodiment;
[0025] Figure 5 A schematic bar chart showing computing power usage in the third cluster mode provided by an embodiment;
[0026] Figure 6 A schematic bar chart showing computing power usage in the fourth cluster mode provided by an embodiment;
[0027] Figure 7 A schematic bar chart showing computing power usage in the fifth cluster mode provided by one embodiment;
[0028] Figure 8 A schematic block diagram of the structure of a control module provided in one embodiment;
[0029] Figure 9 A schematic block diagram of the structure of an array server provided in one embodiment;
[0030] Figure 10 A schematic block diagram of a processing module structure provided in one embodiment;
[0031] Figure 11 A schematic diagram of a network model established by a network module provided in one embodiment. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] An array server is usually provided with a control module and multiple processing modules. The design scheme of applying array servers in the field of server applications is because the existing data flow processing volume is too large, which has exceeded the processing capacity limit of a single processing module. Especially in implementation scenarios such as AI, neural network learning, and image and video rendering, the demand for computing power is extremely large. In order to make full use of the computing power of the processing modules in the array server, this application proposes an array server computing power cluster method. In order to clearly describe the method provided in this embodiment, please refer to Figures 1 to 7 , including steps S110 to S140.
[0034] Step S110: Get the operation information and the data packet including the computing power cluster requirements. The operation information is used to indicate the operation status of the processing module in the array server.
[0035] In one embodiment, the method of this embodiment is applied to a control module on an array server. The control module can control all processing modules and other units within the server, such as network modules or data transmission modules. When data packets are transmitted to the server, they are often temporarily stored in the corresponding receiving unit. The receiving unit can generate a computing power cluster requirement based on the received data packet, indicating various information and status of the data packet, including but not limited to data size, capacity requirement, and priority. The control module can then simultaneously obtain operating information from all controlled processing modules within the array server. This operating information indicates the operating status of the processing modules in the array server, specifically including but not limited to occupied computing power, remaining computing power, and computing power capacity. This assists in the subsequent determination of the cluster mode to facilitate the division and distribution of data packets. It is understood that in this embodiment, the control module may not actually receive data packets, but only obtain computing power cluster requirements and operating information for judgment. This means that the control module does not actually participate in the flow of data, but only serves as a data transmission coordinator. This reduces the complexity of data transmission and the workload of the control module, thereby improving transmission efficiency and the efficiency of the control module. At the same time, for the processing module, it can be an X86 architecture or ARM architecture chip circulating on the market. There is no specific restriction on this, as long as it is a chip that can be set in the array service to provide corresponding computing power.
[0036] Step S120: Determine the cluster mode based on computing power cluster requirements and operation information.
[0037] In one embodiment, determining a clustering mode can be understood as having two meanings. The first is determining the need for clustering across multiple processing modules; the second is how to implement clustering. The first level involves determining how many processing modules are required and how much computing power each module should contribute, from the perspective of the processing modules. This can be determined directly based on the computing power cluster requirements and operational information. For example, processing modules with low occupancy, high processing power, and appropriate processing power can be prioritized. The selected processing module is then selected for online processing, ensuring that the data subsequently segmented or processed can flow between the selected processing modules. This decision must be made based on actual circumstances and is not intended to be a limitation. The determination of how to implement clustering, after determining the first level, involves determining how to segment data packets. This is the meaning of determining a clustering mode in this embodiment. This application defines multiple modes for segmenting data packets based on the computing power cluster requirements and operational information. This will be explained in more detail later and will not be elaborated on here.
[0038] In one embodiment, the computing power cluster requirement includes a computing power size requirement, and the operating information includes the number of idle modules and the remaining computing power of the processing module; the cluster mode is determined according to the computing power cluster requirement and the operating information, including: if the computing power size requirement is greater than the remaining computing power, then matching the first cluster mode, the first cluster mode is used to divide the data packet into multiple first sub-packets, and the first sub-packets are used to preferentially fill the processing module with the largest remaining computing power; if the computing power size requirement is less than or equal to the remaining computing power, and the number of idle modules is less than a preset number, then matching the second cluster mode, the second cluster mode is used to divide the data packet into multiple second sub-packets, and the second sub-packets are used to average the remaining computing power of each processing module; if the computing power size requirement is less than or equal to the remaining computing power, and the number of idle modules is greater than the preset number, then matching the third cluster mode, the third cluster mode is used to divide the data packet into multiple third sub-packets of equal size, and the third sub-packets are used to preferentially fill the processing module with the largest remaining computing power.
[0039] In one embodiment, the computing power cluster requirement may include computing power size requirements, that is, an estimate of how much computing power the data packet will occupy; the operating information includes the number of idle modules and the remaining computing power of the processing module, indicating the number of processing modules that can accept processing work, and the total computing power remaining, or the remaining computing power capacity that each processing module can handle. Based on these two, the corresponding cluster mode can be determined. The cluster mode means that, based on the computing power cluster requirements and operating information of the data packet, it is determined how to process the data packet between the determined multiple processing modules and how to divide the sub-packets to achieve the best processing effect. In this embodiment, the cluster mode can be determined based on the computing power cluster requirements and operating information. There are three cluster modes. For a specific description, please refer to Figures 2 to 5 ,in Figure 2 For the legend, Figures 3 to 5 This is an example of computing power usage in three modes. In addition, Figure 6 、 7 Likewise Figure 2 For example, I will not go into details later. Figures 3 to 5 Each graph shows four bars, each representing one of the four processing modules. The black bar represents the amount of computing power already occupied by the processing module, the white bar represents the remaining computing power, and the diagonal stripes represent the amount of computing power occupied by the processor after the data packet is divided into sub-packets. Figures 3 to 7For the sake of convenience, an embodiment of allocating 4 processing module clusters to realize computing power is adopted. In fact, when determining the cluster mode, the number of processing modules that need to be called can be determined, and this is usually determined based on the computing power cluster requirements and operating information. In this embodiment and subsequent embodiments, for the sake of convenience, an embodiment of calling 4 processing module clusters to realize computing power is adopted, which is not a limitation of the method. When the computing power size requirement is greater than the remaining computing power, that is, the computing power requirement of the currently received data packet is large, when it is determined that the called processing module may not be able to fully meet it, it can be determined as the first cluster mode. The first cluster mode means splitting the data packet into multiple first sub-packets, and the first sub-packets are preferentially input into the processing module with less computing power, that is, the processing module with large remaining computing power. Figure 3 For illustration, assume that the four processing modules are named A1, A2, A3 and A4 from left to right. It can be seen that by inferring the remaining computing power based on the size of the occupied computing power, we can get the result of A3>A1>A2>A4. Therefore, the data packet can be divided into multiple unequal first sub-packets according to the size of the remaining computing power of the four processing modules. Therefore, when inputting the processing module, among the multiple first sub-packets, the processing module with the largest remaining computing power is given priority, that is, the first sub-packet of A3 is input first, and then it is sequentially assigned to the processing module with the largest remaining computing power. It is worth noting that the inequality mentioned here and later is only for the following. Figure 3 This is an explanation of the embodiments. In practice, equal distribution may occur. These are merely illustrations of specific embodiments, not limitations. The first clustering mode allows for the utilization of less-occupied processing modules when processing modules are busy, maximizing the computing capacity of each processing module using the split first sub-packets. This fully utilizes the processing capabilities of the processing modules, maximizing clustering effectiveness and processing efficiency.
[0040] In one embodiment, if the computing power requirement is less than or equal to the remaining computing power, and the number of idle modules is less than a preset number, the second cluster mode is matched, and the second cluster mode is used to divide the data packet into multiple second sub-packets, and the second sub-packets are used to average the remaining computing power of each processing module. Figure 4 , it can be seen that before the sub-packets are split and input into the processing modules. According to the operation information, it can be determined that the processing modules selected for the cluster are generally idle and can meet the computing power requirements. In this case, in order to balance the processing capacity between the processing modules, the data packet can be split into multiple second sub-packets. The second sub-packets are input into the corresponding processing modules, which can balance the processing capacity of each processing module. That is, refer to Figure 4In this case, after a data packet is split into multiple, unequally distributed second sub-packets and input into the processing modules, the remaining computing power is equalized across the processing modules that receive the second sub-packets. This clustering mode balances the load across processing modules when clustering them. If the processing modules have the same processing power, each sub-packet is guaranteed to be output at the same time, facilitating result aggregation and making the data clustering process more controllable.
[0041] In one embodiment, if the computing power requirement is less than or equal to the remaining computing power, and the number of idle modules is greater than a preset number, the third cluster mode is matched, and the third cluster mode is used to divide the data packet into multiple third sub-packets of equal size, and the third sub-packets are used to preferentially fill the processing module with the largest remaining computing power. For this embodiment, please refer to Figure 5 For example, if all processing modules selected for clustering are idle, the data packet can be directly divided equally into multiple third sub-packets of uniform size and input into the processing modules. This embodiment is preferably applicable to machine learning scenarios. For machine learning, even if a data set is input into a predetermined model for training, since it is implemented through a processor cluster, each processing module has a sub-model, which is trained based on the corresponding sub-packet, and the training results are summarized after training. If the training conditions and scenarios within each processing module are inconsistent, it may lead to significant differences in training results, causing unnecessary trouble. To eliminate these differences, the consistency of training conditions can be controlled, that is, for example, to ensure the consistency of sub-packet sizes. Therefore, when dividing the data packet, the data packet can be divided into multiple third sub-packets of equal size for training. Therefore, based on the clustering mode provided by this embodiment, the conditions and circumstances for processing sub-packets in each processing module can be the same or similar, thereby minimizing the differences caused by differences in conditions and operating environments and ensuring the consistency or stability of the processing output.
[0042] In one embodiment, the computing power cluster requirement also includes a priority, and the operating information includes the remaining computing power of the processing module; when multiple data packets are obtained, the cluster mode is determined according to the computing power cluster requirement and the operating information, including: if the remaining computing power is higher than the idle threshold, then the fourth cluster mode is matched, and the fourth cluster mode is used to divide the high-priority data packets into the first processing module and the low-priority data packets into the second processing module, the first processing module is the processing module with the largest remaining computing power in the array server, and the second module is the remaining processing modules in the array server except the first processing module; if the remaining computing power is less than or equal to the idle threshold, then the fifth cluster mode is matched, and the fifth cluster mode is used to divide and process the high-priority data packets first, and then divide the low-priority data packets.
[0043] In one embodiment, the first to third cluster modes mentioned above are cluster modes for a single data packet. In actual application scenarios, the array server often receives multiple data packets requiring processing module clustering. For how to divide within the data packet, please refer to the description of the first to third cluster modes above. That is to say, the first to third cluster modes and the fourth to fifth cluster modes are in a parallel relationship, and the two respectively divide the data packet from two perspectives: within the data packet and between data packets. Therefore, in one embodiment, the fourth or fifth cluster mode can be determined among multiple data packets first, and then one of the first to third cluster modes can be determined for the data packet separately. This embodiment is proposed for how to cluster and split between data packets. The computing power cluster requirement also includes priority, that is, there is a priority order between the corresponding data packets; the same operating information includes the remaining computing power of the processing module. If the remaining computing power is higher than the idle threshold, the fourth cluster mode is matched. For specific implementation, please refer to Figure 6 .like Figure 6 As shown in the figure, the remaining computing power of the processing modules selected for the computing power cluster is relatively high, which meets the condition of being above the idle threshold. Specifically, it means that the processing modules selected for the cluster can realize the computing power cluster at the same time. For this, the data packets with high priority can be arranged for division first, and the data packets with slightly lower priority can be divided later. Among them, the data packets with high priority can be input into the processing modules with low occupancy and more remaining computing power first, that is, Figure 6 The two processing modules on the right of the figure; and the data packets with slightly lower or lower priority are divided into the processing modules with high occupancy and less computing power in the subsequent steps, that is, Figure 6 Therefore, based on this embodiment, when all processing modules are relatively idle, high-priority data packets can be processed first, and the control division input can be directed to the processing modules with low occupancy and more remaining computing power, so that high-priority data packets can be output earlier, reducing the processing waiting time of important data and protecting the importance of the data.
[0044] In one embodiment, if the remaining computing power is less than or equal to the idle threshold, the fifth cluster mode is matched. Figure 7The situation shown. It is worth noting that the idle thresholds used in both the fourth cluster mode and the fifth cluster mode are pre-set. Compared with the remaining computing power, in addition to the size comparison, it can also include but not precede the amount of remaining computing power that meets the idle threshold conditions, the processing capacity of the processing module, etc. In this way, a comprehensive judgment is made to achieve that the data packet can be matched to the processing module corresponding to its priority, optimize the allocation capacity, and improve the processing efficiency. Therefore, when the remaining computing power is less than or equal to the idle threshold, it means that the processing modules of the cluster that can be implemented are relatively busy and cannot process the cluster task of several data packets at the same time. For this purpose, the data packets with high priority can be divided first, and then the data packets with low priority can be divided. As Figure 7 As shown, within a processing module, the lower portion contains occupied computing power, the middle portion contains high-priority sub-packets, and the upper portion contains low-priority sub-packets. Once the occupied computing power in the lower portion is processed, the high-priority sub-packets are processed first, followed by the low-priority sub-packets. This corresponds to processing high-priority data packets first, followed by low-priority sub-packets. Therefore, based on this embodiment, even when all processing modules are busy, data packets can be processed separately in order of priority, allowing high-priority packets to be processed as quickly as possible, reducing processing wait time for important data and protecting the importance of the data.
[0045] Step S130: Divide the data packet into multiple sub-packets according to the cluster mode control, and send them to multiple processing modules for processing.
[0046] In one embodiment, the computing power cluster requirement includes at least one computing power capability requirement, which indicates the processing task that the data packet needs to perform, and the computing power capability requirement includes one of the data processing requirement, graphics processing requirement, video processing requirement, and neural network processing requirement; the processing module includes multiple computing power processors, and the computing power processors are used to complete the corresponding processing tasks. The computing power processors include: central processing unit, graphics processing unit, video processor, and neural network processor; the cluster mode is determined according to the computing power cluster requirement and operation information, including: determining the computing power processor to which the data packet will be distributed according to the computing power capability requirement to determine the cluster mode.
[0047] In one embodiment, the data packets have corresponding task requirements, which may be one or more of image rendering, video decoding, data calculation, etc. At the same time, the corresponding processing module also includes multiple computing power processors, which may specifically include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), a video processor (VPU), and a neural network processor (NPU). Therefore, when determining the cluster mode, it is also necessary to clarify the computing power requirements of the data packet, that is, to which computing power processor the data packet should be split and distributed so that it can be distributed accordingly in the end. For example, for the process of neural network training, its capability requirements may include data processing requirements, graphics processing requirements, and neural network processing requirements, but there is no video processing requirement. Therefore, when determining the cluster mode, it is necessary to clarify that the distributed sub-packets are sent to the CPU, GPU, and NPU of the processing module respectively.
[0048] Step S140: Summarize and output the processing results of the processing modules.
[0049] In one embodiment, the data processed in each processing module are all sub-packets of the data packet, and the processing results of the sub-packets are finally summarized to meet the needs of the data. For example, for video editing, the video may be split into multiple segments and handed over to multiple processing modules for encoding, decoding and rendering. However, what needs to be output in the end is the complete video, not individual segments. Therefore, the processing modules can also have a molecular-parent relationship, with the sub-processing module implementing the data processing of the sub-packets, and the processing results being handed over to the parent processing module for summary. The performance indicators such as performance, frequency, bandwidth, cache size, etc. of the parent processing module can be higher than those of the sub-processing modules, thereby achieving more efficient summary processing.
[0050] In one embodiment, when the computing power requirement includes a neural network processing requirement, the data packet is divided into multiple sub-packets according to a cluster mode and sent to multiple processing modules for processing, including: controlling the processing module to perform gradient calculations based on a preset model and sub-packets; summarizing the processing results of the processing modules and outputting them, including: obtaining the parameter gradients obtained by each processing module; summarizing the parameter gradients through a gradient aggregation operation for updating the parameters of the preset model and conducting the next round of training.
[0051] In one embodiment, neural network processing requires large-scale models and training data, making training impossible using a single card. In large-scale AI training clusters, data parallelism is often used to complete training. Data parallelism means that each device uses the same model but different training samples, and the gradient data calculated by each processing module is aggregated before parameter updates are performed. The core of data parallelism is to split the dataset by sample and distribute it to different processing modules. Each processing module calculates its own gradient based on the assigned data. Gradient aggregation ensures the consistency of the computational logic. After the gradient calculation is completed, an operator is used to implement the gradient aggregation operation between the processing modules, summing all the gradients obtained, typically by adding the gradients of all processing modules. Gradient aggregation during the parameter update phase allows the models of each processing module to enter the parameter update phase simultaneously with the same gradient value, and then proceed to the next round of training on the new data. This clustering of processing modules improves the efficiency of large-scale parallel training, thereby reducing the computational and storage pressure on individual processing modules.
[0052] Therefore, this application can divide the data packets into cluster modes according to the needs of the data packets and the specific operation conditions of the array server, and distribute them to the processing modules for processing, thereby making full use of the computing power of the processing modules, improving the processing efficiency of large-scale data in parallel, and reducing the computing and storage pressure on a single processing module.
[0053] In one embodiment, the present application proposes a control module comprising a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps: Step S110: Retrieve operational information and a data packet including computing power cluster requirements. The operational information is used to indicate the operational status of the processing modules in the array server. Step S120: Determine a cluster mode based on the computing power cluster requirements and the operational information. Step S130: Divide the data packet into multiple sub-packets based on the cluster mode control and send them to multiple processing modules for processing. Step S140: Summarize and output the processing results of the processing modules.
[0054] Figure 8 FIG1 shows an internal structure diagram of a control module in an embodiment. The computer device can be a terminal or a server. Figure 8As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the array server computing power cluster method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can implement the age recognition method. It will be understood by those skilled in the art that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0055] Figure 9 A schematic block diagram of the structure of an array server in one embodiment is shown. Array server 90 includes a control module 910 and at least one processing module 920. Control module 910 is configured to execute the method described above. Processing module 920 is configured to receive sub-packets sent by control module 910, process the sub-packets, and feed back the processing results to control module 910.
[0056] In one embodiment, the processing module 920 includes multiple computing processors, which are used to complete corresponding processing tasks. The specific structure can be referred to Figure 10 , Figure 10 This is a schematic block diagram of the structure of the processing module 920 provided in one embodiment. The computing power processor may specifically include but is not limited to: a central processing unit 921, a graphics processing unit 922, a video processor 923, and a neural network processor 924.
[0057] In one embodiment, the array server 90 further includes a network module, which is used to establish a local area network to connect the control module 910 and all processing modules 920 in the array server 90; when multiple processing modules 920 are selected for the computing power cluster, the network module establishes a subnet between the selected processing modules 920 so that the subnet is isolated from other processing modules 920. For a schematic diagram of the network model established by the network module, please refer to Figure 11 ,like Figure 11As shown, control module 910 and processing modules 1-5 are integrated within a local area network (LAN). The network module can be a 10G network switch. Two subnets are established between processing modules 1-3 and processing modules 4-5, respectively. The two subnets are isolated from each other, and no data is transferred between them. Specifically, the subnets are implemented by configuring virtual local area networks (VLANs) within the network module, thereby dividing an originally large LAN into several smaller ones. Broadcast storms within each subnet and other useless traffic are therefore confined to their own subnets.
[0058] In one embodiment, the present application further proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the aforementioned method.
[0059] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0060] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0061] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for clustering computing power of an array server, applied to a control module in an array server, characterized in that: The steps include: Obtaining operation information and a data packet including computing power cluster requirements, wherein the operation information is used to indicate the operation status of the processing modules in the array server; Determine a cluster mode according to the computing power cluster requirements and the operation information; Dividing the data packet into a plurality of sub-packets according to the cluster mode control, and sending the sub-packets to a plurality of the processing modules for processing; Summarize and output the processing results of the processing modules; The computing power cluster requirement includes computing power requirements, and the operation information includes the number of idle modules and the remaining computing power of the processing module; The determining of the cluster mode according to the computing power cluster requirement and the operation information includes: If the computing power requirement is greater than the remaining computing power, a first cluster mode is matched, where the first cluster mode is used to divide the data packet into a plurality of first sub-packets, and the first sub-packets are used to preferentially fill the processing module with the largest remaining computing power; If the computing power requirement is less than or equal to the remaining computing power, and the number of idle modules is less than a preset number, a second clustering mode is matched, where the second clustering mode is used to divide the data packet into a plurality of second sub-packets, and the second sub-packets are used to average the remaining computing power of each of the processing modules; If the computing power requirement is less than or equal to the remaining computing power, and the number of idle modules is greater than a preset number, a third cluster mode is matched. The third cluster mode is used to divide the data packet into multiple third sub-packets of equal size. The third sub-packets are used to preferentially fill the processing module with the largest remaining computing power.
2. The array server computing power clustering method according to claim 1, wherein: The computing power cluster requirement includes at least one computing power capability requirement, where the computing power capability requirement indicates a processing task that the data packet needs to perform, and the computing power capability requirement includes one of a data processing requirement, a graphics processing requirement, a video processing requirement, and a neural network processing requirement; The processing module includes multiple computing processors, which are used to complete corresponding processing tasks. The computing processors include: a central processing unit, a graphics processing unit, a video processor, and a neural network processor; The determining of the cluster mode according to the computing power cluster requirement and the operation information includes: The computing processor to which the data packet will be distributed is determined according to the computing power capability requirement to determine the cluster mode.
3. The array server computing power clustering method according to claim 2, wherein: When the computing power requirement includes the neural network processing requirement, The step of dividing the data packet into a plurality of sub-packets according to the cluster mode and sending the sub-packets to the plurality of processing modules for processing includes: Controlling the processing module to perform gradient calculation according to a preset model and the sub-package; The summarizing and outputting the processing results of the processing modules includes: Obtaining parameter gradients processed by each processing module; The parameter gradients are aggregated through a gradient aggregation operation to update the parameters of the preset model and perform the next round of training.
4. The array server computing power clustering method according to claim 3, wherein: The computing power cluster requirement also includes a priority, and the operation information includes the remaining computing power of the processing module; when multiple data packets are obtained, The determining of the cluster mode according to the computing power cluster requirement and the operation information includes: If the remaining computing power is higher than the idle threshold, a fourth cluster mode is matched, wherein the fourth cluster mode is used to classify the high-priority data packets into a first processing module and the low-priority data packets into a second processing module, wherein the first processing module is the processing module with the largest remaining computing power in the array server, and the second processing module is the remaining processing modules in the array server except the first processing module; If the remaining computing power is less than or equal to the idle threshold, the fifth cluster mode is matched, and the fifth cluster mode is used to first divide and process the high-priority data packets, and then divide and process the low-priority data packets.
5. An array server, characterized in that: comprising a control module and at least one processing module; The control module is configured to execute the method according to any one of claims 1 to 4; The processing module is used to receive the sub-packets sent by the control module, process the sub-packets and feed back the processing results to the control module.
6. The array server according to claim 5, wherein: The array server also includes a network module, which is used to establish a local area network to connect the control module and all the processing modules in the array server; when multiple processing modules are selected for the computing power cluster, the network module establishes a subnet between the selected processing modules to form a network isolation between the subnet and the other processing modules.
7. The array server according to claim 5, wherein: The processing module includes multiple computing processors, which are used to complete corresponding processing tasks. The computing processors include: a central processing unit, a graphics processing unit, a video processor, and a neural network processor.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Task scheduling method and device, equipment, storage medium and computer program product
CN113742068A
Resource allocation method and device, storage medium and electronic equipment
CN115495235A